Hedra
  • Enterprise
  • Pricing
  • Blog
  • Creators
Log inSign Up
Open Hedra
Your account
    Explore
  • Home
  • Enterprise
  • Pricing
  • Blog
  • Creators
    Log inSign Up
    Open Hedra

Kling O1

All video models
Video modelKling
OverviewImage → VideoFirst & last frame → VideoReference → Video

Overview

Kling O1 is a unified multimodal video model developed by Kuaishou. Built on a Multimodal Visual Language framework, it consolidates text-to-video creation, reference-based generation, and native video editing into a single engine. The model is highly effective for precise video-to-video transformations, allowing users to swap subjects, modify backgrounds, or restyle existing footage using natural language prompts while maintaining strict temporal consistency.

Kling O1 Image to Video

Image → Video — generates video.

Specifications

Input mode
Image → Video
Accepts
reference image (up to 3), start frame, end frame
Aspect ratios
16:9, 1:1, 9:16
Durations
3s, 4s, 5s, 6s, 7s, 8s, 9s, 10s
Max duration
10s
Native audio
No
Pricing
20 credits / second — longer clips and higher resolutions cost more
Typical generation time
~79s
Free tier
Yes

Image → Video examples

A close-up vertical video frame of a young woman with brown hair and bangs, wearing shimmering silver eyeshadow and a silver chainmail halter top. She poses with her hand near her face against a warm yellow backdrop. This 1076x1928 video was generated using the Kling O1 model on Hedra.Woman in Silver Halter Top — Kling O1A vertical 1080x1916 video frame generated by Kling O1 shows a young woman with dark braids, wearing a pink knit sweater-dress, a black beanie, and black winter boots. She walks on a snowy city sidewalk during winter, with classical buildings, cars, and other pedestrians blurred in the background.Winter Street Walk — Kling O1Vertical video first frame of a man in a grey suit and sunglasses sitting at a wooden desk on a grassy hill, talking on a vintage mobile phone. Behind him is a massive, snow-covered mountain peak under a clear blue sky. Generated with Kling O1 at 1076x1928 resolution.Businessman on Mountain Peak — Kling O1A bearded man with dreadlocks wearing a flat cap sits in a green leather armchair, smoking a cigarette. He is dressed in a white polo, suspenders, and brown trousers against a textured green studio backdrop, generated by the Kling O1 model at 1928x1072 resolution.Man Sitting in Armchair, by Kling O1A high-angle view of six young adults wearing flannel shirts, beanies, and jeans posing around a white classic convertible car with a blue interior. Generated by the Kling O1 model at 1928x1072 resolution, this image-to-video scene captures the group on an asphalt lot under direct sunlight.Group in Flannel with Classic Convertible — Kling O1

Kling O1 First & Last Frame

First & last frame → Video — generates video.

Specifications

Input mode
First & last frame → Video
Accepts
reference image (up to 3), start frame, end frame
Aspect ratios
16:9, 1:1, 9:16
Durations
3s, 4s, 5s, 6s, 7s, 8s, 9s, 10s
Max duration
10s
Native audio
No
Pricing
20 credits / second — longer clips and higher resolutions cost more
Typical generation time
~2 min
Free tier
No

First & last frame → Video examples

A close-up portrait of a model with extremely glossy, wet-look skin and light green eyes, generated using Kling O1 at 1076x1928 resolution. The model has blonde hair, bleached eyebrows, and wears a sheer white high-collar top adorned with pearls under bright studio lighting.Glossy Skincare Portrait — Kling O1A video first frame showing a young man with dark hair, wearing a green suit and dress shoes, reclining on a bed of fluffy white clouds against a clear blue sky. Generated with the Kling O1 model at 1928x1076 resolution, this image-to-video asset captures a low-angle perspective.Man Reclining on Clouds — Kling O1A vertical video frame shows a person wearing an oversized bright orange satin coat and patterned trousers, standing in a classical courtyard. Their face is obscured by a long, hanging beaded veil integrated into their dark braided hair. This 1080x1920 video was generated using the Kling O1 model.Model in Orange Satin, Generated by Kling O1A first frame of a 1928x1076 video generated by Kling O1, showing a young man in an olive-green suit lying on a dense bed of white clouds. The low-angle perspective highlights his black shoes in the foreground, set against a clear blue sky.Man Floating in Clouds, by Kling O1

Kling O1 Reference to Video

Reference → Video — generates video.

Specifications

Input mode
Reference → Video
Accepts
reference image (up to 3), start frame, end frame
Aspect ratios
16:9, 9:16, 1:1
Durations
3s, 4s, 5s, 6s, 7s, 8s, 9s, 10s
Max duration
10s
Native audio
No
Pricing
20 credits / second — longer clips and higher resolutions cost more
Free tier
Yes

What is Kling O1 best used for?

Kling O1 is Kuaishou's first unified multimodal video model, making it exceptionally good at native video-to-video editing. Instead of just generating clips from scratch, it allows you to modify existing footage using natural language prompts—such as swapping backgrounds, changing a character's clothing, or shifting the lighting from day to night—without complex masking. Because of its fast generation speed and high prompt adherence, the community has dubbed it the "Nano Banana of AI video."

When was Kling O1 released and what is its lineage?

Kuaishou officially launched Kling O1 on December 1, 2025, kicking off its "Omni" lineup of unified multimodal models. It arrived shortly after the 2.x generation, such as Kling 2.6 Pro, and represented a major architectural shift from pure generation to a hybrid generation-and-editing engine. It was eventually succeeded by the 3.0 generation in February 2026, which includes standard models like Kling V3 Pro and the next-generation omni model, Kling O3 Pro.

How can I get the most consistent character edits with Kling O1?

To maintain strict character consistency while modifying footage, leverage Kling O1's multi-element reference capabilities by uploading reference images alongside your video. When prompting, use a conversational approach rather than traditional keyword stuffing. Because the model processes text, image, and video in a shared semantic space, clear instructions like "remove all cars from the street" or "change the protagonist's outfit to a red dress" yield much better results than comma-separated tags.

Similar models

Kling O3 ProKlingKling O3 StandardKlingHappy HorseAlibabaKling V3 ProKlingKling V3 StandardKlingHedra OmniaHedra

Prompt tips

  • Use numbered references: Explicitly define relationships between your uploaded images by using @Image1, @Image2, etc., directly in your text prompt.
  • Anchor with real footage: For video-to-video edits, start with high-quality stock video as your foundation to provide structural detail and maintain a realistic look.
  • Leverage the Elements feature: Upload a clear, frontal image of your subject to lock in a 3D-consistent actor across multiple scenes.
  • Use the constraint sandwich: Place your most critical constraints (like character identity or specific actions) at both the beginning and end of your prompt to prevent the model from drifting.
Illustration of a laptop with the Hedra spark logo in front of a city skyline at sunset

What Will You Create?

Sign up for free

Product

StudioCommunityFeedbackUse CasesModels

Legal

Privacy PolicyTerms of useAcceptable useCookie PolicyBiometric data policy

Company

AboutTeamChangelogCareersCreatorsSupportAlternatives
LinkedinInstagramDiscord
support@hedra.comHedra 2026 — All rights reserved