Turn any image into video with AI.

Animate a product shot, a portrait, or a frame you've already signed off — and the subject comes through intact. For ads, catalogs, localization, and the video features inside your own app.

Image in
Amber honey frozen mid-pour from an unlabeled glass jar onto a ceramic spoon, backlit against deep brown

"The pour resumes — honey ribbons onto the spoon and over the edge."

PNG · 1280 × 720

Your frame, in motion.

0:15 · 1080P

The frame you approved is the frame that moves.

Your still isn't a suggestion to the model — it's the frame the clip opens on. Product geometry, a face, the light someone spent a day getting right: all of it survives into motion. That's the difference between footage a brand can run tomorrow and a lookalike somebody has to explain.

“Golden hour to blue hour, lights coming up across the city”KLING V3 · START + END FRAME

Two frames and it knows the shot.

Say where the shot starts and where it ends and the model fills the middle. Hand it a reference clip instead and it copies the move; hand it reference images and one subject holds from shot to shot. In Studio that's a second file dropped in, in a request it's another URL — either way the picture is how you stop guessing.

“Abrimos hace tres años con una sola máquina.”VEO 3.1 · 0:08 · SPANISH

Hand the photo a voice.

Point a portrait at audio — a script you typed, a voice you cloned, a recording off your phone — and the face performs it in whatever language the market speaks. One photo covers ten of them, and the take holds for as long as the audio runs.

The right model for every still.

Nearly every model in the catalog starts from a picture — some animate a start frame, some fill the gap between two, some hold one subject across shots from reference images. Let the agent route each job or name the model yourself: in Studio it's a dropdown, in a request it's a string, and it's priced per second of output either way.

Barista laughing with a regular at a sunlit Roman espresso barKLING V3
Man with blond dreadlocks eating noodles in a red-lit restaurant, cards suspended mid-airWAN 3.0
Retro cassette player with headphones on a sunlit dresser, dust in the lightMINIMAX H3
Man in a powder-blue suit loading plush carnival animals into a cream sedanVEO 3.1
Rainforest canopy after rain with layered palm fronds and hanging mangoesSEEDANCE 1.5
Man standing on the Great Wall of China at golden hourLTX-2.3
Hand with chrome-blue nails pouring a pink drink over ice on pink tilesVIDU Q3
Man in a green velvet jacket lying among cushions covered in sleeping catsSEEDANCE 2.0
Close-up of a canned drink being poured against pink tileworkFLUX.3
Red velvet listening room with a turntable built into a sports-car wheelKLING O3
Runners crossing a marathon finish line beneath a FINISH bannerPIXVERSE V6
Skier in a mustard jacket carving through deep powder against a cobalt skyHAILUO 2.3

Built for visual inference.

Serving video is a different problem from serving text — a single request can saturate a GPU, and none of the tricks that made language models cheap apply. Hedra's engine was built for exactly that workload.

Whichever model you choose — ours or anyone else's — it runs on the same infrastructure, tuned for visual inference and priced per second of output.

Generate with API or Agent

Build with the API

One key, every model, priced per second of output.

Start building now

Animate it now.

No code — drop in an image, describe the motion, and export the result.

Create with Agent

FAQs