Image to Video AI
Upload a photo or a drawing, say what should move, pick a camera move, and get a 5 second clip that starts on your picture.
Updated
Example outputs
From an AI-made anime still. Prompt: "the rain grows heavier and the wind lifts her hair and the yellow raincoat; she pulls the hood up over her head and keeps watching the city lights". Camera: slow push in. 5 s at 768P, 50 credits.
From a tall photo of an AI-made person, so the clip is tall too. Prompt: "she notices something outside the window, smiles softly and lifts the coffee cup toward her lips". Camera: static. 5 s at 768P, 50 credits.
From an AI landscape photo. Prompt: "the mist drifts slowly across the water, small ripples pass under the boat and the sunlight on the trees grows warmer". Camera: slow pan to the right. 5 s at 768P, 50 credits.
From an AI product photo. Prompt: "the shoe turns slowly on the pedestal while a soft band of light sweeps across it from left to right". Camera: static. The light sweep came through; the turn is small. 5 s at 768P, 50 credits.
How it works
Upload a photo, a drawing or an AI image (JPG, PNG or WebP, up to 10 MB). The clip keeps its shape: a tall picture gives a tall video.
Say what should move in one or two sentences, and pick a camera move: a slow push in, a still camera, a pull back, a pan or an orbit.
MiniMax H3 Max Turbo makes a 5 second clip at 768P (50 credits), usually in one to three minutes. Download the MP4.
Ideas and uses
| Camera move | What happens | Good for | How to frame the picture |
|---|---|---|---|
| Slow push in | The camera moves closer to the subject | A realization, a reveal, product shots | Start one size wider than the shot you want to end on |
| Static camera | The frame holds still; only the subject moves | Faces, small gestures, rain, steam, wind | Any picture with one clear subject |
| Slow pull back | The camera moves away and shows more of the place | Endings, loneliness, scale | Keep the subject sharp and give the edges something to reveal |
| Pan left or right | The camera turns sideways across the scene | Landscapes, streets, a crowd | Leave open space on the side the camera turns toward |
| Orbit | The camera circles the subject | Products, a hero moment, a statue | A subject in the middle with space all around it |
| Handheld follow | A loose, moving camera behind or beside the subject | Walking, running, a documentary feel | Leave room in front of the subject for where they go |
What image to video AI does
Image to video AI starts from your picture and invents what happens next. Your image becomes the first frame, and the model draws the frames after it: the hair that moves, the steam that rises, the camera that drifts closer. It does not cut out layers or slide the picture around like a slideshow. It draws new frames, so a person can turn their head and a wave can break.
It is at its best with one subject and one clear motion: a face that changes expression, cloth and hair in the wind, water, smoke, rain, light that shifts, a slow camera move through a landscape. It struggles with big actions that need many steps, hands that grab small objects, readable text, and people who appear from outside the frame. Plan for those in the edit instead.
Write the motion like a director
The picture already shows who and where. The prompt only needs the moment that follows it. Three rules from animation storyboards help a lot:
- One action and one camera move per shot. "She looks up and smiles" works. "She looks up, runs to the door and opens it" asks for three shots.
- Start from the instant before the action. A sword still in its sheath, a breath before a laugh. The model animates the release, so the picture should hold the wind-up.
- Name the small motion. Hair, a scarf, leaves, rain, dust in the light. Secondary motion is what makes a clip feel alive.
Weak: "make it move, cinematic, epic". Better: "the wind pushes her scarf to the right, she narrows her eyes against it and takes one slow breath".
Pick a picture that wants to move
- Sharp and well lit. A blurry or dark picture gives a blurry or dark clip.
- Leave room for the move. For a push in, start wider than the shot you want. For a pan, leave space on the side the camera turns toward.
- Off center, with space where the subject looks. A subject on a third line with open space ahead of them gives the motion somewhere to go.
- A clean first frame. No motion blur, speed lines, captions or watermarks in the picture. Add titles later in the edit.
- Faces big enough to read. A tiny face far away in a crowd will not keep its features.
From one clip to a whole scene
A 5 second clip is one shot. A scene is several shots that cut together: a wide that shows the place, a closer shot of the person, a detail, a reaction. In Ziztu Studio the Director plans those shots from one idea, keeps your characters the same with character sheets, and puts the clips, music and captions on one timeline.
Frequently asked questions
Is this image to video AI free?
You can try it free. A new account gets 40 free trial credits. On the trial a clip is made at 480p for 31 credits and carries a small watermark. With a plan, a 5 second clip at 768P costs 50 credits and has no watermark. Plans start at $12 a month for 1,000 credits.
How long is the video?
Each clip is 5 seconds long. In Ziztu Studio one shot can run up to 15 seconds, and a film is built from many shots on a timeline.
Does the video have sound?
Yes. The model adds sound that fits the picture, such as rain, wind, traffic or a room tone. It is quiet in calm scenes, and there is no voice-over or music. Add those in an editor or in Ziztu Studio.
Can I animate a drawing, an anime picture or an AI image?
Yes. Any picture works as the first frame: a photo, a scan of a drawing, a manga panel or an image you made with AI. Clean line art and a clear subject give the most stable motion. Pictures with motion blur or speed lines are harder to animate.
Will the face and the style stay the same?
The first frame is your picture, and the prompt asks the model to keep the faces, colors and look. Big movements and turned heads can still change a face a little. For a character who must look the same in many shots, Ziztu Studio sends character sheets with every clip.
What should I write in the prompt?
One main action and the small motion around it. For example: "she turns toward the window and smiles while the wind lifts her hair". Describe the moment that starts from the picture, not a whole story. The camera move comes from the option above.
Which pictures can I upload?
JPG, PNG or WebP up to 10 MB that you have the right to use. Our terms ask for the consent of anyone in the photo and do not allow sexual content. Prompts with sexual words are blocked, and the model has its own safety filter.
Is my picture stored?
Your uploaded picture is deleted from our server within 24 hours. To make the video, it is sent to our AI provider, fal, which keeps its own copy under its own retention rules.
Which AI model makes the video?
MiniMax H3 Max Turbo image to video. Ziztu Studio also has Kling, Seedance and Grok Imagine video models for the same job.