Skip to main content

Image to Video

Image to Video animates a still image into a moving clip. You provide a first frame (and optionally a last frame), describe the motion you want, and generate. It's ideal when you already have an image and want to bring it to life.

L
Written by LX

Image to Video animates a still image into a moving clip. You provide a first frame (and optionally a last frame), describe the motion you want, and generate. It's ideal when you already have an image and want to bring it to life.

What does Image to Video do?

It takes a source image as the starting frame of a video and animates it according to your description. A hint suggests "We recommend using an image generated by APOB, but feel free to choose any you like." You can also add a last frame so the clip ends on a specific image.

Paid users start with UltraS 720P selected. Free users keep the Fast 480P starting option.

How do I add my source (first frame) image?

In the reference area, you have two buttons:

  • Select content — pick an existing image from your own generations (followed by curated inspiration) or the community feed.

  • Upload image — upload your own file.

Until you add one, the area shows "Your selected reference image will appear here once you select content or upload an image." The first frame "sets how your video starts." If you're not logged in, choosing Select content opens the sign-in dialog first.

What is the last frame and how do I add one?

The last frame "sets how your video ends." Click Add last frame to attach an image as the ending frame. A hint warns: "Last frame sets how your video ends. If it doesn't connect naturally to the first frame, the result might feel random." To remove it, click Remove last frame.

Can I generate a last frame instead of uploading one?

Yes. The last-frame area has a Generate last frame from first frame option that opens a dialog where you create an ending image from your first frame. Once generated and selected, it's attached as the last frame and the view scrolls back to the top.

Why is "Add last frame" greyed out?

Last frame is only supported on certain quality modes and only for single-shot videos. Hovering the disabled button explains which quality you need. You'll also see related notes like "Last frame is only supported in Ultra mode" and "Last frame is only supported for single-shot video generation." Clicking the Generate last frame from first frame option will switch you to a supporting quality automatically.

Can I use a last frame and multiple shots together?

No. Adding a last frame disables multi-shot, and vice versa. The hint reads "Multi-shot generation is unavailable when a last frame is added. Remove the last frame to use multi-shot generation."

What is the description box for?

The description tells the model how to animate the image — for example a pose or movement. The placeholder reads "Describe how the image should move. E.g., The subject slowly turns toward the camera, smiles, adjusts their jacket, and looks out the rain-streaked window while their hair moves gently and the camera gradually zooms in." In No Model mode, it instead reads "Describe how the character in your image should move. E.g., A young woman with warm olive skin, shoulder-length wavy auburn hair, green eyes, and light freckles slowly turns toward the camera, smiles, adjusts her jacket, and looks out the window while her hair moves gently and the camera gradually zooms in." On a quality tier that doesn't support descriptions, the box is replaced by a prompt to switch: it shows "Description" then "only available under" and a button with the supporting quality names you can tap to switch.

What does the Template button do?

Template is the pill button at the right end of every Description editor's bottom button row on desktop, on the same line as the + Element / + Voice model / + Camera triggers. On mobile it sits on its own right-aligned row above those shortcuts, separated from them by a grey divider (the divider hides while the Elements panel is expanded). Each shot's pill opens the same picker. It opens a picker of videos made with Image to Video — public ones under Community, your own creations under Generation. The picker always opens on the Community tab (templates are inspiration-first, even when you have generations of your own); switch to Generation to pick from yours. Every card shows the video's cover and auto-plays its muted video preview (one card at a time in feed order on mobile, while hovered on desktop), with a preview of its prompt near the bottom edge (trimmed with an ellipsis when it's long) and an info icon in the card's top-right corner. Clicking the info icon opens a preview dialog: the video autoplays on a loop beside the full prompt — its text with its style and element pills — and a Use template button. The player has no native controls; the sound button in its bottom-right corner toggles sound (muted by default — the on/off choice is shared with the app's other video players).

Clicking a card — or Use template in the dialog — applies that video as a template, and this completely replaces what you currently have in the composer: the source (first frame) image, the last frame (restored from the template, or cleared when it has none), every shot's description, camera movement, the filled element slots, Quality, resolution, duration, and the Native audio / voice-model state are all overwritten with the template's stored generation settings. (A template whose voice model isn't supported on the restored quality drops the voice model and tells you so.) A "Content settings applied." toast confirms it and the picker closes — so save anything you want to keep before picking a template.

A card that is still generating shows "Generating. Can't select right now." and can't be picked. If a template's details fail to load, a "Failed to load content detail…" error appears and nothing in your composer changes.

What is "Native audio" / the audio section?

On desktop and mobile, the audio card pairs a Native audio toggle with the caption "Synchronized audio-video generation." Turning the toggle on generates sound synced to the video. It's required on the highest tier (you'll see "Native audio is required for UltraS quality" if you try to turn it off there). Image to Video keeps voice-model selection with the Description references instead of embedding it in this card.

What is a voice model and how do I add one?

When native audio is on, you can attach a Voice model so spoken lines use a specific voice. Voice models are available on Ultra and UltraS; on Fast, the purple Voice model control is greyed out and hovering it shows "Only available for Ultra / UltraS". Clicking or tapping the greyed-out control switches directly to Ultra, then continues into the Voice model panel or selector without requiring a second click. After a Voice model is selected or mentioned, switching to Fast is blocked until the model and all of its mentions are removed; the warning says "Voice models aren't available for Fast. Remove the voice model and its mentions before switching." Reusing or remixing legacy Fast content automatically drops its saved Voice model and removes Voice model references from every shot. When anything is removed, a toast says "Voice model data was removed because Fast doesn't support Voice models." On desktop, use the Voice model button in the description footer. On mobile, tap Voice model below a Description or use the Voice mode category inside the opened reference panel; both entry points share one selected model and use the same quality-upgrade behavior. While Element, Voice mode, and Camera are all empty, the footer shortcut opens the selector directly and completing a selection opens the categorized panel automatically. Once any of those references exists, every footer shortcut — including Voice model — only reopens that panel. The collapsed and expanded Voice mode cards keep their purple styling but appear dimmed while Audio is off or the current quality does not support Voice models. With an empty panel, tapping the dimmed footer shortcut turns Audio on, shows "Turn on Native audio to add a voice model." when Audio was the missing prerequisite, and opens the shared selector. Once the panel has content, use its dimmed Voice mode control for that warning-and-enable action. The selected Voice mode card can be tapped to insert that voice into its shot once Audio is on, and typing "@" still lets you choose the voice at the cursor. The old selected-voice pill above every mobile description is not shown. You can't turn audio off while a voice model is still referenced — you'll see "A voice model is in use in the description. Remove the voice model mentions before turning off native audio." Clearing the selected model asks before deleting those references.

What are Elements here?

Elements are reusable characters, products, or backgrounds you can mention in your description. Elements are available on the Ultra and UltraS quality tiers — up to 15 Elements on Ultra and up to 25 on UltraS. On desktop, the pink Elements header shows how many you've used out of the current quality's allowance (e.g. "(0/25)"). On mobile, the closed Description footer shows 24px-high Element, Voice model, and Camera shortcuts. While all three categories are empty, each opens its matching picker directly; after the first successful add or selection, the Storyboard-style categorized field opens automatically inside the Description card. Once any reference exists, closing the field and tapping any of the three shortcuts only reopens it instead of launching another picker. A grey Elements close bar appears first, then Element with its allowance count, followed by Voice mode and Camera, each with "(0/1)" or "(1/1)". Use the close icon in the grey Elements bar to hide the categorized field again; the Element category label itself has no close icon. Expanded Voice mode and Camera cards fill the row; compact controls inside the visible field stay 22px high. If the current quality does not support Elements, tapping the greyed Element button switches directly to the lowest supporting quality before opening the Element picker. Reusing or remixing content with saved Elements also opens the field automatically so the restored pool is immediately visible. The Elements category shows every slot the current quality allows, plus one disabled slot whenever a higher tier unlocks more. The disabled slot carries a per-slot hint such as "Element 16 only available for UltraS." Tapping it switches automatically to the lowest quality that supports that slot and reveals that tier's full slot list at once. On an unsupported tier, the category shows one locked slot and no "(0/0)" count. Empty Element slots use a + prefix; selected Element chips use @ before the thumbnail. The chevron uses the same compact style as Text to Video, and its full list floats over the lower part of the Description card without being clipped inside the card or appearing above the handle. In multi-shot mode, opening the field from any shot shows it in every shot. Camera selection and saving or inserting an Element still target the shot where you use the control, and tapping a selected Voice mode card inserts it into that shot, while Element edits, deletions, and the selected voice model remain shared across the generation. See Using elements in prompts and Elements.

What happens if I switch quality after adding elements?

If the new quality supports fewer element slots than the slots you've filled, a "Some elements will be removed" dialog appears: "{quality} supports up to {count} elements. Switching now will remove {removedCount} extra elements and their mentions from the prompt." Choose Continue to switch or Cancel to keep your current quality.

Switching between two element-supporting tiers doesn't clear anything when the filled slots still fit. Going from UltraS (up to 25) down to Ultra (up to 15) with Elements in slots 16-25 asks first, then keeps the first 15 slots and removes the overflow Elements, along with their mentions in your descriptions, only if you continue.

How do multi-shot videos work?

When supported, click Add shot to add shots, each with its own description and camera movement. On mobile, opening the categorized Element / Voice mode / Camera field from any shot shows it in every shot; Camera selection and Element or Voice mode insertion still apply to the shot where you use them, while the saved element pool and selected voice model are shared. The maximum number of shots depends on duration: "{duration}-second videos support up to {count} shots. For more shots, increase duration to {nextDuration} seconds." If multi-shot isn't supported on your quality, the Add Shot button's tooltip points you to a supporting quality.

Which durations can I choose?

Fast offers 4s, 5s, and 10s. Ultra and UltraS offer 5s, 10s, 15s, 20s, 25s, and 30s. On desktop, unavailable durations remain visible but disabled with the quality tier that unlocks them; on mobile, the duration menu only shows options supported by the selected quality.

How much does it cost?

Video is priced per second of output, so the cost scales with your duration. The per-second rate is shown next to Generate and depends on your plan and the quality you pick (it also reflects resolution and options like native audio). See Credits & billing.

Which resolution does Fast use?

Fast offers 480P, 720P, FHD, and QHD. The 480P option is powered by MiniMax H3, supports 4-second, 5-second, and 10-second output, and is the default selection for free users. On desktop, the duration menu still shows unavailable 15s, 20s, 25s, and 30s options as disabled with the quality tiers that unlock them; on mobile, unavailable durations are filtered out, so Fast shows 4s, 5s, and 10s. Hover the Fast quality control on desktop, or tap its info icon on mobile, to see "Powered by MiniMax H3"; MiniMax H3 is emphasized in the tooltip. Paid users keep the established 720P Fast default unless they choose another resolution.

On mobile

Image to Video works on mobile with the same first-frame / last-frame, description, audio, and elements controls arranged for a narrow screen. Its Audio card keeps the dedicated Native audio toggle and "Synchronized audio-video generation" caption; unlike Text to Video, it has no inline voice-model selector. Element, Voice model, and Camera shortcuts live below each Description and open their pickers directly. Completing an add opens their shared categorized field, where the selected voice is shown in Voice mode and can be inserted into that shot; the old duplicate pill above each description editor is hidden. Quality, resolution, and duration live in the bottom bar. The Template pill sits on its own right-aligned row under every Description editor's counter, separated from the Element / Voice model / Camera shortcuts by a grey divider (hidden while the Elements panel is expanded); its picker opens as a full-screen page where tapping a card's info icon opens a preview sized to its content (up to near-full screen) — the playing video above, the full prompt and Use template button below — and tapping the card applies the template.

Did this answer your question?