Text to Video turns a written description into a generated video clip — no source image or video needed. You write a prompt describing the pose, movement, scene, and any elements you want, then generate.
What does Text to Video do?
It creates a video from your words alone. You describe what should happen — for example a pose, a movement, or any element you want in the shot — and the model generates a matching clip. The prompt field shows this placeholder: "Describe your image with @influencer and styles. E.g., A polished editorial portrait set in a Cafe on a Rainy Day, with @influencer wearing a Hot summer clothing walking from the counter toward a rain-streaked window and pausing to look outside, rendered in the Realistic with soft ambient light and camera movement set to Follow Char." Both @influencer references appear as model pills; Cafe, Rainy Day, Hot summer, and Realistic appear as modifier pills with grayscale catalog-image backgrounds, while Follow Char appears as a camera-movement pill. In No Model mode, the placeholder instead reads: "Describe your video with detailed character features, styles, and camera movement. E.g., A poised young woman with warm olive skin, shoulder-length wavy auburn hair, green eyes, and light freckles enters a Cafe wearing a tailored charcoal Suit on a Rainy Day. In the Realistic art style, she walks toward a rain-streaked window, pauses, and looks outside while the Follow Char. camera movement tracks her." It contains no model pill; Cafe, Suit, Rainy Day, Realistic, and Follow Char. retain their catalog-backed pills. Placeholder pills use the same dimensions as pills entered in the editor. The first shot starts with an empty prompt: the example is not written to the editor or submitted with the prompt.
How do I generate a video from text?
Type your description into the prompt editor.
Use the controls on the left (the tag sidebar) to add an art style or a camera movement, and to set up elements.
Pick your quality (Fast, Ultra, or UltraS), resolution, and duration from the workstation's bottom bar.
Click Generate. The credit cost is shown next to the Generate button.
Paid users start with UltraS 720P selected. Free users keep the Fast 480P starting option.
What is the prompt box and is there a length limit?
The prompt box is a rich editor where you type your description and drop in tags (camera movement, art style, elements, a selected voice model, and a person indicator when a real portrait model is selected). In No Model mode the editor does not insert a technical default-model name or avatar. There is a per-shot character limit — when you paste or type past it, a notice tells you your prompt "exceeded the character limit and was truncated," and a live counter appears near the bottom-right of the editor showing how many characters you've used out of the maximum. The visible text isn't blocked, but only the text within the limit is actually used to generate.
What does the Template button do?
Template is the pill button at the bottom-right of every shot's prompt editor (when the character counter is visible, it sits just below it); each shot's pill opens the same picker. It opens a picker of videos made with Text to Video — public ones under Community, your own creations under Generation. The picker always opens on the Community tab (templates are inspiration-first, even when you have generations of your own); switch to Generation to pick from yours. Every card shows the video's cover and auto-plays its muted video preview (one card at a time in feed order on mobile, while hovered on desktop), with a preview of its prompt near the bottom edge (trimmed with an ellipsis when it's long) and an info icon in the card's top-right corner. Clicking the info icon opens a preview dialog: the video autoplays on a loop beside the full prompt — its text with its style and element pills — and a Use template button. The player has no native controls; the sound button in its bottom-right corner toggles sound (muted by default — the on/off choice is shared with the app's other video players).
Clicking a card — or Use template in the dialog — applies that video as a template, and this completely replaces what you currently have in the composer: every shot's description (a multi-shot template restores all of its shots), the art style and camera movement tags, the filled element slots, Quality, resolution, duration, aspect ratio, and the Native audio / voice-model state are all overwritten with the template's stored generation settings. A "Content settings applied." toast confirms it and the picker closes — so save anything you want to keep before picking a template.
Some cards can't be applied: one that is still generating shows "Generating. Can't select right now.", and a video whose settings can't be restored — for example one styled only with art styles while the style catalog is unavailable — shows "This content can't be reused." and the picker stays open for another pick; in that case nothing in your composer changes. If a template's details fail to load, a "Failed to load content detail…" error appears and nothing in your composer changes.
What is "Native audio"?
"Native audio" is a toggle in the Audio section that generates sound together with your video. It's available on every Text to Video quality (Fast, Ultra, and UltraS), but it behaves differently on the highest tier:
On UltraS quality, native audio is required and stays on — if you try to switch it off you'll see "Native audio is required for UltraS quality."
On Fast and Ultra you can freely turn it on or off.
How do I add a voice model?
Turn Native audio on, then click Voice model on desktop or Add voice model on mobile in the Audio section. Voice models are available on Ultra and UltraS, but not Fast (including Fast 480P); native audio itself can still be used on Fast. On Fast, the Voice model control is dimmed and desktop hover text says "Only available for Ultra / UltraS". Clicking or tapping that control switches to Ultra and enables the Voice model flow. Switching to Fast is blocked while a selected Voice model or any Voice model mention remains; the warning says "Voice models aren't available for Fast. Remove the voice model and its mentions before switching." Reusing or remixing historical Fast content removes its saved Voice model and Voice model mentions from every shot, with the toast "Voice model data was removed because Fast doesn't support Voice models."
The voice-model selector reuses the app's standard picker, including filters, No voice model, and Create Voice Model. Once selected, the voice appears as a pill above each shot's description and in the @ mention list; click the pill or choose it after typing @ to insert a voice chip at the cursor. Switching the selected model updates existing voice chips. If a voice chip is still in any shot, native audio cannot be turned off until the chip is removed, and clearing the selected voice asks before deleting those references.
What are art style and camera movement, and why are they in the sidebar?
The left settings card is split into Audio, Presets, and Element sections. Art style sets the overall look and camera movement adds motion like a pan or zoom inside Presets. A hint explains the rule: "Each shot allows one camera movement selection, while art style is set once globally in the first shot only." So you can choose a different camera movement per shot, but the art style applies to the whole video and can only be set on the first shot.
Why can't I change the art style on a later shot?
Art style is global and is taken from the first shot only. If you try to set it on a second or later shot, you'll see "Art style can only be added at the first shot." Select the first shot to change the art style for the whole video.
What are Elements and why are they greyed out?
Elements are reusable people, products, or backgrounds you can reference inside your prompt. They are available on the Ultra and UltraS tiers — up to 15 Elements on Ultra and up to 25 on UltraS. On desktop the pink Element header shows how many Elements you've used out of your quality's allowance (e.g. "(2/25)"), and you see every slot your current quality allows, plus one greyed Add element card whenever a higher tier unlocks more. The greyed card shows a per-slot hint such as "Element 16 only available for UltraS." — hover to read it, or click the card to switch automatically to the lowest quality that unlocks it, which reveals that tier's full slot list at once. Mobile uses the same allowance-plus-one rule instead of rendering all 25 slots at once. Tapping its greyed card shows the same hint as a toast, switches automatically to the lowest unlocking quality, and reveals that quality's allowance plus the next locked card; the card list scrolls vertically when it is taller than the popup. If switching quality would remove filled Element slots — for example dropping from UltraS to Ultra while slots 16-25 contain Elements — a "Some elements will be removed" dialog appears: "{quality} supports up to {count} elements. Switching now will remove {removedCount} extra elements and their mentions from the prompt." Choose Continue to switch or Cancel to keep your current quality. See Elements for the full picture.
How do multi-shot videos work?
When your quality supports it, you can build a video out of several shots, each with its own prompt and camera movement. Click Add shot to add another. The number of shots you can add depends on your video duration — a hint reads "{seconds} second videos support up to {shots} shots. For more shots, increase the video duration." Multi-shot is only offered on the higher tiers ("Multi-shot generation is supported for Ultra and UltraS quality.").
How do I remove a shot?
Each added shot has a small header showing "Shot N." Click that header to remove the shot. The remaining shots renumber automatically.
Why is the sidebar greyed out with "Please select a shot to edit"?
In multi-shot mode the Presets and Element controls apply to whichever shot is active. If no shot is focused, those controls dim and a tooltip says "Please select a shot to edit." Click into a shot's prompt editor to make the controls apply to it.
How much does it cost?
Video is priced per second of output, so the cost scales with your duration. The exact per-second rate is shown next to the Generate button and depends on your plan and the quality you choose (it also reflects options like resolution and native audio). For example, a live capture showed a rate of 160 credits per second, so a 5-second clip cost 800 credits. See Credits & billing.
Which resolution and duration does Fast support?
Fast offers 480P, 720P, FHD, and QHD, with 4s, 5s, and 10s durations. The 480P option is powered by MiniMax H3 and is the default Fast resolution for free users; paid users retain the established 720P Fast default unless they choose another resolution. Hover the Fast quality control on desktop, or tap its info icon on mobile, to see "Powered by MiniMax H3"; MiniMax H3 is emphasized in the tooltip.
Which durations do Ultra and UltraS support?
Ultra and UltraS offer 5s, 10s, 15s, 20s, 25s, and 30s. On desktop, unsupported durations remain visible but disabled with the quality tier that unlocks them; on mobile, the duration menu only shows options supported by the selected quality.
On mobile
Text to Video works on mobile with the same prompt editor and tags, laid out for a narrow screen. The Template pill sits at every description editor's bottom-right corner here too; its picker opens as a full-screen page where tapping a card's info icon opens a preview sized to its content (up to near-full screen) — the playing video above, the full prompt and Use template button below — and tapping the card applies the template. The Audio card and voice-model selector appear above the descriptions, while quality, resolution, and duration live in the workstation's bottom bar.
Related: Quality, resolution & duration · Image to Video · Voice models · Elements · Credits & billing