T2VA
Text only: describe the complete audiovisual timeline from scratch.
Free deterministic prompt builder
H3 generates video and synchronized audio together. A useful prompt therefore describes an ordered timeline—not a pile of visual adjectives—then separates physical sound from audience-only music. Build the official three-field base format or six-section Ref2VA format below.
This page is a writing and formatting tool. It explains all five documented H3 modes, but does not claim that every mode is enabled in Tinulo's current generation runtime. Preset links open existing creation starting points only.
Runs entirely in your browser
Choose the inputs you actually have. The builder adds the correct keyframe preamble and preserves the exact official field order; it does not send your text to a model or rewrite it unpredictably.
T2VA needs no asset. I2VA, FL2VA, and L2VA align keyframes to exact times. Ref2VA is different: it uses six ordered rewrite sections and explicit subject/reference relationships rather than the three base fields.
T2VA
Text only: describe the complete audiovisual timeline from scratch.
I2VA
First frame: Picture 1 is fixed at 0.00 seconds; describe only the motion that follows.
FL2VA
First and last frames: connect Pictures 1 and 2 through one plausible continuous path.
L2VA
Last frame: invent a compatible opening that converges precisely on Picture 1 at the ending timestamp.
Ref2VA
Full reference: define and selectively retain subjects, motion, visual style, video, or audio from labeled assets.
Begin with [Shot 1]. Give later cuts increasing timestamps such as [Shot 2] At 00:05.000. A cut should reveal new information; use camera motion when only framing changes.
Attach one compatible move—push, pull, pan, truck, tilt, arc, track, static, POV—to a visible subject and action. Add amplitude and speed only when they clarify the intended framing.
Put shot-synchronized dialogue and effects in the timeline, summarize diegetic ambience in overall_soundscape, and reserve non_diegetic_music for score the characters cannot hear.
Give speaking people stable IDs such as (S1) only in their voice description, keep spoken words inside a <d> block, and include a language tag. Do not reuse speaker IDs as Ref2VA subject labels.
The guide (S1) says: <d>[English] Keep the camera moving.</d>Silence is not an empty instruction. Write N/A when you want only diegetic ambience and effects; otherwise specify the score's instrumentation, pace, dynamics, and exit point without duplicating physical sounds.
non_diegetic_music: N/A
These editable examples demonstrate the format; they are not official showcase prompts or promises of a particular result. Copy the structured text, replace its subjects and references, and keep the timeline achievable within the selected duration.
Replace appearance-only adjectives with state changes: who moves, what reacts, where the camera goes, and what becomes visible by the end.
Use one camera move per beat, remove simultaneous contradictory directions, and reserve a timestamped cut for a genuinely new composition.
Define each voice once with a stable speaker ID, keep narration distinct from on-screen speech, and do not pack stage directions inside the dialogue tag.
Resolve every <Subject N> and <Picture N> label, specify what each asset contributes, and avoid asking two references to control the same property in incompatible ways.
The field names, section order, and mode semantics come from MiniMax's published materials. The builder wording and examples are original editorial aids; follow the linked source documents for the canonical specification.
Use integrated_multimodal_description first, overall_soundscape second, and non_diegetic_music third. I2VA, FL2VA, and L2VA add an alignment preamble before those same three fields.
I2VA anchors one picture as the first frame and uses the three base fields. Ref2VA can draw selected identity, motion, style, video, or audio properties from labeled references and therefore uses six sections: subject definitions, summary, retention analysis, detailed description, soundscape, and music.
Name one move that serves the action, its subject, and optionally its speed or amplitude. A slow push toward a turning bowl is clearer than stacking push, pan, orbit, zoom, and handheld in one moment.
Yes. Put timed dialogue and synchronized effects in the detailed timeline, then summarize the scene's ambient and physical sounds in overall_soundscape. Use stable speaker IDs and keep dialogue in its original language.
Enter N/A in non_diegetic_music. This explicitly asks for no audience-only score while leaving room for scene ambience, dialogue, and physical effects.
No. The builder formats prompts for all five documented modes, while Tinulo's available runtime models and inputs can change. A related preset link is a creation starting point, not a claim that its page currently exposes every reference mode.
Start from an existing preset, review credit pricing, or read more production notes. These links do not imply that every guide mode is currently exposed by the generator.