Free deterministic prompt builder

Write a MiniMax H3 prompt that controls picture, motion, and sound

H3 generates video and synchronized audio together. A useful prompt therefore describes an ordered timeline—not a pile of visual adjectives—then separates physical sound from audience-only music. Build the official three-field base format or six-section Ref2VA format below.

This page is a writing and formatting tool. It explains all five documented H3 modes, but does not claim that every mode is enabled in Tinulo's current generation runtime. Preset links open existing creation starting points only.

Runs entirely in your browser

MiniMax H3 prompt builder

Choose the inputs you actually have. The builder adds the correct keyframe preamble and preserves the exact official field order; it does not send your text to a model or rewrite it unpredictably.

Prompt inputs

Text only: describe the complete audiovisual timeline from scratch.

Write the chronological shots, composition, subjects, action, camera, dialogue, and synchronized sounds.

Summarize ambience, physical effects, voices, and other sounds that exist inside the scene.

Describe audience-only score, instruments, tempo, and dynamics. Leave blank for an explicit N/A.

Structured prompt
integrated_multimodal_description: [Shot 1] Live-action cinematic medium shot of a ceramicist at a sunlit wheel. She lifts a finished blue bowl, turns it slowly toward the camera, and smiles. The camera pushes in gently while warm window light catches the wet glaze.

overall_soundscape: A quiet wheel hum, soft clay movement, birds outside, and one natural breath.

non_diegetic_music: N/A

The five H3 prompt modes are input contracts

T2VA needs no asset. I2VA, FL2VA, and L2VA align keyframes to exact times. Ref2VA is different: it uses six ordered rewrite sections and explicit subject/reference relationships rather than the three base fields.

T2VA

Text only: describe the complete audiovisual timeline from scratch.

I2VA

First frame: Picture 1 is fixed at 0.00 seconds; describe only the motion that follows.

FL2VA

First and last frames: connect Pictures 1 and 2 through one plausible continuous path.

L2VA

Last frame: invent a compatible opening that converges precisely on Picture 1 at the ending timestamp.

Ref2VA

Full reference: define and selectively retain subjects, motion, visual style, video, or audio from labeled assets.

What belongs inside a strong H3 timeline

Chronological shots

Begin with [Shot 1]. Give later cuts increasing timestamps such as [Shot 2] At 00:05.000. A cut should reveal new information; use camera motion when only framing changes.

Camera with a purpose

Attach one compatible move—push, pull, pan, truck, tilt, arc, track, static, POV—to a visible subject and action. Add amplitude and speed only when they clarify the intended framing.

Two audio layers

Put shot-synchronized dialogue and effects in the timeline, summarize diegetic ambience in overall_soundscape, and reserve non_diegetic_music for score the characters cannot hear.

Dialogue: stable speakers, original language

Give speaking people stable IDs such as (S1) only in their voice description, keep spoken words inside a <d> block, and include a language tag. Do not reuse speaker IDs as Ref2VA subject labels.

The guide (S1) says: <d>[English] Keep the camera moving.</d>

No music is a valid direction

Silence is not an empty instruction. Write N/A when you want only diegetic ambience and effects; otherwise specify the score's instrumentation, pace, dynamics, and exit point without duplicating physical sounds.

non_diegetic_music: N/A

Six original MiniMax H3 prompt examples

These editable examples demonstrate the format; they are not official showcase prompts or promises of a particular result. Copy the structured text, replace its subjects and references, and keep the timeline achievable within the selected duration.

T2VA cinematic movement

A two-shot landscape sequence with motivated tracking, a timed cut, physical ambience, and restrained score.

integrated_multimodal_description: [Shot 1] Live-action wide shot of a lone cyclist crossing a salt flat at blue hour. The camera tracks beside her at a steady pace as the first stars appear in the reflection. [Shot 2] At 00:06.000, cut to a low front angle; the bicycle passes close and the camera tilts toward the violet sky.

overall_soundscape: Fine tires hiss over damp salt, a light crosswind passes the microphone, and the chain clicks softly.

non_diegetic_music: Sparse glass harmonica and a low sustained synth, fading after the cut.
T2VA product-style dialogue

Two stable speaker IDs, language-tagged lines, a simple push-in, and no background music.

integrated_multimodal_description: [Shot 1] Medium two-shot in a quiet greenhouse after rain. The botanist with a calm low voice (S1) points to a new leaf and says: <d>[English] It opened overnight.</d> The camera pushes in slowly. [Shot 2] At 00:05.000, cut to the apprentice with a bright voice (S2), who touches the fogged glass and replies: <d>[English] Then spring arrived early.</d>

overall_soundscape: Water drips from leaves, glass panels creak gently, and distant morning birds call.

non_diegetic_music: N/A
I2VA landscape animation

Preserves the uploaded first frame and limits change to plausible mist, water, vegetation, and camera motion.

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Preserve the mountain lake, dawn palette, shoreline, and framing in <Picture 1>. The camera pushes forward with small amplitude while existing mist drifts right and one ripple widens from the foreground. Add no people, buildings, or weather changes.

overall_soundscape: Small waves touch the shore, reeds move in a light breeze, and one distant bird calls.

non_diegetic_music: N/A
FL2VA keyframe transition

Connects two supplied pictures through one continuous paper-bird motion with an exact 5.00-second endpoint.

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the 5.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Begin with the folded paper bird in Picture 1 on the desk. A breeze lifts it into a smooth clockwise turn; the camera arcs with it and converges naturally on the open-wing pose and window composition in Picture 2 at the final frame.

overall_soundscape: Paper creases flex, a gentle window breeze enters, and the bird lands with a dry tap.

non_diegetic_music: One warm marimba phrase, ending on the final frame.
L2VA final-frame convergence

Builds a plausible ink motion backward from the required final composition and timing.

How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the 10.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Start close on blue ink entering clear water from above. The camera pulls out slowly as branching tendrils descend and curl. Their shape, density, lighting, and position converge continuously on the exact abstract composition in Picture 1 at 10 seconds.

overall_soundscape: A muted underwater bloom, tiny bubbles, and quiet glass-room ambience.

non_diegetic_music: A low airy drone that resolves softly at the final frame.
Ref2VA identity and motion transfer

Defines a referenced subject, separates what to retain from what may change, and uses all six official sections.

subject_definitions: <Subject 1> is the watchmaker from <Picture 1>, retaining face, round glasses, blue apron, and careful hand posture.

summary: A quiet macro portrait of the referenced watchmaker setting a balance wheel into a movement.

retention_analysis: Retain <Subject 1> identity and wardrobe plus the brass-and-blue palette from <Picture 1>. Borrow only the measured hand rhythm from <Video 1>; do not copy its setting.

detailed_description: [Shot 1] Macro over-shoulder view as <Subject 1> lowers the balance wheel with tweezers. The camera trucks left by a few centimeters, racks focus from fingertips to the turning mechanism, then holds.

overall_soundscape: Fine tweezer contact, a soft bench creak, close room tone, and the first precise ticks.

non_diegetic_music: N/A

Troubleshooting weak H3 prompts

The clip feels like a moving still

Replace appearance-only adjectives with state changes: who moves, what reacts, where the camera goes, and what becomes visible by the end.

Cuts and motion fight each other

Use one camera move per beat, remove simultaneous contradictory directions, and reserve a timestamped cut for a genuinely new composition.

Dialogue changes speaker

Define each voice once with a stable speaker ID, keep narration distinct from on-screen speech, and do not pack stage directions inside the dialogue tag.

References are ignored or blended

Resolve every <Subject N> and <Picture N> label, specify what each asset contributes, and avoid asking two references to control the same property in incompatible ways.

Five checks before copying

  • Every requested action fits inside 5 or 10 seconds.
  • Later shots use increasing timestamps and do not exceed the duration.
  • Camera directions are compatible with the composition and subject motion.
  • Dialogue, soundscape, and music are placed in the correct fields.
  • Every reference label resolves consistently; unwanted music is explicitly N/A.

Primary sources and evidence boundary

The field names, section order, and mode semantics come from MiniMax's published materials. The builder wording and examples are original editorial aids; follow the linked source documents for the canonical specification.

MiniMax H3 prompt FAQ

What is the official MiniMax H3 base prompt order?

Use integrated_multimodal_description first, overall_soundscape second, and non_diegetic_music third. I2VA, FL2VA, and L2VA add an alignment preamble before those same three fields.

How is Ref2VA different from I2VA?

I2VA anchors one picture as the first frame and uses the three base fields. Ref2VA can draw selected identity, motion, style, video, or audio properties from labeled references and therefore uses six sections: subject definitions, summary, retention analysis, detailed description, soundscape, and music.

How should I write camera movement?

Name one move that serves the action, its subject, and optionally its speed or amplitude. A slow push toward a turning bowl is clearer than stacking push, pan, orbit, zoom, and handheld in one moment.

Can an H3 prompt contain dialogue and sound effects?

Yes. Put timed dialogue and synchronized effects in the detailed timeline, then summarize the scene's ambient and physical sounds in overall_soundscape. Use stable speaker IDs and keep dialogue in its original language.

What should I enter when I do not want music?

Enter N/A in non_diegetic_music. This explicitly asks for no audience-only score while leaving room for scene ambience, dialogue, and physical effects.

Does this page guarantee every H3 mode works in Tinulo?

No. The builder formats prompts for all five documented modes, while Tinulo's available runtime models and inputs can change. A related preset link is a creation starting point, not a claim that its page currently exposes every reference mode.

Turn the structure into your own shot

Start from an existing preset, review credit pricing, or read more production notes. These links do not imply that every guide mode is currently exposed by the generator.