返回博客

MiniMax H3 Prompt Structure: Five Parts That Control the Shot

A working structure for MiniMax H3 text-to-video prompts — subject, action, camera, light, and style, then the two audio fields — with three copyable examples from Tinulo's own prompt guide.

2026年8月21日Adrian ValeAdrian Vale
MiniMax H3 Prompt Structure: Five Parts That Control the Shot

Most weak H3 prompts fail the same way: they are a pile of visual adjectives with no timeline. "Cinematic, epic, beautiful sunset, 4k, masterpiece" tells the model almost nothing it can act on. H3 generates video and synchronized audio together, so what it actually needs is an ordered description of what happens, in what frame, with what sound.

After a few hundred local generations on our own 5090, the structure that holds up is five visual parts plus two audio fields. This post is that structure, with prompts you can copy.

The five visual parts

Write one shot at a time, and inside each shot cover these in order:

  1. Subject — who or what is in frame, with the details that must survive the clip. "A ceramicist at a sunlit wheel" beats "a person."
  2. Action — the specific motion the subject performs, at a speed a five- or ten-second clip can actually contain. One action per beat. "She lifts a finished blue bowl, turns it slowly toward the camera, and smiles" is three verbs that read as one gesture. Ten verbs read as noise.
  3. Camera — exactly one move, attached to the action: push, pull, pan, truck, tilt, arc, track, static, or POV. "The camera pushes in gently" is a direction. "Dynamic camera angles" is a wish.
  4. Light — where the light comes from and what it touches. "Warm window light catches the wet glaze" gives the model a light source, a direction, and a surface to play against.
  5. Style — one anchoring phrase, placed last so it colors everything above rather than replacing it. "Live-action cinematic medium shot" works. Five style words stacked on each other fight.

For multi-shot prompts, give each shot its own marker — [Shot 1], then [Shot 2] At 00:06.000 — and only cut when the new frame reveals new information. If only the framing changes, that is a camera move, not a cut.

The two audio fields people skip

H3's base format has three sections. The first is the visual timeline above. The other two are where most prompts leave quality on the table:

  • overall_soundscape — the sounds that exist inside the scene: ambience, physical effects, voices. Rain on pavement, a pottery wheel's hum, footsteps. If you leave this blank, the model invents it, and invented ambience is how a quiet studio scene ends up sounding like a crowd.
  • non_diegetic_music — score the characters cannot hear. Write N/A when you want none. Silence is a valid direction; an empty field is not the same thing as a deliberate one.

Three prompts you can copy

These follow the same format as the examples on the MiniMax H3 prompt guide, which also has a browser-only builder that assembles the fields for you without sending your text anywhere.

A single-shot product clip (5 seconds, T2VA):

integrated_multimodal_description: [Shot 1] Live-action cinematic medium shot of a ceramicist at a sunlit wheel. She lifts a finished blue bowl, turns it slowly toward the camera, and smiles. The camera pushes in gently while warm window light catches the wet glaze.

overall_soundscape: A quiet wheel hum, soft clay movement, birds outside, and one natural breath.

non_diegetic_music: N/A

Note what is not here: no style stack, no "masterpiece," no second camera move. Every clause maps to one of the five parts.

A two-shot landscape with a timed cut (10 seconds, T2VA):

integrated_multimodal_description: [Shot 1] Live-action wide shot of a lone cyclist crossing a salt flat at blue hour. The camera tracks beside her at a steady pace as the first stars appear in the reflection. [Shot 2] At 00:06.000, cut to a low front angle; the bicycle passes close and the camera tilts toward the violet sky.

overall_soundscape: Fine tires hiss over damp salt, a light crosswind passes the microphone, and the chain clicks softly.

non_diegetic_music: Sparse glass harmonica and a low sustained synth, fading after the cut.

The second shot carries an explicit timestamp. That is what turns "then something else happens" into an edit the model can schedule.

Animating a still image without breaking it (5 seconds, I2VA):

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Preserve the mountain lake, dawn palette, shoreline, and framing in <Picture 1>. The camera pushes forward with small amplitude while existing mist drifts right and one ripple widens from the foreground. Add no people, buildings, or weather changes.

overall_soundscape: Small waves touch the shore, reeds move in a light breeze, and one distant bird calls.

non_diegetic_music: N/A

Image-to-video prompts have one job beyond the five parts: say what must not change. "Add no people, buildings, or weather changes" is doing as much work as the motion description. Without it, the model treats your first frame as a suggestion.

What this looks like in practice

The four clips on our Explore page were all written this way before they were rendered — the rainy night market walk is "handheld forward walk" (camera) plus "warm lanterns and abstract neon colors reflecting in puddles" (light) plus a soundscape of "steady rain, footsteps, and distant indistinct market ambience." Each one is a single continuous shot, because at five seconds, one well-described move beats three cramped cuts.

If a generation comes back as a moving still, the fix is almost never a longer prompt. It is a missing action verb or a camera field that says "dynamic" instead of naming a move. Add the verb, name the move, and resubmit the same five parts.