Composes H3 prompts with native audio in the mandatory structured format, for full-reference (Ref2VA) and text-to-video (T2VA) only.
Composes H3 prompts with native audio in the mandatory structured format, for full-reference (Ref2VA) and text-to-video (T2VA) only.
Original: minimax-worldbuilder
Hearmeman24 · Source
Compose MiniMax-H3 video-with-native-audio prompts in H3's mandatory structured format — full-reference (Ref2VA) and text-to-video (T2VA) only. Use whenever the target model is MiniMax-H3: the user says MiniMax, MiniMax-H3, H3, ref2va / r2v / t2v, names a minimax_h3 checkpoint, or names the MiniMaxH3ReferenceToVideo / EmptyMiniMaxH3LatentAV / MiniMaxH3SigmaShift ComfyUI nodes. Also use when the request needs H3's own vocabulary — reference tags <Subject N> / <Picture N> / <Video N> / <Audio N>, subject_definitions, retention_analysis, detailed_description, integrated_multimodal_description, overall_soundscape, non_diegetic_music, [Shot N] cut timestamps, (S1) speaker IDs, <d> dialogue blocks. H3 prompts are not free prose: fixed section names in a fixed order, fixed camera-motion and retention enums, and a native-audio pass that goes silent if the sound sections are thin — so use this skill instead of a generic prompt skill any time H3 is the target model. NOT for other generative models (Flux, SDXL, Qwen Image, WAN, Kling, Veo, Sora, Runway, Seedance — use cinematic-prompts), NOT for still images (use cinematic-still-director), and NOT for H3's I2VA / FL2VA / L2VA keyframe paths, which this skill deliberately does not cover.
MiniMax-H3 Worldbuilder — Ref2VA & T2VA Director
H3 is omni-modal: one forward pass produces video and native 32 kHz stereo audio together. Sound is not a garnish — a thin soundscape yields a near-silent clip. And H3's prompt format is a structured intermediate representation, not prose. Named sections, fixed order, fixed enum vocabularies. A paraphrased camera move or an invented retention marker is a defect, not a style choice.
This skill is the director's apparatus on top of that IR: it reads the user's references, picks a look, extracts an inventory, and emits a prompt that conforms to MiniMax's own guides exactly.
Scope
In scope — two paths, and only two:
Out of scope: I2VA, FL2VA, L2VA. Those paths require a first-line image-alignment instruction that this skill does not specify. If the user is running one of those checkpoints/endpoints, say plainly that this skill covers Ref2VA and T2VA only and point them at MiniMax's base prompt-writing guide — do not improvise an alignment line.
"Use this image as the first frame" — which path? Decide by the checkpoint, not by the wording of the request. Running the i2va / fl2va / l2va checkpoint → out of scope, stop and say so. Running the ref2va checkpoint → in scope: full-reference mode has its own way to anchor a frame — a standalone <Picture N> plus the keyframe completion task type in summary, and no alignment instruction line. If you don't know which checkpoint they're on, ask before writing anything.
Unverified: the sources do not say how strictly the ref2va checkpoint honours a frame anchor compared to the dedicated keyframe checkpoints. Say so if the user is leaning on exact first-frame fidelity.
FIXED vs HOUSE — read this before anything else
Two kinds of content live in this skill and they are not interchangeable:
- FIXED — comes from MiniMax's guides. Section names, tag names, enum values, timestamp format, structural rules. Never paraphrase, never extend, never invent a new member. If what you want isn't in a FIXED list, pick the nearest member that is.
- HOUSE — this skill's directorial judgment (look modes, shot-count guidance, word-count guidance for T2VA, pacing). Free English text. Adjust it freely to serve the scene.
Every table below is labelled. When HOUSE guidance and a FIXED rule appear to conflict, the FIXED rule wins and the HOUSE guidance bends.
SESSION OPENER — REFERENCE & CHARACTER GATE
The first time the user asks for an H3 prompt in a session, ask once:
"Any recurring characters or locked references in this batch? If so — do you already have the reference assets, or do we build the look from text?"
Branch:
- Has references → get them, then run the extraction pass below and mirror back the locked spec in plain language for confirmation before composing. This is the Ref2VA path.
- No references / inventing from text → T2VA path. Skip straight to the pre-prompt confirmation.
- Wants recurring identity but has no references → say so directly: T2VA cannot hold identity across separate generations. Recommend generating or sourcing 3–4 varied stills of the character first, then switching to Ref2VA.
Ask once per session. Carry the answer.
Reference asset limits (FIXED): ≤9 images, ≤3 video clips (each 2–15 s), ≤3 audio clips, ≤12 files total. Audio cannot be the sole input. 3–4 varied shots of a character hold identity far better than one — say this out loud when the user offers a single image.
Numbering (FIXED, load-bearing): <Picture N>, <Video N>, <Audio N> are numbered in the order the assets are connected/supplied, and <Video N> / <Audio N> are numbered independently of each other — the same source clip can be <Video 1> and <Audio 2>. If you don't know the connection order, ask. A prompt whose numbers don't match the wiring references the wrong asset.
READING REFERENCES — INVENTORY EXTRACTION (run before composing)
Extract everything visible by visual description only, then compose. Never invent detail that isn't in the reference or in the user's text.
Per character: hair (colour with nuance, length, texture, parting, styling, accessories) · skin and complexion, visible freckles/marks · makeup register if visible · every garment top to bottom (fabric, colour, fit, neckline, sleeve, hem, layering, structural details) · jewellery and accessories · visible piercings, tattoos, nail colour · build and posture · expression register.
Per environment: interior/exterior, architecture, materials, scale · time of day, weather, light direction and colour temperature · set dressing (every object that shapes the world) · dominant palette and contrast structure.
Per audio reference: what role it plays — voice timbre, delivery, music style, ambience, sound-effect texture, beat, or a signal to be copied outright. This decides its retention marker later.
Naming rule. Do not use proper names for characters in the prompt. Ref2VA identifies subjects by <Subject N> plus visual description; T2VA identifies them by visual description alone. No example in either guide names a character.
Age and gender are required, not forbidden. The base guide explicitly asks for character type, age, gender, pitch, timbre, speaking rate and accent when a speaker first appears — write them. (This reverses the age-blind convention used for other platforms.) The one hard floor: never write minors into sexual, suggestive, or violent-victim content.
Brands and on-screen text. Real signage that is genuinely visible in the scene goes in English double quotes, verbatim (see the on-screen text rule). Otherwise describe products generically. If the user is sending to MiniMax's hosted API rather than running locally, keep brand marks out — hosted moderation is stricter than a local checkpoint.
No-invention rule. If a scene needs a detail the reference doesn't carry (a new outfit, a location not shown), ask, or state in your reply that you composed it from the user's text rather than from the reference.
LOOK MODES (HOUSE) — the director's register
H3 has no camera-spec block. The look is carried by (a) the style opening, (b) the enum camera moves you choose, (c) texture and palette words inside the shot descriptions, and (d) the two sound sections. Never append a trailing gear/spec block — every clause in an H3 prompt must correspond to something visible or audible.
Pick one mode. It seeds the style opening and biases the camera choices; it is not a constraint.
Other named styles the guides list and you may open with: 2D-animated, 3D CG, claymation, watercolor. (FIXED list — these plus Cinematic, live-action, vintage film.) You may add free-text look language after the named style.
DURATION AND SHOT TIMING (decide this FIRST)
Cut timestamps must fall inside the duration, so the duration is chosen before a single shot is written.
Model spec (FIXED): 4–15 s, 24 fps, 32 kHz stereo. Aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16. Open weights are 768p only — 2K needs the hosted H3-Regenerate-2K API.
Local ComfyUI frame grid (FIXED): length = 17k + 5, trained range 124–362 frames.
192 frames / 8.000 s is the only whole-second duration in the trained range. Every other length lands on a fraction — write the real number, don't round it into the prompt's timestamps.
Which range wins. The two ranges are not in conflict, they describe different things: 4–15 s is the model's published output range, 124–362 frames is the trained frame range of the local ComfyUI nodes. When the target is local ComfyUI, pick a length from the table — the shortest usable clip is 124 frames ≈ 5.167 s, so a request for "4 seconds" gets rounded up to 124 frames and the user gets told why. When the target is the hosted API, 4 s is available and durations are specified in seconds.
Shot-count guidance (HOUSE): ≤7 s → 1–2 shots. 8–11 s → 2–3 shots. 12–15 s → 3–4 shots. Land the last cut at least ~1.5 s before the end so the closing shot has room to play. Dialogue runs roughly 2.5–3 English words per second — count the words before you promise a line will fit.
FIXED VOCABULARIES
Everything in this section is an enum. Pick a member. Never paraphrase, never invent.
The only angle-bracket tags that may ever appear in an H3 prompt are <Subject N>, <Picture N>, <Video N>, <Audio N>, <d>…</d>, <scenetrans> and <cutoff>. Anything written as {like this} in this skill is a placeholder for you to replace — never copy braces into a prompt, and never invent a new angle-bracket tag.
Camera motion
Motion type: Zoom In · Zoom Out · Push In · Pull Out · Pan Left · Pan Right · Truck Left · Truck Right · Tilt Up · Tilt Down · Pedestal Up · Pedestal Down · Arc Shot · Tracking Shot · Static Shot · Shake Slightly · Shake Strongly · POV · Roll Clockwise · Roll Counterclockwise
Amplitude: with small amplitude · with large amplitude — medium is expressed by omitting amplitude.
Speed: at slow speed · at fast speed — normal is expressed by omitting speed.
Order is fixed: motion type, then amplitude, then speed — "pushes in with small amplitude at slow speed", never "at slow speed with small amplitude". Add amplitude and speed only when they carry meaning; omitting one is how you say "medium" or "normal". Write the move as natural English action inside the shot, conjugated to fit the sentence — never stacked as trailing labels:
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.
Conjugation is allowed (Push In → "pushes in", "pushing in"). Substituting a synonym is not: no "dollies forward", "creeps closer", "whip pan", "crane down", "handheld float". If the move you want isn't in the list, take the nearest member.
Cut verbs
the camera cuts to · the shot cuts to · the shot transitions to · the shot changes to · the shot switches to. Cross-dissolve, fade or wipe only when the user explicitly asks.
A cut must introduce new information — subject, space, state, viewpoint or time. If only the camera distance or a slight angle changes, use camera motion instead of a cut.
Shot headers
[Shot 1] carries no timestamp. Every later shot opens with a strictly increasing cut time inside the duration, in exactly this format:
[Shot 2] At 00:03.500, the camera cuts to ...
MM:SS.mmm — two-digit minutes, two-digit seconds, three-digit milliseconds.
Retention markers
Visual — for <Subject N>, <Picture N>, <Video N>:
Audio — for <Audio N>:
newly_generated does not exist. There is no marker for new content. Newly added actions, backgrounds and plot events are not losses of reference fidelity and get no retention entry at all. Choose a marker only inside the reference role already defined for that label in subject_definitions.
Task-type prefixes for summary
keyframe completion · reference generation · video editing · video continuation · audio reuse · audio reference
Combine with +, never repeat a type: [video continuation + keyframe completion].
- An image/video/audio giving generation guidance for a character, scene, style, action, camera move or storyboard, without being a concrete frame or the source being edited → reference generation.
- video editing only when a source video is directly modified. video continuation only when new content continues or extends it. A reference video that supplies only camera movement, cuts or rhythm is reference generation.
- The mere presence of a video or audio file does not create a task type. Editing a source video whose original audio stays audible adds audio reuse; continuing one while only matching its audible characteristics adds audio reference.
- For a video editing task, the summary starts (after the prefix) with The target video is an edited version of <Video 1>.
Reference labels
The trap: an image used only to define a character, scene, costume or style gets no standalone <Picture N> entry — cite it inside that <Subject N> line instead. Same for a <Video N> that only identifies where a subject came from. A standalone entry exists only when that asset is analysed or used separately later.
One subject may be defined by several assets; one asset may supply several subjects. Content reused from a reference video is still a <Subject N> — <Video N> marks the asset, not the visible content. A reference video does not create an <Audio N> merely because the file has sound. Once assigned, a label keeps the same meaning across every section.
<Picture N> / <Video N> / <Audio N> numbering follows the order the assets are supplied. <Subject N> numbering follows the order you define them in subject_definitions — the guides state no other rule for subjects.
SPEAKERS, DIALOGUE AND SOUND (both paths)
Speaker IDs (FIXED). (S1), (S2), … assigned in the order of actual vocal events in the target video, and stable across shots. Simultaneous speakers get a compound ID: (S1,S2). Characters who never vocalize get no ID. Never write (Sx) in retention_analysis.
The <d> split (FIXED). Identity, action and delivery go outside <d>. Inside <d> goes only [Language] plus the verbatim spoken words. Never translate, never rewrite, never paraphrase.
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>
When a speaker first appears, establish a stable identity outside <d>: character type, age, gender, on-screen or off-screen, pitch, timbre, speaking rate, accent.
Languages with stable dialogue support: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish.
When reference dialogue or lyrics are reused verbatim, preserve the exact source words and original language inside <d>; write [unclear] for unintelligible spans rather than guessing. Standardize punctuation to , . ? ! — strip tildes, emoji, bullets and decorative repetition, and end each complete sentence with ., ? or ! before </d>. When only timbre, rhythm, emotion or delivery is referenced, do not carry the original words into the target video.
Voiceover (FIXED). Use the exact phrase says in an off-screen voiceover, and immediately after the <d> block state that the on-screen character's lips stay closed:
The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.
Speech across a cut. When the same line or lyric is split across a cut, place <scenetrans> at the connecting point in both parts and state explicitly that the audio continues, using one of: continues seamlessly across the cut · continues uninterrupted into the next shot · carries over from the previous shot · remains audible across the transition. Use <cutoff> when speech is truncated by the end of the video.
Source gap — read this. Neither guide shows a worked example of <scenetrans> or <cutoff> in place, so their exact positioning is inferred. The base guide's own Case 1 keeps the whole <d> block inside one shot and lets it ring out over the cut using only the continuity phrase carries over from the previous shot, with no <scenetrans>. Follow that pattern by default: keep each <d> block whole inside one shot and use a continuity phrase. Reach for <scenetrans> only when the user genuinely needs a line split mid-sentence across a cut, and tell them the tag's placement is inferred.