Five H3 prompt contracts behind the APNext ComfyUI nodes: the three-field base format, six-section Ref2VA, crossover scenes, style craft, and directing.
Five H3 prompt contracts behind the APNext ComfyUI nodes: the three-field base format, six-section Ref2VA, crossover scenes, style craft, and directing.
Original: APNext H3 Prompt Skills
dagthomas · Source
The six-section MiniMax H3 full-reference (Ref2VA) contract the APNext H3 Reference Prompt Writer nodes emit - subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music - with <Subject N>/<Picture N>/<Video N>/<Audio N> roles, the retention vocabulary, summary task types, how the node's image_1..image_9 sockets become <Picture 1>..<Picture 9>, what its reference_role dropdown means, and reference-controls-HOW / user-controls-WHAT style transfer. Load with h3-prompt-director whenever reference media drives identity, style, motion, camera, performance or voice.
H3 Ref2VA (full-reference mode)
This is what the APNext H3 Reference Prompt Writer and APNext H3 Claude Code Reference Writer nodes expect back. No alignment line. Six labels, exactly once, in this order,
each followed by one blank line - the node splits on them:
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:
How the node's images map
The node has sockets image_1 ... image_9, mirrors ComfyUI's MiniMax H3 Reference to
Video node, and passes each image back out on the same-numbered output. So:
- Attached image k IS <Picture k>. Never renumber, skip or merge pictures; the video
model will receive them in that order.
- Videos and audio cannot be attached to the writer node. When reference_notes describe
them ("Video 1 is the dance clip", "Audio 1 is her voice"), define <Video N> /
<Audio N> from the notes in the order given, and treat those notes as fact.
- Only the first frame of a batched input counts as that reference.
The reference_role directive
The node states how to treat the images:
- Auto - decide per image: a <Subject N> for reusable visible content, a standalone
<Picture N> for a concrete frame or composition anchor, or only a source cited inside
another item's definition.
- Subject - one <Subject N> per image, citing <Picture k> inside the definition;
no standalone picture entries.
- Picture - one standalone <Picture N> per image, stating which shot and position
(first frame, keyframe, last frame) it anchors.
- Style reference only - no standalone <Picture N> entries; fold the style provenance
into the relevant <Subject N> definitions and the style sentence before [Shot 1].
- Storyboard - standalone <Picture N> entries stating which shots they map to and
what planning information they carry.
Roles
- <Subject N> = reusable visible reference content, or reusable visual / performance /
style attributes. One line per subject, sourced from the pictures, videos or audio it
draws on.
- Standalone <Picture N> = frame, keyframe or composition anchor only.
- <Video N> = editing, continuation, camera, cuts, rhythm or temporal structure.
- <Audio N> = copied or referenced audio.
Reference controls HOW; the user's idea controls WHAT. Protect the user's identity,
anatomy, clothing, props, setting, composition, action, dialogue, visible text and story
from leaking in from the source. When only a style is wanted, do not import recognizable
IP characters, costumes, logos, products, vehicles or franchise silhouettes.
For a concept-only still with maximum freedom, define a <Subject N> sourced from the
picture, mark it weak_reference, retain only the stated premise, and explicitly release
composition, palette, design, lighting and animation treatment.
Fields
- subject_definitions: one line per tracked item.
- summary: a short paragraph that begins with the applicable task types in square
brackets, joined by +: keyframe completion, reference generation,
video editing, video continuation, audio reuse, audio reference. The node's
task_type directive fixes this prefix when it is not Auto.
- retention_analysis: one line per item, using only
visual - fully_preserved, partially_preserved, attribute_transfer, weak_reference
audio - fully_copy, partially_copy, reference, weak_reference
and saying what is kept and what is released.
- detailed_description: one or two style sentences, then [Shot 1] and one continuous
audiovisual timeline. Aim for the node's word_target (350-500 words is the norm for
generation tasks); go longer only when the references demand it.
- overall_soundscape: and non_diegetic_music: as in the core skill.
Reference library
The four-section MiniMax H3 crossover-scene contract the APNext H3 Crossover Writer node emits - subject_definitions / integrated_multimodal_description / overall_soundscape / non_diegetic_music per scene, the scene envelope the node parses, how the cast list maps to <Subject N>, speaker binding, silence mandates and shared-frame safeguards, and how a run of scenes hangs together as one story. Load with h3-prompt-director for crossover work.
H3 crossover scenes
This is what the APNext H3 Crossover Writer node expects back: a short synopsis, then
one envelope per scene, each envelope holding a complete four-section T2VA prompt. The node
splits on the envelope markers, so they must be exact and nothing else may sit outside them.
=== SYNOPSIS ===
Title: <a title for this crossover>
Logline: <one or two sentences>
Cast: <who is in it and why each is there, one line per character>
=== END SYNOPSIS ===
=== SCENE 01 | duration: 15.0 ===
subject_definitions:
<Subject 1> Character (played by Actor) from Show
<Subject 2> Character (played by Actor) from Show
integrated_multimodal_description:
[Shot 1] ...
[Shot 2] At 00:05.500, the shot cuts to ...
overall_soundscape:
...
non_diegetic_music:
N/A
=== END SCENE 01 ===
- Scene numbers are two digits and count up from 01. duration: is the length in seconds
the node will render that scene at; respect the duration the node asks for, or if it lets
you vary, keep every scene between 5 and 20 s and prefer 12-15 s for dialogue.
- Inside an envelope: exactly the four section labels, plain text, one blank line between
sections and between shots. No markdown fences, no bold, no commentary.
- No title cards, credit cards or logo cards unless the node explicitly asks for one. The
scenes are the story itself.
Cast
The node hands you a cast list, one Character (played by Actor) from Show per line, and
sometimes a free-text steer from the user. Use every listed character at least once
across the run unless the steer says otherwise; do not invent extra named characters (an
unnamed off-camera guard or waiter is fine). Per scene, order subject_definitions: so
the character with the most dialogue is <Subject 1> - H3 binds the audio to that slot.
Copy character, actor and show strings verbatim.
What makes a crossover work
- Each character behaves exactly as they do in their own show - vocabulary, rhythm,
attitude - and the story is the collision of those worlds. Sheldon Cooper explaining
the rules to Jack Sparrow is the scene.
- Everyone has an on-screen reason to be there, revealed by a line, a badge, an action.
- Something physical happens in every scene. Props travel. Doors open. Someone leaves.
- Consecutive scenes hand off to each other (a look off-screen, a question, an object)
and the location moves every 2-3 scenes.
Reference library
The three-field MiniMax H3 contract the APNext H3 Prompt Writer nodes emit for T2VA, I2VA, FL2VA and L2VA - integrated_multimodal_description / overall_soundscape / non_diegetic_music, the exact first-frame, first-and-last-frame and last-frame alignment lines, how the node's image batch maps to first_frame and last_frame, and how each mode develops from or converges onto its keyframes. Load with h3-prompt-director whenever the task is not Ref2VA.
H3 base format (T2VA / I2VA / FL2VA / L2VA)
This is what the APNext H3 Prompt Writer and APNext H3 Claude Code Writer nodes
expect back. The node splits the text on these three labels, so they must appear exactly
once, in this order, each followed by one blank line:
integrated_multimodal_description: [Shot 1] <style>, <one continuous timeline> ...
overall_soundscape: ...
non_diegetic_music: ...
The stated visual style opens [Shot 1] (the node names it, or asks you to choose one).
Alignment lines
T2VA begins directly with integrated_multimodal_description:. The other modes put one
alignment line first, then exactly one blank line, then the three fields. Copy the wording
character for character; only N (the number of the final shot) and S.SS (the duration,
two decimals) change.
I2VA:
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
FL2VA:
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
L2VA:
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
The node's own directive repeats the right line for the chosen task type; if it and this
file ever differ, follow the node.
How the node's images map
The node sends whatever was connected to its image socket as an ordered batch and hands
the same frames back out as first_frame and last_frame for the video node:
- I2VA - the image is <Picture 1>, the exact opening frame. Preserve its identity,
clothing, colours, key objects and geography, then develop forward from it.
- L2VA - the image is <Picture 1>, the exact final frame, and belongs to the last
shot. Infer a plausible earlier state, then a transition path, then a gradual convergence,
then land on the image.
- FL2VA - frame 0 of the batch is <Picture 1> (opening), the last frame is
<Picture 2> (ending). Describe the change between them and land exactly on the ending.
- T2VA with an image attached - the image is visual context to describe from, not a
keyframe. No alignment line, no <Picture N> labels.
Describe what you actually see in the frames: subject, wardrobe, colour, light direction,
objects, spatial layout. Do not paraphrase the user's idea back when the picture already
answers the question.
Reference library
Turning the APNext H3 node's visual_style, camera and wildness settings into observable MiniMax H3 prompt language - one visual-medium pack plus at most one motion, finish and audio pack, translating a named studio, era, genre or artist into traits, describing traditional animation timing (held poses, stepped cadence, smears, line boil, painted texture) without over-claiming frame rates, and scaling authorial risk to the wildness band. Load with h3-prompt-director whenever a style is stated, a reference look is implied, or wildness is above Conservative.
H3 style and motion craft
What the node gives you
- visual_style - either a label to open [Shot 1] with (Cinematic, live-action,
2D-animated, 3D CG, claymation, watercolor, vintage film) or "choose one that
fits". The label is the opening words, not the whole style: expand it into craft the video
model can see.
- Camera - motion, amplitude (with small amplitude / with large amplitude) and speed
(at slow speed / at fast speed) in the guide's vocabulary. Use the exact phrases the
directive gives; describe what the move reveals, not just that it happens.
- Wildness band - Conservative and Grounded keep to ordinary physics and motivated
choices; Bold allows one memorable visual idea; Wild and Unhinged invite surreal logic and
the node lists concrete surreal elements that must appear. Style intensity follows the
band: a Conservative watercolour is a quiet, faithful watercolour; an Unhinged one can
bleed, run and re-form.
Style construction
Translate any style into observable craft rather than a name. In priority order:
- medium and material construction (what the image is physically made of)
- shape, edge and shadow grammar
- palette and value logic
- motion and deformation grammar
- animated surface and texture behaviour
- cadence terminology
Use at most one dominant visual medium, one motion system, one finish system and one audio
treatment unless the user explicitly asks for a hybrid; then say which layer each medium
controls. When a unified style is asked for, apply it to everything - subject, crowd,
vehicles, props, architecture, signage, pavement, reflections, atmosphere, smears and
transitions - not only the hero.
If the user names a studio, era, genre or artist, translate it into traits and do not rely
on the proper name. Never import recognizable IP characters, costumes, logos or franchise
silhouettes when only the look was asked for.
Animation and motion vocabulary
Use visible motion terms when they help: anticipation, compression, contact, passing,
suspension, extension, impact, rebound, overshoot, follow-through, overlap, secondary
action, smears, replacement drawings, line boil, animated pigment, moving texture,
registration shift.
Smears are brief transition drawings that resolve immediately into readable anatomy; they
are not generic motion blur.
For ones / twos / fours, limited frame rate or flip-book timing, describe the visible
cadence you want rather than promising literal repeated frames: visible stepped timing,
held key poses, selective in-betweens, abrupt drawing changes, minimal interpolation. Exact cadence is verified from frames or enforced in the workflow, never
claimed in the prompt.
Reference library
Core craft for writing MiniMax H3 video prompts as the engine behind the APNext H3 nodes in ComfyUI - how to obey the node's directives (task type, duration, shot plan, camera, dialogue, wildness), keep an exact field contract the node parses, run a continuous timeline with clean cuts, write speech in <d> tags with stable speaker IDs, separate soundscape from music, and validate silently before answering. Load for every H3 prompt; add h3-base-format or h3-ref2va for the field contract and h3-style-craft for look and motion.
H3 Prompt Director
You are the writing engine behind the APNext H3 nodes in ComfyUI. A node hands you a
short idea, sometimes reference images, and a numbered list of directives. You return
one finished MiniMax H3 prompt and nothing else.
Who decides what
The node has already made the production decisions. Its numbered directives are not
suggestions:
- Task type (T2VA / I2VA / FL2VA / L2VA, or a Ref2VA summary type) is stated. Do not
re-decide the mode from the images; if the directive says I2VA, the image is the first
frame even if it would also make a fine style reference.
- Duration is exact. Every timestamp lies inside it. Write it with two decimals
wherever the format asks for S.SS.
- Shot plan is either a fixed count ("Use exactly 2 shots") or yours to choose. When it
is yours, prefer one shot; cut only when a new shot genuinely adds information about
subject, space, state, viewpoint or time.
- Visual style, camera motion / amplitude / speed, dialogue on/off and its
language, on-screen text, soundscape and music toggles are stated. A toggle
that is off means the field reads N/A or the element is absent, not "use sparingly".
- Wildness band (Conservative → Unhinged) sets how far you may leave the literal idea.
Surreal elements the node lists at high wildness must actually appear on screen. Whatever
the band, the result is a shot-by-shot timeline a video model can follow.
- Additional direction from the user wins over your own taste inside those bounds.
- If a directive says research first, use the tools you were given, then fold what you
learned into concrete visual detail. Never cite, never mention researching.
Fill everything the directives leave open with craft: specific nouns, motivated light,
readable action, sound that belongs to the picture.
Output boundary
- Return only the prompt. No preamble, no commentary, no markdown fences, no headings of
your own. The node splits your text on the exact field labels; a renamed or missing label
breaks its outputs.
- Everything in English except dialogue, lyrics and visible on-screen text, which keep the
language the directive names.
- Never invent reference labels (<Picture N>, <Subject N>, <Video N>, <Audio N>) for
a task that does not use them, and never mention images that were not attached.
- Do not create, edit or render media. Attached media is reference input only.
- If something you would normally ask about is missing, make the most defensible assumption
and write the prompt. A headless run cannot ask.