Builds and debugs H3 ComfyUI workflows with native nodes first, covering every mode, multimodal references, and workflow JSON edits.
Builds and debugs H3 ComfyUI workflows with native nodes first, covering every mode, multimodal references, and workflow JSON edits.
Original: minimax-h3-comfyui
hkhdair · Source
Build and debug MiniMax H3 ComfyUI workflows using native nodes first. Use for H3 T2VA, I2VA, L2VA, FL2VA, Ref2VA, multimodal references, prompt authoring, workflow JSON editing, or optional Qwen3-VL enhancer troubleshooting.
MiniMax H3 in ComfyUI
Overview
Use this skill to build, modify, explain, or troubleshoot native MiniMax H3 workflows in ComfyUI. It applies the prompt-structure, reference-role, timing, routing, and validation knowledge demonstrated by ethanfel/ComfyUI-MiniMax-H3-Guide, while preferring official ComfyUI nodes for newly created workflows.
Treat the target server's live /object_info response as authoritative because node schemas and filenames can change.
Core Policy: Native Nodes First
When asked to create a MiniMax H3 workflow, do not add nodes from ethanfel/ComfyUI-MiniMax-H3-Guide by default. Build the workflow with official/native ComfyUI H3 nodes and manually apply the extension's best practices:
- choose the correct T2VA/I2VA/L2VA/FL2VA/Ref2VA family;
- write the proper H3 prompt structure directly;
- keep one-based <Picture N>, <Video N>, <Audio N>, and <Subject N> labels aligned with zero-based native sockets;
- declare reference roles, retention intent, shot scope, timing, dialogue, soundscape, and music explicitly in the prompt;
- use native auto-growing image/video/audio sockets and correct paired-audio routing;
- validate duration, media counts, frame-grid rounding, and native model selection.
Treat the custom-node repository as a knowledge and troubleshooting reference, not the default runtime dependency. Use its extension nodes only when one of these conditions is true:
- The user explicitly asks for the Prompt Guide, Prompt Enhancer, Reference Sheets, Target Timing, or other extension nodes.
- The user asks to modify or troubleshoot a workflow that already contains them.
- The user explicitly prefers automated Qwen3-VL prompt rewriting or visual-reference analysis inside ComfyUI and accepts the extra model/runtime cost.
If extension nodes are proposed, state why native nodes alone are insufficient for that request. Never silently introduce the custom-node dependency into a new workflow.
When to Use
- Building native H3 text-, image-, or reference-to-video workflows.
- Translating the extension's H3 prompting practices into native workflow prompts.
- Wiring multiple images, videos, paired soundtracks, or standalone audio.
- Extending the inputs visible in an H3 template.
- Editing a saved H3 workflow without running it.
- Troubleshooting existing Prompt Guide or Qwen3-VL Prompt Enhancer workflows.
- Diagnosing missing nodes, models, sockets, labels, or enhancer output.
Choose the Correct H3 Family
H3-Base-FL2VA accepts at most two endpoint images. H3-Base-Ref2VA accepts up to 9 images, 3 videos, and 3 audio clips, with at most 12 media files in total. Reference video/audio clips must each be 2–15 seconds; total duration per media type is at most 15 seconds. Audio cannot be the only Ref2VA input.
Do not choose endpoint mode merely because an image is present. Appearance/style guidance belongs to Ref2VA; exact first/last frames belong to I2VA/L2VA/FL2VA.
Native-First Workflow Recipe
For a new workflow, start from the official ComfyUI MiniMax H3 template or assemble only official nodes:
Model/CLIP/VAE loaders
↓
Native MiniMax H3 Image to Video or Reference to Video
↓
SamplerCustomAdvanced
├─→ VAEDecode video
└─→ VAEDecodeAudio
↓
CreateVideo → SaveVideoThen apply these extension-derived practices manually:
- Resolve the family: decide endpoint vs Ref2VA semantics before wiring media.
- Route native media: connect exact endpoints or auto-growing Ref2VA sockets directly.
- Inventory references: map every zero-based socket to its one-based H3 label and semantic role.
- Author Context-IR-style text: write the correct section structure directly in the native prompt widget or a standard string node.
- Validate timing: use 4–15 seconds, 24 FPS, and a valid 17k+5 frame count.
- Validate limits: image/video/audio counts, per-clip duration, total duration, and 12-file mixed cap.
- Run a small test first: low resolution, short duration, and conservative steps before production settings.
For T2VA/I2VA/FL2VA/L2VA, manually produce exactly:
integrated_multimodal_description:
...
overall_soundscape:
...
non_diegetic_music:
...
For Ref2VA, manually produce these six sections in order:
subject_definitions:
...
summary:
...
retention_analysis:
...
detailed_description:
...
overall_soundscape:
...
non_diegetic_music:
...
Native prompt-writing rules:
- Define reusable visible content as <Subject N> and cite its source <Picture N> or <Video N>.
- Keep standalone <Picture N> only for concrete keyframe/composition/storyboard roles.
- Keep standalone <Video N> for direct editing, continuation, or whole-video temporal structure.
- Give each tracked reference exactly one retention row using fully_preserved, partially_preserved, attribute_transfer, or weak_reference; audio uses fully_copy, partially_copy, reference, or weak_reference.
- Start [Shot 1] without a timestamp. Start later shots as [Shot N] At MM:SS.mmm with strictly increasing times inside the duration.
- Put spoken/sung words only inside <d>[Language] ...</d> and keep stable speaker IDs (S1), (S2).
- Put physical/ambient sounds in overall_soundscape; put audience-only score in non_diegetic_music; use N/A when there is no score.
- Preserve exact reference numbering and never invent an unsupplied asset.
Completion criterion: the workflow contains no extension node types unless the user explicitly requested them, while the native prompt and routing still satisfy the same H3 role, label, timing, and audio rules.
Verified Model Set
For the reference-to-video workflow:
models/diffusion_models/
└── minimax_h3_ref2va_pruned_int8_convrot.safetensors
models/text_encoders/
├── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
└── MiniMax-H3/
└── qwen3vl_32b_h3_instruct_generation_tail_50_63_nvfp4_awq.safetensors
models/vae/
├── minimax_h3_video_vae_fp16.safetensors
└── minimax_h3_audio_vae_fp32.safetensorsThe normal H3 Qwen3-VL encoder contains language layers 0–49. It conditions H3 but cannot reliably generate rewritten text alone. The matching generation tail temporarily supplies layers 50–63, final normalization, and the LM head.
Verified NVFP4/AWQ tail:
filename: qwen3vl_32b_h3_instruct_generation_tail_50_63_nvfp4_awq.safetensors
bytes: 5396902102
sha256: 40a9a12feb0822b6eaa39bfb03f7f41415fc1f0fd041315391ea968627e01480
source: https://huggingface.co/ethanfel/Qwen3-VL-32B-Ultra-Heretic-H3-ComfyUI-INT8-ConvRot
Do not pair this 0–49 encoder with an arbitrary tail when the matching Instruct NVFP4/AWQ tail exists.
Optional Extension: Install and Discover
Skip this section when creating a normal native H3 workflow. Use it only when the user explicitly requests the extension or an existing workflow already depends on it.
Guide repository:
https://github.com/ethanfel/ComfyUI-MiniMax-H3-Guide
Install under:
ComfyUI/custom_nodes/ComfyUI-MiniMax-H3-Guide
After adding nodes or models:
- Check /queue before restarting.
- Refuse restart when running/pending work exists unless interruption is explicitly accepted.
- Restart ComfyUI.
- Verify /system_stats and /object_info.
- Hard-refresh the user's browser.
Prove node availability through /object_info, not filesystem presence alone:
import json, urllib.request
with urllib.request.urlopen("http://127.0.0.1:8188/object_info") as r:
info = json.load(r)
for name in (
"MiniMaxH3ReferenceToVideo",
"MiniMaxH3PromptGuide",
"MiniMaxH3GenerationTailLoader",
"MiniMaxH3PromptEnhancer",
):
print(name, "PRESENT" if name in info else "MISSING")Optional Extension: Prompt Guide
Do not add this node to newly created workflows unless explicitly requested. Use the behavior below as a reference for writing equivalent native H3 prompts manually.
Simple Ref2VA wiring:
Prompt Guide.h3_prompt ─→ Native H3.prompt
Prompt Guide.h3_length ─→ Native H3.length
Guide outputs:
- h3_prompt: deterministic structured H3 draft.
- rewrite_request: instructions for an external LLM.
- mode_report: family, references, checkpoint, and warnings.
- h3_length: duration rounded to H3's 17k+5 frame grid at 24 FPS.
Use h3_prompt, not rewrite_request, as the included Prompt Enhancer's manual_prompt.
Optional Extension: Qwen3-VL Prompt Enhancer
Do not add the enhancer or generation tail to a newly created workflow by default. Prefer a carefully written native H3 prompt that applies the same section ordering, label discipline, reference roles, timing, dialogue, soundscape, and music rules. Add this path only for explicit in-ComfyUI automated rewriting or when maintaining an existing extension-based workflow.
Add:
MiniMax H3 Generation Tail Loader
MiniMax H3 Prompt Enhancer (Qwen3-VL)
Required wiring:
Prompt Guide.h3_prompt ─────────────→ Prompt Enhancer.manual_prompt
Prompt Guide.mode_report ───────────→ Prompt Enhancer.mode_report
CLIP Loader.CLIP ───────────────────→ Prompt Enhancer.clip
Generation Tail Loader.clip_tail ───→ Prompt Enhancer.clip_tail
Prompt Enhancer.enhanced_prompt ────→ Native H3.prompt
Prompt Guide.h3_length ─────────────→ Native H3.length
CLIP Loader.CLIP ───────────────────→ Native H3.clip
The CLIP output fans out to both enhancer and native H3. The tail loader is a descriptor; the enhancer loads and unloads the tail during prompt generation.
Safe first-test settings:
max_new_tokens: 1000
sampling: deterministic
temperature: 0.7
top_k: 64
top_p: 0.95
min_p: 0.05
repetition_penalty: 1.05
presence_penalty: 0.0
seed: 0
thinking: false
offload_after_generation: true
For richer output after validation, use sampling=sample, max_new_tokens=1200–1400, and temperature 0.6–0.7.
If the 50-layer CLIP lacks a tail, the enhancer safely returns the manual prompt unchanged. Inspect enhancer_report; do not mistake fallback text for successful enhancement. Connect enhanced_prompt, never llm_prompt, to native H3.
Text Enhancement vs Visual Analysis
Images connected directly to native H3 are not automatically visible to Qwen3-VL. A Guide+CLIP-only enhancer rewrites text.
To let Qwen inspect pixels, add one MiniMax H3 Enhancer Visual Reference per image or reference-video frame batch:
Image 1 ─→ Visual Reference 1.media
Image 2 ─→ Visual Reference 2.media
Visual Reference 1.reference_context
└─→ Visual Reference 2.previous_context
Final Visual Reference.reference_context
├─→ Prompt Guide.reference_context
└─→ Prompt Enhancer.reference_contextKeep every original media output separately connected to native H3 according to routing_report. Visual context informs Qwen; it does not replace native media routing.
For video visual references, put MiniMax H3 Target Timing upstream:
Target Timing.timing_context ─→ Prompt Guide.timing_context
Target Timing.h3_length ──────→ each video Visual Reference.h3_length
Target Timing.h3_length ──────→ Native H3.length
Never feed Prompt Guide.h3_length backward into a Visual Reference whose final context feeds the same Guide; that creates a cycle.
Auto-Growing Native Inputs
MiniMaxH3ReferenceToVideo uses COMFY_AUTOGROW_V3. A template can show only three image sockets although the live limit is nine.
Connect the final empty socket to create the next automatically:
ref_image_0 → <Picture 1>
ref_image_1 → <Picture 2>
ref_image_2 → <Picture 3>
...
ref_image_8 → <Picture 9>
When ref_image_2 is connected, ref_image_3 appears. Repeat until ref_image_8. Do not duplicate the H3 node to add references. Video and audio groups use the same auto-grow behavior up to their limits.
With many images, start with:
ref_image_size: match
max retains larger references for identity fidelity but can be substantially slower because reference tokens participate in every sampling step.
Reference Videos and Audio
ref_video_audios.ref_video_audio_0 is the synchronized soundtrack paired with ref_videos.ref_video_0. It is not a general audio socket.
Correct chain:
Load Video.VIDEO ─→ Get Video Components.video
Get Video Components.images ─→ ref_videos.ref_video_0
Get Video Components.audio ─→ ref_video_audios.ref_video_audio_0
Indexes must match:
ref_video_0 ↔ ref_video_audio_0
ref_video_1 ↔ ref_video_audio_1
ref_video_2 ↔ ref_video_audio_2
Leave paired audio disconnected when only motion, action, camera, cuts, or visual style is wanted. Use standalone audio sockets for a separate recording:
Load Audio.AUDIO ─→ ref_audios.ref_audio_0
Reference Labels and Roles
Socket indexes are zero-based; prompt labels are one-based. Keep labels stable and describe assets in actual downstream order:
Picture 1: nighttime fire, smoke, orange lighting, and environment
Picture 2: main subject identity, face, hair, glasses, and clothing
Video 1: walking motion and camera rhythm to transfer
Audio 1: voice timbre and delivery reference
Use explicit roles:
- Identity or appearance.
- Object, prop, clothing, interface, or effect.
- Scene or environment.
- Visual style.
- Concrete keyframe or composition anchor.
- Storyboard or shot planning.
- Motion or action.
- Camera, cuts, or rhythm.
- Source video to edit or continue.
Safe Workflow-JSON Editing
When modifying editor-format workflow JSON:
- Read the target file and discover actual node IDs.
- Confirm live schemas through /object_info.
- Create a timestamped backup.
- Prefer frontend-serialized node shapes.
- For every link, update the top-level link record, source output links, and target input link.
- Remove obsolete links from all three locations.
- Update last_node_id, last_link_id, and revision.
- Write atomically.
- Load the file into an isolated ComfyUI frontend graph and assert no missing node types.
- Confirm /queue matches the user's run/no-run instruction.
Do not POST to /prompt when the user asked not to run. Frontend graph loading plus structural validation is enough for editor compatibility.
Warn users with an already-open tab to hard-refresh and reopen the saved workflow before saving. A stale browser tab can overwrite server-side changes.
Verification
curl -s http://127.0.0.1:8188/queue | python3 -m json.tool
curl -s http://127.0.0.1:8188/object_info/MiniMaxH3ReferenceToVideo | python3 -m json.tool
curl -s http://127.0.0.1:8188/object_info/MiniMaxH3GenerationTailLoader | python3 -m json.tool
After a requested render, verify the real file with ffprobe. One verified L40S baseline was approximately 367.48 seconds for a 608×352, 24 FPS, 4.459-second H.264/AAC output. Treat it as a benchmark, not a guarantee.
Common Pitfalls
- Only three image sockets appear: connect the final empty socket; the next appears automatically.
- Qwen ignores image content: native H3 connections do not expose pixels to the enhancer; add Visual Reference nodes.
- Enhancement changes nothing: inspect enhancer_report; the 50-layer encoder may lack its matching tail.
- Wrong output: native H3 receives enhanced_prompt, not llm_prompt or rewrite_request.
- Wrong audio group: same-video sound goes to ref_video_audio_N; unrelated audio goes to ref_audio_N.
- Labels drift: sockets are zero-based while <Picture N>, <Video N>, and <Audio N> are one-based.
- Role contradicts family: endpoints use I2VA/L2VA/FL2VA; reference relationships use Ref2VA.
- Timing cycle: Target Timing must precede video Visual References.
- Unexpected slowdown: use ref_image_size=match before max.
- Server edit disappears: a stale browser tab overwrote it; hard-refresh and reopen.
- Monitor fails but render succeeds: inspect /history, /queue, service logs, and output files independently.
Verification Checklist
- New workflow uses official/native H3 nodes only unless extension nodes were explicitly requested and justified.
- Native prompt manually applies H3 section order, reference-role, retention, timing, dialogue, and audio practices.
- Correct H3 family selected.
- Native H3 nodes exist in /object_info; custom nodes are checked only for an explicitly requested extension workflow.
- Required native diffusion model, encoder, and VAEs are in the correct folders.
- If the optional enhancer was requested, its matching tail appears in the loader after restart and its wiring is valid.
- Length routing has no graph cycle.
- Original media remains connected to native H3.
- Auto-growing sockets and media limits are respected.
- Video audio is paired with the same-numbered video.
- Prompt labels match socket order and roles.
- Saved JSON backlinks are consistent.
- Workflow loads with no missing nodes.
- Queue state matches the user's instruction.