Character references and dialogue: assign image, video and audio roles — MiniMax H3 tutorial
Use Ref2VA to assign identity and sound references, then review dialogue and identity retention before trying camera edits.
Original: ComfyUI MiniMax H3 Text-to-Video, Image-to-Video, and Reference-to-Video Workflows
ComfyUI · Source
MiniMax H3 Reference to Video (R2V)
Generate videos that lock in a character, style, motion, camera move, or voice from any mix of reference images, videos, and audio.
Model downloads
Model storage
ComfyUI/
├── 📂 models/
│ ├── 📂 diffusion_models/
│ │ └── minimax_h3_ref2va_pruned_int8_convrot.safetensors
│ ├── 📂 text_encoders/
│ │ └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
│ ├── 📂 vae/
│ │ ├── minimax_h3_video_vae_fp16.safetensors
│ │ └── minimax_h3_audio_vae_fp32.safetensors
│ ├── 📂 loras/
│ │ └── minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors
│ └── 📂 embeddings/
│ └── minimaxh3_art_is_explosion.safetensors
Prompting tips
- Reference by tag: Reference each input by tag in the exact order it was connected, for example , , ``
- Assign each reference a job: State which reference drives which part of the shot (identity, style, motion, camera, voice). Explicit assignments tend to work much better
- Limits: Up to 9 reference images, 3 reference videos (each can carry its own soundtrack), and 3 standalone reference audio clips
- ref_image_size: match scales references down to the generation resolution for speed; max keeps up to a 2048px short edge for stronger identity fidelity at the cost of speed
- Note: R2V uses the ref2va diffusion model, a different set of weights from the fl2va model used by the T2V and I2V workflows
- Turbo mode (optional): The workflow generates at 20 steps by default; raise the step count (for example to 25) for better motion quality. Enable the Lightning LoRA checkbox to use the 4-step turbo LoRA (minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16) for much faster generation, with slightly lower audio and motion quality
The official full-reference-mode prompt writing guide is summarized in the prompt guide.
Audience: Creators with a working single-clip H3 setup who want reference media to guide identity and dialogue.
Hardware: ComfyUI already running H3, with separate Ref2VA weights. More complex inputs increase compute requirements.
Prerequisites
- A usable character image or video and any required sound references
- A clear decision on what to retain and what to change
Steps
- Open the R2V template and select Ref2VA instead of the FL2VA weights used for T2V/I2V.
- Connect image, video and audio references as shown in the template and record their input order. Start with only the assets the task requires.
- Use the reference guide to assign identity, action, background or sound roles and specify what should be retained or changed.
- Place one spoken line in a short scene, with the official speaker and language format. Check that every reference corresponds to an uploaded input.
- Inspect identity, motion, dialogue and sound separately. Reduce conflicting references for unstable identity and shorten incomplete dialogue.
- Once the baseline is satisfactory, try a camera or action change while specifying what stays the same. Do not replace all references at once.
Caveats
- Do not expect pixel-level reconstruction; fast motion and occlusion can still drift.
View original source · ComfyUI · 2026-09-06
In depth · Source checked: 2026-09-06 · Site testing: No generation test performed
Assign a role to each reference: identity, motion, camera or sound. The guide makes the intended use of each input explicit.
Local generation uses your own hardware, RAM and storage. See the original model and tool terms.
Links lead to the original workflows, models or examples. Supply your own reference assets where needed. This generation workflow has not been independently run by this site.
Applicable versions
ComfyUI 0.30+ · native H3 templates
Resources and demonstrations
Troubleshooting
Identity drift
Reduce conflicting inputs and separate identity from action references. Compare within the same short scene.
Truncated dialogue or unexpected sound
Shorten the line and check language, speaker tags and whether audio is referenced or reused.
Examples and techniques
Learn next