人物参考与对白:分配图像、视频和声音 — MiniMax H3 教程
用 Ref2VA 指定人物与声音参考,检查生成对白及身份保持;机位编辑作为后续变化。
原文: ComfyUI MiniMax H3 Text-to-Video, Image-to-Video, and Reference-to-Video Workflows
ComfyUI · 出处
MiniMax H3 Reference to Video (R2V)
Generate videos that lock in a character, style, motion, camera move, or voice from any mix of reference images, videos, and audio.
Model downloads
Model storage
ComfyUI/
├── 📂 models/
│ ├── 📂 diffusion_models/
│ │ └── minimax_h3_ref2va_pruned_int8_convrot.safetensors
│ ├── 📂 text_encoders/
│ │ └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
│ ├── 📂 vae/
│ │ ├── minimax_h3_video_vae_fp16.safetensors
│ │ └── minimax_h3_audio_vae_fp32.safetensors
│ ├── 📂 loras/
│ │ └── minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16.safetensors
│ └── 📂 embeddings/
│ └── minimaxh3_art_is_explosion.safetensors
Prompting tips
- Reference by tag: Reference each input by tag in the exact order it was connected, for example , , ``
- Assign each reference a job: State which reference drives which part of the shot (identity, style, motion, camera, voice). Explicit assignments tend to work much better
- Limits: Up to 9 reference images, 3 reference videos (each can carry its own soundtrack), and 3 standalone reference audio clips
- ref_image_size: match scales references down to the generation resolution for speed; max keeps up to a 2048px short edge for stronger identity fidelity at the cost of speed
- Note: R2V uses the ref2va diffusion model, a different set of weights from the fl2va model used by the T2V and I2V workflows
- Turbo mode (optional): The workflow generates at 20 steps by default; raise the step count (for example to 25) for better motion quality. Enable the Lightning LoRA checkbox to use the 4-step turbo LoRA (minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16) for much faster generation, with slightly lower audio and motion quality
The official full-reference-mode prompt writing guide is summarized in the prompt guide.
适用人群: 已能生成单段 H3 视频,希望用参考素材约束人物特征并组织对白的创作者。
硬件要求: 已能运行 H3 的 ComfyUI;需另备 Ref2VA 权重,输入越复杂计算负担越高。
前置条件
- 可使用的人物图或视频,以及需要的声音参考
- 知道哪些内容要保留、哪些要改变
执行步骤
- 打开 R2V 模板,确认模型切换到 Ref2VA;不要继续使用 T2V/I2V 的 FL2VA 权重。
- 按模板连接参考图、视频和音频,记下输入顺序;第一次只保留任务必需的少量素材。
- 对照官方参考指南声明各素材的任务:身份、动作、背景或声音;明确保留和修改的部分。
- 在同一短场景里安排一句对白,按官方格式标出说话人和语言;检查所有引用都对应已上传素材。
- 生成后分别检查人物特征、动作、台词与声音。人物不稳先减少互相冲突的参考,台词不完整先缩短对白。
- 基线满意后再尝试改变机位或动作,并写清哪些内容保持不变;不要同时替换所有参考。
注意事项
查看原始来源 · ComfyUI · 2026-09-06
深度精选 · 来源核对: 2026-09-06 · 本站实测: 未进行生成实测
从“每个参考负责什么”入手,区分人物、动作、机位和声音来源,避免把素材堆进去却没有明确任务。
本地生成占用自己的硬件、内存和磁盘;模型与工具的使用条件请看原始说明。
链接提供原作者工作流、模型或示例;自选参考素材需自行准备。本站没有另行实测这套生成流程。
适用版本
ComfyUI 0.30+ · native H3 templates
材料与演示
故障排查
人物漂移
减少冲突参考,明确哪个素材提供身份、哪个提供动作;保留同一短场景做对比。
对白被截断或声音不符合预期
缩短台词,检查语言与说话人标记,以及音频是引用还是复用。
案例与技巧
下一步学习