ComfyUI zero to your first H3 video with audio — MiniMax H3 tutorial
Start with a clean ComfyUI setup, place all five model files correctly, and render a playable T2VA MP4 with picture and synchronized audio.
Original: ComfyUI MiniMax H3 Video Generation Guide
ComfyUI · Source
How to use open-weights MiniMax H3 in ComfyUI: text-to-video, image-to-video, and reference-to-video workflows with native stereo audio, prompt writing tips, and Sage Attention speedups.
MiniMax H3 is MiniMax's general-purpose, omni-modal generation model, now available as open weights. It jointly understands text, images, video, and audio in a single context, and generates video with native stereo audio: voice, sound effects, and music are modeled together in a single forward pass instead of being layered on afterward. Output is up to 2K resolution, 24fps, and about 15 seconds.
ComfyUI natively supports MiniMax H3. The documentation is split across five pages:
- Overview (this page): model capabilities, workflow index, output resolution, and speedups
- Native workflows: Text-to-Video, Image-to-Video, Reference-to-Video, plus advanced native-node techniques
- Multiframe Reference: anchor reference frames at specific points along the output timeline
- Fun ControlNet Union: drive H3 with a control video, or run video inpainting with a mask
- Prompt guide: official MiniMax prompt writing guides, general tips, and prompt embeddings
H3's open weights let you run the model locally. Commercial use of locally generated outputs requires a MiniMax commercial license, available through Comfy, the only official reseller. Generations on Comfy Cloud already include commercial rights.
Key features
- Native stereo audio: Dialogue, sound effects, and music are generated together with the video, synced in one MP4
- Multimodal context: Text, images, video, and audio references can be combined in one generation
- Reference-driven generation: Lock a character's identity, a style, a motion, a camera move, or a voice from reference materials
- Instruction following: Describe the relationship between references and the target shot in natural language
- Accurate text rendering: Spelled-out text and brand elements render cleanly
- Open weights: Run locally in ComfyUI with full control over every parameter
Getting started
MiniMax H3 is supported in ComfyUI with open weights. To get started:
- Update ComfyUI to version 0.30.0 or later
- Go to Template Library > Video > choose any MiniMax H3 workflow
- Follow the pop-up to download models and run the workflow
The model files are hosted on Hugging Face in the Comfy-Org/MiniMax-H3 repository.
Workflow index
The template library currently ships with five example workflows. They are example templates, not an exhaustive list: the model supports more generation modes through the native MiniMax H3 nodes, and you can build additional workflows with them.
- Text to Video (T2V) — Generate videos from text prompts with native stereo audio
- Image to Video (I2V) — Generate videos from an input image, with optional first/last-frame control
- Reference to Video (R2V) — Lock in a character, style, motion, camera move, or voice from reference images, videos, and audio
- Multiframe Reference — Anchor reference frames at specific points along the output timeline with chained Add Guide nodes
- Fun ControlNet Union — Drive H3 with a Canny, Depth, HED, MLSD, or Pose control video, or run video inpainting with a mask
Underlying node modes: first/last-frame image-to-video (fl2va) via the MiniMaxH3ImageToVideo node, and reference-driven generation with images, videos, and audio (ref2va) via the MiniMaxH3ReferenceToVideo node.
For prompt writing resources (official MiniMax guides, general tips, and prompt embeddings), see the prompt guide.
Setting the output resolution
Each workflow uses a Resolution Selector node to control the overall output size. The node computes width and height from three settings, and its outputs connect directly to the width and height inputs of the MiniMax H3 node:
- Aspect ratio: Pick a preset such as 16:9 (Widescreen), 9:16 (Portrait Widescreen), or 1:1 (Square)
- Megapixels: Target total pixel count for the output. Higher values give larger frames; lower values run faster
- Multiple: The computed resolution is rounded to the nearest multiple of this number. Keep it at 32 to match H3's resolution grid
The template ships with a fast preview size. For full-quality output at 16:9, set the Resolution Selector's Megapixels to 0.98 for H3's native canvas (a 768px short edge, 1344x768 at 16:9), or enter 1344 x 768 directly in the MiniMax H3 node's width and height inputs (its default). Skip the 1.0 Megapixel step: it yields 1376x768, above the model's 768x1344 pixel area cap.
Speeding up generation with Sage Attention
The example workflows use the standard attention implementation. You can roughly double the generation speed with Sage Attention, with minimal quality loss. Sage Attention is an optional dependency, so you need to install it yourself:
- Install the sageattention Python package. Download the wheel that matches your PyTorch and CUDA versions from the SageAttention releases page, then install it with pip install <wheel-file>.
- Install the KJNodes custom nodes, which provide the Patch Sage Attention KJ node. Use the ComfyUI Manager, or clone the repository into ComfyUI/custom_nodes/ and restart ComfyUI.
- Add a Patch Sage Attention KJ node to the workflow and connect it between the UNETLoader and the BasicGuider node: its model input receives the model from the UNETLoader, and its model output feeds the model input of the BasicGuider. Set sage_attention to auto.
- Run the workflow as usual. Only the guider needs the patch; the scheduler only generates the sigmas and can stay as is.
Notes:
- Sage Attention requires float16 or bfloat16 tensors. MiniMax H3 runs some layers in other dtypes, so you may see "Input tensors must be in dtype of torch.float16 or torch.bfloat16, using pytorch attention instead" messages in the console. These are expected; the affected layers fall back to standard attention and generation still works.
- Alternatively, you can enable Sage Attention globally by launching ComfyUI with the --use-sage-attention flag instead of adding the node.
Audience: First-time NVIDIA users who want a guided local MiniMax H3 run instead of reverse-engineering a node graph.
Hardware: Windows or Linux with an NVIDIA GPU and ample disk space. Start at the template preview size when VRAM is limited instead of jumping to 1344×768.
Prerequisites
- ComfyUI 0.30 or newer, starting without errors
- NVIDIA drivers compatible with your ComfyUI installation; use the official installer for Desktop dependencies
- Access to model files and enough disk space for downloads and outputs
Steps
- Install or update ComfyUI. Windows beginners should follow the linked official Desktop installer and launch the app after NVIDIA dependencies finish. For Linux/manual installs, first follow the manual guide to clone ComfyUI, activate its isolated Python environment and install the matching CUDA PyTorch build. Only that route uses the pip and main.py commands below, from the ComfyUI folder containing requirements.txt. Confirm the UI opens without CUDA errors before continuing.
- Open Template Library → Video → MiniMax H3 T2V. Use the model-download popup when available; preserve the exact filenames for manual downloads.
- Place minimax_h3_fl2va_pruned_int8_convrot.safetensors in ComfyUI/models/diffusion_models/. T2V and I2V use FL2VA; only R2V switches to Ref2VA weights.
- Place qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors in ComfyUI/models/text_encoders/. Put both the video and audio VAEs in ComfyUI/models/vae/. A silent render is not a successful audio-video setup.
- Place minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors in ComfyUI/models/loras/. It powers the optional turbo_mode; keep Turbo off and use the default 20 steps for the first baseline.
- Restart ComfyUI or refresh the model list. Confirm every Loader finds the diffusion model, text encoder, both VAEs, and LoRA. Fix null, undefined, or red nodes before queuing.
- Render a short clip at the template preview size, 16:9, and multiple=32. Describe the scene first, then shots and camera, then dialogue, effects, or music. Add no custom nodes or stacked accelerators yet.
- Queue the graph and wait through VAE decode and MP4 assembly. Open the file and verify picture, duration, audio, and event or lip synchronization. Save this untouched graph as the baseline.
- After the baseline passes, move to the native 1344×768 canvas or 0.98 MP in Resolution Selector. Do not choose 1.0 MP because it produces 1376×768, above the model area cap.
- For first/last frames, connect first_frame or last_frame in the same FL2VA graph. For identity, style, motion, camera, or voice references, open the R2V template and switch to Ref2VA weights instead of mixing model families.
Commands
python --version
python -m pip install -r requirements.txt
python main.py
ComfyUI/models/diffusion_models/minimax_h3_fl2va_pruned_int8_convrot.safetensors
ComfyUI/models/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
ComfyUI/models/vae/minimax_h3_video_vae_fp16.safetensors
ComfyUI/models/vae/minimax_h3_audio_vae_fp32.safetensors
ComfyUI/models/loras/minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
Caveats
- Model files are large; storage, interrupted downloads, and wrong folders are the most common first-run failures.
- T2V and I2V use FL2VA while R2V uses Ref2VA; mixing weight families causes missing models or bad output.
- H3 resolutions must be multiples of 32, and higher resolution or duration sharply increases VRAM and wait time.
- Templates and filenames evolve; verify the current reference before following an old screenshot.
View original source · ComfyUI · 2026-09-06
In depth · Source checked: 2026-09-06 · Site testing: No generation test performed
The native template gives a picture-and-audio baseline before introducing third-party nodes.
Local generation uses your own hardware, RAM and storage. See the original model and tool terms.
Links lead to the original workflows, models or examples. Supply your own reference assets where needed. This generation workflow has not been independently run by this site.
Applicable versions
ComfyUI 0.30+ · native H3 T2V template
Resources and demonstrations
Troubleshooting
A model is missing from a Loader
Check the original filename and folder for all five files, refresh the model list, and fully restart ComfyUI.
The output is silent or out of sync
Confirm the audio VAE is loaded and reproduce a short 20-step render without third-party nodes.
Examples and techniques
Learn next