MiniMax Director: multi-shot video with sound on a timeline — MiniMax H3 tutorial
Arrange shots, recurring subjects and dialogue in ComfyUI, inspect the compiled prompt, and render an MP4 with sound.
Audience: Creators who have already rendered an H3 clip with the official ComfyUI template and want to work with multiple shots and recurring subjects.
Hardware: Rendering requirements come from H3 and the selected workflow; the plugin primarily compiles prompts. In the submission, the author reports roughly 30 minutes for 8 seconds on Linux with an NVIDIA H200 141GB. This is neither a minimum requirement nor a site benchmark; other GPUs were not timed.
Prerequisites
- Render one clip with the official ComfyUI H3 template first; use our ComfyUI setup guide if needed.
- Use ComfyUI 0.31.0 or newer and Python 3.10 or newer.
- Obtain the four Comfy-Org/MiniMax-H3 weights without renaming them. The author's September 2 snapshot totals about 39.6 GiB; allow additional space for outputs and runtime files.
Steps
- Run the clone command below from ComfyUI/custom_nodes, then restart ComfyUI.
- Place minimax_h3_ref2va_pruned_int8_convrot.safetensors in ComfyUI/models/diffusion_models/.
- Place qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors in ComfyUI/models/text_encoders/. Put minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors in ComfyUI/models/vae/.
- Open Workflow → Browse Templates → minimax-director, or drag in the plugin's example_workflows/minimax-director.json. Check every loader and leave Turbo and Upscale off initially.
- Select the shot block under the playhead and describe its picture, camera movement and sound. Arrange subsequent shots on the timeline; duration snaps to an accepted H3 frame count.
- Create WHO & WHAT cards for recurring people or objects with reference assets and descriptions. Click a card in the shot's SUBJECTS row to insert <Subject n> into the shot text.
- Inspect the actual model input in Prompt and resolve explicit errors in Report before queueing. Check the resulting MP4 for picture, duration and sound.
- For a faster preview, obtain the author's documented 4-step Turbo LoRA, place it in models/loras/ and enable Turbo in Speed. Use its matching four steps, strength 1.0 and simple scheduler. Disable Turbo if sound or fast motion deteriorates.
Commands
cd ComfyUI/custom_nodes
git clone https://github.com/imbutus/ComfyUI-MiniMaxDirector.git
Caveats
- At 24 fps, accepted H3 lengths satisfy length % 17 == 5. Timeline snapping does not enable unlimited duration.
- Reference cards and dialogue compilation cannot guarantee subject consistency or perfect lip sync in every generation.
- The four-step Turbo path is a preview option with author-reported audio and fast-motion limitations.
- Upscale is off by default. Enabling it requires the extra node pack and weights in the author's instructions and adds computation.
View original source · imbutus · 2026-09-08
Author submission: imbutus · Submission
In depth · Source checked: 2026-09-08 · Site testing: No generation test performed
The author provides installation, a complete graph, shot editing and troubleshooting in one tutorial, with an English walkthrough and a Chinese dub.
The open-source plugin requires no separate service purchase. Rendering consumes local or rented GPU resources; we have not paid for a generation test.
Applicable versions
MiniMax Director 0.17.14 (11fd464) · ComfyUI >= 0.31.0 · Python >= 3.10
Resources and demonstrations
Video chapters
Troubleshooting
Cannot find the example workflow
Use example_workflows/minimax-director.json or Browse Templates. The examples/ path in the initial submission is outdated.
Empty weight dropdown
Check original filenames and the directories above, then refresh or restart. Also confirm the download completed.
Turbo picture looks fine but audio is distorted
Confirm ComfyUI >= 0.31.0 with ModelSamplingAV. Disable Turbo if the problem persists; the author also notes four-step audio limitations.