RTX 4080 16GB: pruned plus INT8 low-VRAM setup — MiniMax H3 tutorial
Run a ten-second local H3 test within 16GB VRAM.
Audience: Users around 16GB VRAM who need a reduced-memory path.
Hardware: RTX 4080 16GB as reported by the source; roughly ten minutes for a ten-second clip.
Prerequisites
- The native ComfyUI workflow
- Matching pruned and INT8 weights
- Correct model directories
Steps
- Treat the source's RTX 4080 16GB result as a hardware snapshot, not a universal ten-minute promise.
- Install DiffSynth-Studio from source with its quant extras in a separate Python environment so the ComfyUI environment remains untouched.
- Open MiniMax-H3-Int8-ConvRot-Pruned-FL2VA.py under model_inference_low_vram and verify model IDs, origin_file_pattern, vram_limit, and output path.
- Run the example first at its short duration and small canvas. Monitor peak VRAM, system memory, downloads, and VAE decode before increasing duration or resolution.
- After an MP4 with video and audio succeeds, hold the seed constant and increase frames or canvas one variable at a time. Keep the last successful configuration before OOM.
- Compare quantized and native outputs with the same prompt, seed, and output specification; fitting in memory is not a quality verdict.
Commands
git clone https://github.com/modelscope/DiffSynth-Studio.git
cd DiffSynth-Studio && pip install -e .
pip install "diffsynth[quant]"
python examples/minimax_h3/model_inference_low_vram/MiniMax-H3-Int8-ConvRot-Pruned-FL2VA.py
Caveats
- Quantization and pruning may affect quality; reproduce first, then evaluate.
View original source · Shuohao · 2026-08-23
Guide · Source checked: 2026-08-23 · Site testing: No generation test performed
Check the prerequisites, then follow the author for the full demonstration and updates.
Follow the original source for its materials. Missing assets or settings are not reconstructed here.
Applicable versions
DiffSynth-Studio main @ 2026-08-23 · INT8 ConvRot pruned FL2VA low-VRAM example
Troubleshooting
comfy-kitchen or the quant backend will not import
Install diffsynth[quant] inside the isolated DiffSynth environment and match Torch/CUDA to the GPU.
The model loads but generation OOMs
Return to the example's default canvas and frames, close competing GPU processes, and scale one step at a time.