Low-VRAM Maestro: pruned 20B plus Turbo — MiniMax H3 tutorial
Reduce VRAM pressure with the pruned 20B model, six-step Turbo, and optional First Block Cache.
Audience: Maestro users who cannot comfortably run the full model.
Hardware: A lower-VRAM NVIDIA setup; use current Maestro documentation for the actual minimum.
Prerequisites
- A current Maestro release
- Pruned 20B weights
- Turbo weights
Steps
- Select the pruned 20B model and six-step Turbo.
- Run a fixed-input baseline and record VRAM.
- Enable First Block Cache only if still needed, then inspect quality again.
Caveats
- Cache is another variable; do not claim lossless stacking with Turbo without testing.
View original source · cocktail peanut · 2026-08-23