原文: h3-metal
antirez · 出处
Native MiniMax-H3 inference for Apple Silicon. The project is being built as a
sequence of working vertical slices: deterministic host/model metadata first,
then portable Metal block parity, prompt encoding, prompt-to-video/audio, and
first/last-frame conditioning and then ordered references.
Prompt-to-video/audio, first/last-frame conditioning, and ordered Ref2VA
image/video/audio references work end to end. The current work is incremental
H3-specific Metal performance and memory optimization on M3 Max and M5 Max.
Tutorial
1. Build and inspect the model
The examples assume that the Hugging Face snapshot is in ./MiniMax-H3 and
that FFmpeg and FFprobe are available on PATH.
make -j8
mkdir -p outputs
./h3 --info -d ./MiniMax-H3
--info checks the model layout and prints the selected Metal device without
mapping all weights or generating media. Run ./h3 --help for the complete CLI
reference.
Without -p, the same binary starts an Iris-style interactive session:
./h3 -d ./MiniMax-H3 --width 512 --height 512 --steps 6
Type a prompt to generate a numbered video. The session keeps the exact BF16
prompt conditioning, prepared DiT, and video decoder in memory, so repeating a
prompt with another seed avoids loading and encoding them again. Useful commands
are !status, !seed random, !seconds 2, !show, !save output.mp4, and
!cache. Use !help for the full, short list.
First/last-frame conditioning is persistent in the session:
h3> !first opening.png
h3> !last ending.png
h3> The camera moves slowly around the subject.
Use !first clear or !last clear to remove an anchor. Generated videos are
written to the session directory printed at startup.
For a general Ref2VA conditioning image, use !ref-image PATH instead. Images
are appended in order and exposed to the model as <Picture 1>, <Picture 2>,
and so on; filenames have no meaning to the model.
h3> !ref-image person.png
h3> Make the person shown in Picture 1 wave to the camera.
!refs lists the current order, !ref-remove N removes one entry, and
!refs clear removes them all. Ref2VA references cannot be mixed with
!first/!last anchors.
2. Make a first fast video
Start with the validated balanced preset. It generates 22 frames at 24 fps
(about 0.92 seconds), displays the evolving middle-video frame after every
denoising transition in a supported graphical terminal, and prints phase
timings:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A red fox walks through fresh snow in a pine forest. Medium tracking shot, natural winter light, realistic fur, soft footsteps and wind." \
--width 512 --height 512 \
--frames 22 --steps 20 \
--layers 45 --reuse 2 \
--show \
-o outputs/fox-fast.mp4
This is deliberately not the most aggressive configuration:
- --steps 20 performs the default 20 denoising passes.
- --reuse 2 computes 11 fresh denoiser velocities instead of all 20 and
extrapolates the skipped transitions.
- --layers 45 runs 45 of the 50 transformer blocks, reducing both time and
unified-memory use.
- --show is optional. It supports Kitty/Ghostty and
iTerm2/WezTerm/Konsole graphical protocols. It loads a resident preview VAE,
displays one representative middle-video frame after every Euler transition,
and then displays all final frames. Display dimensions default to 2x so the
image has its intended logical size on macOS Retina screens; use --zoom 1
on a non-HiDPI display. This adds preview decode time and roughly 10 GiB of
temporary model residency; runs without --show are unchanged.
- --profile is optional and does not select a different generation path.
The first process invocation also pays model loading and filesystem-cache
costs. Compare performance using repeated runs, and alternate variants when
the machines are warming up because this workload is sensitive to thermal
throttling.
For a very short iteration, request four denoising passes directly:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A red fox walks through fresh snow in a pine forest. Medium tracking shot, natural winter light, realistic fur." \
--width 512 --height 512 --frames 22 \
--steps 4 --layers 50 --reuse 1 \
--show \
-o outputs/fox-four-step.mp4
--steps N always means exactly N denoising passes. Four through seven passes
use the same schedule that won the low-budget comparison; increasing from 4
to 7 progressively improves detail and motion. Keep --reuse 1 at such small
budgets so every requested pass runs the model. --show displays one preview
after each pass.
Several tail-heavy schedules were evaluated because most visible cleanup
happens late in a long run. They preserved too few early composition updates
and produced woven texture, weak motion, or clipped colors. The retained mode
uses the released linear base grid with one terminal point. On the 512-square,
22-frame fox test, the selected four-pass result had 0.556 full-video SSIM
against a 29-pass reference; an independent surfer test measured 0.547. The
four-pass denoise took about 3.5 seconds on M5 Max, versus 26.4 seconds for the
reference.
For a low-memory run, add --ssd-streaming:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A red fox walks through fresh snow in a pine forest." \
--width 512 --height 512 --frames 22 --steps 20 \
--layers 50 --reuse 1 --ssd-streaming \
-o outputs/fox-ssd.mp4
This uses the original BF16 checkpoint without conversion or quantization. It
keeps two DiT blocks in memory and reads the next block from SSD while the GPU
runs the current one. On M5 Max, tracked DiT storage fell from about 36.5 GiB to
2.0 GiB at 512 square and 2.1 GiB at 864x480. A warm 50-block forward measured
1.35 versus 2.49 seconds at 512 square (84% slower), and 2.14 versus 2.68
seconds at 864x480 (26% slower). These are comparisons against the same
full-residency BF16 path, and the results were byte-identical in both checks.
The 2.0--2.1 GiB figure is the DiT's tracked tensor storage, not total system
RAM. Prompt encoding and the two VAEs run in separate phases rather than adding
their full peaks to it; the OS, media buffers, and output resolution still need
headroom. --show keeps a preview VAE resident and adds roughly 10 GiB, so omit
it for the lowest-memory run.
SSD streaming is an explicit memory/speed tradeoff and is not the default. It
cannot be combined with --use-int8-row-fc2. In an interactive session, use
!ssd-streaming on.
3. Move toward reference quality
Change one control at a time when evaluating quality. First restore all layers,
then all denoiser evaluations, and finally raise the default 20-pass schedule
to the slower 50-pass reference:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A red fox walks through fresh snow in a pine forest. Medium tracking shot, natural winter light, realistic fur, soft footsteps and wind." \
--width 512 --height 512 \
--frames 22 --steps 50 \
--layers 50 --reuse 1 \
-o outputs/fox-close.mp4
The defaults are --steps 20 --layers 50 --reuse 1; keep --steps 50
explicit for this close path. It performs 50 complete 50-block denoiser
forwards and is much more expensive than the default, but is the right oracle
when a fast mode changes the subject, anatomy, motion, or composition.
Numerical pixel identity with MLX is not expected because the random-number and
execution engines differ; the depicted content and motion should agree.
4. Choose a speed/quality preset
These controls are independent unless noted otherwise:
On M5, --use-int8-row-fc2 uses one activation scale per FC2 row and a single
full-width TensorOps product. It is optional because it is less numerically
conservative than grouped int8. It reduced complete denoiser forwards by about
2.6% in reciprocal tests. Matched four-step fox and surfer videos kept the same
subjects, setting, and motion (full-video SSIM 0.919 and 0.828). In the
interactive session, use !int8-row-fc2 on.
--reuse and --core-reuse are mutually exclusive. Layer thinning can be
combined with either one.
To make the first command faster while keeping its output resolution, add
token reduction:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A surfer riding inside a sharp blue ocean wave, one rider and one white board, realistic spray." \
--width 512 --height 512 --frames 22 --steps 20 \
--layers 45 --reuse 2 --token-reduction \
-o outputs/surfer-fast.mp4
At the validated 512 square shape, token reduction cut the 45 layers + reuse 2 denoise profile from 16.69 to 12.60 seconds on the IT M5 Max. Independent
fox and surfer renders stayed coherent, but composition can diverge more from
the close path.
For an aggressive preview, render internally at 320 square and upscale to the
requested 512 square output:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A red fox walking through snow, realistic, tracking shot." \
--width 512 --height 512 \
--render-width 320 --render-height 320 \
--frames 22 --steps 20 --layers 40 --reuse 3 \
-o outputs/fox-aggressive.mp4
This combination produced a clean, recognizable 22-frame fox in validation,
but loses fine detail and can change framing. Do not add --token-reduction
to both --layers 40 and --reuse 3: that tested combination produced color
ringing, outlines, and ghosted limbs.
As an alternative to whole-velocity reuse, this keeps the timestep-dependent
patch and output heads fresh at every transition:
./h3 --profile \
-d ./MiniMax-H3 \
-p "A surfer riding a blue ocean wave." \
--width 512 --height 512 --frames 22 --steps 20 \
--layers 45 --core-reuse 4 \
-o outputs/surfer-core-reuse.mp4
Use --core-reuse 6 only as an aggressive preview. Values above 6 are not
exposed because validation lost subject fidelity.
5. Pick resolution and duration
Width and height must each be multiples of 32, at least 32, and their product
must not exceed 768 * 1344 pixels. Those are mechanical limits, not a promise
that every tiny canvas has good model quality. H3-Base is a 768p model.
For a fast native 256-square preview:
./h3 -d ./MiniMax-H3 \
-p "A red fox walks through fresh snow in a pine forest." \
--width 256 --height 256 \
--frames 22 --steps 20 \
--layers 50 --reuse 1 \
-o outputs/fox-256.mp4
At 256 square, H3 has only an 8x8 effective spatial-token grid, so it has less
room for fine detail and complex composition. H3 automatically halves spatial
RoPE coordinates at exactly 256 square. This removed repeating lattice
artifacts in long fox renders and stayed coherent on an independent portrait,
without adding tokens or runtime. Use --use-reference-rope to restore the
released/MLX coordinates for parity checks. Keep token reduction off at this
size. Native 128 square remains unsupported: its 4x4 token grid did not
recover a recognizable subject even with adjusted RoPE.
--render-width and --render-height must be set together, must have the same
aspect ratio as the output, and cannot exceed the output dimensions. The model
and VAE use the internal size; terminal frames and the encoded video retain the
requested output size.
H3 emits 24 fps and aligns frame requests upward to 5 + 17*n:
Use --seconds N for a duration-oriented request, or --frames N for direct
frame control; the two options are mutually exclusive. Fractional seconds are
accepted. Seconds are converted at 24 fps and then rounded upward to the next
legal H3 temporal shape, so --seconds 10 produces 243 frames (10.125 seconds).
Short clips are useful for development. The released workflow is intended for
roughly 4–15 second videos. A request such as --frames 23 is rounded up to 39
frames rather than producing an arbitrary temporal shape.
6. Improve the prompt
A short prompt works, but the released system expects a Context-IR-like
description. State the subject, action, setting, camera, lighting/style, and
desired sound. For example:
Scene: a single red fox in a snow-covered pine forest at dawn.
Action: the fox walks steadily left to right and looks toward the camera once.
Camera: medium-height lateral tracking shot, 50 mm lens, stable framing.
Look: photorealistic fur, cold blue ambient light, warm sunrise rim light.
Audio: soft footsteps in snow, light wind through pine branches, no music.
Keep identity and object counts explicit when they matter. --seed N controls
the native random stream; the default is 42. Compare options with the same
prompt, seed, resolution, frame count, and step count.
7. Preview frames and diagnose performance
- --show displays a representative frame after every denoising transition,
followed by all frames from the completed video. Like Iris, it advertises 2x
display dimensions by default for Retina terminals; --zoom N changes that
factor without resizing the generated video or the encoded terminal image.
- --frames-dir DIR writes final callback frames as PPM files. Intermediate
--show previews are not written there.
- -o '' disables MP4 encoding; combine it with --frames-dir when FFmpeg is
unavailable.
- --profile reports phase wall time, Metal encoding/wait time, peak live
tensor storage, cumulative allocation, and dispatch counts.
For example:
./h3 --profile -d ./MiniMax-H3 -p "A hummingbird hovering over red flowers." \
--width 512 --height 512 --frames 22 --steps 20 \
--layers 45 --reuse 2 --frames-dir outputs/hummingbird-frames \
-o ''
8. Add image, video, and audio references
First/last-frame anchors select the FL2VA path:
./h3 -d ./MiniMax-H3 -p "The fox keeps walking through the snow." \
--width 512 --height 512 --frames 22 --steps 20 \
--layers 45 --reuse 2 \
--first-frame fox.png --last-frame fox-later.png \
-o outputs/fox-anchored.mp4
Ordered references select the distinct Ref2VA checkpoint. Use the flag matching
the media semantics:
# One image reference.
./h3 -d ./MiniMax-H3 -p "Use the animal and setting in the reference." \
--width 512 --height 512 --frames 22 --steps 20 \
--ref-image fox.png -o outputs/fox-reference.mp4
# Continue a clip but ignore its soundtrack.
./h3 -d ./MiniMax-H3 -p "Continue the motion in this clip." \
--width 512 --height 512 --frames 22 --steps 20 \
--ref-silent-video fox.mp4 -o outputs/fox-video-reference.mp4
# Preserve the clip's embedded audio.
./h3 -d ./MiniMax-H3 -p "Continue this audiovisual scene." \
--width 512 --height 512 --frames 56 --steps 20 \
--ref-video fox-with-audio.mp4 -o outputs/fox-video-audio.mp4
# Replace a video's soundtrack explicitly.
./h3 -d ./MiniMax-H3 -p "Continue the scene with the supplied music." \
--width 512 --height 512 --frames 56 --steps 20 \
--ref-video-audio silent-fox.mp4 replacement.wav \
-o outputs/fox-replaced-audio.mp4
# An ordered image plus standalone audio reference.
./h3 -d ./MiniMax-H3 -p "Use the animal and music from the references." \
--width 512 --height 512 --frames 56 --steps 20 \
--ref-image fox.png --ref-audio music.wav \
-o outputs/fox-image-audio.mp4
Reference flags may be repeated and their command-line order is preserved.
Standalone audio must accompany an image or video reference. Audio references
must be 2–15 seconds; at most three audio inputs are accepted and their total
decoded duration is capped at 15 seconds.
Tests and runtime requirements
make test
make parity
make test runs the deterministic host suite and, when the ignored MLX fixture
is installed under misc/fixtures/, compiles the Metal source at runtime and
checks a complete toy H3 block against named MLX outputs. Runtime compilation is
intentional: it follows Iris and does not require Xcode's optional offline Metal
toolchain. The test covers both an F32 diagnosis path and the production BF16
storage path; wide BF16 matrix products and SDPA use cached MPSGraph graphs, with
direct Metal correctness fallbacks. make parity runs only those Metal/MLX
checks.
FFmpeg and FFprobe must be available on PATH for media inputs and MP4 output
(H3_FFMPEG and H3_FFPROBE may select explicit executables). Generated RGB24 and
32 kHz stereo F32 PCM are fed through concurrent pipes; no intermediate
uncompressed media file is created.
Implementation and performance notes
The remainder documents the implementation behind the tutorial presets and the
environment variables retained for exact A/B diagnosis.
Sampler and DiT controls