原文: H3MLX
RobZombAI · 出处
High-performance video and image generation toolkit for H3XML (MiniMax H3 Metal 4 NAX accelerated engine). Features calibrated aspect ratio presets, fast iteration modes, text-to-image and text-to-video pipelines on Apple Silicon (M-series with Unified Memory).
H3XML Engine & Presets Guide
H3XML is the native Metal 4 NAX inference engine for MiniMax H3 (33B parameters) on Apple Silicon (M-series with Unified Memory). It provides full mathematical compatibility with antirez/h3.c while delivering accelerated GPU denoise passes and zero-overhead model residency.
Calibrated Resolution & Aspect Ratio Configurations
1. 768x512 Balanced Widescreen (3:2)
- Command Flags: --width 768 --height 512 --frames 90 --steps 20 --layers 45 --reuse 2 --use-int8-row-fc2
- Passes: 10 GPU passes calculated
- Performance (4.0s / 90F): Denoise = 113.69s, Total = 144.91s
- Characteristics: Balanced token density for 3:2 landscape compositions.
2. 864x480 Standard Wide 16:9 (Panavision)
- Command Flags: --width 864 --height 480 --frames 90 --steps 40 --layers 50 --reuse 6 --use-int8-row-fc2
- Passes: 8 GPU passes calculated
- Performance (4.0s / 90F): Denoise = 104.43s, Total = 137.82s
- Characteristics: Standard 16:9 cinematic aspect ratio with full 50 layers.
3. 864x480 Standard Wide Balanced 16:9
- Command Flags: --width 864 --height 480 --frames 90 --steps 20 --layers 45 --reuse 2 --use-int8-row-fc2
- Passes: 11 GPU passes calculated
- Performance (4.0s / 90F): Denoise = 126.03s, Total = 158.20s
- Characteristics: 45 layers with step reuse 2 for 16:9 widescreen.
4. 512x512 Master Square (1:1) — Fast Portrait
- Command Flags: --width 512 --height 512 --frames 90 --steps 40 --layers 50 --reuse 6 --use-int8-row-fc2
- Passes: 8 GPU passes calculated
- Performance (4.0s / 90F): Denoise = 48.82s, Total = 75.73s
- Characteristics: Square 1:1 format focusing token budget on close-up facial features.
5. 512x512 Balanced Square (1:1)
- Command Flags: --width 512 --height 512 --frames 90 --steps 20 --layers 45 --reuse 2 --use-int8-row-fc2
- Passes: 11 GPU passes calculated
- Performance (4.0s / 90F): Denoise = 59.86s, Total = 86.08s
- Characteristics: Fast square generation with 45 layers.
6. 768x768 High-Res Square (1:1)
- Command Flags: --width 768 --height 768 --frames 90 --steps 20 --layers 45 --reuse 2 --use-int8-row-fc2
- Passes: 11 GPU passes calculated
- Performance (4.0s / 90F): Denoise = 237.12s, Total = 276.90s
- Characteristics: 2304 spatial tokens for high spatial detail in 1:1 framing.
Fast Turnaround Configurations (1.0s / 22 Frames)
For rapid prompt exploration and lighting checks:
7. Fast Mode: 512x512 (22 Frames / ~1s)
- Command Flags: --width 512 --height 512 --frames 22 --steps 20 --layers 45 --reuse 2 --use-int8-row-fc2
- Performance: Denoise = 8.79s, Total = 28.38s
8. Fast Mode: 768x512 Widescreen (22 Frames / ~1s)
- Command Flags: --width 768 --height 512 --frames 22 --steps 40 --layers 50 --reuse 6 --use-int8-row-fc2
- Performance: Denoise = 15.01s, Total = 36.65s
Text-to-Image (T2I) Snapshot Mode
For 1-frame image generation:
- Command Flags: --width 768 --height 512 --frames 5 --steps 20 --layers 45 --reuse 1 --use-int8-row-fc2
- Generates an instant high-resolution frame via causal frame slicing ($T=5 \to t_0$).
Execution & Environment Setup
Native Metal 4 NAX environment flags:
export H3_PROFILE=1
export H3_NAX=1
export H3_CPU_SAMPLER=1
export H3_ZERO_COPY_WEIGHTS=1
export H3_REUSE_MPS_COMMAND=1
export H3_DIT_COMMAND_BLOCKS=0
export H3_SOLVER=euler
export OMP_NUM_THREADS=18
Fast Master execution toolkit for MiniMax H3 / H3-Max video generation on Apple Silicon M-series (Unified Memory). Combines 8-step / 5-step exact DPM++ 2M Trailing Flow, 50 full layers, dynamic int8 FC2 quantization, single-chunk and multi-chunk causal temporal lattice (T = 17n + 5), and zero-loss optical macro definition.
MiniMax H3 Fast Master: Architecture & Execution Guide
This skill specifies the engineering configuration for video generation with MiniMax H3 / H3-Max on Apple Silicon (M-series / Unified Memory), calibrated for photorealistic optical fidelity and high GPU throughput (DiT denoise pass in ~10s on 1s clips, ~44s on 4s clips on M5 Max).
💎 Architecture & Pipeline
graph TD
subgraph Fast_Master_Pipeline ["Fast Master Pipeline"]
P["Macro Optical Prompting (35mm f/1.4, micro-textures)"] --> Q["Text Encoder Qwen 3-VL (4.5s)"]
Q --> D["H3 DiT (50 Full Layers, 100% Spatial Tokens)"]
D --> S["DPM++ 2M Trailing Flow Solver + INT8-Row-FC2"]
S --> C1["GPU Evaluation Mode:<br/>1. Exact Mode (--reuse 1): 8 evals (Maximum Quality)<br/>2. Turbo Mode (--reuse 2): 5 evals (Fast Turnaround)"]
C1 --> V["3D Causal Video VAE (3x3 Tiling with 32px Overlap)"]
V --> MP4["Direct Native MP4 Output"]
end⚡ Key Configuration Parameters
🚀 CLI Execution Example
#!/bin/bash
# Fast Master Runner (Apple Silicon Native)
export H3_PROFILE=1
export H3_NAX="qkv-attn"
export H3_ZERO_COPY_WEIGHTS=1
export H3_REUSE_MPS_COMMAND=1
export H3_GPU_SAMPLER=1
export OMP_NUM_THREADS=18
export METAL_DEVICE_WRAPPER_TYPE=0
export MTL_DEBUG_LAYER=0
export MTL_SHADER_VALIDATION=0
export METAL_CAPTURE_ENABLED=0
MODEL_DIR="./models/MiniMax-H3-PDD-8Step"
PROMPT="Cinematic close-up portrait of a woman in natural light. Crisp optical definition, detailed radial iris fibers and specular reflections, natural skin texture, soft warm smile, wavy hair catching rim lighting. Soft blurred cafe background, authentic 35mm f/1.4 lens bokeh."
# 1 Second High-Quality (22 Frames, Denoise ~10s):
./h3-lora-lab/h3 --profile \
-d "$MODEL_DIR" \
-p "$PROMPT" \
--width 640 --height 640 \
--frames 22 \
--steps 8 \
--layers 50 \
--reuse 1 \
--use-int8-row-fc2 \
--seed 333 \
-o outputs/fast_master_1s.mp4
# 4 Seconds Master (90 Frames, Denoise ~44s):
./h3-lora-lab/h3 --profile \
-d "$MODEL_DIR" \
-p "$PROMPT" \
--width 640 --height 640 \
--frames 90 \
--steps 8 \
--layers 50 \
--reuse 2 \
--use-int8-row-fc2 \
--seed 333 \
-o outputs/fast_master_4s.mp4
⏱️ Telemetry Reference Metrics (M5 Max 128GB UMA)
- 1 Second ($640 \times 640$, 22 frames, 8 exact steps):
- DiT GPU Denoise: $10.46\text{ seconds}$
- VAE Decoder: $8.73\text{ seconds}$
- Total Cold Latency: $36.0\text{ seconds}$ (Warm run: $\approx 19.5\text{ s}$).
- 4 Seconds ($640 \times 640$, 90 frames, 8 steps reuse-2):
- DiT GPU Denoise: $44.66\text{ seconds}$
- VAE Decoder: $43.89\text{ seconds}$
- Total Cold Latency: $107.0\text{ seconds}$ (Warm run: $\approx 88.0\text{ s}$).
Baseline high-fidelity guide and execution toolkit for MiniMax H3-Max video generation, SGLang miles RL + LoRA SFT training, PDD 6-step/8-step high-fidelity inference, and Apple Silicon Metal 4 NAX native acceleration. Activate whenever the user mentions MiniMax H3, H3-Max, Hailuo 3, SGLang miles, or asks for high-fidelity 6-step/8-step cinematic video generation with synchronized 48 kHz audio.
MiniMax H3-Max (Golden High-Fidelity Standard Suite)
This skill provides full technical instructions, training recipes, and native Apple Silicon M5 Max execution commands for MiniMax H3-Max with 6-step and 8-step high-fidelity DiT denoise.
1. Golden Standard Execution Baseline (6-8 Steps / 50 Layers / 73 Frames)
cd /Users/robzomb/Documents/antigravity/cool-hopper/h3-lora-lab
export H3_PROFILE=1
export H3_NAX=1
export H3_ZERO_COPY_WEIGHTS=1
export H3_REUSE_MPS_COMMAND=1
export H3_GPU_SAMPLER=1
export OMP_NUM_THREADS=12
caffeinate -dimsu nice -n -20 ./h3 --profile \
-d "/Users/robzomb/h3-models/MiniMax-H3-PDD-8Step" \
-p "YOUR_CINEMATIC_PROMPT" \
--width 960 --height 544 \
--frames 73 \
--steps 8 \
--layers 50 \
--reuse 2 \
--use-int8-row-fc2 \
--seed 333 \
-o "outputs/raw_video.mp4"
Golden Rules:
- Canvas Geometry: Always use $960 \times 544$ (Width $\ge$ Height) to preserve 3D-RoPE rotary coordinate alignment and prevent body/limb stretching.
- Causal Chunking: Always use $73$ frames ($17 \times 4 + 5$) for exactly 32 VAE tiles (cuts VAE decode to ~40s with zero padding).
- Social Vertical Reels (9:16): Generate natively at $960 \times 544$ and master to $1080 \times 1920$ via FFmpeg hardware center crop.
- Forbidden Flags: Never use --use-reference-rope (causes ghosting) or --sol-attn/--sol-cache (causes CPU locks).
2. 10-Bit Main10 Apple Native Mastering
# 16:9 Landscape Master (1920x1080)
ffmpeg -y -i raw.mp4 \
-filter_complex "[0:v]scale=1920:1080:flags=lanczos+accurate_rnd+full_chroma_int+full_chroma_inp,cas=0.30,format=yuv420p10le[v];[0:a]aresample=48000,loudnorm=I=-14:TP=-1.0:LRA=11[a]" \
-map "[v]" -map "[a]" \
-c:v hevc_videotoolbox -profile:v main10 -pix_fmt p010le -b:v 80M -tag:v hvc1 -r 24 \
-c:a aac -b:a 320k -ar 48000 -movflags +faststart master_1080p.mp4
# 9:16 Vertical Reel Master (1080x1920)
ffmpeg -y -i raw.mp4 \
-filter_complex "[0:v]crop=ih*9/16:ih:(iw-ih*9/16)/2:0,scale=1080:1920:flags=lanczos+accurate_rnd+full_chroma_int+full_chroma_inp,cas=0.30,format=yuv420p10le[v];[0:a]aresample=48000,loudnorm=I=-14:TP=-1.0:LRA=11[a]" \
-map "[v]" -map "[a]" \
-c:v hevc_videotoolbox -profile:v main10 -pix_fmt p010le -b:v 80M -tag:v hvc1 -r 24 \
-c:a aac -b:a 320k -ar 48000 -movflags +faststart master_reel_9x16.mp4
Ultra-fast ComfyUI co-designed execution toolkit for MiniMax H3 / H3-Max Turbo video generation. Features SLA-Attention (Sparse Local Attention in middle DiT blocks 14-36), Predictive Euler Step Reuse (reuse 2), 4-step PDD distillation, Motion Context temporal video chaining, and sub-90s generation on Apple Silicon M5 Max. Activate whenever the user asks for fast H3 generation, ComfyUI H3 workflows, SLA-attention, motion context, or sub-90s / turbo video generation.
MiniMax H3-Turbo (ComfyUI Co-Design & Fast Suite)
This skill provides ultra-fast inference recipes and continuous sequence chaining based on the ComfyUI Cloud graph architecture and Sol-Engine optimizations.
1. Fast Turbo Engine (4 Steps / 45 Layers / Reuse 2 / SLA-Attention)
cd /Users/robzomb/Documents/antigravity/cool-hopper/h3-lora-lab
# SLA-Attention Token Sparsity (Blocks 14-36)
export H3_TOKEN_REDUCTION=1
export H3_TOKEN_REDUCTION_BLOCKS="14:36"
export H3_TOKEN_REDUCTION_SCALE="1.0"
export H3_PROFILE=1
export H3_NAX=1
export H3_ZERO_COPY_WEIGHTS=1
export H3_REUSE_MPS_COMMAND=1
export H3_GPU_SAMPLER=1
caffeinate -dimsu nice -n -20 ./h3 --profile \
-d "/Users/robzomb/h3-models/MiniMax-H3-PDD-8Step" \
-p "YOUR_PROMPT" \
--width 960 --height 544 \
--frames 107 \
--steps 4 \
--layers 45 \
--reuse 2 \
--token-reduction 1 \
--use-int8-row-fc2 \
-o "outputs/turbo_raw.mp4"
2. Motion Context (Temporal Sequence Chaining)
To extend a video continuously across multiple shots without visual jumping:
# Pass the last frame of Clip 1 as the first frame of Clip 2
./h3 --profile \
-d "/Users/robzomb/h3-models/MiniMax-H3-PDD-8Step" \
-p "Continuation prompt..." \
--first "outputs/clip1_last_frame.jpg" \
--frames 107 --steps 4 --layers 45 --reuse 2 \
-o "outputs/clip2_raw.mp4"
3. Automated Unified Runner
You can also run directly via:
./h3_max_suite/inference/h3_max_engine.sh "YOUR_PROMPT" "outputs/my_video.mp4" [STEPS=4] [WIDTH=960] [HEIGHT=544] [LAYERS=45] [FRAMES=107] [REUSE=2]
Autonomous AI agent execution skill for high-speed, photorealistic MiniMax-H3 video and 48kHz audio generation on Apple Silicon (Metal 4 / MLX / Pure C). Includes Fast Master (8-step) and FastVideo (4-step) presets, hardware auto-profiling, and broadcast mastering.
H3MLX: AI Agent Execution Skill for MiniMax-H3 on Apple Silicon
This skill enables any autonomous agent (such as Hermes, Antigravity, or LangChain agents) to programmatically invoke, benchmark, and orchestrate photorealistic video and synchronized 48kHz audio generation using the high-performance C/Metal 4 MiniMax-H3 engine on Apple Silicon.
⚡ Agent Capability Matrix
🛠️ Autonomous Agent Invocation Protocol
When an AI agent needs to generate or master video assets:
Step 1: Tool Execution via Bash
# Execute the Master CLI with designated preset, prompt, and optional conditioning
./h3_master_cli.sh [preset_id] "[descriptive prompt]" [width] [height] [frames] [first_frame_path]
Step 2: Causal Temporal Alignment Rule
Agents must align frame counts to the causal temporal lattice formula:
$$\text{Frames} = 17n + 5 \quad (n \ge 1)$$
- $n=1 \to 22 \text{ frames}$ ($0.9\text{s}$ @ 24fps)
- $n=2 \to 39 \text{ frames}$ ($1.6\text{s}$ @ 24fps)
- $n=3 \to 56 \text{ frames}$ ($2.3\text{s}$ @ 24fps)
- $n=5 \to 90 \text{ frames}$ ($3.8\text{s}$ @ 24fps)
- $n=8 \to 141 \text{ frames}$ ($6.0\text{s}$ @ 24fps)
- $n=11 \to 192 \text{ frames}$ ($8.0\text{s}$ @ 24fps)
Step 3: Structured Output Parsing
The script outputs structured JSON metadata upon completion:
{
"status": "success",
"preset": "champion",
"frames": 39,
"resolution": "640x640",
"raw_output": "outputs/raw_champion_39f_TIMESTAMP.mp4",
"master_output": "outputs/master_champion_39f_TIMESTAMP.mp4",
"audio_spec": "48000Hz stereo AAC (-14 LUFS EBU R128)",
"metrics": {
"gpu_denoise_sec": 12.55,
"vae_decode_sec": 9.88,
"total_latency_sec": 44.92,
"throughput_fps": 3.11
}
}🧠 Environment & Subagent Tool Equipping
To equip Hermes or subagents with this skill:
- Place the repository in the agent's active workspace or skills folder (.agents/skills/h3mlx or ~/.agents/skills/h3mlx).
- The agent reads SKILL.md and discovers available execution binaries (./h3_master_cli.sh, ./h3).
- The agent can trigger generation, inspect outputs via ffmpeg, and report status back to user conversations seamlessly.