原文: h3lite
Rimagination · 出处
Use when configuring, repairing, planning, or running MiniMax H3 locally on a Windows NVIDIA computer, especially when installation compatibility, component sets, resolution, generation-time budget, path, low-VRAM risk, or H3 prompt mode must be chosen from hardware and user requirements.
MiniMax H3 Local Video
Use this skill to turn a user's local computer into a reproducible MiniMax H3 audio-video workstation and to generate short clips. The validated primary route is Windows + NVIDIA + ComfyUI; macOS Apple Silicon is a community/experimental alternative, not an equivalent tested backend. Treat platform selection, path selection, hardware-aware planning, prompt writing, execution, timing, and verification as one workflow. The default is the validated fast route; other modes must be chosen explicitly or justified by a time budget.
Platform scope and routing
Detect the operating system and accelerator before giving installation commands or model links. Do not give CUDA, Windows virtual-environment, .bat, or Windows path instructions to a macOS user.
For a non-Windows user, preserve the useful cross-platform guidance (prompt structure, 32-pixel canvas alignment, low-resolution preview, disk/log/output checks), but label every resource and timing as platform-specific. A community implementation that can run H3 on a Mac is evidence for feasibility, not evidence that this ComfyUI skill supports Metal.
The Mac alternative described in the project notes uses MLX/Metal and a third-party mmh3turbo package. If a Mac user explicitly chooses that route, point them to the author's community bundle and package (uvx mmh3turbo), and state that these are external resources with independent licensing, updates, and validation. It may reduce the download footprint with GGUF/4-bit bundles, but it is outside this skill's tested component sets. Do not silently install it, mix its weights with ComfyUI models, or present its 30–43 minute 5-second 720p timings as Windows benchmarks.
Agent production workflow
For complex creative requests, use the compact workflow contract in
references/agent-workflow.md:
intent route → reference/identity anchors → prompt enhancement → execute → verify
Route from the user's input and acceptance criteria, not from a style adjective.
Use I2VA for a specified opening frame, FL2VA/L2VA for endpoint anchors,
and Ref2VA only after its model, text encoder, and workflow are confirmed.
For recurring characters or multi-shot work, define stable subject/reference
labels, what must be retained, what may change, and what drift is forbidden
before writing the timeline. This workflow pattern is local and does not add a
cloud service, MCP dependency, second model, or second inference pass.
When a creative brief is vague (for example, it only says “more cinematic”
or “make it a nice 3D animation”), optionally read
references/prompt-assist.md. It adapts the
public Higgsfield prompt structure—stable style/identity lock, one clear scene
action, physical camera motion, audio, and compact anti-drift constraints—to
H3's native fields. Use it as a writing aid only: do not call Higgsfield, copy
model-specific flags or capabilities, or let a web lookup change the local
route, resolution, component set, or verification rules. If a live lookup is
needed, follow the host web-access skill and use public pages only.
Operating rules
- Hot path first: for an ordinary text-to-video request on an already validated installation, run scripts/h3_fastpath.py once. It combines /system_stats, fresh-cache reuse, in-process planning/preflight, queue submission, and one bounded completion watch. Do not issue repeated one-shot status calls, reread the full reference set, run --help, or ask nonessential questions during this path. If the command yields a running terminal cell, wait on that cell; do not start another monitor.
- Keep cold work out of the hot path: model download manifests, hash or size verification, Torch/custom-node repair, repository checks, browser workflow discovery, and full recursive doctor scans belong to installation, migration, repair, or first-run validation. A normal generation on an unchanged machine must not pay those costs.
- Cold path can be heavier when it prevents hour-scale waste: during installation or repair, verify download source, target folder, expected size/hash when available, runtime imports, and model-role mapping before queueing. Store the result so later prompts reuse it instead of repeating it.
- Inspect before changing anything. Run scripts/h3_doctor.py --json and locate the target ComfyUI directory before installing packages, nodes, or weights.
- On first contact, run a platform/accelerator check before the Windows doctor. If the machine is not Windows + NVIDIA, stop the CUDA installation branch and route the user using the platform matrix above.
- Prefer an isolated ComfyUI directory when no installation is supplied. Never overwrite an existing installation or silently replace model files.
- Keep the deployment path configurable. Do not copy paths from another computer into scripts or workflows.
- Report required disk space before large downloads. Use resumable downloads and verify file size or hash when a source provides one.
- Before a long generation or multi-shot batch, check free space and pagefile headroom and keep per-shot logs. A pipeline that filters away the process exit code or traceback is not a successful run; preserve the full log and stop on the first failed shot.
- If the user can only download from the public internet, run a cold-path download plan before fetching multi-GB files: test candidate raw URLs with a small ranged download, choose the fastest stable source, estimate wall-clock time, then use resumable .part downloads. Do not pretend scripts can beat the user's real bandwidth.
- Before downloading large assets, run the doctor compatibility probe. Stop on a Torch import error; treat a comfy-kitchen/Torch mismatch as a repair decision, not a post-download surprise. Do not silently substitute model files or start unlimited parallel downloads.
- Treat the diffusion checkpoint, text encoder, Turbo LoRA, workflow, and node revisions as one component set. Read references/component-sets.md during installation, migration, model replacement, or kernel repair. Never construct an unvalidated set from individually plausible filenames.
- Use --component-set auto for one unambiguous installed set, or explicitly select A/validated-low-vram-a or B/portable-16gb-b when both sets are installed. Record the selected set in the run manifest; never resolve a partial set role by role.
- Prefer the maintained Baidu package for the registered A/B component sets. Keep the selected set atomic, and respect the licenses of model weights and third-party nodes when using either the package or upstream sources.
- Prefer the ComfyUI HTTP API with an API-format workflow JSON. Use browser/CDP capture only as a recovery path when no reusable workflow JSON exists.
- Check http://127.0.0.1:8188/system_stats before starting anything. If ComfyUI is already healthy, reuse it and do not restart it or rediscover its workflow history.
- Preserve MiniMax H3's audio path and flow/sigma-shift handling when the user wants native audio. Do not remove audio VAE, audio conditioning, or the H3 sampling node merely to make a graph look simpler.
- Zero-inference optimization constraint: hardware compatibility checks, timing calibration, face routing, and media QA may run before or after generation, but must not add sampling steps, extra generation models, or a second video inference pass. Keep the selected graph unchanged unless the user explicitly requests a different quality profile.
- Face-quality routing: if the user needs a recognizable or speaking human face, do not treat low-VRAM W4A8 T2VA at 640x352 as a final-quality route. Prefer I2VA with a clear first-frame reference; prefer Ref2VA when identity must persist across shots. Read references/face-quality.md, confirm MiniMaxH3ReferenceToVideo through /object_info, and confirm the matching reference-capable text encoder/projection and workflow before selecting that route. A registered node alone is not enough; the bundled Ref2VA templates are an experimental local path until a complete run passes media and manual identity QA.
- Anchor before prompt: for multi-shot or identity-sensitive requests, first create an internal anchor sheet with stable subject/picture labels, retention rules, allowed changes, and forbidden drift. Use the same labels in the prompt, output prefix, and run manifest; read references/agent-workflow.md for the compact contract.
- Assist vague creative briefs without inventing facts: when the request lacks a concrete camera, action, sound, or finish, read references/prompt-assist.md and use its bounded defaults or ask one targeted question if the omission changes the route or acceptance criteria. A public Higgsfield lookup is optional and pattern-only; fall back to the local references when browsing is unavailable.
- On current ComfyUI builds, the API class MiniMaxH3SigmaShift is the native ModelSamplingMiniMaxH3 node and uses the merged ModelSamplingAV video/audio schedule fix. Detect it by /object_info or the local source before adding a custom dual-clock sampler; do not duplicate the fix merely because the API class keeps its compatibility name.
- Run the read-only planner before a non-trivial generation. It must report selected mode, resolution, steps, cache policy, paths, and an estimated time range. Do not present an estimate as a guarantee.
- Run the read-only preflight after the doctor and planner. Treat low available RAM/VRAM as a caution, but stop when the pagefile is critically low, required assets are missing, or the doctor recommends an alternative backend.
- Do not perform a full recursive doctor scan for every prompt. Cache the environment report under <ComfyUI>/user/h3lite_runs/_environment/; reuse it for a normal session (normally no older than 30 minutes), invalidate it after ComfyUI/model/node/driver changes or a failed run, and use h3_preflight.py --refresh-runtime for volatile resource fields.
- Do not revalidate large model files before every prompt. Trust the cached download/component manifest unless the file is missing, has a different size/mtime than recorded, the user changed components, or the previous run failed with a model/node/loader error.
- For registered Set B files, require the recorded SHA-256 on first use or after a size/mtime change. A same-size corrupted W4A8 checkpoint produced colored mosaic frames, so byte count alone is not proof of integrity; reuse the cached integrity result on unchanged files.
- Treat every submission as an auditable run: save the effective prompt, mutated API workflow, configuration fingerprint, queue ID, actual execution time, and verified output in the run manifest.
- For identity-sensitive or multi-shot runs, the runtime also writes anchors.json beside manifest.json and records advisory anchor_qa comparisons; these signals support manual continuity review but are not face recognition.
- Keep agent-facing status compact: omit ComfyUI's full history graph by default; use verbose history only when diagnosing a failure.
- Never submit an identical configuration while its manifest is submitting, queued, or running. Return the existing prompt ID instead; use --allow-duplicate only when the user explicitly asks for a second identical run.
- Treat low-VRAM timing as an empirical estimate. The first run can be much slower because kernels compile and weights move between system RAM and VRAM.
- For expensive renders, use a cheap preview pass first: validate the complete prompt/shot list at the smallest supported canvas (for example 256p or the local fast bucket), then promote only approved shots to the requested resolution. This is especially important for multi-shot work; it is a planning optimization, not a second quality-generation pass for a single requested clip.
- Prefer NORMAL_VRAM when a validated 16 GB system can keep Set B resident. In a same-model/workflow/prompt/seed 640x352 comparison, an RTX 4060 Ti 16 GB run took 77.08 seconds versus 591.22 seconds on an RTX 4070 Laptop 8 GB using LOW_VRAM; treat dynamic loading/offload as the main operational explanation, not as a pure GPU benchmark or a promise.
- When launching ComfyUI as a background process, redirect stdout and stderr to persistent files. A detached pipe can become invalid after the launching session is cleaned up, leaving ComfyUI alive but causing tqdm/logger writes to fail with OSError: [Errno 22] Invalid argument. On that signature, restart ComfyUI with persistent logs; do not redownload models or rerun a full doctor unless the restart exposes another error.
- Keep media verification attached to the selected ComfyUI root. The verifier searches system PATH, H3LITE_FFPROBE, and common locations in or beside <ComfyUI> for ffprobe; both h3_generate --watch and standalone h3_status must receive or infer that root. Standalone status may infer the parent only when --output-dir points exactly <ComfyUI>\output; otherwise pass --comfyui explicitly. Treat ffprobe_not_found as a missing verifier, not evidence that generation failed, and do not requeue the video until the existing output has been inspected.
- Treat run-history cleanup as explicit maintenance, never hot-path work. Use scripts/h3_cleanup.py in dry-run mode first and require --apply before deleting eligible run snapshots. Preserve _environment, _hotpath, _workflows, _experiments, prompt folders, timing data, and generated output files.
Preferred component download source
For installation or repair, use the maintained Baidu package before assembling
the set from multiple upstream repositories. Select one complete package after
the hardware check; do not ask the user to download both sets or mix their
exclusive files.
Guide the user to open the matching link, enter the code, and download the
whole package. If the baidu-drive skill or a Baidu Drive connector is
available, use it for the download; otherwise give the link and code directly
and continue after the user places the files locally. Merge the package's
models and
custom_nodes folders into the selected <ComfyUI> root, then import or copy
the packaged workflows and keep component-manifest.json with the install.
Run the doctor after the merge.
Set A contains the INT4 text encoder and optional low-VRAM acceleration nodes.
Set B contains the FP8 text encoder and validated compatibility workflows. Both
packages include their own shared ClipProj and VAE files, so a user only needs
one link. If the Baidu package is unavailable or the user explicitly requests
upstream downloads, use the exact sources, filenames, sizes, and hashes in
references/component-sets.md.
Installation target contract
Before installing, downloading, or moving any component, establish one explicit
deployment target and state it to the user:
Install mode: reuse-existing | current-project | dedicated-folder
ComfyUI: <absolute path>
Models: <ComfyUI>\models
Custom nodes: <ComfyUI>\custom_nodes
Output: <ComfyUI>\output
Use these rules:
- reuse-existing: use the exact existing ComfyUI path supplied by the user or discovered and confirmed by the user. Do not clone, reinstall, or create a second model directory.
- current-project: keep everything under the active workspace in <workspace>\.h3lite\ComfyUI so the project-scoped choice is unambiguous and does not scatter models across the repository.
- dedicated-folder: use the user's absolute path, preferably a non-repository path such as D:\AI\MiniMax-H3\ComfyUI or F:\MiniMax-H3\ComfyUI. Put the venv, custom nodes, models, user data, and output under this ComfyUI root.
- If no existing installation and no target path are available, recommend dedicated-folder and ask the user to confirm the absolute path before downloading large files. Never silently choose a drive or install into the current project root.
- If the user says “当前项目” without naming the workspace, resolve and display the active workspace path before proceeding. If the user gives a path ending in ComfyUI, use it directly; if they give a parent install folder, append ComfyUI and display the resulting path for confirmation.