原文: MiniMax-H3-Codex-Drama
chiphoton · 出处
Create, enhance, diagnose, and optionally execute MiniMax H3 audio-only prompts through the verified fixed 32x32 ComfyUI proxy. Use for prompt-generated dialogue, narration, announcements, ambience, foley, robot sounds, or animal calls when no reference media is required; use R2V for reference-controlled voices or sounds.
MiniMax H3 Audio
Create native H3 audio without retaining a video. The bundled workflow gives the joint FL2VA sampler a disposable 32x32 visual latent, decodes the audio branch, and saves lossless FLAC through SaveAudio.
Route the request
- Use this skill for audio invented entirely from text: one or more voices, narration, announcements, ambience, foley, robot sounds, or animal vocalizations.
- Use minimax-h3-reference-to-video when an image, video, or audio reference must control identity, performance, voice, or sound. This audio proxy is prompt-only and is not a voice-cloning workflow.
- Keep requested duration between 0 and 15 seconds. Split longer material into coherent clips.
- Never expose visual size as a creative control. The proxy must remain fixed at 32x32.
Write the prompt
Return a production-ready prompt when generation was not requested. Preserve user-supplied spoken words exactly inside H3 dialogue tags.
Use a compact structure:
[Integrated audio description: speaker or source, delivery, mood, acoustic setting, event order, and timing]
Dialogue: <d>[English] Exact spoken line.</d>
Soundscape: [diegetic ambience and effects, or "clean studio recording; no music"]
Non-diegetic music: [music direction, or "none"]
For people, specify only attributes that affect sound: adult/child/older speaker, vocal register or timbre, pace, intensity, mood, and scenario such as documentary narration, news report, or airline announcement. Do not rely on vague labels like “woman voice” alone.
For speech, fit the script to the duration. As a starting estimate, allow roughly two to three spoken words per second, then leave additional room for pauses, ambience, and sound effects. Exact wording is stochastic and must be checked after generation.
For non-speech audio, state the source, event count and order, distance, room or outdoor acoustics, and whether overlap is allowed. Say no intelligible human speech when applicable. For animals, describe audible behavior rather than human emotions or dialogue.
Execute through ComfyUI
When the user asks to generate, test, submit, render, or retrieve the audio, follow ../minimax-h3-comfyui/SKILL.md in audio mode. Preserve the finished prompt verbatim.
- Use Turbo by default. Use standard sampling only when the user asks for quality-first generation or supplies [turbo=false].
- Prepare audio-turbo.api.json or audio.api.json through the shared preparer; do not hand-edit node IDs.
- Save the result as FLAC. The prepared workflow has no video decode, mux, or video-save branch.
- If the ComfyUI skill delegated prompt enhancement back to this skill, return the prompt to that workflow and do not recursively start another execution pass.
Verify the result
Treat the output as generated source audio, not a mastered deliverable.
- Confirm the exact ComfyUI prompt ID and history entry before claiming success.
- Confirm the expected duration, decodable FLAC, and native 32 kHz stereo stream.
- For speech, transcribe the result and compare it with the locked line. Distinguish punctuation or number formatting from missing or changed words.
- For non-speech audio, listen or inspect a spectrogram for event order, separation, and absence of unwanted dialogue.
- Measure integrated loudness, true peak, and clipped samples. H3 can generate hot peaks; flag or normalize downstream rather than describing raw output as broadcast-ready.
Read references/verified-behavior.md when reporting capability, limitations, or test evidence.
Return
State the mode and variant, duration and seed, prompt ID, output filename, transcript or event check, and any peak/clipping warning. Return the playable local audio when available.
Prepare, validate, submit, monitor, and retrieve MiniMax H3 text-to-video, first/last-frame image-to-video, reference-to-video, and verified 32x32 audio-only workflows on a local or self-hosted ComfyUI instance. Use when a user explicitly asks to run, generate, execute, queue, preview, or fetch MiniMax H3 work through ComfyUI. Uses pinned Comfy-Org templates and bundled Turbo variants, and patches known fields deterministically; it does not adapt arbitrary custom workflows.
MiniMax H3 ComfyUI
Run a finished MiniMax H3 prompt through ComfyUI using bundled T2V, I2V, R2V, or audio-only workflows. Turbo is the default variant; preserve the selected pinned graph unless the user explicitly overrides a declared field.
Interpret control flags
Recognize these case-insensitive controls anywhere in the request:
For every <boolean> below, accept true, on, yes, or 1 as true and false, off, no, or 0 as false.
- [return=<boolean>]: wait for completion and fetch results when true; submit and return the prompt_id immediately when false. Default true.
- [prompt_enhance=<boolean>], [pe=<boolean>]: improve the prompt through the matching MiniMax H3 prompt specialist when true; preserve it when false. Default false.
- [preview=<boolean>]: show or hide the ComfyUI browser after loading the prepared workflow. This has no effect when load_workflow resolves to false. Default true.
- [load_workflow=<boolean>]: replace the live canvas with the prepared workflow when true. When false, stay headless: do not initialize or open a browser, and do not announce preview or canvas behavior. Default false.
- [turbo=<boolean>]: use the 6-step MiniMax H3 Turbo LoRA workflow when true; use the original 20-step workflow when false. Default true. In particular, [turbo=off], [turbo=no], and [turbo=false] disable Turbo.
Strip control flags before sending text to MiniMax H3. Treat unambiguous natural-language name/value execution settings, such as "use seed 42" or "at 8 sampling steps," as controls and remove only those clauses from the prompt payload. Preserve ambiguous operational wording as prompt text; do not invent a setting from phrases such as "the original sampler."
Match every boolean flag name and value case-insensitively. If the same flag appears more than once, use its last occurrence. Treat any other value as invalid control syntax and stop before workflow preparation.
Without an enhancement flag, treat the prompt as finished: do not silently rewrite, expand, translate, or “improve” it. If enhancement is enabled, load and apply exactly one matching specialist:
- T2V: ../minimax-h3-text-to-video/SKILL.md
- I2V: ../minimax-h3-frame-to-video/SKILL.md
- R2V: ../minimax-h3-reference-to-video/SKILL.md
- Audio only: ../minimax-h3-audio/SKILL.md
Select the bundled workflow
- Explicit audio-only output with no controlling media: assets/workflows/audio-turbo.api.json by default; audio.api.json when Turbo is false. Keep its disposable visual latent fixed at 32x32.
- No controlling media and video output: assets/workflows/t2v-turbo.api.json by default; t2v.api.json when Turbo is false.
- One literal first frame, or literal first and last frames: assets/workflows/i2v-turbo.api.json by default; i2v.api.json when Turbo is false.
- Media used for identity, style, motion, camera, performance, voice, music, or rhythm: assets/workflows/r2v-turbo.api.json by default; r2v.api.json when Turbo is false.
Use the .api.json files for validation and execution. Use each variant's matching .ui.json file for provenance and browser preview. The Turbo copies preserve the original mode graph while inserting MiniMaxH3TurboLoRA after the diffusion loader, replacing the sampler selector with MiniMaxH3TurboSampler, and setting simple scheduling with 6 steps. Reject arbitrary attached workflow JSON at runtime; custom-workflow adaptation is intentionally deferred.
Execute the workflow
Read references/runtime.md before using ComfyUI. Follow it in order:
- Parse and strip controls, resolve configuration, select the mode and variant, and validate request-derived settings without connecting or uploading. Turbo rejects sampler or scheduler overrides and steps outside 4 through 8; stop before side effects on an invalid setting.
- Test reachability. A successful default connection check is silent.
- Resolve installed models and, for Turbo, the custom nodes and Turbo LoRA conservatively; never install or download without explicit permission.
- Inspect attached assets and upload them through the matching ComfyUI media tool.
- Run scripts/prepare_workflow.py to patch only the manifest-declared fields, then validate the prepared API graph.
- If explicitly requested, load or preview the matching prepared UI graph.
- Submit once. Respect the return behavior above. For awaited runs, prefer the live sampler ETA from ComfyUI's log stream to fixed-interval polling, while binding completion and errors to the exact prompt_id.
- On completion, fetch the output audio or video to a temporary or user-selected directory and return it with the prompt_id.
Do not submit a workflow while required media, a compatible model choice, or validation errors remain unresolved.
Preserve selected defaults
Unless explicitly supplied by the user or non-empty configuration, preserve the selected workflow's prompt-independent defaults for resolution, seed, scheduler, steps, denoise, reference sizing, model filenames, Turbo LoRA, and output prefix. Turbo pins its custom sampler, simple scheduler, LoRA strength 1.0, low-VRAM mode off, and 6 steps. A Turbo step override must remain from 4 through 8; disable Turbo to select another sampler or scheduler.
Audio mode is the exception to resolution configuration: it always patches the disposable visual latent to 32x32 and rejects any other explicit size. It accepts no first/last frame or reference media; route reference-controlled sound or voice work to R2V.
Allowed deterministic patches are:
- prompt
- uploaded media filenames and reference connections
- width and height together
- duration, converted to the H3 17k+5 frame grid at 24 fps
- seed
- compatible model filenames
- compatible Turbo LoRA filename
- output filename prefix
- named sampler and scheduler overrides for the standard workflow
- 4–8 steps and denoise overrides for Turbo, or steps and denoise overrides for the standard workflow
- R2V reference-size overrides
Never patch fields by searching for example text or relying on UI coordinates. Use assets/workflows/manifest.json and the preparer script.
Return a concise execution report
For completed jobs, return the selected mode, relevant settings, prompt_id, and fetched audio or video. For asynchronous jobs, return the mode and prompt_id plus how to request status or results later. For failures, return the failed node/error, what was checked, and the smallest next action.
Do not claim the job completed merely because it left the queue; confirm its exact history entry or fresh output file.
Advise, enhance, diagnose, and route MiniMax H3 video or audio prompts across official H3 style workflows, text-to-video, first/last-frame, multimodal reference-to-video, precise video-editing, verified 32x32 audio-only generation, and optional ComfyUI execution. Use when a user has a media idea, draft prompt, failed generation, uncertain input strategy, asks which MiniMax H3 workflow to use, or explicitly asks to run it through ComfyUI. Grill one decision at a time unless the user requests fast mode, use an official h3style overlay only on a clear match, and invoke ComfyUI only on explicit execution intent.
MiniMax H3 Adviser
Turn an idea, draft, failure report, or asset set into the right MiniMax H3 workflow and a finished prompt. Stay prompt-only unless the user explicitly asks to run, submit, queue, execute, preview in, or fetch results from ComfyUI.
Start by classifying the request
Choose one entry path:
- Build: turn an idea into a prompt.
- Enhance: preserve the user's intent while improving a draft prompt.
- Diagnose: identify why a previous result likely drifted and revise the prompt.
- Recommend: choose a workflow, template, vocabulary, or input strategy.
Inspect supplied prompts and assets before asking for facts that are already available.
Detect fast mode
Enter fast mode when the user uses any case-insensitive phrase below or clearly asks for an immediate answer:
- use your best judgement or use your best judgment
- help me handle the rest
- skip the grilling
- [mode=fast]
- answer immediately
- give prompt immediately
In fast mode:
- Ask no more questions.
- Make conservative creative assumptions.
- Label only assumptions that could materially change the result.
- Route to a specialist and finish the prompt immediately.
Grill in guided mode
Ask exactly one question per turn and wait for the answer. Include a recommended answer with each question. Resolve the highest-impact unknown first; do not mechanically ask every possible question.
Use this order when relevant:
- Clarify the intended viewer experience or edit outcome.
- Inventory existing text, images, videos, audio, and first/last frames.
- Resolve what each asset controls and what it must not influence.
- Resolve duration, aspect ratio, and shot structure.
- Resolve the action timeline, camera, look, and audio.
- Resolve must-preserve details and likely failure modes.
Stop grilling once the specialist can produce a coherent prompt. Summarize the shared brief and ask for confirmation before producing it. Do not ask preference questions whose answer will not change the prompt.
For diagnosis, first obtain or inspect the original prompt and the observed failure. Ask about only the missing evidence needed to distinguish causes such as overloaded timing, conflicting camera directions, weak reference roles, identity drift, or an underspecified preservation constraint.
Apply an official style workflow when it materially helps
The separate h3style plugin contains Codex adapters for the official MiniMax-H3 skills. Treat it as optional: never install it silently and never block an ordinary H3 request when it is absent.
Choose at most one official style workflow, and only when the user explicitly selects it or the request clearly matches its specialty:
Use the minimalist product skill only when the clean premium product-film grammar is central. Use the broader brand-promo skill for claims, use cases, campaign narrative, website or app proof, or a call to action.
When a matching skill is available, load and apply it in adviser overlay mode: extract its confirmed facts, style DNA, narrative or beat grammar, asset roles, preservation constraints, negative direction, and QC checks into a compact style brief. Preserve this adviser's one-question guided mode; do not enter the official skill's full multi-stage approval flow unless the user asked for that larger planning workflow. Ignore MiniMax Hub-only canvas or hub_* operations—the h3style adapter owns their Codex translation.
If the adviser was itself loaded by an active h3style skill, or the request already contains a style brief marked with the same h3style:<skill>, do not reload that style skill. Continue with the supplied brief so the skills cannot recurse.
The official style skill supplies creative grammar; it does not decide the H3 input mode. Continue routing by the actual job of the supplied media.
Use h3style:h3-prompt-writing separately, as a final formatter, only when the user explicitly requests the official structured H3 schema or the confirmed target accepts fields such as integrated_multimodal_description, subject_definitions, or retention_analysis. Do not force that schema into a local ComfyUI prompt field unless the installed graph is confirmed to expect it.
Route by production scope and input role
If the user requests a finished multi-shot film rather than one prompt or one generated clip, carry the official style brief into ../minimax-h3-drama-producer/SKILL.md. Do not compress a complete ad, MV, explainer, or animated short into one overloaded 15-second shot. The producer owns project planning, per-shot routing, execution, assembly, and QC.
For one prompt or one clip, use references/workflow-map.md for the complete routing table and provider-neutral starting settings, then choose one prompt specialist: