Image to video: first frames, last frames and modes — MiniMax H3 tutorial
Animate a still image or constrain the opening and ending with two images.
Original: ComfyUI MiniMax H3 Text-to-Video, Image-to-Video, and Reference-to-Video Workflows
ComfyUI · Source
MiniMax H3 Image to Video (I2V)
Generate videos from an input image, with optional first/last-frame keyframes.
Model downloads
Model storage
ComfyUI/
├── 📂 models/
│ ├── 📂 diffusion_models/
│ │ └── minimax_h3_fl2va_pruned_int8_convrot.safetensors
│ ├── 📂 text_encoders/
│ │ └── qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
│ ├── 📂 loras/
│ │ └── minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
│ ├── 📂 vae/
│ │ ├── minimax_h3_video_vae_fp16.safetensors
│ │ └── minimax_h3_audio_vae_fp32.safetensors
│ └── 📂 embeddings/
│ └── minimaxh3_art_is_explosion.safetensors
Prompting tips
- Keyframes: The first_frame and last_frame inputs are optional; the model generates the motion between them
- Prompt: Describe the shots, motion, and the accompanying audio (dialogue, SFX, music) in one block
- Resolution: H3's native canvas is a 768px short edge, which is 1344x768 at 16:9, and resolutions are rounded to a multiple of 32
- Duration: The duration input snaps to the model's 17-frame-per-block (17k+5) grid at 24fps
- Turbo mode: Enable turbo_mode on the MiniMax H3 node to switch to the turbo LoRA and generate in turbo_steps (default 8) instead of 20 steps. turbo_model_strength controls the LoRA strength (default 1.0).
The official base-mode prompt writing guide is summarized in the prompt guide.
Audience: Users seeing H3's multiple ComfyUI workflows for the first time.
Hardware: Local or cloud ComfyUI already running H3; image-to-video reuses FL2VA weights.
Prerequisites
- One first-frame image; add a last-frame image only when the ending must be constrained
- Complete the first-run guide so the model and both VAEs are available
Steps
- Open the official I2V JSON or MiniMax H3 I2V template and confirm the FL2VA model.
- Upload the opening image and connect first_frame. Leave last_frame empty when animating a single still.
- For boundary-frame control, upload the ending image to last_frame. Choose two images that a clear action can connect.
- Choose an aspect ratio that suits the composition and begin at preview size. Follow the official example to describe motion starting from the input.
- After rendering, compare the opening and ending with the inputs and inspect unwanted subject changes. Adjust either composition or action at a time.
Caveats
- Input modes use different weights and wiring; do not merely swap input ports.
View original source · ComfyUI · 2026-09-06
In depth · Source checked: 2026-09-06 · Site testing: No generation test performed
Distinguish first frames, last frames and general references, then connect inputs in the native template. A practical route from a still image.
Local generation uses your own hardware, RAM and storage. See the original model and tool terms.
Links lead to the original workflows, models or examples. Supply your own reference assets where needed. This generation workflow has not been independently run by this site.
Applicable versions
ComfyUI 0.30+ · native H3 templates
Resources and demonstrations
Troubleshooting
Unexpected image crop
Crop the input to the target aspect ratio, then inspect the workflow preview dimensions.
A character reference became a literal first frame
Use the Ref2VA guide when you need identity reuse rather than a fixed opening composition.
Examples and techniques
Learn next