MiniMax H3 / Hailuo 3.0 · Ref2VA

Character, Audio, and Lyrics in One Performance

The original creator published the complete generation prompt for character, audio, and lyrics in one performance on X. Native-video output: 10s · 2560×1440. The creator published the complete prompt, preserved here verbatim.

Generation mode
Ref2VA
Model
MiniMax H3 (creator-labeled)
Output
10s · 2560×1440 · landscape
Tags
Music Video
Prompt provenance
Creator, verbatim

Original Prompt

Verbatim prompt as published by the source; not translated, rewritten, or completed.

10-second dark emotional K-pop music video. Realistic young East Asian woman from the character reference, natural face, real skin texture, messy hair. 
Important rules: 
- The girl is always framed in close-up or medium shots only. No wide or distant shots at all. 
- The girl performs continuously for the full 10 seconds without any cuts on her. Her movement is natural and fluid — she can shift weight, make small dance-like gestures, turn her head, move her shoulders and hands, but the camera stays on her in consistent medium/close framing. 
- Typography and background change with hard cuts according to the storyboard order, while the girl keeps performing through them.  

Sequence of text (hard cuts): 
1. “THE TABLE’S SET AGAIN” 
2. “NOBODY SITS DOWN” 
3. “YOUR WORDS ARE HARD GLASS” 
4. “FALLING TO THE FLOOR NOW” 
5. “I KEPT THE DOOR OPEN” 
6. “YOU KEPT YOUR HANDS CROSSED” 
7. “SAME ROOM, SAME SCAR” 
8. “SAME OLD COST” 
9. Final rose  The girl sings the entire song with clear, accurate lip-sync. 

Mouth and teeth must stay natural and undistorted at all times — no deformation, no melting, no artifacts. Emotional facial performance. Typography stays sharp and reacts (cracks, shifts, particles). Changing typography every 1–1.5 seconds. Heavy red atmosphere, cinematic lighting, photoreal

View original source · Original X post · @Kiber_Alla