Most bad AI video comes from good English. People write a lovely sentence describing a scene, the model renders a lovely sentence describing a scene, and the result is a moving stock photo — pretty, static, and useless.
The fix is a shift in posture. Stop describing. Start directing. A director does not tell the crew "a beautiful, cinematic kitchen scene". They call a shot size, a lens move, a light source, and one action. That is what Hailuo responds to.
The seven-part formula
Work through these in order. You will not always need all seven, but the first four are close to mandatory.
- Subject — who or what, described specifically enough to stay the same across takes.
- Action — one beat. One.
- Setting — where, and when.
- Camera — shot size plus movement.
- Light and style — the source of the light and the look it produces.
- Sound and dialogue — what you want to hear, with spoken lines in quotes.
- Constraints — what to keep out of frame.
Part by part
Subject
Vague: a woman. Better: a woman in her fifties with short grey hair and a navy work apron.
Specific nouns give the model something to hold on to between shots. If a character has to survive across several generations, describe them once, save the sentence, and reuse it verbatim — or better, attach a reference image and stop describing them at all.
Action
This is where most prompts break. "She picks up the mug, walks to the window, and looks out at the rain" is three beats. In a six-second clip the model will attempt all three, rush all three, and land none.
Pick the one that matters: She lifts the mug to her mouth and pauses. Then generate the walk as a second shot if you need it.
Setting
Place and time of day, briefly. A cluttered home office at dusk. A wet street at 2am. Time of day does an enormous amount of work because it implies the light before you have said anything about lighting.
Camera
Two components: how close, and how it moves.
| Instead of | Say |
|---|---|
| cinematic | Medium close-up, slow push in |
| dynamic | Handheld, tracking behind the subject |
| epic | Wide shot, slow crane up |
| dramatic angle | Low angle, static, subject fills the left third |
| beautiful shot | Over-the-shoulder, shallow depth of field |
The left column is a mood. The right column is an instruction. Only one of them can be executed.
Light and style
Name the source, not the vibe. Hard midday sun through vertical blinds. Single practical lamp, everything else falling into shadow. Overcast, flat, no shadows. If you want a film look, name a stock or an era rather than writing "cinematic" again.
Sound and dialogue
Hailuo generates stereo music, ambient effects, and character speech in the same pass as the picture, synced to the action — so audio is a prompt input, not an afterthought.
- Ambient: Room tone, distant traffic, a fridge hum.
- Score: Sparse piano, slow. Or, just as usefully, no music.
- Dialogue: put the line in quotes and say who says it. She says, "You're early."
Keep spoken lines short. A ten-word line in a six-second clip will be rushed.
Constraints
Say what you do not want, especially the things models like to add uninvited: No on-screen text. No lens flare. No crowd in the background. No music.
A prompt, rewritten
Before:
- A cinematic shot of a man drinking coffee in a cafe, beautiful lighting, highly detailed, 4k, masterpiece
After:
- Medium close-up of a man in his thirties in a charcoal jacket, sitting at a window table in a near-empty cafe. He lowers his coffee cup and glances off-frame left. Camera holds static, shallow depth of field. Morning sun through the window rakes across the table, warm highlights, deep shadow on the far side of his face. Ambient: espresso machine hiss, low murmur, no music. No on-screen text.
Same idea, but the second version specifies a shot size, a single action, an eyeline, a light source and direction, an audio bed, and an exclusion. Every one of those is something the model would otherwise decide for you.
Multi-shot prompts
For sequences, write each shot as its own block with its own camera and action, and keep the subject description identical across all of them. Give the sequence an arc — the shots should escalate rather than restate.
- Shot 1 — Wide. The workshop, lights off, one window. Static.
- Shot 2 — Medium. The same woman in the navy apron flicks the bench light on. Slow push in.
- Shot 3 — Close-up on her hands as she sets a tool down. Static, shallow focus.
Three shots, one beat each, escalating from empty room to hands. That reads as a scene. Three variations on "she works in the workshop" reads as three attempts at the same shot.
References beat adjectives
Words set the look; references set the specifics. Hailuo 03 takes up to 50 references in a single generation, mixed however you like — images, video clips, and audio.
If you have rewritten a description three times and the face, the product label, or the brand colour still drifts, that is the signal to attach a file rather than reach for another adjective.
Five mistakes that waste credits
- Adjective stacking. "Cinematic, hyper-detailed, award-winning, 8k" adds nothing and pushes real instructions further down the prompt.
- More than one action. The single biggest cause of mangled motion.
- Describing a photo, not a shot. If your prompt would work as a caption for a still image, it has no camera in it.
- Fighting drift with words. Consistency problems are a reference problem, not a wording problem.
- Drafting at full resolution. Iterate at 768p on a fast model, then run the winning prompt at 2K.
FAQ
How long should a Hailuo prompt be?
Long enough to cover subject, action, setting, and camera — typically two to four sentences. Length past that usually means adjectives, not information.
Can I write dialogue in a prompt?
Yes. Put the line in quotes and attribute it. Dialogue is generated with the picture and lip-synced, so keep lines short enough to fit the clip.
Do negative instructions work?
Yes. "No on-screen text", "no music", and "no lens flare" are worth including whenever a model keeps adding something you do not want.
Why does my character look different in every clip?
Because words cannot pin a face. Attach the same image reference to every generation in the sequence.
Should I prompt differently for different models?
The formula holds across Hailuo 03, 2.3, 2.3 Fast, and 02. What changes is headroom: 03 takes far more references and a longer clip, so it can carry a denser multi-shot prompt.
What resolution should I prompt at?
Prompt the same way at any resolution. Draft at 768p to save credits, then rerun the prompt that worked at 2K.


