Why AI Short Drama Prompts Break — and How to Rewrite Them Like a Shot List
Most AI short drama prompts break because they are written like a novel paragraph — full of mood, metaphor, and stacked actions — when video models need a shot list: one subject, one action, one move, one look, per shot.
If your AI short drama clips come back with drifting faces, changed outfits, impossible camera moves, or characters doing three actions at once, the problem is usually not the model. It is the prompt language.
Video models do not read prose the way a human reader does. They do not infer "she is angry, so she slams the door, then turns, then speaks" as three separate beats. They compress. They blend. They pick one visual thread and run with it. The fix is to stop writing prompts like a scene in a novel and start writing them like instructions on a shot list.
What "shot list language" actually means
A shot list is a production document. Each line answers a small set of concrete questions: who is in frame, what are they doing, where are they, how is the camera moving, what does it look like, and what must not happen.
Shot list language is:
- Specific about the visible action. "She slams the door with her right hand" beats "she is furious."
- Specific about the camera. "Slow push-in on her face" beats "cinematic shot."
- Specific about constraints. "No logo, no extra people, face stays consistent" beats nothing.
- Modest about motion. One continuous action beats a five-beat action sequence.
- Disciplined about reference. If a character has a reference image, the prompt does not re-describe their clothes or face.
Novel language is the opposite. It piles on atmosphere, internal state, metaphor, and chained actions. That works for a reader. It confuses a video generator.
The eight things every shot prompt needs
Every usable AI short drama shot prompt covers the same eight blocks. You do not need to label them in the final prompt, but you should be able to point to each one.
| Element | What it answers | Bad version | Good version |
|---|---|---|---|
| 1. Precise subject | Who is in frame, how many | A person | Female lead in medium close-up |
| 2. Action detail | What exactly happens | She reacts | She freezes, then slowly lowers the phone |
| 3. Scene environment | Where we are | A room | Dim apartment hallway at night |
| 4. Light and color | What it looks like | Cinematic lighting | Cool blue practical light from a phone screen |
| 5. Camera movement | How the shot moves | Dynamic camera | Static shot, slight handheld |
| 6. Visual style | The look of the world | Realistic | Modern urban live-action realism |
| 7. Image quality | Technical floor | HD | Sharp focus, stable face, clean frame |
| 8. Constraints | What must not happen | — | No logo, no extra characters, no outfit change |
The first six tell the model what to generate. The last two tell it where the floor is. Skipping constraints is how you get twins, watermarks, and random extras in the background.
When to use the three-part structure
Simple shots can be one paragraph. Complex shots — especially dialogue, reverses, reveals, or action with multiple beats — work better in three parts.
- Overall setup. Location, time of day, lighting, style, who is present.
- Shot-by-shot beats. One short paragraph per camera angle or action beat, numbered as Shot 1, Shot 2, Shot 3.
- Constraint pack. Quality, face stability, no logos, no extra people, style anchor, and any reference rules.
A three-part prompt for a short drama scene might look like this:
Overall: A narrow apartment hallway at night. Warm yellow wall lamp on the left, dark front door at the end. Female lead stands near the door holding a phone. Modern urban live-action realism, vertical 9:16 frame.
Shot 1: Medium shot from the waist up. She stares at the phone screen, thumb hovering. Her hand trembles slightly.
Shot 2: Slow push-in to close-up on her face. Her eyes widen. She presses her lips together and takes one small step back.
Constraints: Sharp focus, stable face, no logo, no extra people, no outfit change, smooth motion, no sudden camera shake.
Notice what is missing: backstory, inner monologue, metaphors about her heart sinking, and three actions crammed into one sentence. The prompt describes only what the camera can see.
Camera movement: one move per shot
One of the most common failure modes is stacking camera moves. A prompt that says "push in, pan left, then crane up and zoom out" asks the model to do four things at once. The result is usually a drifting, warped frame.
Use one camera move per shot. Good options for vertical short drama:
- Static shot
- Slow push-in
- Slow pull-back
- Slow pan left / right
- Slow tilt up / down
- Slight handheld
- Over-the-shoulder angle
- Low angle, locked off
If the scene needs two moves, split it into two shots. That is how real coverage works anyway.
Reference images change what the prompt should say
Reference images are the strongest tool for character consistency, but they also change how you write. When a character is bound to a reference image, the prompt should not re-describe their face, hair, or outfit. That creates a conflict: the reference says one thing, the text says another, and the model picks randomly.
The rule is simple:
- With a character reference: describe action, expression, and visible injury only.
- Without a character reference: describe appearance briefly, but expect less consistency.
- With a scene reference: describe only what changes in this shot, not the whole room again.
- With a prop reference: name the prop once and describe how it is used.
This is why professional AI short drama pipelines build a reference sheet for every shot in a fixed order: scene first, then props, then characters. The prompt then writes around those references instead of fighting them.
Action writing: small, continuous, quantified
Video models handle small continuous motion better than big explosive motion. "She slowly tightens her grip on the letter" is more stable than "she screams, throws the letter, turns, runs down the hall, and slams the door."
When you write action:
- Break chained actions into separate shots.
- Use one verb per beat where possible.
- Quantify degree: "slight smile," "one step back," "slowly turns her head."
- Prefer low-speed motion over fast motion.
- Put dialogue and sound cues in a consistent format so they do not get read as visual description.
For dialogue and sound, keep the format clean. Put spoken lines, sound effects, and music cues in distinct markers so the prompt stays visually readable. The exact symbols can vary by tool, but the principle is the same: do not bury audio inside visual prose.
How long should a single AI short drama shot be?
For vertical short drama, aim for about 10 seconds per shot block as a starting point. That is long enough for one line, one reaction, or one small action beat, and short enough to stay stable.
If a scene in the script runs long, split it. A 40-second scene is not one prompt. It is three or four shot blocks:
- Establishing action
- Character reaction
- Line delivery
- Cliffhanger beat
This also makes reshoots practical. If one beat fails, you regenerate that shot instead of throwing away an entire minute of footage.
A simple rewrite workflow
When a prompt comes back broken, rewrite it in this order before blaming the model:
- Cut novel prose. Remove internal state, metaphor, and backstory.
- Name the subject precisely. Who is in frame, at what shot size.
- Reduce to one action. If there are three verbs, split into three shots.
- Assign one camera move. Delete stacked moves.
- Add concrete light and location. Replace "cinematic" with something visible.
- Respect references. Remove outfit and face descriptions for referenced characters.
- Add the constraint pack. Stability, no logos, no extras, no outfit change.
- Shorten the shot. If it still breaks, cut duration and split the beat.
Most "model failures" disappear after steps 1 through 5.
What this approach does not fix
Shot list language improves stability and consistency, but it is not magic.
- It does not guarantee perfect lip sync by itself.
- It does not replace good character and scene references.
- It does not make a 10-second shot hold five plot beats.
- It does not remove the need for human review and reshoots.
- It does not replace editing; shots still need to be cut together into a finished episode.
Think of the prompt as a shot list, not a spell. The better the instruction, the more predictable the take — but the final cut still belongs to the person making the show.
For a full walkthrough of the AI short drama production pipeline, from idea through script, look development, shot blocking, and rendering, see the complete guide to AI short drama production. For character-specific failures such as face swaps and outfit changes, see the guide to AI short drama character consistency.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com