AI Film Shot Prompt Standard: Eight Elements and the Three-Part Structure

Maosika Editorial | Last updated

Most AI video prompts fail because they read like a novel, not a shot list. A usable prompt names one subject, one action, one camera move, and one set of constraints — and it separates setup from shot instructions from guardrails.

Why shot prompts break

When an AI video model produces a character whose face changes mid-shot, a camera that both dollies in and pans left at the same time, or a scene that ignores the script entirely, the problem is rarely the model alone. The prompt usually asked for too many things at once, mixed story description with camera instructions, or forgot to tell the model what *not* to do.

A shot prompt is a production instruction for one camera set-up, not a paragraph of prose. The goal is not to be poetic. The goal is to be unambiguous enough that the model produces a usable take, and structured enough that a human editor can read it in 10 seconds and know what will come out.

This is the standard used inside the vertical short drama production pipeline at Maosika (猫斯卡). It is built for 9:16 vertical micro-dramas cut into roughly 10-second shot blocks, but the same structure works for landscape narrative work.

The eight elements every shot prompt must contain

Every shot prompt, no matter how short, should account for eight elements. Missing any one of them is where drift begins.

#ElementWhat it actually controlsCommon mistake
1Precise subjectWho or what is on screen, named consistently with the scriptWriting "a man" when the script names a specific character
2Action detailThe single physical action happening in this shotStacking three actions ("she turns, runs, and shouts")
3Scene environmentWhere the shot takes place, interior/exterior, time of dayVague locations like "a nice room"
4Light and colorLighting direction, quality, and overall color toneSkipping it entirely and letting the model guess
5Camera movementOne move, or staticCombining dolly, pan, and zoom in one shot
6Visual styleThe look path shared with character and scene artSwitching style between shots in the same scene
7Image qualityResolution, sharpness, stability expectationsAssuming the model defaults to cinematic quality
8ConstraintsWhat must not appear, and stability guardrailsForbidding nothing, then being surprised by watermarks or extra limbs

These eight elements are the minimum viable shot brief. A prompt that has all eight is not guaranteed to be great, but a prompt missing one is guaranteed to be a coin flip.

One-line vs. three-part structure

Not every shot needs a long prompt. Use the complexity of the scene to decide the format.

Simple scenes — one paragraph is enough.

A single character performing a single action in a known location can be written as one block that walks through the eight elements in order. Example shape:

Subject + action + environment + light + camera + style + quality + constraints.

This covers static dialogue shots, reaction close-ups, insert shots of objects, and establishing shots with no cast.

Complex cinematic scenes — use the three-part structure.

When a shot involves multiple characters, a choreographed move, a visual effect, or a tonal shift across the shot, split the prompt into three clearly labeled sections:

  1. Overall setup — environment, time of day, light, color tone, visual style, and the spatial arrangement of characters. This is everything that is true for the whole shot.
  2. Shot-by-shot instructions — one numbered beat per camera idea. Beat 1, Beat 2, Beat 3. Each beat names the subject, the action, and the single camera move for that beat. Use shot numbers, not absolute timestamps like "0–3s".
  3. Constraint pack — the hard guardrails: image quality, facial stability, no watermark or logo, no on-screen text, no twins or duplicates in multi-character shots, and any style anchor for non-realistic looks.

The three-part structure exists because models read linearly. If you bury a constraint in the middle of a paragraph describing action, it gets treated as flavor text. If you put constraints in their own block at the end, they are applied as rules.

The rules that keep takes stable

Structure alone is not enough. These are the writing rules that turn a prompt from a description into something the model can execute.

One shot, one camera move

Do not write "dolly in while panning left and slowly zooming". Pick one move. If the scene truly needs multiple moves, split it into multiple shot beats in part two, or split the scene into separate shot blocks at the editorial level.

Compound camera moves are where AI video most often produces wobble, geometry errors, and character warping.

Use shot numbers, not timestamps

Write "Shot 1", "Shot 2", "Shot 3" inside the prompt. Do not write "0–2s", "3–5s". Different models and different renders interpret absolute seconds differently, and timestamp instructions frequently get ignored or mis-parsed. Shot numbers are a language model construct and travel reliably across providers.

Quantify action, don't dramatize it

Instead of "she is overwhelmed with emotion and dramatically collapses", write "she grips the doorframe, her shoulders drop, she slowly slides down to a seated position".

Prefer low, slow, continuous motion. High-velocity actions — fast punches, spinning kicks, rapid turns, crowds running — are where generation quality collapses first. If the script demands a burst of motion, stage it as the tail of a shot and keep the lead-in simple.

Reference images override text description

When a character has an approved reference image, the prompt must not re-describe their clothes, face, or appearance in text. Text only describes action, expression, and temporary state — wounds, sweat, tears, disheveled hair from running.

This is the single most important consistency rule. If the prompt says "wearing a red dress" but the reference shows a black coat, the model has to choose, and it chooses differently every take. The result is the infamous AI costume swap.

Mark dialogue, sound, and music with symbols

Use a consistent notation so post-production can read the prompt without guessing:

  • {Dialogue line here} for spoken lines
  • <door creaks> for sound effects
  • (low tense strings) for music cues

These markers separate what the audience *hears* from what the camera *sees*, and keep audio instructions from contaminating the visual generation.

Feed only this shot's assets

The prompt for shot 7 must only see the characters, props, and location that appear in shot 7. Do not pass in the whole cast list or the whole scene catalog. Models will happily pull in a character from another scene if their name appears in the context window.

This is a prompt engineering discipline, not a model feature: no cross-scene contamination in the input, no cross-scene contamination in the output.

Always include the fallback pack

End every prompt with a short constraint block. The exact wording can be templated, but it must cover:

  • Image quality and sharpness
  • Facial stability across the shot
  • No watermark, no logo, no on-screen text
  • For multi-character shots: no duplicate faces, no twins, no extra limbs
  • For non-realistic styles: style anchor named explicitly, so the shot does not drift toward photorealism mid-render

The fallback pack is boring. It is also where 30% of bad takes are prevented.

Sanitization before the prompt ships

Before a prompt is sent to a video model, two cleanup steps save a lot of failed renders:

  • Strip specific IP references. Remove named films, shows, or directors as direct style references. Keep the *technique* — "soft side lighting, shallow depth of field, muted teal-and-amber grade" — instead of "in the style of [specific title]". This reduces copyright and content-policy blocks.
  • Flag真人 photo references. If a reference image is detected as an actual person's photograph rather than a generated character look, route it back to the art pipeline for a drawn look-dev pass before shooting. Feeding real photos as character references is a common source of policy rejects and likeness drift.

A worked example, before and after

Weak prompt (novel prose):

A beautiful woman in a red dress is standing in a rainy alley at night, she looks really sad and then suddenly starts running as the camera moves dramatically, cinematic, 4K, high quality.

Problems: two actions (standing, then running), two camera moves implied, costume contradicts the reference, "beautiful" and "really sad" are un-filmable adjectives, no constraint pack, no light direction.

Strong prompt (three-part structure):

Overall setup. Exterior narrow brick alley, night, heavy rain. Cool blue practical light from a distant streetlamp, wet pavement reflections. Visual style: urban live-action realism. The character Li Na stands alone in frame, three-quarter view toward camera.

Shot 1. Static wide. Li Na stands still, head lowered, rain dripping off her hair. Her right hand slowly tightens into a fist at her side.

Shot 2. Slow dolly in to medium close-up. She lifts her head, eyes wet, jaw set. No other camera move.

Constraints. Sharp focus, facial features stable across the shot, no watermark, no logo, no on-screen text, no extra people in frame. Do not alter clothing or appearance from the reference image; describe only action and expression.

The second version is longer. It is also the version that produces a usable take on the first or second render, instead of the eighth.

What this standard does not do

Be honest about the boundaries so you don't blame the prompt for things it cannot fix:

  • A prompt cannot enforce perfect character identity on its own. Consistency comes from the look-dev and reference image pipeline, not from clever wording alone.
  • A prompt cannot fix bad asset preparation. If the character reference is weak, the scene art is missing, or the script beat is unclear, the best prompt in the world will still produce a mediocre take.
  • A prompt is not a replacement for editorial. Roughly 10-second shot blocks are an engineering heuristic, not frame-accurate editing; assembly, pacing, and final cut still belong in post.
  • The system teaches the model to speak in shot-list language. It does not replace the director's taste in choosing which take to keep.

Good prompts are not magic. They are the form of discipline that matches how the model actually reads — one subject, one action, one move, one block of rules, repeated for every shot in the show.

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It turns a single idea into a full production pipeline — story bible, script batches, look-dev, shot blocks, engineered prompts, and multi-take rendering — with human review at every decision point.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com