AI Vertical Short Drama Production: The Industrial Order From Brief Lock to Final Take

Maosika Editorial | Last updated

Most AI short drama failures are not model failures — they are order failures. Teams try to generate video before locking the brief, or shoot scenes before the cast is in wardrobe. The fix is an industrial sequence that mirrors a real set.

Why "order" matters more than prompts

When a live-action crew shows up on day one without a locked script, without cast fittings, and without a shot list, nobody blames the camera. They blame the production manager. AI short drama production is no different — the toolchain is new, but the failure modes are old.

A vertical short drama production pipeline is a staged workflow where each stage produces a reviewable artifact before the next stage is allowed to consume it. The artifact is the handoff: a locked brief, a cast lineup, a continuity bible, a scene breakdown, a set of reference images, engineered shot prompts, and per-take renders.

This is the core difference between a toy workflow and a production workflow. In a toy workflow you type a premise and get a video. In a production workflow you cannot start shooting until the brief is signed, the cast is approved, and the references are mapped — because every later stage multiplies the cost of an early mistake.

The seven gates, in order

Below is the sequence that holds up under real production pressure. The order is not arbitrary; each gate exists to prevent a specific, expensive failure downstream.

GateWhat you produceThe failure it prevents
1. Intake & creative evaluationA scored intake; a guided fill of missing dimensionsVague premises that drift episode to episode
2. Creative brief lockLogline, core conflict, arc, hook rhythm, platform, audience, episode outlineRewrites cascading into 50+ episodes
3. Cast & visual confirmationNamed cast with visual fields aligned to the briefCharacters who change face, gender, or age mid-series
4. Story archive (continuity bible)Character states, relationships, open/resolved threads, per-episode appearance, prop visual notes"The model forgot" plot holes and dropped threads
5. Batched scriptwritingBeat sheets first, then dialogue, then rule-based QC, then archive write-backCliffhangers that don't cliffhang; placeholder filler; missing cast lines
6. Look dev & asset readinessStyle manual selection; character / scene / prop sheets; pre-shoot kit checkStyle breaks between characters and background; missing references mid-render
7. Scene-blocked shooting~10-second scene blocks, per-block reference maps, engineered shot prompts, multi-take outputOne giant prompt per episode that cannot be reviewed, re-shot, or edited

Gate 1–2: Intake and brief lock

The intake is scored 0–100 on completeness. Above 80, you can go straight to brief; 40–79, you fill only the missing dimensions; below 40, you walk through the full guided flow. The point is not bureaucracy — it is that the brief is the contract every later stage reads from.

A locked brief for vertical short drama includes at minimum: the logline, the core conflict, the story direction, the ending direction, the hook and payoff rhythm, the target platform and audience, the episode-by-episode outline, and notes for the writer. The lead character entry must include the character arc, not just a job title and a hair color.

Vertical format has its own dimensions that get pinned here: genre, protagonist, conflict, trajectory, episode count, per-episode length (typically 1–2 minutes, capped around 3–4), tone, payoff density, ending direction, platform, and audience. These are not creative flourishes; they are constraints the writer and the shot engine both need.

Gate 3–4: Cast confirmation and the story archive

Cast confirmation is a hard gate, not a suggestion. The lineup must be complete, names valid, visual fields filled, and leads aligned to the brief. Until that passes, there is no next button.

The story archive is a structured continuity record that travels with the production for its entire run: who each character is, their stable traits, their current mutable state (injuries, revealed identities, changed allegiances), relationships, which threads are open and which are resolved, an appearance table by episode, batch-by-batch plot summaries, and visual notes for props. It is the digital equivalent of a writers' room continuity bible.

The writing stage does not read the entire history every time. It reads a slice: current character states, unresolved threads, recent batch summaries, and the current batch's main line. This is how a show can run to long episode counts without the model "forgetting" what happened in episode 3.

Gate 5: Batched scriptwriting with rules built in

Scripts are written in batches of several episodes, not as one long dump. Each batch follows the same internal order: beat sheet and end-of-episode cliffhanger first, then the actual scenes, then a rules-based quality check, then a write-back into the archive before the next batch starts.

From the second batch onward, four things are pinned before writing begins: the batch's plot direction, the focal characters, the threads and conflicts in play, and the end-of-batch hook. These are treated as hard user-confirmed constraints; if they conflict with the archive, the confirmed creative intent wins.

The writing rules themselves are the same ones that working vertical-drama writers use:

  1. Golden 3 seconds: The first scene must open on conflict or suspense. No slow pan, no lore dump, no voiceover explaining the world.
  2. Per-episode shape: Opening hook (one scene) → rising conflict → end-of-episode cliffhanger (the final scene).
  3. Payoff density: At least one small payoff per episode (a reveal, a reversal, an identity hint, evidence landed); a larger payoff every few episodes.
  4. Dialogue discipline: Short lines, typically under 20 characters in the original writing rhythm; no essay-style speeches, no explanatory narration.

A rule-based QC pass catches concrete failures: wrong episode titles, mismatched scene counts, missing cast lines, too little dialogue, placeholder text like "to be continued." Failing scripts are sent back for rewrite with the error attached, up to a fixed retry cap. Half-finished scripts do not get handed to the user.

Gate 6: Look dev and the pre-shoot kit check

Style is chosen from a shared style manual library covering 2D, 3D, and photoreal directions — urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink-painting style, xianxia, 3D donghua, stop-motion clay, cyberpunk-Chinese fusion, and more. The manual covers character sheets, scene and prop sheets, and video style tags together, so the characters, the world, and the footage all live in the same visual language.

Character sheets are produced in two stages: first the archive description is polished into a sheet-ready prompt using the style manual and hard gender rules (this step is human-overridable), then the image is generated and versioned in history. Video only consumes sheets whose status is "complete" and whose file is readable — half-rendered drafts are never silently bound as references.

Before any prompt is generated for a scene, a pre-shoot kit check lists any missing scene, character, or prop references and routes you back to draw them first. You are allowed to skip and shoot from text only — but the system tells you this is a quality tradeoff, not the default path.

Gate 7: Scene-blocked shooting and multi-take discipline

The script is cut into scene blocks targeting roughly 10 seconds each, with a soft cap on body length and re-splitting rules for longer beats. Crowd and generic characters are separated out from the drawable main cast so they do not consume reference slots. You do not see "episode 3, one big video"; you see episode 3 → scene 1, scene 2, scene 3… each with its own prompt, its own reference set, its own output, and its own take history.

Default delivery is 9:16 vertical — native, not cropped after the fact. Other ratios are supported, but vertical is the first-class output.

Each scene builds its own reference table in a fixed order: scene image first, then the scene's props, then character sheets. Slots are only filled when an image exists; nothing is invented to fill a gap. The highest-priority rule is explicit: for any character with a reference image, the prompt must not re-describe their clothing or appearance in text. The reference is the source of truth; text only describes action, expression, and injury state. This single rule eliminates the most common AI video failure — text and reference fighting each other and producing a costume or face swap mid-shot.

The shot prompt itself follows an engineered structure, not freeform prose:

  • Eight components: precise subject, action detail, scene environment, lighting and tone, camera movement, visual style, image quality, constraint pack.
  • Complexity routing: simple scenes use one block; cinematic scenes use a three-part structure (overall setup → shot 1/2/3… → constraint pack).
  • One move per shot: no push-pull-pan stacking inside a single shot.
  • Shot numbers, not timestamps: prompts reference "shot 2," not "0–3s."
  • Mandatory fallback pack: quality, face stability, no watermark or logo; multi-person scenes add twinning / duplicate-body fallbacks; non-realistic styles get an explicit style anchor.
  • Action principle: fine-grained limb movement with quantified intensity; slow continuous motion preferred over high-dynamic action that breaks.
  • Notation: dialogue in {}, sound effects in <>, BGM in ().

In plain terms: the system is teaching the model to read a shot list, not to write a novel.

After prompts are generated, you edit them like a director revising shot notes, swap references like a wardrobe change, bind names manually for nicknames or guest roles, include or exclude props to control visual focus, switch style manuals, choose a model tier for the quality/cost/speed tradeoff, and shoot multiple takes per scene. The worker renders using the prompt you confirmed and the references that were locked at the time — changing the prompt and resubmitting is a new take, exactly like a real set.

Where the industrial order actually saves you

Three failure patterns show up on almost every AI short drama production that skips the order:

  1. Brief drift. The premise changes three times in the first ten episodes because nothing was locked. Every later asset — cast, props, sets, edits — has to be redone. Locking the brief before look dev is cheaper, not slower.
  2. Face and costume swap. Characters change appearance scene to scene because each shot re-describes them in text instead of referencing a locked character sheet. The fix is not a better model; it is the "reference wins over text" rule.
  3. Unre-editable output. When an episode is one long generated video, a bad line in scene 4 means regenerating the whole episode. When it is scene blocks with per-scene takes, you re-shoot scene 4.

The production queue is built for this reality: renders run on an isolated lane separate from writing and art tasks, in-flight scenes block parallel submissions to avoid double-charging and state confusion, credits are pre-deducted and released on failure, hung tasks time out and become retryable, and already-issued vendor jobs are polled for continuation rather than re-created. This is operational infrastructure, not a demo script.

What this pipeline does not do

Honest boundaries matter — especially for working producers and directors who have seen too many "one-click hit" claims.

  • There is no automatic quality scoring or auto-pick of the best take. Final judgment sits with the creator and the producer; the system gives you multiple takes and the tools to re-shoot with revised prompts.
  • Reference images are not a hard blockade. Missing references trigger a warning, but you can still shoot from text-only — quality is usually worse, which is why the professional path is to lock looks before shooting.
  • Reference markers in prompts rely on the spec, not a second hard-merge pass. Creators should still glance at the prompt to confirm references are present.
  • Shot grammar follows the built-in camera spec; the full style manual drives look dev, while video gets the style tags.
  • The ~10-second scene block is an engineering heuristic, not a timecode-precise cut. Long scenes get re-split, but some blocks run long; final trimming and stitching belong in a later edit pass.
  • There is currently no productized cross-scene auto-extend workflow. The unit of work is "single scene, multi-modal references → single clip."
  • Character consistency relies on the look-dev asset chain, not a face-embedding verification pass. Period/costume routing is rule-based; final look still depends on sheet quality and on the prompt obeying the "don't re-describe referenced characters" rule.

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. Its position is not "no humans needed, zero-error output." It is: take the parts of production that are repetitive, drift-prone, and easy to lose control of, and turn them into a constrained pipeline — while aesthetic judgment stays with the person in the director's chair.

If you are planning a vertical short drama slate, the single highest-leverage decision you can make is not which video model to use. It is whether your production order forces a locked brief before a single frame is rendered, a locked cast before a single scene is shot, and a scene-blocked, multi-take shooting discipline instead of one long prompt per episode. High-quality short drama is never one long generation cut up after the fact. It is a stack of controllable units, each reviewable, each re-shootable, each built on the artifacts that came before it.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com