AI Vertical Short Drama Production: The Industrial Order That Prevents Mid-Shoot Chaos

Maosika Editorial | Last updated

AI short drama production breaks down when teams skip order: they shoot before the brief is locked, generate video before characters are designed, and write episodes without a continuity record. The fix is a strict industrial sequence with verifiable checkpoints at every stage.

Why "just generate" fails for vertical short dramas

The common failure pattern in AI micro-drama production is not bad prompts—it is bad sequencing. A team writes a loose logline, jumps straight into video generation, and then discovers by episode 8 that the lead's face has drifted, props from episode 2 vanished, and the plot contradicts itself. Re-shooting at that point costs far more than building the pipeline correctly up front.

AI vertical short drama production is the process of turning a single idea into a finished vertical-form episode through a locked sequence of verifiable stages: idea evaluation, creative brief, script in batches, character and environment look development, scene blocking, engineered shot prompts, multi-modal generation, and review with re-takes.

The core principle is simple: every stage produces an artifact you can inspect, reject, or lock before the next stage consumes it. You do not ask the model to remember your world—you give it a structured record to read.

The industrial sequence, stage by stage

The order below is not a suggestion. It mirrors how a real crew runs a production, with the machine handling the repetitive, drift-prone work and humans making the审美 calls at gates.

StageWhat gets producedHuman gateWhat goes wrong if you skip it
1. Idea intake & evaluationCompleteness score, missing-dimension listDecide whether the idea is ready to briefVague concepts get expanded into 60 episodes of nothing
2. Creative brief (locked)Logline, core conflict, arc, hook rhythm, platform, episode outlineSign off; no script starts until this is frozenScripts wander; later episodes can't be reconciled
3. Story archive (continuity bible)Characters, relationships, open/resolved threads, per-episode cast, prop descriptionsConfirm core cast and arcsModel "forgets" who people are by episode 20
4. Batch script writingBeat sheets first, then dialogue, then rule-based QCApprove each batch before the nextCliffhangers don't land; filler episodes pile up
5. Style & look-devStyle manual selection; character, scene, prop sheetsLock visual identity"Anime face, live-action body" mismatches
6. Scene blocking~10-second scene blocks per episodeRe-cut blocks that are too longOne giant 2-minute prompt that the model can't hold
7. Engineered shot promptsPer-block prompts with reference mapsRead and edit prompts before shootReference images and text contradict each other
8. Shoot & re-takeMultiple takes per block; select bestFinal quality callOne-and-done outputs with no recovery path

Stage 1 — Idea intake: score before you brief

Before any writing starts, the idea is scored on completeness. A mature system routes the creator differently based on how filled-in the concept is: near-complete ideas skip most hand-holding and go straight to brief; partial ideas get prompted for the missing dimensions; thin ideas go through a full guided intake.

Think of this as the development meeting. You would not greenlight a pitch in a real room that cannot name the protagonist, the core conflict, or the ending direction—there is no reason to let AI do it either.

Stage 2 — Lock the creative brief

The brief is the contract with every downstream stage. It must fix, at minimum: the one-sentence logline, the central conflict, the story direction, the ending shape, the hook and payoff rhythm, target platform and audience, episode-by-episode outline, and notes for the writer. The protagonist entry must include a character arc, not just a job title.

For vertical short dramas, the intake also pins the vertical-specific dimensions: genre, lead, conflict, trajectory, episode count, episode length (typically 1–2 minutes, capping at 3–4), tone, beat density, ending direction, and platform/audience.

Once the brief is locked, it does not move. If you want to change the core conflict in episode 30, you go back and amend the brief—you do not whisper a new direction into a single episode prompt and hope it propagates.

Stage 3 — Build the story archive before you write

The story archive is a structured continuity bible that travels with the production for its entire run.

It records: character identities, stable traits, current mutable state (injuries, revealed identities, changed relationships), relationship maps, every plot thread with an open/resolved status, a per-episode appearance table, batch-by-batch plot summaries, and visual descriptions of recurring props.

When writing a new batch, the system reads only the relevant slice of this archive—current character states, unresolved threads, recent batch summaries, and the current batch's main line. After each batch is approved, the archive is updated. Continuity comes from structure, not from model memory.

Stage 4 — Write scripts in batches, beat sheets first

Vertical short drama script writing has hard rules that must be enforced, not suggested:

  1. Golden 3 seconds: the first scene of every episode opens on conflict or悬念. No slow establishing shots, no exposition dumps.
  2. Single-episode shape: opening hook (one scene) → rising conflict → end-of-episode cliffhanger (final scene).
  3. Payoff density: at least one small payoff per episode (a reveal, a reversal, a status hint, evidence in hand); a larger payoff every few episodes.
  4. Dialogue: short lines, generally under 20 words. No essay-speak; no narrator dumping backstory.

The writing order inside each batch matters too. Beat sheets and the end-of-episode cliffhanger are planned before any dialogue is written. Past the halfway point of the series, ending constraints are injected so the story actually converges. From the second batch onward, the system confirms four things before writing: this batch's plot direction, focal characters, threads and conflicts to advance, and the end-of-episode hook. Those become hard constraints.

Every finished batch goes through a rule-based QC pass that catches wrong episode titles, mismatched scene counts, missing character lines, too-little dialogue, and placeholder text like "to be continued." Failing batches are sent back for rewrite with the specific error attached. Nothing half-baked reaches the creator's desk.

Stage 5 — Lock the visual identity before you shoot

AI productions famously drift visually because characters, scenes, and the video model are all pulling in different style directions. The fix is a shared style path: pick one style manual up front, and have character art, scene art, prop art, and video prompts all consume that same path.

A mature production system ships with multiple style manuals covering 2D, 3D, and photoreal directions—urban live-action, period live-action, mature urban romance animation, 90s anime, Chinese ink, xianxia, 3D donghua, stop-motion clay, cyberpunk-Chinese fusion, and more. Each manual includes how to draw faces consistently, how scenes and props should look, and the video style tags to inject at shoot time.

Character approval is a hard gate: the cast must be complete, names valid, visual fields filled, and leads aligned with the brief. You do not start generating video with a half-cast.

Stage 6 — Block scenes into ~10-second units

A finished episode script is not fed to the video model as one lump. It is cut into scene blocks, each targeting roughly 10 seconds of screen time, with a soft cap on the body text per block. Over-long blocks are split further along action beats and sentence boundaries.

This is the editing-room clip list, not a single render. Each block has its own prompt, its own reference set, its own output, and its own history of takes. Crowd characters and generic extras are separated from the named principal cast so they do not consume design slots meant for leads.

The default deliverable is 9:16 vertical, not a landscape video cropped after the fact.

Stage 7 — Engineer the shot prompt, don't freestyle it

A shot prompt for an AI short drama block is not a paragraph of prose. It is a structured document built from eight elements:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color tone
  5. Camera movement (one move per shot, never stacked)
  6. Visual style
  7. Image quality
  8. Constraints (face stability, no watermarks, no twins in multi-character shots, style anchor for non-realistic looks)

Complex scenes use a three-part structure: overall setup, shot-by-shot breakdown using shot numbers (never timecodes like "0–3s"), and a constraint package. Dialogue is wrapped in {}, sound effects in <>, and BGM in ().

The single most important consistency rule at this stage: for any character with a reference image, the prompt must not re-describe their clothing or appearance in text. Text describes only action, expression, and injury. The image is the authority on how they look. When text and reference fight, you get outfit swaps and face swaps.

Reference maps are built automatically per block in a fixed order—scene first, then props for that block, then character sheets—with empty slots marked as text-only rather than invented. Manual binding handles cases where the script calls a character "the officer" but the archive knows them as "Li Qiang." Props can be manually included or excluded to control visual focus.

Before prompts are generated, a readiness check lists any missing scene, character, or prop references. You can skip and shoot from text only—but you should know that quality typically drops when you do, which is why the professional order is: design first, shoot second.

Stage 8 — Shoot, review, re-take

The standard shoot loop is: open the episode's video workspace, confirm references are present (or deliberately skipped), generate the shot prompts, read and edit them, pick model tier / ratio / resolution / duration, submit, wait in the render queue, and review the resulting takes.

Crucially, the system does not re-invent prompts at render time. It uses the prompt you approved and the reference map that was locked when you submitted. Editing the prompt and submitting again produces a new take—exactly like revising a shot list and rolling camera again.

There is no automatic "best take" picker. Final quality judgment sits with the creator and the producer. The system gives you multiple takes, the ability to edit prompts, swap references, rebind characters, toggle props, switch styles, and choose model tiers—but it does not pretend to score artistic merit for you.

The honest boundaries

A production-grade system is honest about what it does not do:

  • It does not auto-score finished footage or auto-select the best take. Human eyes still call the cut.
  • Reference images are not a hard blockade. You can skip them and shoot text-only; expect weaker consistency.
  • Reference markers in prompts rely on the drafting规范; creators should still read prompts before submitting.
  • Shot grammar follows the system's built-in camera rules; full art manuals apply to the look-dev side, not frame-by-frame direction.
  • The ~10-second block is an engineering heuristic, not a timecode-precise cut. Final assembly and trimming still belong in editing.
  • There is no cross-block automatic video continuation workflow. The unit of work is a single block with its own references, producing a single clip.
  • Character consistency depends on the look-dev asset chain, not on facial embedding verification. Era-specific costumes are picked by rule, and final look still depends on the quality of the character sheets and on respecting the "don't describe what the image already shows" rule.

What this order buys you

When the sequence is respected, a single operator can move an idea through a full production run without the classic AI failure modes: faces that change mid-series, props that appear and vanish, plot threads that hang, and episodes that feel like they were written by a stranger. The machine does the repetitive, drift-prone work—maintaining the archive, drafting beat sheets, mapping references, building structured prompts, queuing renders. The human keeps the审美 judgment: which brief to greenlight, which take to keep, when a line is wrong, when a performance doesn't land.

That is the bet Maosika (猫斯卡) makes: it does not try to replace the crew with one big generate button. It enforces, in software, the order that working crews already follow—story first, then continuity, then look, then blocks, then shots, then multiple takes—and leaves the taste calls on the human side of the screen.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com