How Vertical Short Dramas Actually Get Made: An 11-Stage AI Production Pipeline
AI short drama production fails not because the model is weak, but because teams skip the order of operations. The reliable path is the same one live-action crews use: lock the brief, build continuity, approve looks, block scenes, then shoot.
Most AI short drama workflows still behave like a demo: type a premise, press generate, then patch the result. That works for a single clip. It falls apart the moment you need 60 episodes with the same lead, the same clothes, the same unresolved subplot, and a cliffhanger every episode.
The production-grade approach is different. It treats AI not as a magic button, but as a pipeline with handoffs, approvals, and versioned assets at every stage. Below is the full sequence, in the order it has to happen.
The core principle: order matters more than tooling
A vertical short drama is not one long generation cut into pieces. It is a stack of controlled units: a locked brief, a structured story record, approved character and environment art, scene blocks, engineered shot prompts, rendered takes, and review passes.
A production pipeline for vertical short drama is a staged workflow where every stage produces a reviewable artifact before the next stage is allowed to start.
Skipping stages is the root cause of the familiar failures: faces change between episodes, props vanish, plots forget their own setup, episodes open with exposition instead of conflict, and the visual style drifts from 2D to realistic halfway through a scene.
The 11 stages, in order
The table below is the spine of the pipeline. Each row answers three questions: what gets produced, who approves it, and what breaks if you skip it.
| # | Stage | Deliverable | Gate before next stage | What breaks if skipped |
|---|---|---|---|---|
| 1 | Idea intake & scoring | Intake score across premise, character, conflict, platform, tone | Score determines how much guidance is needed | Vague brief, rewrites explode |
| 2 | Creative guidance | Filled gaps in premise, protagonist arc, hook rhythm, ending direction | All required dimensions covered | Writer invents missing pieces |
| 3 | Brief lock | One-line logline, core conflict, arc, hook cadence, audience, episode outline | Brief is frozen; later conflicts resolve to it | Drift across batches |
| 4 | Cast & visual confirmation | Named cast with visual fields aligned to the brief | Cast complete and fields valid | Wrong protagonist, unnamed leads |
| 5 | Story archive / continuity bible | Character states, relationships, open/resolved threads, per-episode appearance, batch notes | Archive exists and is current | Amnesia across episodes |
| 6 | Batch script writing | Beat sheets first, then dialogue, then rule check, then archive update | Batch passes rule-based quality check | Cliffhangers missing, format errors |
| 7 | Style & look selection | Chosen visual style with shared rules for characters, scenes, props, video | One style path locked for all assets | Style fracture across assets |
| 8 | Look dev: characters, scenes, props | Approved reference art for cast, locations, key items | Video only consumes finished, readable art | Face swaps, costume drift |
| 9 | Scene blocking | Episode cut into ~10-second scene blocks with cast and locations parsed | Each block has its own reference set | Long unmanageable clips |
| 10 | Engineered shot prompts | Per-block prompts with subject, action, environment, lighting, camera, style, quality, constraints | Prompt reviewed and editable before render | Model writes fiction, not shots |
| 11 | Shoot, review, re-shoot | Multiple takes per block; re-shoot after changing prompt or references | Best take selected per block | No controlled iteration |
Stage 1–3: Intake, guidance, brief lock
The intake stage scores how complete the idea is. A high-completeness premise can go almost straight to brief. A thin one needs a structured conversation that fills in the ten dimensions vertical drama actually requires: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook and payoff rhythm, ending direction, and target platform.
Episode length for vertical drama mostly lives in the 1–2 minute range, sometimes extending to 3–4 minutes. The brief is not a formality. It is the document that later batches answer to when the model would otherwise improvise.
A locked brief includes:
- A one-line logline
- The core conflict
- Story direction and ending shape
- Hook and payoff rhythm
- Target platform and audience
- Episode-by-episode outline
- Notes for the writer
- A protagonist entry that includes character arc, not just appearance
The analogy is simple: you would not call action on a set before the production meeting. The same rule applies here.
Stage 4–5: Cast confirmation and the story archive
Before a single scene is written, the cast must exist as named entities with complete visual fields, and the lead must match the brief. This is a hard gate, not a suggestion.
Then the story archive is built. This is the single most important structure for long-running short dramas.
A story archive is a structured continuity record that tracks character identity, stable traits, current state (injuries, revealed identity, changed status), relationships, open and resolved plot threads, per-episode appearance, batch summaries, and prop visual notes.
The archive solves the problem that language models are bad at: remembering. Instead of asking the model to "keep everything in mind," the writing process only receives the relevant slice of the archive for the current batch: current character states, unresolved threads, recent batch summary, and the batch's main line. After each batch, the archive is updated. This is how episode 40 still knows what happened in episode 3.
Stage 6: Batch script writing
Scripts are written in batches of several episodes, not as one endless run. Each batch follows a strict order:
- Beat sheet first, including the episode-end cliffhanger
- Then full scene dialogue
- Then a rule-based quality check
- Then archive update
The writing rules embedded in the process are not arbitrary; they are vertical drama mechanics:
- Golden 3 seconds: the first scene must open on conflict or suspense, no slow setup
- Episode shape: opening hook → escalation → cliffhanger
- Payoff density: at least one small payoff per episode (a reveal, a reversal, a piece of evidence, a status shift); a larger payoff every few episodes
- Dialogue: short lines, generally under ~20 words; no lecture-style monologue
From the second batch onward, four things are locked before writing starts: the batch's plot direction, focus characters, threads and conflicts, and end-of-episode hook. These become hard constraints. If they conflict with older archive notes, the newly confirmed creative intent wins.
The quality check is rule-based, not taste-based. It catches wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, and placeholder text like "to be continued." Failing scripts are sent back for rewrite with the specific error attached. Half-finished scripts do not move forward.
Stage 7–8: Style selection and look dev
AI drama consistency breaks most often at the visual stage because character art, scene art, props, and video prompts are generated on separate paths with no shared rules. The fix is a single style source.
A production-grade style system includes a shared manual for character rendering (face anchors, material, temperament, view consistency), scene and prop rules, and video style tags. Once a style is chosen, characters, locations, props, and shot prompts all follow the same visual path.
Look dev happens in two steps:
- Text polish: the archive description is turned into an image-generation prompt using the chosen style's character manual and hard gender rules. The creator can override this.
- Image generation: the final prompt is rendered, optionally with a reference image, and saved as a versioned asset.
Video only consumes finished, readable art. It does not grab half-rendered drafts.
The reference image rule that prevents most face swaps
This single rule eliminates a large class of AI video failures:
When a character has an approved reference image, the prompt must not re-describe that character's clothing or appearance in text. The reference image is the source of truth for appearance; text only describes action, expression, and injury.
When text and reference fight, the model invents a compromise. That compromise is the face swap, the costume change, the age shift. The rule is not "add more reference images." It is "stop contradicting the reference in text."
Reference sets are built per scene in a fixed order: scene → scene-specific props → characters. Empty slots are marked as text-only rather than filled with invented bindings. Creators can manually bind a script name to the right character, include or exclude props, swap to an older approved version of an asset, and route period-correct looks for time-travel or flashback scenes.
Scene art follows one discipline: empty scenes contain no people. Props are extracted per episode from the script's own wording, then merged into a show-wide catalog so nothing is lost across episodes.
Before shooting, the pipeline checks what reference art is missing and shows a list. You can skip it and shoot from text only — but quality is usually worse, which is why the professional path is to finish look dev first.
Stage 9: Scene blocking into ~10-second blocks
Episodes are not rendered as one long video. They are cut into scene blocks targeting roughly 10 seconds each, with a soft cap on scene text length and further splitting on action beats, paragraph breaks, and sentence endings when needed.
This is the equivalent of a clip list on an editing timeline. Each block has its own prompt, its own reference set, its own rendered outputs, and its own history of takes. Crowd characters and generic extras are separated from the named cast so they do not consume reference slots meant for leads.
The default deliverable is 9:16 vertical, not a landscape video cropped after the fact. Other ratios are supported, but vertical is native.
Stage 10: Engineered shot prompts
Prompts are where most teams quietly lose control. A novel-style prompt like "a dramatic confrontation in a rainy alley" asks the model to invent coverage. A shot-list prompt tells it exactly what to do.
An engineered prompt covers eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color tone
- Camera movement
- Visual style
- Image quality
- Constraints
Complex cinematic scenes use a three-part structure: overall setup, then shot-by-shot instructions, then a constraint pack. The rules are deliberately conservative:
- One camera move per shot; no stacking push, pull, pan, and tilt together
- Shot numbers instead of absolute timestamps
- A mandatory constraint pack for quality, face stability, and no watermark or logo
- Multi-person scenes add a "no twins / no duplicates" guard
- Non-realistic styles anchor the style explicitly
- Actions are specific and quantified; slow continuous motion is preferred over high-dynamic bursts
- Dialogue, sound effects, and background music use a consistent notation so they are not confused with visual instructions
- The prompt only sees the current scene's assets; it does not pull characters or props from other scenes
Before delivery, prompts are cleaned of specific copyrighted IP references while preserving technique and aesthetic description, reducing downstream blocking risk. If a reference image is detected as a real-person photo, the workflow guides the creator to use platform-rendered art instead.
The simplest way to say it: the system is teaching the model to speak in shot-list language, not novel language.
Stage 11: Shoot, review, re-shoot
The shoot stage is where the pipeline behaves like a real set, not a render farm.
The creator opens an episode's video workspace, confirms reference art (or deliberately skips it), generates the shot prompt, reads and edits it, chooses model tier, aspect ratio, resolution, and duration, then submits. The job enters a dedicated queue separate from writing and art tasks. Re-submitting the same block while it is running is blocked to avoid duplicate charges and state confusion. Credits are reserved on submit and released on failure; stalled jobs time out and can be retried.
Crucially, rendering does not invent a new prompt. It uses the prompt you confirmed or edited, with the reference set locked at that moment. If you change the prompt and render again, that is a new take — exactly like revising shot notes and calling action again.
The professional control points map cleanly to on-set roles:
| Control point | On-set equivalent |
|---|---|
| Editing the shot prompt | Director revising shot notes |
| Swapping character / scene / prop art | Changing a look or location board |
| Manual character binding | Fixing name mismatches and cameos |
| Including or excluding props | Controlling visual focus in the frame |
| Switching style | Unifying the visual language |
| Choosing model tier | Trading quality, cost, and speed |
| Reviewing historical takes per block | Multi-take selection |
There is no automatic "choose the best take" engine. Final quality judgment stays with the creator and producer. The pipeline gives you multiple takes and a controlled way to re-shoot after changing something; it does not pretend to replace the director's eye.
Where Maosika fits in this pipeline
[Craftdoc note: tier 艺 — explain the industrial solution first, then state what Maosika solidifies.]
Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not replace the writer, director, or producer. It solidifies the handoffs described above into an enforced sequence: intake scoring, brief lock, cast confirmation, story archive, batch writing with rule checks, shared style path, look dev for characters/scenes/props, ~10-second scene blocks, engineered shot prompts, queued rendering, and multi-take re-shoots.
The 18 digital specialists in the system mirror real crew roles — archivist, producer, writer, script supervisor, casting director, art director, costume and makeup, props, storyboard artist, director of photography, director, camera operator, editor, and VFX supervisor — and their work is visible as a streaming production log rather than hidden inside a black box.
It is honest about what it does not do:
- It does not auto-score or auto-pick the best finished take; final review is human.
- Reference art is not a hard gate; missing art warns you but allows text-only shooting, with the expected quality tradeoff.
- It does not build a cross-scene automatic video continuation workflow; the unit of work is one scene block to one clip.
- Character consistency depends on the look-dev asset chain, not on facial embedding magic; period routing is rule-based, and final look still depends on art quality and prompt discipline.
That honesty is the point. Maosika turns the repetitive, drift-prone, failure-prone parts of short drama production into a constrained pipeline, while keeping aesthetic judgment on the human side of the table.
High-quality short drama is never one long generation cut up. It is a sequence of controlled units, each reviewable, each re-doable, each anchored to assets that were approved before the camera rolled.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com