How AI Vertical Short Dramas Actually Get Made: A 12-Stage Production Pipeline From Idea to Final Cut
AI short drama production is not one text-to-video button. It is a staged pipeline with checkpoints between each stage—brief lock, character look lock, scene-by-scene shot instructions, then rendering and retakes. Skipping those checkpoints is where most teams lose consistency.
The mental model: a pipeline, not a prompt
Most teams new to AI short drama start the wrong way: they write a script, drop it into a video model, and hope the cast stays the same across 60 episodes. It does not. The cast changes faces, costumes drift, props disappear, and episode 12 forgets a secret episode 3 buried.
A production-grade AI vertical short drama pipeline treats the show the way a real crew does: you lock the brief before you write, you lock the cast looks before you shoot, you block scenes before you call action, and you keep takes so you can pick the good one. The AI does not replace the crew—it executes the repeatable parts once a human has approved the direction.
Below is the 12-stage pipeline as it is implemented in Maosika (猫斯卡), an AI production operating system for vertical short dramas. The stages are presented in the order they must happen; rearranging them is almost always the cause of consistency failures.
Stage 1 — Intake and creative assessment
Before any writing begins, the raw idea is scored for completeness across the dimensions that actually matter for vertical drama: genre, protagonist, core conflict, story arc, episode count, episode length, tone, hook rhythm, ending direction, and target platform / audience.
Ideas that score high enough can move straight to a locked brief. Ideas in the middle range get prompted only for the missing dimensions. Ideas that are too thin go through a full guided intake. The point is not gatekeeping—it is preventing the team from writing 80 episodes on top of a premise that was never actually decided.
Stage 2 — Creative brief lock
The brief is the first hard gate. Until it is locked, nothing downstream starts. A usable brief contains:
- A one-sentence logline
- The core conflict
- Story direction and ending direction
- Hook and payoff rhythm
- Target platform and audience
- Episode-by-episode outline at the arc level
- Notes to the writer about what must not change
- A protagonist entry that includes their character arc, not just their name and job
This is the equivalent of a greenlight meeting. You would not call a real crew to set without one; the same discipline applies here.
Stage 3 — Character lineup and visual confirmation
Characters are assembled as a lineup before any art is generated. The system checks that the cast is complete, names are valid, visual fields are filled in, and the leads match what the brief promised. This is a hard gate—an incomplete lineup blocks the next stage.
The reason is simple: if you discover in episode 20 that the second male lead was never properly defined, every scene he is in becomes a continuity repair job.
Stage 4 — Story archive (the continuity bible)
The story archive is a structured record that travels with the show for its entire run. It tracks:
- Character identities and stable traits
- Variable current state (injuries, revealed identities, changed allegiances)
- Relationship maps
- Every plot thread, marked open or resolved
- Per-episode appearance tables
- Per-batch plot summaries
- Visual descriptions of recurring props
The archive is what prevents "model amnesia." When a new batch of episodes is written, the writer is not relying on memory of everything that came before—it is fed the relevant slice of the archive: current character states, unresolved threads, recent batch summaries, and the current batch's main line. After each batch, the archive is updated before the next batch starts.
Stage 5 — Batched script writing with beat sheets first
Scripts are written in batches of a few episodes, not all at once and not one at a time. For each episode, the beat sheet and the end-of-episode cliffhanger are planned before any dialogue is written.
Vertical short drama has hard rhythm rules that are enforced at this stage:
- Golden 3 seconds: the first scene must open on conflict or a strong hook—no slow setup, no exposition dump.
- Single-episode structure: opening hook → escalating conflict → cliffhanger on the final scene.
- Payoff density: at least one small payoff per episode (a face-slap, a reversal, an identity hint, evidence obtained); a larger payoff every few episodes.
- Dialogue: short lines, generally under ~20 characters in the original Chinese writing convention (translate to tight, spoken English), no essay-speak, no long narrator explanations.
From the second batch onward, four things are confirmed before writing starts: the batch's plot direction, focus characters, threads and conflicts to advance, and the end-of-episode hooks. These become hard constraints that override archive drift.
Stage 6 — Script rule check
After writing, each episode passes through a rule-based quality check. It catches things like:
- Wrong episode titles
- Scene count mismatches
- Missing character lines
- Too little dialogue
- Placeholder text like "to be continued"
Failures are sent back for an automatic rewrite with the specific error attached, up to a capped number of attempts. Nothing that fails this check is delivered as a finished script. This is not an AI "opinion" about quality—it is a format and completeness gate, the way a script coordinator would flag a missing scene heading.
Stage 7 — Look development and style selection
The show picks one visual style from a built-in library of style manuals covering 2D, 3D, and live-action-realistic directions (urban realistic, period realistic, mature urban romance animation, 90s Japanese anime, Chinese ink-and-brush, xianxia, 3D donghua, clay stop-motion, cyber-Chinese, and others).
Each style manual includes at least:
- A character sheet guide (face anchors, materials,气质, view consistency)
- A scene and prop guide
- Video style tags
Once a style is chosen, character art, scene art, prop art, and video prompts all follow the same style path. This is what prevents the common AI failure where characters look like anime but the video output flips to photorealism.
Stage 8 — Character, scene, and prop art (look lock)
Character art is generated in a two-step process: first the archive description is polished into an art-ready prompt using the chosen style manual and hard gender rules (which the creator can override manually), then the final prompt is rendered. Historical versions are kept so you can roll back.
Scene art follows a strict rule: no people in establishing shots. Prop art is extracted episode by episode from the script using a prop-master's eye—pulling props by the exact name the script uses, then merging them into a show-wide catalog so nothing is lost across episodes.
Only art that is finished and readable is consumed downstream; half-rendered previews are never bound as references for shooting.
Stage 9 — Scene blocking into ~10-second blocks
Each episode script is cut into video scene blocks, targeted at roughly 10 seconds each, with a soft cap on body length. Longer blocks are split along action beats, paragraph breaks, and sentence endings. Crowd and generic characters are separated from the drawable main cast so they don't consume character reference slots.
What the creator sees is not "one big video for episode 3" but "episode 3 → scene 1, scene 2, scene 3…" Each scene has its own prompt, its own reference set, its own output, and its own history of takes. This is the editing-room clip list, not a single render.
The default deliverable is 9:16 vertical, not a horizontal video cropped after the fact. Other aspect ratios are supported, but vertical is the native shape.
Stage 10 — Engineered shot prompts
Each scene's prompt is built from eight components:
| Component | What it covers |
|---|---|
| Precise subject | Who or what is in frame |
| Action detail | What they are doing, with quantified motion |
| Scene environment | Where it happens |
| Lighting and color | Mood, time of day, palette |
| Camera movement | One move per shot, no stacked pans/zooms |
| Visual style | Pulled from the chosen style manual |
| Image quality | Resolution and stability anchors |
| Constraints | What must not happen |
Simple scenes use a single-block prompt; complex cinematic scenes use a three-part structure (overall setup → shot 1 / shot 2 / shot 3 → constraint pack). Shots are numbered rather than timed in absolute seconds. A mandatory fallback pack is always appended: image quality, face stability, no watermark or logo, twin/duplicate prevention for multi-person scenes, and style anchoring for non-realistic looks.
The highest-priority consistency rule at this stage: for any character that has a reference image, the prompt must not describe clothing or appearance in text again. The reference image is the source of truth; text only describes action, expression, and injury state. This single rule eliminates the majority of face-swap and costume-drift failures.
Reference images are bound in a fixed order per scene: scene art → props in this scene → character look art. Slots are only filled if an image exists; empty slots are marked as text-only rather than filled with a fabricated match. Creators can manually bind a script name ("the officer") to a specific character's locked look, manually include or exclude props, swap in a historical version of an image, and—for time-travel or flashback stories—let the system route to the period-correct look based on scene keywords.
Before prompts are finalized, a readiness check lists any missing scene / character / prop references and offers to route the creator back to art generation. Skipping is allowed (the system will shoot text-only), but the professional path is to finish look lock first.
Stage 11 — Multi-modal rendering and the take system
Shooting follows a clear path: open the episode's video workspace, confirm references (or deliberately skip), generate the shot prompts, read and edit them as a director would edit a shot list, choose model tier / aspect ratio / resolution / duration, submit, wait in the render queue, then review.
The things a creator can directly control map cleanly to on-set roles:
| Control | On-set equivalent |
|---|---|
| Edit the video prompt | Director revising the shot list |
| Swap character / scene / prop references | Changing a look or a location plate |
| Manual character binding | Fixing name mismatches and cameos |
| Include / exclude props | Controlling visual focus in the frame |
| Switch style | Unifying the visual language |
| Choose model tier | Trading quality / cost / speed |
| Review historical takes per scene | Picking the best take |
A new take is only created when the prompt or references change and you submit again—exactly like changing the shot note and rolling camera again. The render queue is isolated from writing and art tasks, scenes cannot be double-submitted while in flight, credits are pre-deducted and released on failure, hung jobs time out and are retryable, and existing vendor task IDs are reused rather than double-created. Finished videos are faststart-processed and stored with measured duration rather than trusting the vendor's reported length.
Stage 12 — Review, revise, re-render
The final stage is human. There is no automatic scoring engine that picks the "best" take, and no auto-reshoot loop that decides quality for you. The creator or showrunner watches each scene, picks takes, asks for prompt edits or reference swaps, and re-renders. Cross-scene automatic video extension is not part of the workflow—the product unit is "one scene, with multi-modal references, producing one clip." Final assembly, fine cutting, and stitching across scenes belong in a downstream edit.
Where the boundaries are
It is worth being explicit about what this pipeline does not do, because those boundaries are exactly what make it usable as a production tool rather than a demo:
- No automatic quality scoring or auto-pick of the best take—final judgment stays with the creator.
- Reference images are not a hard gate; you can skip them and shoot text-only, but quality is usually worse, which is why the professional path is look-lock first.
- Reference tokens in prompts rely on the prompt spec; creators should still verify them when reviewing.
- The ~10-second scene block target is an engineering heuristic, not a timecode-precise cut; long scenes get re-split but some blocks still run long.
- Character consistency depends on the look-development chain, not on a face-embedding verification step; period routing is rule-based, and final look quality still depends on the art and on respecting the "don't describe looks in text when a reference exists" rule.
Maosika's position is that of an AI production operating system for vertical short dramas: it hard-encodes the verified quality-control order of short-drama production—story and continuity first, then look lock, then scene blocking, then shot instructions, then multi-take selection—rather than claiming it can produce hit shows with one click or replace a crew outright. The repeatable, drift-prone, failure-prone parts become a constrained pipeline; aesthetic judgment still sits with the person making the show.
If you are planning a show and want to see how the staged pipeline behaves end to end, you can explore the workflow at Maosika.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com