The AI Vertical Short Drama Pipeline: 12 Stages From Idea to Finished Vertical Cut
Most AI short drama failures happen not because the model is bad, but because the production order is wrong. A reliable vertical short drama pipeline forces a strict sequence: lock the brief, build continuity, approve looks, then shoot scene by scene.
Why this guide exists
This is not a "one-click hit" tutorial. It is a production workflow map for teams that already know short dramas and want to understand how an AI-assisted pipeline should be structured before they start spending render credits.
The core idea is simple: high-quality short dramas are not produced as one long generation that gets cut up later. They are built as controllable units, each with an approvable deliverable.
A vertical micro-drama pipeline is a staged production system in which every stage produces a reviewable artifact—brief, character lineup, continuity bible, beat sheet, script, look references, scene blocks, shot prompts, and final takes—so creators can correct drift before it compounds across episodes.
The 12-stage production pipeline
The table below lays out the full sequence. The order matters more than the tool used at each step.
| Stage | Artifact produced | Who decides | Common failure if skipped |
|---|---|---|---|
| 1. Idea intake | Scored brief completeness | Creator + producer | Vague premise that drifts by episode 3 |
| 2. Creative guidance | Filled gaps in premise, audience, hooks | Creator | Writing starts before the core conflict is clear |
| 3. Brief lock | Approved logline, conflict, arc, ending direction, hook rhythm | Creator / showrunner | Constant rewrites mid-production |
| 4. Character lineup & visual confirmation | Named cast with arcs and visual fields | Creator | Characters change face, name, or motive |
| 5. Story archive / continuity bible | Structured record of identities, states, relationships, open threads | System + creator | Episode 40 forgets episode 5 |
| 6. Batch script writing | Beat sheets first, then dialogue, then rule checks | Writer + system | Weak cliffhangers, filler episodes |
| 7. Style / look selection | Unified look manual for characters, scenes, props | Art lead | Characters look 2D while scenes look realistic |
| 8. Look development assets | Character sheets, scene plates, prop catalog | Art lead + creator | Reference gaps force text-only guesswork |
| 9. Scene blocking | ~10-second scene blocks per episode | Editor / director logic | Long unwieldy clips that cannot be re-shot cleanly |
| 10. Shot prompt engineering | Per-scene reference mapping + camera-language prompts | Director / prompt editor | Inconsistent motion, wrong costumes, mixed eras |
| 11. Multimodal shoot | Rendered vertical takes per scene | Render queue | Silent failures, duplicate costs, lost takes |
| 12. Review and reshoot | Selected takes, edited prompts, re-renders | Creator / editor | Teams accept first bad take and move on |
Stage 1–3: Intake, guidance, and brief lock
The first three stages exist to prevent the single most expensive mistake in AI production: writing and shooting before the show is actually defined.
A strong intake covers the vertical-short-drama dimensions: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, platform, and target audience. Episode length usually lands in the 1–2 minute range for vertical, with an outer ceiling around 3–4 minutes.
A useful gate here is completeness scoring. If the brief is too thin, the system should force a longer guidance path; if it is already detailed, it can move straight to a locked brief. This is the AI equivalent of a development meeting: no one calls "action" until the logline, character arc, conflict escalation, and cliffhanger rhythm are signed off.
Stage 4–5: Lineup and the continuity bible
Character consistency does not start at render time. It starts when the cast is defined.
The lineup should include every named character, with a stable identity, personality anchors, and a clear arc for leads. Visual fields must be complete before art begins.
The continuity bible is a structured story record that tracks character identities, stable traits, current mutable states (injuries, disguises, revealed identities), relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual notes.
The point is not to rely on model memory. The point is to give every future writing pass only the slice of archive it needs: current character states, unresolved threads, recent batch summary, and the current batch's main line.
Stage 6: Batch writing with hard rules
Scripts should be written in batches of several episodes, not one endless generation. Each batch follows the same internal order:
- Beat sheet for the batch, including episode-end cliffhangers
- Full dialogue draft
- Rule-based quality check
- Archive update after approval
Vertical short drama writing has rules that are non-negotiable enough to be encoded:
- Golden 3 seconds: the first scene must open with conflict or suspense, never exposition
- Episode shape: opening hook → escalation → closing cliffhanger
- Payoff density: at least one small beat per episode; a larger payoff every few episodes
- Dialogue: short lines, generally under ~20 characters in Chinese or the equivalent tight line in English; no lecture-like monologue
A rule checker should catch structural failures: missing character lines, too little dialogue, mismatched scene counts, placeholder text like "to be continued," or malformed episode titles. Failing drafts should be sent back with specific feedback rather than handed to the user half-finished.
From the second batch onward, the pipeline should lock intent before writing: where this batch goes, which characters carry it, which threads advance, and how each episode ends. Those locked intentions override stale archive entries when the two conflict.
Stage 7–8: Look development before shooting
One of the most visible AI failures is style drift: characters rendered in one visual dialect, backgrounds in another, props in a third.
A robust pipeline uses unified look manuals. A typical system includes around 17 style presets spanning 2D, 3D, and realistic directions—urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink style, xianxia, 3D donghua, stop-motion clay, cyber-Chinese fusion, and so on. Each manual covers character rendering rules, scene and prop guidance, and video style tags so that stills and moving clips share one visual language.
Asset generation is a two-step process:
- Text polishing: turn archive descriptions into render-ready prompts using the chosen style manual, with hard rules such as gender preserved
- Image generation: produce the final asset, with optional reference images, and store version history
The video side should only consume completed, readable assets. It should never try to shoot using half-finished character sheets.
Stage 9: Scene blocks, not one long video
Each episode should be cut into scene blocks targeting roughly 10 seconds each, with a soft text ceiling around 200 characters before further splitting. The split follows action beats, paragraph breaks, and sentence endings rather than arbitrary timestamps.
This is the editing-room view: instead of "Episode 3, one big file," the team sees Episode 3 as Scene 1, Scene 2, Scene 3, and so on. Each scene has its own prompt, its own reference set, its own render history, and its own takes.
This matters because reshoots in AI production are scene-level. If one line delivery is wrong, you do not regenerate the whole episode; you regenerate the block.
Default delivery is 9:16 vertical, not a horizontal master that gets cropped afterward. Other ratios can be supported, but vertical is the native output shape.
Stage 10: Shot prompts written like a shot list
The biggest prompt mistake is writing prose when the model needs a shot list.
A production-grade shot prompt contains eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color tone
- Camera movement
- Visual style
- Image quality anchors
- Constraint / anti-failure package
Simple scenes can be written as one block. Complex cinematic scenes use a three-part structure: overall setup, numbered shots, and a constraint package. One shot gets one camera move—no mixing push, pull, pan, and tilt inside a single shot. Number shots instead of stamping absolute timestamps like "0–3s."
The constraint package is not optional. It should include image quality, facial stability, no watermark or logo, and anti-twinning rules for multi-character scenes. Non-realistic styles need an explicit style anchor. Motion should favor slow, continuous actions over high-impact bursts that models often break.
There is one reference rule worth treating as law: if a character has an approved reference image, the prompt must not re-describe that character's clothing or appearance in text. Text only describes action, expression, and injury state. This single rule eliminates a huge share of costume swaps and face changes.
Reference mapping follows a fixed order per scene: scene plate → scene props → character sheets. Slots are only filled when an image exists; the system should never invent a binding to fill empty space. Creators can manually bind names (for example, a script that says "the officer" maps to the character "Li Qiang"), include or exclude props, swap in older approved asset versions, and route era-specific looks for time-travel or flashback scenes.
Before prompts are finalized, IP and specific title names should be stripped while keeping the technical and aesthetic direction, reducing downstream copyright blocking risk.
Stage 11–12: Shoot, review, reshoot
The shooting flow should be familiar to anyone who has worked on a set:
- Open the episode's video workspace
- Confirm reference assets are present, or deliberately choose text-only
- Generate shot prompts
- Read and edit the prompts
- Choose model tier, aspect ratio, resolution, and duration
- Submit to render
- Review historical takes and choose the best one
The controls map directly to real production roles:
| Control | Production equivalent |
|---|---|
| Edit shot prompt | Director revising shot notes |
| Swap character / scene / prop reference | Changing a look or location plate |
| Manual character binding | Fixing nickname or cameo references |
| Include / exclude props | Controlling visual focus |
| Switch style manual | Unifying art direction |
| Choose model tier | Balancing quality, speed, and cost |
| Review historical takes | Multi-take selection |
A production queue, not a toy script, should handle reliability: separate lanes for video, art, and writing tasks; no parallel submissions for the same in-progress scene; pre-deducted credits with automatic release on failure; timeout recovery for stuck jobs; and duration validation on output rather than trusting vendor metadata.
Where human judgment stays in the loop
It is important to be honest about what this kind of pipeline does not do.
- There is no automatic quality score that magically picks the best take. Final review still belongs to the creator.
- Reference images are not a hard gate. Teams can skip them and shoot text-only, but quality usually drops, so professional process should approve looks first.
- Reference tags in prompts depend on disciplined formatting; creators should still read prompts before rendering.
- The ~10-second scene block is an engineering heuristic, not a frame-accurate edit. Final trimming and assembly still happen in post.
- The product unit is one scene with multimodal references producing one clip. There is no mature cross-scene automatic video extension workflow.
- Character consistency depends on the asset chain, not a magic face-lock. Era routing is rule-based, and final look still depends on asset quality and prompt discipline.
In short, the pipeline does not remove filmmakers from the process. It removes the parts of the process that are repetitive, drift-prone, and easy to make routine, so that aesthetic decisions stay with the people who own the show.
How Maosika fits in
Maosika (猫斯卡) is built around this staged logic. It is an AI production operating system for vertical short dramas, not a single text-to-video button. The 18 digital expert roles—five core roles for creative and script work, plus 13 video-production roles covering prep, visual design, shooting, and post QC—mirror a real crew structure, and users can follow the streaming work logs the way a producer reads a daily production report.
The system does not claim to replace writers, directors, or editors. It enforces the sequence that professional short drama production already relies on: story and continuity first, then look development, then scene breakdown, then shot-level instructions, then multiple takes with deliberate reshoots.
If you are building your own AI short drama workflow, use this 12-stage order as your checklist. If you want the sequence already hardened into a product, you can explore Maosika at https://www.maosika.com.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com