How AI Vertical Short Dramas Actually Get Made: A 12-Stage Production Pipeline From Idea to Deliverable
Most AI short drama failures are not model failures; they are pipeline failures. A usable episode is produced by locking decisions in order—idea, brief, bible, cast, look, scene, shot, take—not by pressing one big generate button.
The core insight: production order is the product
If you treat AI short drama as a single prompt-to-video problem, you get the familiar failure modes: the lead changes face between scenes, props vanish, episode 8 forgets a secret set up in episode 2, and the opening wastes ten seconds on setup instead of a hook. These are not bugs in one model. They are what happens when there is no enforced production order.
A professional AI vertical short drama pipeline is a 12-stage chain with checkpoints. Each stage produces a reviewable artifact. You do not move forward until the artifact is locked. The human keeps the审美 calls; the system enforces that those calls are carried into every episode and every scene.
AI vertical short drama production is the discipline of turning one idea into a sequence of lockable artifacts—brief, bible, script batch, look, cast art, scene block, shot prompt, take—so that drift is caught at the stage where it is cheapest to fix.
This guide walks through that chain in the order it actually runs, written for producers, writers, and directors who need to ship vertical episodes, not admire demo reels.
---
Stage 1: Idea intake and completeness scoring
Everything starts with a one-line idea, but a one-line idea is not a brief. The first job of the pipeline is to measure how much of the brief is actually present.
A complete intake for a vertical short drama covers ten dimensions:
| Dimension | What it must answer |
|---|---|
| Genre / trope | Is this revenge, hidden identity, contract marriage, xianxia, rebirth? |
| Protagonist | Who is the lead, and what is their arc? |
| Core conflict | What is the engine that runs episode after episode? |
| Story direction | Rags to power? Slow reveal? Dual identity? |
| Episode count | How many episodes total? |
| Episode length | Typically 1–2 minutes per vertical episode, up to 3–4 at most |
| Tone | Gritty? Glossy? Melodramatic? Deadpan? |
| Hook and payoff rhythm | How often does a beat land? |
| Ending direction | Open, closed, twist, season-two hook? |
| Platform and audience | ReelShort-style? Domestic platforms? Age and gender skew? |
The intake is scored 0–100. Above 80, you can go almost straight to brief lock. In the 40–79 range, the system only asks for the missing dimensions. Below 40, you go through a full guided development path. The point is not gatekeeping; the point is that a brief that has not answered these questions will be answered randomly by the model later, and you will pay for it in reshoots.
---
Stage 2: Guided creative development
When the intake is thin, the pipeline runs a structured development conversation rather than asking the model to "be creative." This is the equivalent of a development room asking the hard questions before the writer goes to draft.
The output of this stage is not a script. It is a set of decisions:
- What the protagonist wants, what stands in their way, and what they are willing to do
- Where the story starts in media res (vertical episodes do not have time for slow world-building)
- What the first hook is, before the first line of dialogue is written
- What the audience is supposed to feel at the end of episode 1, episode 3, episode 10
This stage exists because vertical short drama writing has hard rules that do not come naturally to language models trained on long-form prose:
- The golden 3 seconds: scene 1 must open on conflict or suspense, never on a slow pan or exposition.
- Per-episode structure: opening hook → escalation → end-of-episode cliffhanger.
- Payoff density: at least one small payoff per episode (a reveal, a face-slap, a piece of evidence landing, an identity hint); a larger payoff every few episodes.
- Dialogue discipline: short lines, generally under twenty words; no essay-style monologues; no narrator explaining what the scene should show.
These are not stylistic preferences. They are the format constraints of vertical drama, where the swipe is one thumb-move away.
---
Stage 3: Creative brief lock
The brief is the first hard gate. It does not move to script until it is locked.
A locked brief contains:
- Logline
- Core conflict
- Story direction
- Ending direction
- Payoff and hook rhythm
- Platform and audience
- Episode-by-episode synopsis
- Notes to the writer
- Lead character fields, including character arc
The brief matters because everything downstream—script, bible, art, shot prompts—treats it as a hard constraint. If you change the brief after art is locked, you are not "tweaking"; you are restarting a production. In live action, you would not change the logline after sets are built. The same discipline applies here.
---
Stage 4: Cast lineup and visual confirmation
Before a single scene is written in final form, the cast is lined up and the visual fields are confirmed. The gate checks four things:
- The lineup is complete (no unnamed "friend" or "boss" carrying major scenes)
- Names are valid and consistent
- Visual description fields are filled
- The leads match the brief
This is the stage where consistency starts. If the cast is vague here, every later stage will invent details on the fly, and that is how you get three different faces for the same character across twelve episodes.
---
Stage 5: Story bible (continuity bible) creation
The story bible is a structured record that travels with the production from first episode to last. It is not a prose document; it is a structured continuity file.
The story bible is the production's single source of truth for who characters are, what they look like, what they know, what they have done, which plot threads are open, which are resolved, and which items and locations carry across episodes.
A working bible tracks:
- Character identities and stable traits
- Mutable current state (injuries, revealed identities, changed allegiances)
- Relationship maps
- Plot threads, marked open or resolved
- Per-episode appearance tables
- Batch-by-batch plot summaries
- Visual descriptions of recurring props
The key mechanism: when writing, the system does not rely on the model's "memory" of earlier episodes. It pulls a bible slice for the current batch—current character states, unresolved threads, recent batch summaries, and the current batch's main line. After each batch is written, the bible is updated before the next batch begins.
This is how a 100-episode drama still remembers what was set up in episode 3.
---
Stage 6: Batch script writing with beat sheets first
Scripts are written in batches of several episodes, not one giant run. Each batch follows a strict internal order:
- Beat sheet first: for each episode, the scene-by-scene beats and the end-of-episode cliffhanger are planned before any dialogue is written.
- Then the draft: dialogue and action are written on top of the locked beats.
- Then rules-based quality check: an automated checker catches episode title errors, wrong scene counts, missing character lines, too little dialogue, and placeholder text like "to be continued."
- Then bible update: the new batch is written back into the continuity record.
If the quality check fails, the batch is rewritten with the specific error fed back, up to a fixed cap. Half-finished work that does not pass does not get delivered.
From the second batch onward, the pipeline locks four things before writing: where this batch goes in the plot, which characters carry it, which threads and conflicts are active, and what the end-of-batch hook is. These become hard constraints. If they conflict with the older bible, the newly confirmed creative intent wins—because the writer or showrunner just decided it.
Past the midpoint of the total run, the system also injects ending constraints, so the final episodes do not drift into a new plot instead of paying off the one that was set up.
---
Stage 7: Art style selection and look development
Vertical AI drama fails hardest when the art direction drifts between assets: the character is 2D anime, the background is photoreal, the video looks like oil paint. The solution is to lock one style path and use it everywhere.
A production-ready style library covers multiple directions—2D, 3D, live-action realism, period, xianxia, cyberpunk, and others—with each style providing:
- Character art rules (face anchors, material, tone, view consistency)
- Background and prop rules
- Video style tags
Once a style is chosen, character art, scene art, prop art, and video prompts all share the same style path. The look is decided once, then enforced downstream.
---
Stage 8: Character, scene, and prop art production
Art is produced in a two-step process:
- Text polish: the bible's character descriptions are rewritten into art-ready prompts using the chosen style's character rules and hard gender constraints. This step is human-overridable.
- Image generation: the final prompt is rendered, with optional reference images, and saved as a versioned asset.
Three disciplines prevent the classic AI art failures:
- Scene art discipline: scene backgrounds must not contain characters. Empty frames stay empty.
- Prop catalog discipline: props are extracted episode by episode from the script using the original names the script uses, then merged into a show-level catalog so nothing is lost between episodes.
- Content authority rule: the script and the creator's descriptions override the style manual's subject-matter restrictions. The manual controls how things are drawn, not what is allowed to exist.
Before shooting begins, a readiness check flags any missing scene, character, or prop references and lists them. You can choose to proceed with text-only prompts, but the system tells you clearly that this usually produces weaker results; the professional path is to finish look-dev before rolling camera.
---
Stage 9: Reference map assembly (the consistency engine)
This is the single most important mechanism for character consistency, and it is worth understanding in detail.
For each scene, the pipeline builds a reference table in a fixed order:
- Scene art
- Props appearing in this scene
- Character look-dev art for characters in this scene
Slots are only filled if an image actually exists. Nothing is invented to fill a gap.
Then a hard rule fires:
For any character that has a reference image, the prompt must not describe clothing or appearance in text. The reference image is authoritative. Text only describes action, expression, and visible injury.
This rule directly attacks the most common AI video failure: text describing one outfit while the reference image shows another, which causes the model to re-invent the character mid-shot.
Additional mechanisms support this:
- Manual character binding: if the script calls someone "Officer," you bind that reference to the named character's look-dev art.
- Manual prop include / exclude: you control which props actually occupy reference slots in a given scene.
- Version swapping: you can swap in a different historical version of the same named asset.
- Era routing: for time-travel or flashback stories, the system uses scene keywords to pick modern or period look-dev art for the same character.
Consistency is not achieved by asking the model nicely. It is achieved by controlling which assets are handed to the model, and by forbidding text that fights those assets.
---
Stage 10: Scene blocking into ~10-second chunks
A finished episode script is not sent to video as one long piece. It is cut into scene blocks targeting roughly ten seconds each.
The rules:
- Target length is about ten seconds per block.
- Body text has a soft cap around two hundred characters; longer scenes are split on action beats, paragraph breaks, or sentence endings.
- Extras and generic unnamed roles are separated from the drawable main cast so they do not consume reference slots.
- You can re-block an entire episode.
What you see in production is not "episode 3 as one video." You see episode 3 broken into scene 1, scene 2, scene 3—each with its own prompt, its own reference set, its own output, and its own history of takes. This is exactly how a clip list looks on an editing table.
Default delivery is 9:16 vertical, not a landscape video cropped after the fact. Other ratios are supported, but vertical is native.
---
Stage 11: Engineered shot prompts
At this stage the pipeline is no longer writing prose. It is writing shot instructions.
A shot prompt is built from eight elements:
| Element | Function |
|---|---|
| Precise subject | Who or what is in frame |
| Action detail | What they are doing, with quantified motion |
| Scene environment | Where they are |
| Light and color | Time of day, palette, mood |
| Camera movement | One move per shot, never stacked |
| Visual style | Injected from the chosen style path |
| Image quality | Baseline quality and stability flags |
| Constraints | Negative constraints and stability safeguards |
Simple scenes use a single-block prompt. Complex cinematic scenes use a three-part structure: overall setup → shot 1 / shot 2 / shot 3 → shared constraint pack.
Hard prompt-writing rules:
- One camera move per shot. No "push in while panning while tilting."
- Shots are numbered. No absolute timecodes like "0–3s."
- A mandatory fallback pack is added for quality, face stability, and no watermark; multi-person scenes add safeguards against twin/duplicate artifacts; non-realistic styles get explicit style anchors.
- Motion favors slow, continuous action over high-energy bursts, which are far more likely to break.
- Dialogue is wrapped in
{}, sound effects in<>, BGM in(). - Only assets for the current scene are fed in; nothing from another scene leaks across.
- The current scene's script body is the highest-priority source.
- Generation settings are conservative, biased toward stability over wildness.
Before delivery, prompts are cleaned of specific copyrighted IP or title references, keeping the technical and aesthetic language but reducing downstream blocking risk.
In one sentence: the system teaches the model to speak the language of a shot list, not the language of a novel.
---
Stage 12: Shoot, review, retake
The actual shooting path is deliberately close to a real set:
- Open the episode's video workspace.
- Confirm references are present (or consciously skip them).
- Generate the shot prompts.
- Read and edit the prompts directly.
- Choose model tier, ratio, resolution, and duration.
- Submit to render.
- Review historical takes and pick the best one.
The professional control points map cleanly onto real production roles:
| Control point | Live-action equivalent |
|---|---|
| Editing the shot prompt | Director revising shot notes |
| Swapping character / scene / prop reference | Changing a look or location board |
| Manual character binding | Fixing a name mismatch or offscreen reference |
| Including / excluding a prop | Controlling visual focus in the frame |
| Switching style | Unifying the visual language |
| Choosing model tier | Trading quality, cost, and speed |
| Reviewing historical takes per scene | Multi-take selection |
Crucially, hitting generate does not invent a new prompt. The renderer uses the prompt you confirmed or edited, with the reference map locked at that moment. If you change the prompt and generate again, that is a new take—exactly like revising shot notes and rolling camera again.
Production-grade reliability lives here too: videos are faststart-processed and stored with actual measured duration; video tasks run in a separate queue from writing and art; the same scene block cannot be double-submitted mid-run; credits are pre-deducted and released on failure; stalled tasks time out and can be retried; and existing vendor task IDs are reused rather than double-created.
---
What this pipeline does not do
It is important to be explicit about the boundaries, because honest limits are what make a production tool trustworthy.
- There is no automatic scoring or auto-pick of the best take. Final quality judgment stays with the creator and producer; the system gives you multiple takes and the tools to reshoot.
- Reference images are not a hard gate. You can skip them and shoot text-only, but quality is usually worse; the professional workflow locks art first.
- Reference tags in prompts rely on the prompt spec, not hidden hard-patching. You should still read the prompt and confirm references are present before shooting.
- Shot grammar follows the built-in lens spec. The full style manual drives art; video gets style tags, not the entire art book.
- The ~10-second block is an engineering heuristic, not timecode-accurate editing. Long scenes are re-split, but blocks can still run long; final cutting and stitching are a later stage.
- There is no cross-scene automatic video continuation as a product flow. The unit of production is one scene block with its references, producing one clip.
- Character consistency depends on the look-dev asset chain, not face-embedding verification. Era routing is rule-based; final look still depends on asset quality and on respecting the "do not describe clothed appearance when a reference exists" rule.
---
Where Maosika fits
Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not promise one-click hits or fully autonomous crews. It encodes the twelve-stage pipeline above—intake scoring, brief lock, story bible, batch writing with beat sheets, look-dev, reference mapping, scene blocking, engineered shot prompts, queued rendering, and multi-take selection—so that the order professional productions already follow is enforced by the product, rather than left to willpower.
The philosophy is simple: the parts of production that are repetitive, drift-prone, and easy to get wrong become a constrained pipeline; the审美 judgment stays with the human.
If you are planning an AI vertical short drama slate, the question is not which model looks most impressive in a demo. It is whether your pipeline locks decisions in the right order, produces reviewable artifacts at each stage, and lets you retake a bad scene without restarting the whole episode.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com