The AI Vertical Short Drama Pipeline, Step by Step: Why Order Matters More Than Any Single Prompt
Most AI short drama failures are not model failures—they are order failures. Teams skip the brief, skip the look-dev, and ask the video model to invent story, faces, and props in one shot; the fix is a staged pipeline with verifiable handoffs between stages.
The core claim: pipeline order is the product
If you ask a video model to "make a 60-episode vertical short drama about a betrayed heiress," you will get one pretty clip and 59 episodes of strangers. The reason is not that the model is bad; it is that you asked it to do six different jobs at once—story structure, continuity, casting, art direction, shot design, and performance—without giving it stable assets to anchor any of them.
A production-ready AI short drama pipeline is defined by one rule: each stage produces a lockable artifact before the next stage starts. You do not film a script that has not been signed off. You do not bind reference images for a character whose look has not been approved. You do not write shot prompts for a scene whose props have not been cataloged.
This is not a theoretical ideal. It is the same order a live-action set runs—greenlight, writers' room, casting, look-dev, shot list, shoot, dailies—rebuilt around what generative models are actually good and bad at.
The eight stages, in order
Below is the industrial order, mapped to what you should be able to point at and say "this is done" at each gate.
| Stage | Lockable artifact | What goes wrong if you skip it |
|---|---|---|
| 1. Idea intake & scoring | A completeness score on the core creative dimensions | Vague briefs get "filled in" by the model differently every episode |
| 2. Guided development & locked brief | A signed-off creative brief (logline, conflict, arc, hook rhythm, platform, audience) | The story drifts episode to episode with no single source of truth |
| 3. Cast & visual confirmation | Approved character lineup with visual fields complete and aligned to the brief | Characters rename themselves, swap genders, or change personalities mid-series |
| 4. Continuity bible (story archive) | A structured record of identities, relationships, open/resolved plot threads, per-episode appearance, prop descriptions | The show "forgets" its own setup by episode 20; buried threads never resurface |
| 5. Batch scriptwriting with rule checks | Beat sheets per episode, cliffhangers, dialogue passes, rule-based QA rewrites | Weak cold opens, flat cliffhangers, placeholder lines like "to be continued" leaking through |
| 6. Look selection & asset creation | Style pack chosen; character / scene / prop art generated and versioned | Characters rendered in one style, scenes in another; props reinvented per shot |
| 7. Scene blocking & shot prompts | ~10-second scene blocks, each with bound references and engineered shot prompts | One giant 2-minute prompt per episode; faces and outfits change between cuts |
| 8. Shoot, review, reshoot | Per-scene takes, editable prompts, reference swaps, best-take selection | You accept first-pass output and cannot iterate without regenerating everything |
The order is not cosmetic. Each row in the table feeds the next; removing a stage forces a later stage to invent information that should have been inherited.
Stage 1–2: Intake, scoring, and the locked brief
A creative brief for a vertical short drama is the document everyone—writers, the digital crew, the model, the editor—agrees to before a single frame exists. It contains at minimum: the logline, the central conflict, the story arc direction, the ending direction, the beat of payoffs and hooks, the target platform and audience, episode-by-episode outlines, and notes to the writer. The protagonist entry must include their character arc, not just their name and job.
The intake step scores how complete the incoming idea is across the vertical-format dimensions: genre, protagonist, central conflict, arc direction, episode count, episode length, tone, payoff/hook rhythm, ending direction, platform, and audience. Ideas that score high can move straight to a brief; ideas in the middle get only the missing dimensions filled in; ideas that score low go through a full guided development pass. The point is that the routing is enforced by the system, not left to a model to say "sounds good."
Think of it as a greenlight meeting. You would not let a show start shooting because the writer "sort of knows" who the lead is. The same discipline applies here.
Stage 3–4: Cast lock and the continuity bible
Continuity bible is a structured, series-spanning record: character identities, stable traits, variable current state (injuries, revealed identities, shifting allegiances), relationships, open and resolved plot threads, per-episode appearance tables, batch-by-batch plot summaries, and visual descriptions of recurring props. In a traditional writers' room this is the document the script coordinator yells at you for violating; in an AI pipeline it is the only thing preventing the model from improvising backstory on episode 47.
The cast confirmation step is a hard gate: lineup complete, names valid, visual fields filled, leads aligned to the brief. Until that passes, nothing downstream runs. This is deliberately rigid because the single most common failure in AI dramas is not bad writing—it is a character who is a 28-year-old doctor in episode 1 and a 45-year-old detective by episode 6, and nobody noticed until edit.
When writing, the system should consume only a slice of the bible relevant to the current batch: current character states, unresolved threads, recent batch summaries, this batch's main line. That prevents both "context amnesia" (the model forgetting earlier events) and "context noise" (the model getting distracted by irrelevant details from 30 episodes ago). After each batch, the bible is updated; before the next batch, it is reloaded. Continuity comes from structure, not from hoping the model remembers.
Stage 5: Batch scriptwriting with vertical-format rules
Vertical short drama scripts live or die by a small set of hard rules that are cheap to enforce and expensive to fix in post:
- The golden 3 seconds. The first scene must open on conflict or suspense—no slow pan, no exposition, no backstory monologue.
- Single-episode structure. Hook in scene 1, escalating conflict through the middle, cliffhanger in the final scene.
- Payoff density. At least one small payoff per episode (a reveal, a comeback, a piece of evidence landing); a larger payoff every few episodes.
- Dialogue discipline. Short lines, generally under twenty characters in the original-language phrasing; no essay-style dialogue, no narrator dumping exposition.
The writing process itself is staged: beat sheet and episode-end cliffhanger first, then the actual pages. A rule-based QA pass catches structural failures—wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, placeholder text like "to be continued"—and triggers an automatic rewrite with the specific error fed back. Half-finished drafts that fail QA do not get handed to the user.
From the second batch onward, four things are locked before writing starts: this batch's plot direction, the characters in focus, the threads and conflicts in play, and the episode-end hook. Those become hard constraints. If they conflict with something in the bible, the freshly confirmed creative intent wins—this is how showrunners course-correct mid-series without the model reverting to an older plan.
Stage 6: Style, art, and the look-dev chain
A style pack is a unified visual language that covers character sheets (facial anchors, materials, mood, view consistency), scene and prop sheets, and video style tags. When a style is selected, character art, scene art, prop art, and video prompts all draw from the same style path. The alternative—characters rendered in 90s anime, scenes in live-action realism, props in 3D—is the visual schizophrenia viewers can smell within ten seconds.
Character art is produced in two steps: a text pass that turns the bible's character description into an art-ready prompt using the style pack's character rules (with gender as a hard constraint, not a suggestion), then an image generation pass that supports reference images and stores every version. Only finished, readable art is consumed downstream; half-rendered previews never get bound as references.
Scene and prop handling follows set discipline:
- Scene headers are parsed from the script into interior/exterior, location, and day/night—no manual form-filling.
- Scene art must be empty of people; reference plates with characters baked in cause contamination.
- Props are extracted episode by episode from the script's own wording, then merged into a series-wide catalog so nothing is lost between batches.
- When user or script descriptions conflict with a style pack's world rules, the script wins—the style pack governs how things are drawn, not what is allowed to exist.
Before any shot prompt is generated, a readiness check flags missing scene, character, or prop references and surfaces a list. You can skip and go text-only, but you should know you are skipping—silent fallback is how you get three episodes of a character wearing different suits because the model improvised.
Stage 7: Scene blocks and engineered shot prompts
A scene block is the unit of production. Episodes are cut into roughly 10-second chunks, with a soft cap on body length and re-splitting rules for long passages. Each block is independent: its own prompt, its own bound references, its own output, its own history of takes. You do not see "episode 3, one big video"; you see "episode 3 → scene 1, scene 2, scene 3…" This is the editing-room view, and it is non-negotiable for iteration.
Default delivery is 9:16 vertical, not a landscape master cropped in post. Length, model tier, resolution, and audio are selectable per block.
Shot prompts follow an engineered format, not freeform prose:
- Eight core elements: precise subject, action detail, scene environment, lighting and color, camera movement, visual style, image quality, constraint pack.
- Complexity routing: simple scenes in one block; cinematic scenes in three parts (overall setup → shot-by-shot → constraint pack).
- One camera move per shot. No push-pan-zoom stacks inside a single shot.
- Shot numbers, not timestamps. "Shot 1," "Shot 2," never "0–3s."
- Mandatory fallback pack: quality, facial stability, no watermark or logo; multi-person scenes get twin/duplicate fallbacks; non-realistic styles get an explicit style anchor.
- Action principle: fine-grained limb movement, quantified intensity, slow continuous motion preferred over high-energy bursts that break.
- Symbol conventions: dialogue in curly braces, sound effects in angle brackets, BGM in parentheses.
Crucially, prompts are built only from this scene's materials—no cross-scene bleeding, no pulling characters from the next episode. The system is teaching the model to read a shot list, not to write a novel.
Before delivery, prompts are cleaned of specific copyrighted IP names while retaining technique and aesthetic, reducing downstream content blocks. If a reference image is flagged as a real-person photo, the pipeline returns a clear instruction to use platform-generated art instead of binding a real face.
Stage 8: Shoot, review, reshoot—the takes that actually matter
The shoot path is intentionally boring, because boring is what makes a production tool usable: open the episode's video workspace, confirm references (or deliberately skip), generate shot prompts, read and edit them, pick model tier / ratio / resolution / length, submit, wait in the queue, review, pick the best take.
The professional controls map directly to on-set equivalents:
| Control in the tool | Equivalent on a real set |
|---|---|
| Editing the shot prompt | Director revising the shot list |
| Swapping character / scene / prop references | Changing a look or a location plate |
| Manual character binding | Fixing name mismatches and off-screen references |
| Including / excluding props | Controlling what is in frame |
| Switching style packs | Unifying the visual language |
| Choosing model tier | Trading quality, cost, and speed |
| Browsing historical takes per scene | Dailies and pickups |
When you press generate, the worker uses the prompt you confirmed and the references that were locked at that moment. Changing the prompt and regenerating creates a new take, exactly like revising a shot list and calling "action" again. The system does not secretly rewrite your prompt between takes.
On the operational side this means: faststart-processed files with actual measured duration, a dedicated video queue isolated from script and art tasks, no parallel submissions on the same scene block to avoid double-charging, pre-authored credits that release on failure, hung-task timeouts, and resumable polling when a vendor job ID already exists. It is a production queue, not a toy script.
The reference-image rule that fixes most consistency problems
If you internalize one rule from this guide, make it this one:
When a character has a bound reference image, the prompt must not re-describe their clothing or appearance in text. The reference is authoritative; text describes only action, expression, and injury state.
This single rule eliminates the most common AI video failure: text and reference fighting each other, producing a mid-shot wardrobe change or a face swap between cuts. The reference-mapping order is also fixed: scene plate first, then this scene's props, then character looks. Slots are only filled when an image exists; empty slots are marked "text only" rather than silently invented. Manual binding lets you map a script reference like "the officer" to the correct character in the bible, and era-based routing picks the right look for time-travel or flashback scenes.
Where the human still sits (and must sit)
It is worth being explicit about what this pipeline does not do, because the honesty is the point:
- There is no automatic quality-scoring engine that picks the "best" take for you. Final judgment stays with the creator and the producer; the system gives you multiple takes and the tools to reshoot.
- Reference images are not a hard gate. You can skip them and go text-only—quality will usually be worse, which is why the professional flow locks looks first.
- Reference tags inside prompts rely on the format being respected; creators should still glance at the prompt to confirm tags are present.
- The ~10-second scene block is an engineering heuristic, not a timecode-precise cut; long scenes get re-split but some blocks run over, and final assembly still belongs in an edit.
- There is no cross-scene automatic video continuation as a product workflow today. The unit of work is "single scene with references → single clip."
- Character consistency depends on the look-dev chain, not on facial-embedding verification; era routing is rule-based, and final look still depends on the quality of the art and on prompts respecting the "do not re-describe appearances" rule.
The goal of an AI vertical short drama pipeline is not to remove the crew. It is to take the parts that are repetitive, drift-prone, and easy to get wrong, and turn them into an enforceable assembly line—while keeping every aesthetic decision in human hands.
Maosika (猫斯卡) is an AI production operating system for vertical short dramas that encodes this staged order: from idea scoring and locked brief, through a structured continuity bible and look-dev chain, to ~10-second scene blocks with engineered shot prompts and multi-take review. It does not promise one-click hits; it promises that when you say "that's the look, that's the story, that's the shot," the next episode inherits exactly those decisions instead of reinventing them.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com