How AI Vertical Short Dramas Actually Get Made: A 12-Stage Production Pipeline From Idea to Deliverable
Most AI short drama failures are not model failures — they are pipeline failures. The fix is not a bigger "generate" button, but a locked sequence of intermediate deliverables that a real crew would recognize.
The pipeline is the product, not the prompt
If you have tried making vertical short dramas with AI tools, you have probably seen the same failure pattern: the first clip looks promising, the third clip changes the lead's face, the fifth clip forgets a key prop, and by episode ten the plot contradicts itself. Teams often blame the video model. The model is rarely the root cause.
The root cause is that most tools treat a short drama as one long generation task. A real production does not work that way. A real production is a sequence of handoffs: idea becomes brief, brief becomes outline, outline becomes script, script becomes cast and look, look becomes shot list, shot list becomes filmed takes, takes become selected cuts. Each handoff produces something you can look at, reject, or lock before money is spent on the next stage.
An AI vertical short drama pipeline is the same idea, with digital crew roles enforcing the handoffs. The claim is not "AI replaces the crew." The claim is that the repeatable, drift-prone, failure-prone parts can be turned into a constrained assembly line, while taste and final judgment stay with the human producer.
Stage 1 — Creative intake and completeness scoring
Production starts before writing. The raw idea is scored for completeness across the dimensions that actually matter for vertical drama: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, target platform, and audience.
If the idea is already well-formed, the system can route straight to a locked brief. If key dimensions are missing, it asks only for what is missing. If the idea is too thin, it runs a fuller guided development flow. This is not a chatbot being polite. It is a forced gate: an under-specified brief is the single largest source of mid-production rewrites.
Stage 2 — Guided creative development
For ideas that do not yet pass intake, the development stage fills the gaps. The goal is not to write the script here. The goal is to make the inevitable decisions early: who the lead is, what they want, what is blocking them, what the audience is supposed to feel every few episodes, and how the ending is shaped.
A useful rule from vertical drama craft: the first episode does not have time to set up the world slowly. The first scene needs conflict or a strong悬念 — a question the viewer must see answered. If the brief cannot support that, the brief is not ready.
Stage 3 — Lock the creative brief
The creative brief is the first hard lock. Nothing downstream moves until it is confirmed. A production-ready brief includes:
- Logline
- Core conflict
- Story direction
- Ending direction
- Payoff and hook rhythm
- Target platform and audience
- Episode-by-episode synopsis
- Notes for the writer
- Lead character fields including character arc
Think of this as the development meeting ending with a signed-off brief. You can change it later, but not by accident.
Stage 4 — Build the story archive
The story archive is a structured continuity record that follows the whole drama. It tracks character identities, stable traits, current state — injuries, revealed identities, changed relationships — relationship maps, open and resolved plot threads, episode appearance tables, batch-level plot notes, and prop visual descriptions.
The story archive is the answer to "how does episode 47 remember what happened in episode 3?" It does not rely on model memory. It relies on a structured file. Before each batch of episodes, the writer is fed only the relevant slice: current character state, unresolved threads, recent batch summary, and the current batch's main line. After each batch, the archive is updated.
Stage 5 — Confirm the cast and visual fields
Before any image is generated, the cast must be complete: names are valid, visual fields are filled, and leads match the brief. This is another hard gate. A missing or vague character description here turns into face drift later, and fixing it after shooting is far more expensive than fixing it now.
This is also where you decide which named roles are main cast that need look development, versus extras or generic characters that should not consume character reference slots.
Stage 6 — Write scripts in batches, with a beat sheet first
Script writing happens in batches of several episodes, not one giant 80-episode generation. For each batch after the first, the system first confirms four things: the batch's plot direction, key characters, threads and conflicts, and the episode-end hooks. Those become hard constraints.
Within each episode, the order matters:
- Beat sheet for the episode
- Episode-end cliffhanger defined
- Full scene-by-scene script
- Rule-based quality check
- Archive update
Vertical drama writing rules that should be enforced, not suggested:
- Golden 3 seconds: the first scene must open with conflict or suspense, no slow exposition.
- Episode structure: opening hook, escalating conflict, final-scene cliffhanger.
- Payoff density: at least one small payoff per episode; a larger payoff every few episodes.
- Dialogue: short lines, generally under twenty characters in Chinese source pacing; no lecture-like monologues.
The quality check catches concrete failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, or placeholder text like "to be continued." Failing drafts are sent back for rewrite with the error feedback attached. Half-finished scripts should not reach the director or the renderer.
Stage 7 — Select the visual style
Look development is not a single image filter. A usable style pack needs to cover character sheets, scene sheets, prop sheets, and video style tags under one consistent visual path. Otherwise you get a common AI failure: characters drawn in one style, video rendered in another.
A production system should offer multiple style manuals covering 2D, 3D, and live-action-realistic directions — for example urban realism, period realism, mature urban romance animation, 1990s anime, Chinese ink-fantasy, xianxia, 3D donghua, clay stop-motion, and cyber-Chinese styles. Once chosen, the same style path feeds character art, scene art, prop art, and video prompts.
Stage 8 — Generate locked character, scene, and prop art
Art generation is a two-step process. First, the archive description is polished into an image-generation prompt using the chosen style manual and hard gender rules; this step can be manually overridden. Second, the final prompt generates the image, with optional reference input, and the result is saved as a versioned asset.
Two disciplines prevent later pain:
- Scene art must be empty plates — no people in scene references.
- Video production only consumes completed, readable art assets, never half-finished ones.
Props are extracted episode by episode from the script using the script's own wording, then merged into a show-wide catalog. This prevents prop amnesia: the jade pendant in episode 2 should still look like the same jade pendant in episode 22.
Stage 9 — Cut each episode into shot blocks
An episode is not rendered as one long clip. It is cut into shot blocks, each targeting roughly ten seconds of screen time, with a soft cap on script length per block. Longer passages are split further by action beats, paragraph breaks, or sentence boundaries.
This is the equivalent of a clip list on an editing timeline. Each block gets its own prompt, its own reference set, its own rendered output, and its own history of takes. You are not looking at "episode 3, one big video." You are looking at "episode 3, scene 1, scene 2, scene 3…" each independently reviewable and reshootable.
Default delivery is vertical 9:16, not a landscape video cropped after the fact. Other aspect ratios can be supported, but vertical is the native output shape for short drama platforms.
Stage 10 — Bind references and build shot prompts
This stage is where most consistency problems are either solved or created. For each shot block, the system builds a reference table in a fixed order: scene reference first, then prop references for that shot, then character references. Slots are only filled if an asset exists; the system does not invent bindings to fill empty slots.
The highest-priority consistency rule can be stated plainly:
For any character with a reference image, the prompt must not re-describe clothing or appearance in text. The reference image is authoritative. Text describes only action, expression, and injury state.
This rule directly attacks the classic AI failure where text says "black suit" but the reference shows a white dress, and the model compromises by generating neither correctly.
Manual controls belong here too: bind a script name like "the officer" to the correct cast card, include or exclude props to control visual focus, swap in an older version of an asset, or route period-versus-modern looks for time-travel or flashback stories based on scene keywords.
Before prompts are finalized, there is a readiness check: missing scene, character, or prop references are surfaced as a list. You can skip them and shoot with text only, but you are explicitly choosing a lower-consistency path. Professional workflow locks art first, then shoots.
The shot prompt itself should follow cinematic structure, not novelistic prose. A strong prompt covers eight elements: precise subject, action detail, scene environment, light and color, camera movement, visual style, image quality, and constraints. Complex shots use a three-part structure: overall setup, shot-by-shot instructions, and a constraint pack. One shot gets one camera move — no push-pan-zoom stacks. Camera instructions use shot numbers, not absolute timestamps. A mandatory fallback pack covers image quality, face stability, no watermark, twin/duplicate prevention in multi-character shots, and style anchoring for non-realistic looks.
Before delivery, prompts are cleaned of specific copyrighted work or IP names, keeping the technique and aesthetic descriptors. This reduces downstream copyright blocking risk.
Stage 11 — Render through a production queue
When a shot is submitted, the renderer uses exactly the prompt you confirmed and the reference map that was locked at submission time. Editing the prompt and resubmitting creates a new take, mirroring real-set practice: change the shot note, then shoot again.
Production-grade queue behavior matters more than it sounds:
- Video tasks run in a separate queue lane from writing and art tasks.
- The same shot block cannot submit parallel jobs while one is running, preventing duplicate charges and state confusion.
- Credits are pre-deducted and released on failure.
- Stuck jobs time out and can be retried.
- Existing vendor task IDs are polled rather than creating duplicate jobs.
- Final videos get faststart processing and are stored with measured runtime, not only the vendor's reported duration.
This is what makes the system an operating system rather than a demo script.
Stage 12 — Review, retake, and hand off
After rendering, you review takes per shot block. You can keep one, reshoot with edited prompts, swap a reference image and reshoot, change the model tier for speed-versus-quality tradeoffs, or resplit a block that came out too long. The unit of retake is the shot, not the whole episode.
There is no automatic "choose the best take" engine. Final quality judgment sits with the creator or supervising producer. The system provides multiple takes and the tools to reshoot with controlled changes; it does not pretend to replace the director's eye.
Final assembly across shots, transitions, sound balancing, and episode-level polish still belongs in a finishing stage. The pipeline delivers reliable shot-level units; a skilled editor still assembles them into the release cut.
What this pipeline does not claim
Honest boundaries matter, especially in a category full of overpromising:
- There is no automatic scoring or auto-retake engine that guarantees a good shot. Quality review is human.
- Reference images are not a hard block. You can shoot without them; results are usually worse, which is why professional workflow locks art first.
- Reference tokens in prompts rely on disciplined structure, not invisible magic. The creator should still check that references are present when reviewing prompts.
- The roughly ten-second shot block is an engineering heuristic, not frame-accurate timecode editing. Longer blocks can still occur.
- There is currently no cross-shot automatic continuation or video extension workflow. The product unit is "single shot with multimodal references in, single clip out."
- Character consistency depends on the art asset chain, not a face-embedding verification lock. Final look depends on asset quality and on respecting the rule that referenced characters are not re-described in text.
Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not promise one-click hits or zero-failure output. It encodes a professional sequence — story and continuity first, then locked look, then shot planning, then shot-level prompts, then multiple takes — so that production can scale without the usual AI drift.
If you are planning a vertical short drama slate, the question to ask any tool is not "how good is your first sample clip?" It is "what locked intermediate deliverables do I get to review before the next stage spends money?" That is the difference between a demo and a production pipeline.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com