AI Vertical Short Drama Production, Step by Step: The 12-Stage Pipeline From Idea to Deliverable Cut
AI short dramas fail not because the model is weak, but because teams skip the order of operations. The reliable path is a staged pipeline with sign-off gates between each step—from locked brief to batched scripts to approved look dev to field-blocked shots to multi-take delivery.
Why most AI short drama runs collapse in the middle
A typical AI micro-drama project starts with enthusiasm and ends in chaos: the lead's face changes in episode 4, the costume shifts between shots, a prop from episode 2 vanishes, episode 8 forgets the cliffhanger from episode 7, and the team burns credits regenerating the same scenes over and over.
This is not a model-quality problem. It is a production-management problem. Traditional crews solve it with call sheets, continuity bibles, storyboards, and dailies. AI-native crews need the same discipline, translated into a pipeline where every stage produces a reviewable artifact before the next stage begins.
A production pipeline for AI vertical short drama is a staged sequence of work where each stage has a defined input, a reviewable intermediate deliverable, and a go/no-go gate before anything moves downstream. The goal is not to remove human judgment; it is to make sure human judgment is applied at the right checkpoints, not wasted re-prompting to fix drift.
The 12-stage pipeline, in order
The stages below are written the way a real production office would run them. Each stage names what you should have in hand before you move on.
| # | Stage | Reviewable artifact | Gate before next stage |
|---|---|---|---|
| 1 | Idea intake & completeness scoring | Scored intake across premise, lead, conflict, arc, platform, episode count, episode length, tone, hook rhythm, ending direction | Score high enough to route into briefing |
| 2 | Creative briefing | Locked one-line logline, core conflict, story direction, ending direction, hook and payoff rhythm, target platform and audience, episode-by-episode outline, writer notes | Brief locked; no script starts without it |
| 3 | Cast & character lock | Named cast with visual fields complete; lead includes character arc | Cast complete and aligned to brief |
| 4 | Continuity bible build | Structured record of identities, stable traits, current state, relationships, open/resolved plot threads, per-episode appearance table, batch summaries, prop visual notes | Bible initialized and referenced by writing |
| 5 | Batched script writing | Beat sheet per episode first, then dialogue, then cliffhanger; scripts written in batches of several episodes | Batch passes rule-based quality check |
| 6 | Look development & style lock | Chosen visual style; shared style path for characters, scenes, props, and video prompts | Style locked across asset types |
| 7 | Character, scene, and prop look-dev art | Final approved character turnarounds, scene plates, and prop art; scene plates contain no people | Assets readable and complete for scheduled scenes |
| 8 | Episode-to-shot blocking | Each episode cut into roughly 10-second field blocks; each block is an independent shooting unit | Blocks reviewed; long blocks re-split |
| 9 | Reference mapping per shot | Per-block reference table in fixed order: scene → props → characters; manual binding for nicknames and cameos | Missing references flagged; creator decides to fill or skip |
| 10 | Engineered shot prompts | Prompt per block covering subject, action, environment, lighting, camera move, style, quality, constraints; one camera move per shot | Prompt reviewed and editable before submission |
| 11 | Shooting, queue, and takes | Submitted render jobs in an isolated queue; multiple takes per block; failed jobs visible and retryable | Creator selects preferred takes |
| 12 | Review, reshoot, and handoff | Selected cuts, reshoots by editing prompts or swapping references, final faststart-ready files | Deliverable cut handed to editing and ops |
The order matters. Locking look dev before the script is finished wastes art. Generating video before references are mapped guarantees drift. Writing episode 20 without a continuity bible guarantees amnesia.
Stages 1–4: Before a single script page is written
Most teams want to jump straight to scripts. That is the single most expensive mistake.
Stage 1 — Idea intake
A raw idea like "a disgraced heiress returns for revenge" is not shootable. Intake converts it into a structured brief by forcing decisions across ten dimensions: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook and payoff rhythm, ending direction, and target platform and audience.
For vertical short drama, episode length generally lives in the 1–2 minute range, with an outer ceiling around 3–4 minutes. The intake should also score completeness; a half-baked idea gets routed through deeper guided development, while a well-formed idea can move straight to briefing. Treat this like a greenlight meeting: no shooting starts until the brief is locked.
Stage 2 — Creative briefing
The brief is the contract between the creator and everything downstream. It must include:
- One-line logline
- Core conflict
- Story direction
- Ending direction
- Hook and payoff rhythm
- Target platform and audience
- Episode-by-episode outline
- Notes for the writer
- Lead character with a stated arc
If the lead has no arc, the brief is not locked.
Stage 3 — Cast & character lock
Before art or scripts, name the cast and fill in their visual fields. The gate here is mechanical: cast must be complete, names valid, visual fields populated, and leads aligned to the brief. AI video does not forgive vague protagonists; if "the male lead" is still a placeholder when you start generating, every shot will invent a new person.
Stage 4 — Continuity bible
The continuity bible is a structured, living record of everything the show must remember across episodes: character identities, stable traits, current and changing states (injuries, revealed identities, status changes), relationships, open and resolved plot threads, a per-episode appearance table, batch-level plot summaries, and visual notes for props.
This is the opposite of asking a model to "remember" earlier episodes. Memory drifts; a structured bible does not. Writing pulls only the slice it needs—current character states, unresolved threads, recent batch summaries, and the current batch's main line—then writes back to the bible after each batch.
Stages 5–7: Scripts and look dev
Stage 5 — Batched script writing
Vertical short drama scripts follow tight rules that should be enforced, not hoped for:
- Golden 3 seconds: The first scene must open on conflict or suspense; no slow setup, no expository voiceover.
- Per-episode structure: Opening hook (one scene), escalating conflict, end-of-episode cliffhanger (final scene).
- Payoff density: At least one small payoff per episode (a reveal, a counter-strike, an identity hint, evidence secured); a larger payoff every few episodes.
- Dialogue: Short lines, generally under twenty characters in the original-language phrasing intent; no essay-speak, no paragraph-long narration explaining the plot.
The writing order inside each batch matters. Beat sheets and cliffhangers come first; dialogue comes after. From the second batch onward, lock four things before writing: the batch's plot direction, focal characters, threads and conflicts to advance, and the end-of-batch hook. These become hard constraints that override stale bible state when the creator has clearly changed intent.
Every batch passes through a rule-based quality check that catches malformed episode titles, mismatched scene counts, missing character lines, too-thin dialogue, and placeholder text like "to be continued." Failing batches are sent back for rewrite with the error attached. Half-finished scripts should never reach the shooting stage.
Stage 6 — Style lock
Choose one visual style path and apply it everywhere. A style manual should cover character art rules (face anchors, materials, mood, view consistency), scene and prop rules, and video style tags. Characters, scenes, props, and video prompts all consume the same style path; otherwise you get anime faces cut into live-action footage, the most common tell of an amateur AI production.
A production-ready system typically offers a library of style manuals spanning 2D, 3D, and photoreal directions—urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink-and-brush, xianxia, 3D donghua, stop-motion clay, cyberpunk-Chinese fusion, and similar lanes.
Stage 7 — Look-dev art
Generate and approve three asset classes:
- Character art: Final turnarounds per character, with gender treated as a hard rule and style manual applied. Historical or time-travel stories route to period-appropriate looks by scene.
- Scene plates: Clean location art with no people in frame; these become the scene reference for shots set there.
- Props: Pulled from scripts using the script's own naming, then merged into a show-wide catalog so nothing is lost between episodes.
Only completed, readable assets move downstream. Shooting should never consume half-rendered art.
Stages 8–10: From episode to shootable shot
Stage 8 — Shot blocking into field blocks
An episode is not one big render. It is a list of roughly 10-second field blocks, each with its own prompt, its own references, its own takes, and its own history. Think of the editing room bin: clip 1, clip 2, clip 3, each addressable on its own.
The soft target is about two hundred characters of script body per block; longer scenes are re-split on action beats, paragraph breaks, or sentence endings. Crowd characters and generic extras are separated from the named, drawable cast so they do not consume character reference slots. You can always re-block an entire episode if the rhythm is wrong.
Stage 9 — Reference mapping
This is the single most important mechanism for character consistency. For every field block, build a reference table in a fixed order:
- Scene reference
- Props appearing in this block
- Characters appearing in this block
The rule is strict: if a character has a reference image, the prompt must not describe that character's clothing or appearance in text. Text describes only action, expression, and injury. When text fights the reference image, you get costume swaps and face changes. This rule is what separates consistent AI drama from random generations.
Reference mapping also needs manual overrides for real production problems:
- Binding a script nickname ("the officer") to the correct character card ("Li Qiang")
- Including or excluding props so irrelevant items do not steal reference slots
- Swapping to a different approved version of the same asset
- Routing period looks for flashback or cross-time scenes
Before prompts are generated, run a completeness check that lists missing scene, character, or prop references. You can choose to skip and shoot on text-only prompts, but you should know going in that quality usually drops; the professional move is to finish look-dev first.
Stage 10 — Engineered shot prompts
A good shot prompt is written in the language of a storyboard, not a novel. It covers eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color tone
- Camera movement
- Visual style
- Image quality
- Constraints
For simple scenes, one paragraph is enough. For complex cinematic scenes, use a three-part structure: overall setup, then shot-by-shot beats numbered by shot (not by absolute timestamps like "0–3s"), then a constraint pack. One camera move per shot; do not stack push, pull, pan, and tilt into a single shot. Actions should be specific and quantified, favoring slow continuous motion over explosive motion that models tend to break on.
The constraint pack is non-negotiable: quality level, face stability, no watermark or logo, anti-twinning for multi-person scenes, and style anchoring for non-photoreal looks. Dialogue, sound effects, and music use consistent notation so downstream editing can parse them. Prompts are built only from the current block's material—never leak references or characters from other scenes—and generation settings err on the conservative, stable side rather than the wildly creative side.
Before prompts go out, strip specific copyrighted IP or title references while keeping the technique and aesthetic descriptors, to reduce downstream blocking risk.
Stages 11–12: Shooting, retakes, and handoff
Stage 11 — Shooting and takes
The shooting workflow mirrors a real set:
- Open the episode's video workspace.
- Confirm references are complete (or deliberately note gaps).
- Generate the shot prompt.
- Read and edit it—this is the director revising the storyboard note.
- Choose model tier, aspect ratio, resolution, and duration.
- Submit to the render queue.
- Review historical takes per block and pick the best.
Default delivery is 9:16 vertical, not a landscape render cropped in post. Duration can be intelligent or fixed, typically in the 5–15 second range. Multiple model tiers give you the quality/cost/speed tradeoff a real production chooses between setups.
Production-grade queue behavior matters more than people expect: video jobs run in isolated lanes separate from writing and art, you cannot submit parallel jobs for the same block while one is running, credits are pre-deducted and released on failure, hung jobs time out and become retryable, and existing vendor task IDs are polled rather than re-created. Record actual delivered duration after faststart processing rather than trusting the vendor's reported length. This is what makes the pipeline operable day after day, not a demo script.
Stage 12 — Review, reshoot, handoff
Reshoots work like they do on a real set:
- Edit the prompt and reshoot = revise the storyboard note and shoot another take.
- Swap a character, scene, or prop reference = change the wardrobe or location plate.
- Rebind a character manually = fix a casting mismatch.
- Include or exclude a prop = control what is in frame.
- Switch style = change the visual language show-wide.
- Pick a different historical take = choose the better performance from dailies.
There is no automatic "best take" scorer. Final quality judgment sits with the creator and the producer. The system's job is to give you clean takes, editable prompts, and a reliable queue—not to declare a shot finished for you.
What this pipeline deliberately does not do
Being honest about limits is what makes a production tool trustworthy. Seven boundaries are worth stating plainly:
- No automatic quality scoring or auto-retake engine. Final selection stays with the creator; the system provides multiple takes and prompt-driven reshoots.
- References are not a hard gate. Missing art triggers a reminder, but you can shoot text-only; expect lower quality and plan accordingly.
- Reference tokens in prompts rely on the specification, not hidden re-injection. Read the prompt before you submit and confirm references are present.
- Shot grammar follows the built-in camera specification. Style manuals inject style tags into video; the full art manual applies on the image side.
- The ~10-second block is an engineering heuristic, not a timecode-precise cut. Over-long blocks are re-split, but some variance remains; final trimming and stitching belong in editing.
- There is no cross-shot automatic continuation or video extension workflow. The unit of work is one field block with its own references, producing one clip.
- Character consistency depends on the look-dev asset chain, not face-embedding verification. Period routing is rule-based, and final look still depends on asset quality and on following the "don't describe what the reference already shows" rule.
The point of the pipeline is not to promise zero-drift, human-free output. It is to turn the parts of production that are repetitive, drift-prone, and easy to break into a constrained, reviewable assembly line, while keeping aesthetic judgment firmly on the human side of the screen.
Where Maosika sits in this
Maosika (猫斯卡) is an AI production operating system for vertical short dramas that hard-codes the staged pipeline above into a working production chain: idea intake and routing, locked creative brief, cast and continuity bible, batched scripting with rule-based checks, shared style path across art and video, look-dev art for characters/scenes/props, roughly 10-second field blocks, per-block reference mapping, engineered shot prompts, queued multi-take shooting, and prompt- or reference-driven reshoots. It uses 18 digital specialists modeled on real crew roles across the writing and video chains, with streaming work logs you can read like a producer reading a daily production report.
It does not replace writers, directors, or producers. It enforces the order of operations that professional crews already follow, so that the decisions made at each gate actually carry through to every episode and every shot.
If you are building an AI short drama slate and want a pipeline that treats production like production rather than like a single generate button, you can explore the workflow at Maosika.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com