The 12-Step AI Vertical Short Drama Pipeline: From Idea to Finished Episode

Maosika Editorial | Last updated

Most AI short drama failures aren't model failures — they're pipeline failures. The fix isn't a better prompt; it's a locked order of operations: brief first, then continuity, then look-dev, then shots, then takes.

If you've watched AI-generated micro-dramas, you already know the symptoms: a lead whose face changes between scenes, costumes that reset mid-episode, props that appear and vanish, cliffhangers that forget what the last episode promised, and a third act that drifts because no one was holding the story together.

These aren't isolated bugs. They happen when a team treats "make an episode" as one big generation instead of a sequence of controlled handoffs. Professional crews don't shoot a movie before the script is locked, and they don't design costumes on the day of the shoot. AI workflows need the same discipline — just reorganized around what machines are good at (consistency, repetition, structured handoffs) and what humans are still needed for (taste, story judgment, final take selection).

Below is the 12-step pipeline that a production-grade AI vertical short drama system should enforce. It's written for producers, writer-directors, and studios who want to ship episodes at scale, not toy with one-off generations.

Why order matters more than the model

A single text-to-video button is a demo. A pipeline is a factory. The difference is that every stage in a pipeline produces a verifiable intermediate artifact that the next stage consumes. If a stage is broken, you fix it there — you don't try to patch it three steps later with a stronger prompt.

The core principle is simple:

High-quality short dramas are never one long generation cut up; they're stacks of controllable units.

For vertical (9:16) micro-dramas — typically 1–2 minutes per episode, sometimes stretching to 3–4 — this matters even more. Episodes are short, hooks are merciless, and audiences drop off in seconds. A consistency break isn't a nitpick; it's a retention killer.

The 12-step production pipeline

The steps below must run in this order. Skipping ahead is exactly what causes the failures listed at the top of this article.

#StepWhat gets lockedWho decides
1Idea evaluation & intakeWhether the concept is complete enough to proceedSystem scores, human confirms
2Creative guided intakeMissing dimensions filled inHuman, guided
3Creative brief lockLogline, conflict, arc, hooks, audience, episode breakdownHuman sign-off
4Cast & visual confirmationCharacter lineup aligned with the briefHuman sign-off
5Story archive builtContinuity bible: characters, threads, relationships, stateSystem maintains, human reviews
6Batch script writingBeat sheets → drafts → rule checks → archive updateHuman approves batches
7Style selectionOne visual language for the whole showHuman picks
8Look-dev: characters, scenes, propsFinal approved art assetsHuman approves each
9Episode split into scene blocks~10-second shootable unitsSystem proposes, human can re-split
10Per-scene binding & shot promptsReference map + engineered shot instructionsHuman reviews and edits
11Multi-modal shootRaw takes per scene blockSystem renders, human picks takes
12Review, revise, re-shootLocked episodeHuman final call

Let's walk through each step with the actual work that happens, and the failure mode it prevents.

Step 1 — Idea evaluation

Not every idea is ready to become a show. A good intake scores the concept across the dimensions vertical drama actually needs: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, platform, and target audience.

A complete concept can skip most hand-holding. A half-formed one needs guided questions. A thin one gets a full, structured intake. The point is: the system should refuse to pretend a vague idea is a finished brief, the same way a real development room would.

Step 2 — Creative guided intake

This is where missing pieces get filled in, not by the model guessing, but by asking the creator. For vertical drama, the intake is tuned to the form: hook density, cliffhanger placement, episode length, and platform expectations are first-class inputs, not afterthoughts.

Step 3 — Creative brief lock

The brief is the first hard gate. Until it's locked, nothing after it should run. A locked brief contains the logline, core conflict, story direction, ending direction, beat and hook rhythm, target platform and audience, episode-by-episode synopsis, and notes for the writer. The protagonist's field must include their character arc — not just their name and job.

Think of this as the development meeting that ends with everyone signing the same document before the cameras roll.

Step 4 — Cast & visual confirmation

Characters are confirmed against the brief: lineup complete, names valid, visual fields filled, leads aligned with what the brief promised. This is a hard gate, not a suggestion. If the cast list doesn't match the brief, the system should refuse to move forward.

Step 5 — Story archive (the continuity bible)

A story archive is a structured, show-long record of who characters are, how they relate, what they currently look like, which plot threads are open or resolved, who appears in which episodes, batch-by-batch plot notes, and how props are described. It's the digital equivalent of a writers' room continuity bible — the thing that makes sure episode 200 still remembers what episode 3 set up.

This is the single most important structure for long-running shows. Models don't "remember" — they get fed a slice of the archive for the batch they're writing: current character states, unresolved threads, recent batch summaries, and the current batch's main line. After each batch, the archive updates itself.

Don't rely on model memory. Rely on a structured archive.

Step 6 — Batch script writing

Scripts are written in batches of several episodes, never one giant run. The order inside each batch is itself a quality gate:

  1. Beat sheet first — every episode starts with a scene-by-scene beat sheet and the episode-end cliffhanger, *before* any dialogue is written.
  2. Then the draft.
  3. Then a rule-based quality check that catches: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, and placeholder cop-outs like "to be continued."
  4. Failures auto-rewrite with error feedback, up to a capped number of attempts; nothing half-baked reaches the user.
  5. Archive update after the batch passes.

Vertical drama has hard writing rules baked into this stage:

  • Golden 3 seconds: the first scene must hit hard conflict or strong suspense — no slow opens, no exposition dumps.
  • Per-episode structure: opening hook (1 scene) → rising conflict → end-of-episode cliffhanger (final scene).
  • Payoff density: at least one small payoff per episode (a comeuppance, a reveal, an identity hint, evidence secured); a bigger payoff every few episodes.
  • Dialogue: short sentences, generally under ~20 characters in the original Chinese writing discipline; no lecturing, no voiceover explaining the plot.

From the second batch onward, before writing starts, four things are pinned down: this batch's plot direction, focus characters, threads and conflicts, and end-of-episode hooks. These become hard constraints that override archive drift when the creator's intent and the archive disagree.

Step 7 — Style selection

A show needs one visual language. A production system should ship with multiple style packs covering 2D, 3D, and photoreal directions — urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink-and-brush, xianxia fantasy, 3D donghua, clay stop-motion, cyberpunk-Chinese fusion, and so on.

Each pack includes a character sheet guide (facial anchors, materials, vibe, multi-view consistency), a scene/prop guide, and video style tags. Once chosen, character art, scene art, prop art, and video prompts all share the same style path — so you don't end up with anime characters in a photoreal world.

Step 8 — Look-dev: characters, scenes, props

Art is produced in a two-step process:

  1. Text polish: the archive description is turned into an art-ready prompt using the chosen style pack and hard gender rules. The creator can override this by hand.
  2. Image generation: the final prompt is rendered, reference images are supported, and every version is saved to a browsable history.

Only finished, readable art assets are passed downstream. The video pipeline never shoots against half-rendered concept art.

A few look-dev disciplines separate a serious pipeline from a toy:

  • Scene art must be empty of people. No cheating by generating a scene with a stand-in character baked in.
  • Props are extracted per episode from the script using the script's own wording, then merged into a show-wide catalog so nothing is forgotten between episodes.
  • Creator/script descriptions override style-pack world rules. The style pack controls *how* things are drawn, not *what* is allowed to exist.
  • Gender is a hard rule at generation time, not left to the model.

Step 9 — Episode split into scene blocks

An episode isn't one big video. It's sliced into scene blocks targeted at roughly 10 seconds each, with a soft cap on script length and re-splitting on action beats, paragraph breaks, and sentence endings when a block runs long. Extras and generic characters are separated from the drawable main cast so they don't consume character reference slots.

What the creator sees isn't "Episode 3, one file." It's "Episode 3 → Scene 1, Scene 2, Scene 3…" — each with its own prompt, its own reference images, its own output, and its own history of takes. This is the editing room's clip list, not a single export.

Step 10 — Per-scene binding & shot prompts

This is where most of the "AI face swap" problem is actually solved. For each scene, a reference table is built in a fixed order:

Scene art → scene props → character look-dev art

Slots are only filled when an image exists; gaps are marked as text-only rather than silently fabricated. The highest-priority rule is:

When a character has a reference image, the prompt is forbidden from describing their clothes or appearance in text. The reference image is the source of truth; text only describes action, expression, and injury state.

That one rule eliminates the classic failure where the text says "black suit" and the reference shows a leather jacket, and the model compromises by generating a third thing. Additional controls include manual character binding (when the script says "the officer" but the archive name is "Li Qiang"), manual prop inclusion/exclusion to control visual focus, swapping in a different historical version of the same asset, and era-based routing (time-travel/flashback scenes pick period-matching character looks by keyword).

Before prompts are generated, a readiness check lists any missing scene/character/prop references and offers to route the creator back to draw them. Skipping is allowed — but it's an explicit choice, not a silent failure.

The shot prompt itself follows an engineered, eight-element structure:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting & color
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Constraints

Complex cinematic scenes use a three-part structure (overall setup → shot 1/2/3… → constraint pack). One shot = one camera move — no stacking pans, tilts, and zooms into a single shot. Shots are numbered; absolute timestamps like "0–3s" are forbidden. A mandatory fallback pack covers image quality, facial stability, and no watermarks/logo; multi-person scenes add twin/duplicate fallbacks; non-realistic styles pin the style anchor. Actions are described with quantified, continuous motion rather than explosive, hard-to-render dynamics. Dialogue, sound effects, and BGM use dedicated symbol conventions so they're not confused with visual direction.

Crucially, prompts are built only from this scene's materials — no bleeding in characters or props from other scenes. The script text for this scene is the highest-priority source. Generation parameters are conservative, favoring stability over wild creativity.

The system teaches the model to speak in shot-list language, not novel prose.

Before delivery, prompts are cleaned of specific copyrighted work/IP names (keeping the technique and aesthetic descriptors) to reduce downstream copyright blocking. If a reference image is flagged as a suspected real-person photo, the system returns clear guidance to re-route through the platform's art pipeline rather than shooting against a real photo. Topic intake runs against a compliance filter that tracks platform and regulatory restrictions as they evolve.

Step 11 — Multi-modal shoot

The standard shoot path is: open the episode's video workspace → confirm references are ready (or deliberately skip) → generate shot prompts → read and edit them by hand → pick model tier, aspect ratio, resolution, and duration → submit → queue renders → review the take history and pick the best.

Aspect ratio defaults to 9:16 vertical (other ratios are supported). Duration can be smart or fixed at 5–15 seconds. Model tiers trade quality, cost, and speed; resolution follows each model's allowlist. Audio is on by default; watermarking is off by default. Vertical is the *default deliverable*, not a post-hoc crop — that's a meaningful distinction for platform-native delivery.

The professional control points map directly to real crew roles:

Control pointOn a real set, that's…
Editing the shot promptThe director revising the shot list
Swapping character/scene/prop referencesChanging a costume or switching the location board
Manual character bindingFixing name mismatches and cameo references
Including/excluding propsControlling what's in frame
Switching style packsUnifying the visual language
Choosing model tierTrading quality / cost / speed
Per-scene take historyShooting multiple takes and picking the best

When you hit generate, the worker uses exactly the prompt you confirmed (or edited) and the reference map locked at that moment. Change the prompt and hit generate again, and that's a new take — exactly like revising shot notes and rolling camera again.

Step 12 — Review, revise, re-shoot

Finished videos get faststart treatment and are stored with their actual measured runtime, not just the vendor's reported duration. Video jobs run on an independent queue lane, isolated from script and art tasks, with concurrency slots. A scene block already in flight can't be double-submitted, preventing duplicate charges and state confusion. Credits are pre-deducted and settled; failures auto-release. Stuck tasks time out and are recoverable and retryable; if a vendor task ID already exists, polling resumes rather than spinning up a second job and double-charging.

This is what makes the pipeline operable as production infrastructure, not a toy script.

Where the human still sits

It's worth being explicit about what this kind of system does *not* do, because the honest boundary is where trust is built:

  1. There is no automatic scoring or auto-pick of the best take. Final quality judgment stays with the creator and showrunner; the system provides multiple takes and the tools to re-shoot with revised prompts.
  2. Reference images are not a hard gate. Missing images trigger a reminder, but you can skip and go text-only — quality will usually be worse, which is why the professional flow locks look-dev before shooting.
  3. Reference tags in prompts rely on the spec, not a second hard injection. Creators should still confirm tags are present when reviewing prompts.
  4. Video shot grammar follows the platform's built-in shot spec. The style pack injects video style tags; the full art manual is applied on the look-dev side.
  5. The ~10-second scene block is an engineering heuristic, not timecode-precise editing. Overlong scenes get re-split, but blocks can still run long; final trimming and stitching belong in a later edit stage.
  6. There is currently no cross-scene auto-extend or continuation workflow at the product level. The unit of work is "single-scene multi-modal reference → single clip."
  7. Character consistency depends on the look-dev asset chain, not face-embedding verification. Era routing is rule-based; final look still depends on art quality and whether the prompt respects the "don't describe what the reference shows" rule.

The positioning this leads to is straightforward:

The role of an AI production operating system is to hard-code the quality-control order that professional short drama already proved works — story and continuity first, then look-dev, then scene breakdown, then shot instructions, then multi-take selection. It doesn't claim to remove humans or ship zero-error episodes; it turns the repetitive, drift-prone, loss-of-control parts of the job into a constrained pipeline, while taste and judgment stay on the human side of the table.

Good stories are never scarce. What's scarce is the road that gets them to screen, episode after episode, without the show falling apart on episode 12.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com