How AI Vertical Short Dramas Actually Get Made: A 12-Stage Production Pipeline From Idea to Deliverable

Maosika Editorial | Last updated

Most AI short drama failures are not model failures — they are pipeline failures. The fix is not a better "generate" button, but a locked sequence: brief first, then continuity, then look, then shots, then takes.

The pipeline is the product

If you have watched enough AI-generated vertical dramas, you already know the failure modes by heart: the lead changes face between episodes, the costume drifts scene to scene, episode 7 forgets a clue episode 3 planted, the opening spends 40 seconds on backstory, and one episode ends without a cliffhanger because the model wandered.

These are not separate bugs. They are symptoms of the same root cause: the production treated the project as a single long generation, instead of a sequence of controlled handoffs.

A production-ready AI vertical short drama pipeline is a staged workflow in which each stage produces a reviewable artifact, that artifact is locked before the next stage begins, and the downstream stages only consume approved assets from upstream. The point is not to remove the filmmaker — it is to remove the parts of the process that drift, forget, or contradict themselves.

This article walks through that pipeline the way a real production room would: stage by stage, with the deliverable, the gate, and the common failure at each step.

Why "one big prompt" fails for vertical drama

Vertical short drama is unforgiving in a way standalone AI video clips are not:

  • It is episodic, so continuity compounds across episodes.
  • It is 9:16 native, so framing and pacing cannot be fixed by cropping later.
  • It lives or dies on the first three seconds, so structure is not optional.
  • It is asset-heavy — recurring characters, recurring locations, recurring props.
  • It is a deliverable, not a demo — queues, retries, and version history matter.

A single long prompt cannot enforce any of that. It has no memory of what was approved, no concept of a locked brief, no difference between a character description and a character asset, and no notion of a "take" you can compare against another.

The fix is industrial, not literary: break the job into stages where each stage has one job, one artifact, and one gate.

The 12 stages, in order

The order matters. Skipping a stage, or doing them out of order, is where most teams quietly lose quality.

#StagePrimary artifactGate before moving on
1Idea intake & evaluationScored brief readinessMissing dimensions are filled, not guessed
2Creative guidanceCompleted creative dimensionsGenre, lead, conflict, arc, platform, tone are explicit
3Brief lockOne-page creative briefLogline, conflict, arc, hook rhythm, ending direction signed off
4Cast & visual confirmationCharacter lineup with visual fieldsCast is complete and aligned to the brief
5Story archive (continuity bible)Structured show bibleCharacters, relationships, open threads, prop notes are recorded
6Batch script writingBeat sheets → scripts → rule checkEach episode passes structural rules before delivery
7Style selectionChosen look with shared style pathOne style governs characters, scenes, props, and video
8Look development: characters, scenes, propsApproved reference artOnly finished, readable assets move to shooting
9Episode splitting into scene blocks~10-second shot blocksEach block has its own script slice and asset map
10Shot promptingEngineered shot prompts per blockPrompts reference assets, not prose descriptions of them
11Multi-modal shootingRendered takes per blockTakes are reviewable, retryable, and versioned
12Review, reshoot, selectFinal selected takes per blockHuman picks; system supports re-word and re-shoot

Below is what each stage actually does in practice.

1. Idea intake & evaluation

Not every idea arrives ready to write. A pitch that says "a CEO revenge drama, 80 episodes, female lead" is missing most of what a writer would need to start.

The job of this stage is to score how complete the incoming idea is, then route it: ideas that are already well-formed go straight to brief; partial ideas get prompted for the missing dimensions; thin ideas go through a full guided intake. The key principle is that missing information is filled by asking the creator, not invented by the model.

2. Creative guidance

Vertical drama has a specific set of dimensions that must be explicit before writing starts: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook and payoff rhythm, ending direction, and target platform / audience.

Episode length for vertical drama typically lands in the 1–2 minute range, with an outer bound around 3–4 minutes. This is not an arbitrary preference — it changes how beats are distributed, how often cliffhangers hit, and how much dialogue can fit per scene.

3. Brief lock

The creative brief is the first hard gate. Nothing proceeds to writing until it is locked.

A locked brief includes the logline, core conflict, story direction, ending direction, payoff and hook rhythm, target platform and audience, episode-by-episode outline, and notes for the writer. The lead character entry must include the character arc — not just a job title and a mood.

Think of this as the greenlight meeting. Once it is locked, downstream stages treat it as a constraint, not a suggestion.

4. Cast & visual confirmation

Before any art is generated, the cast must be complete and aligned to the brief. Names must be valid, visual fields must be filled, and the leads must match who the brief says the story is about.

This is a boring step, and it is where a surprising number of productions quietly break. If "the second male lead" has no entry in the cast list, every downstream stage will improvise him differently.

5. Story archive (continuity bible)

The story archive is a structured record that travels with the show for its entire run: character identities, stable traits, current state (injuries, revealed identities, changed relationships), relationship maps, open and resolved plot threads, episode appearance tables, batch-by-batch plot summaries, and visual descriptions of recurring props.

In traditional writers' rooms this document is called a continuity bible. Its job is to make sure episode 60 still remembers what episode 4 established. In an AI pipeline, the rule is simple: do not rely on model memory; rely on a structured archive that gets read before writing and updated after each batch.

6. Batch script writing

Writing happens in batches of several episodes, not as one 80-episode generation. Each batch follows the same internal order:

  1. Beat sheet first — list the scenes and the episode-end cliffhanger before writing dialogue.
  2. Then the script itself.
  3. Then a rule-based quality check.
  4. Then archive update, so the next batch inherits the new state.

Vertical drama has hard-earned structural rules that belong in the check, not in the vibes:

  • Golden 3 seconds: the first scene must open on conflict or suspense; no slow backstory opens.
  • Episode shape: opening hook → escalation → end-of-episode cliffhanger.
  • Payoff density: at least one small payoff per episode; a larger payoff every few episodes.
  • Dialogue: short lines, generally under ~20 characters in Chinese or the equivalent tight line in English; no essay-like monologue, no narrator explaining what we could watch.

The rule check catches mechanical failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, or placeholder text like "to be continued." Scripts that fail are sent back for rewrite with the specific error attached. Half-finished scripts do not get handed to the shooting stage.

7. Style selection

A consistent look is not a matter of typing the same style phrase over and over. It requires a shared style path that characters, scenes, props, and video prompts all use.

A production-grade style library covers multiple directions — urban live-action realism, period realism, mature urban romance animation, 90s anime, Chinese ink style, xianxia, 3D donghua, stop-motion clay, cyberpunk-Chinese fusion, and others — with each style defining how faces are anchored, what materials look like, what mood the lighting carries, and which video style tags apply. Once a style is chosen, everything downstream inherits it. This is how you avoid the common failure where the character art is anime but the video output turns photoreal.

8. Look development: characters, scenes, props

Before a single scene is shot, the recurring assets must exist as finished, readable art:

  • Character lookdev turns archive descriptions into image prompts using the chosen style's character rules, then generates the art. Gender is a hard rule, not a suggestion.
  • Scene art is generated as empty plates — no people in the scene reference.
  • Props are pulled episode by episode from what the script actually calls them, then merged into a show-wide catalog so nothing is lost between episodes.

Only finished, readable assets move on. The shooting stage does not consume half-drawn placeholders.

9. Episode splitting into scene blocks

Each episode is split into scene blocks targeting roughly 10 seconds each, with a soft cap on script length per block. Longer beats get split again on action, paragraph breaks, or sentence boundaries.

This is one of the most important shifts in mindset from "AI video demo" to "AI drama production." You do not think of "episode 3" as one big video. You think of it as a clip list: episode 3, scene 1; scene 2; scene 3 — each with its own prompt, its own reference set, its own output, and its own history of takes.

Background extras and generic characters are separated from the drawable main cast, so they do not steal reference slots from people who actually recur.

10. Shot prompting

This is where most prompt-writing advice for generic AI video is actively wrong for drama. A shot prompt is not a paragraph of novelistic prose. It is a shot list entry.

A well-engineered shot prompt covers eight elements:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color tone
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Negative / stability constraints

A few rules separate production prompts from hobby prompts:

  • One shot, one camera move — do not stack push, pull, pan, and tilt into a single shot.
  • Use shot numbers, not absolute timestamps like "0–3s".
  • Always include a constraint package for face stability, image quality, and no watermark; add twin / duplicate prevention for multi-character shots.
  • Keep actions low and continuous rather than explosive — fast chaotic motion breaks more often.
  • Mark dialogue, sound effects, and music with consistent notation.
  • Only feed this scene's assets into this scene's prompt — no cross-scene leakage.

Most importantly: if a character has a reference image, the prompt must not re-describe their clothes or appearance in text. The reference is the source of truth. Text only describes action, expression, and injury state. This single rule eliminates a huge share of "the character changed outfit mid-scene" failures.

Before prompts are handed off, IP names are stripped out — style and technique descriptions stay, specific copyrighted titles do not.

11. Multi-modal shooting

Shooting a scene block follows a clear path: open the episode's video workspace, confirm references are present (or intentionally skipped), generate the shot prompt, read and edit it, choose model tier / aspect ratio / resolution / duration, submit, wait in the render queue, then review the resulting takes.

The default deliverable is 9:16 vertical — not a horizontal video cropped after the fact. That changes composition from the first prompt.

A real production needs production-level controls, not just a generate button:

ControlWhat it is equivalent to on a real set
Editing the shot promptThe director revising the shot list
Swapping a character / scene / prop referenceChanging a lookbook plate or location
Manual character bindingFixing a nickname or off-screen reference in the script
Including / excluding a propControlling what is in frame for this shot
Switching styleUnifying the visual language across units
Choosing model tierTrading quality, cost, and speed
Reviewing historical takes per sceneMulti-take selection

The queue itself has to behave like production infrastructure: separate lanes for script, art, and video tasks; no parallel submissions for the same block; pre-deducted credits that release on failure; timeout recovery for stuck jobs; and no duplicate task creation when polling an existing vendor job.

12. Review, reshoot, select

There is no magic "auto-pick the best take" engine. The final quality call sits with the creator. What the pipeline provides is the ability to compare takes, change the wording of the prompt, swap a reference, and reshoot — exactly like on a real set, where you adjust the shot and roll again.

This is also where you accept an honest limitation: the ~10-second scene block is an engineering heuristic, not a timecode-precise edit. Some blocks run longer, and stitching blocks into a finished episode still belongs in the edit.

Where humans still decide

A common misunderstanding is that an AI production pipeline replaces the filmmaker. It does not, and it should not pretend to. The human is the decision-maker at every point where taste, story, and audience matter:

  • Locking the brief.
  • Approving the cast.
  • Signing off on lookdev.
  • Setting batch intent before each new batch of episodes (direction, focus characters, threads to advance, end hook).
  • Reading and editing shot prompts.
  • Choosing between takes.
  • Deciding when a reshoot is needed.

What the pipeline does is take the parts that are repetitive, drift-prone, and self-contradictory — remembering continuity, applying the same style, enforcing shot structure, routing assets, managing the queue — and turn them into a constrained process. The aesthetic judgment still sits with a person.

What this pipeline does not do

It is worth being explicit about the boundaries, because teams that trust the hype waste the most time:

  1. There is no automatic scoring or auto-reshoot engine that picks the best take for you. Final review is human.
  2. Reference images are not a hard gate. You can skip them and shoot from text-only — quality will usually be worse, which is why the disciplined path is to lock lookdev first.
  3. Reference tags in prompts rely on disciplined structure, not silent magic. The creator should still read the prompt and confirm references are present.
  4. Shot grammar follows the built-in camera spec; the full art manual applies to lookdev. The style system feeds video style tags, not the entire art book.
  5. ~10-second blocks are heuristic, not precision editing. Long scenes may still need to be split further, and final assembly is still an edit job.
  6. There is no cross-scene automatic video continuation workflow today. The unit of work is one scene block, with its own references, producing one clip.
  7. Character consistency depends on the lookdev asset chain, not a face-embedding lock. Era-specific looks are routed by rule, but final quality still depends on the reference art and on prompts obeying the "do not re-describe clothed appearance" rule.

These are not apologies. They are the boundaries of a tool that is meant to be used by professionals, not a demo that promises zero-human hit videos.

The shape of a reliable AI drama workflow

If you take one thing away from this breakdown, it is this: high-quality vertical drama is not generated in one long shot. It is assembled from controlled units, each of which you can see, change, and roll back.

The order is the discipline:

Idea → brief → archive → scripts → style → lookdev → scene blocks → shot prompts → takes → human selection.

Every time a team skips a step — writes before the brief is locked, shoots before characters are designed, prompts from prose instead of from assets, treats episodes as single giant generations — they re-introduce the exact failures AI drama is famous for.

Maosika (猫斯卡) is built around this sequence. Rather than offering a single "generate drama" button, it hardcodes the staged handoffs, the reviewable artifacts, the continuity archive, the asset chain, and the take-based shooting workflow — while leaving the creative decisions firmly in the creator's hands.

If you are planning your own AI vertical drama production, building or choosing a tool, the question to ask is not "how good is the video?" on a single cherry-picked clip. The question is: does the pipeline enforce the right order, and can you inspect and intervene at every handoff? That is what separates a demo from a deliverable.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com