How Vertical Short Dramas Actually Get Made: An 11-Stage AI Production Pipeline

Maosika Editorial | Last updated

AI short drama production fails not because the model is weak, but because teams skip the order of operations. The reliable path is the same one live-action crews use: lock the brief, build continuity, approve looks, block scenes, then shoot.

Most AI short drama workflows still behave like a demo: type a premise, press generate, then patch the result. That works for a single clip. It falls apart the moment you need 60 episodes with the same lead, the same clothes, the same unresolved subplot, and a cliffhanger every episode.

The production-grade approach is different. It treats AI not as a magic button, but as a pipeline with handoffs, approvals, and versioned assets at every stage. Below is the full sequence, in the order it has to happen.

The core principle: order matters more than tooling

A vertical short drama is not one long generation cut into pieces. It is a stack of controlled units: a locked brief, a structured story record, approved character and environment art, scene blocks, engineered shot prompts, rendered takes, and review passes.

A production pipeline for vertical short drama is a staged workflow where every stage produces a reviewable artifact before the next stage is allowed to start.

Skipping stages is the root cause of the familiar failures: faces change between episodes, props vanish, plots forget their own setup, episodes open with exposition instead of conflict, and the visual style drifts from 2D to realistic halfway through a scene.

The 11 stages, in order

The table below is the spine of the pipeline. Each row answers three questions: what gets produced, who approves it, and what breaks if you skip it.

#StageDeliverableGate before next stageWhat breaks if skipped
1Idea intake & scoringIntake score across premise, character, conflict, platform, toneScore determines how much guidance is neededVague brief, rewrites explode
2Creative guidanceFilled gaps in premise, protagonist arc, hook rhythm, ending directionAll required dimensions coveredWriter invents missing pieces
3Brief lockOne-line logline, core conflict, arc, hook cadence, audience, episode outlineBrief is frozen; later conflicts resolve to itDrift across batches
4Cast & visual confirmationNamed cast with visual fields aligned to the briefCast complete and fields validWrong protagonist, unnamed leads
5Story archive / continuity bibleCharacter states, relationships, open/resolved threads, per-episode appearance, batch notesArchive exists and is currentAmnesia across episodes
6Batch script writingBeat sheets first, then dialogue, then rule check, then archive updateBatch passes rule-based quality checkCliffhangers missing, format errors
7Style & look selectionChosen visual style with shared rules for characters, scenes, props, videoOne style path locked for all assetsStyle fracture across assets
8Look dev: characters, scenes, propsApproved reference art for cast, locations, key itemsVideo only consumes finished, readable artFace swaps, costume drift
9Scene blockingEpisode cut into ~10-second scene blocks with cast and locations parsedEach block has its own reference setLong unmanageable clips
10Engineered shot promptsPer-block prompts with subject, action, environment, lighting, camera, style, quality, constraintsPrompt reviewed and editable before renderModel writes fiction, not shots
11Shoot, review, re-shootMultiple takes per block; re-shoot after changing prompt or referencesBest take selected per blockNo controlled iteration

Stage 1–3: Intake, guidance, brief lock

The intake stage scores how complete the idea is. A high-completeness premise can go almost straight to brief. A thin one needs a structured conversation that fills in the ten dimensions vertical drama actually requires: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook and payoff rhythm, ending direction, and target platform.

Episode length for vertical drama mostly lives in the 1–2 minute range, sometimes extending to 3–4 minutes. The brief is not a formality. It is the document that later batches answer to when the model would otherwise improvise.

A locked brief includes:

  • A one-line logline
  • The core conflict
  • Story direction and ending shape
  • Hook and payoff rhythm
  • Target platform and audience
  • Episode-by-episode outline
  • Notes for the writer
  • A protagonist entry that includes character arc, not just appearance

The analogy is simple: you would not call action on a set before the production meeting. The same rule applies here.

Stage 4–5: Cast confirmation and the story archive

Before a single scene is written, the cast must exist as named entities with complete visual fields, and the lead must match the brief. This is a hard gate, not a suggestion.

Then the story archive is built. This is the single most important structure for long-running short dramas.

A story archive is a structured continuity record that tracks character identity, stable traits, current state (injuries, revealed identity, changed status), relationships, open and resolved plot threads, per-episode appearance, batch summaries, and prop visual notes.

The archive solves the problem that language models are bad at: remembering. Instead of asking the model to "keep everything in mind," the writing process only receives the relevant slice of the archive for the current batch: current character states, unresolved threads, recent batch summary, and the batch's main line. After each batch, the archive is updated. This is how episode 40 still knows what happened in episode 3.

Stage 6: Batch script writing

Scripts are written in batches of several episodes, not as one endless run. Each batch follows a strict order:

  1. Beat sheet first, including the episode-end cliffhanger
  2. Then full scene dialogue
  3. Then a rule-based quality check
  4. Then archive update

The writing rules embedded in the process are not arbitrary; they are vertical drama mechanics:

  • Golden 3 seconds: the first scene must open on conflict or suspense, no slow setup
  • Episode shape: opening hook → escalation → cliffhanger
  • Payoff density: at least one small payoff per episode (a reveal, a reversal, a piece of evidence, a status shift); a larger payoff every few episodes
  • Dialogue: short lines, generally under ~20 words; no lecture-style monologue

From the second batch onward, four things are locked before writing starts: the batch's plot direction, focus characters, threads and conflicts, and end-of-episode hook. These become hard constraints. If they conflict with older archive notes, the newly confirmed creative intent wins.

The quality check is rule-based, not taste-based. It catches wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, and placeholder text like "to be continued." Failing scripts are sent back for rewrite with the specific error attached. Half-finished scripts do not move forward.

Stage 7–8: Style selection and look dev

AI drama consistency breaks most often at the visual stage because character art, scene art, props, and video prompts are generated on separate paths with no shared rules. The fix is a single style source.

A production-grade style system includes a shared manual for character rendering (face anchors, material, temperament, view consistency), scene and prop rules, and video style tags. Once a style is chosen, characters, locations, props, and shot prompts all follow the same visual path.

Look dev happens in two steps:

  1. Text polish: the archive description is turned into an image-generation prompt using the chosen style's character manual and hard gender rules. The creator can override this.
  2. Image generation: the final prompt is rendered, optionally with a reference image, and saved as a versioned asset.

Video only consumes finished, readable art. It does not grab half-rendered drafts.

The reference image rule that prevents most face swaps

This single rule eliminates a large class of AI video failures:

When a character has an approved reference image, the prompt must not re-describe that character's clothing or appearance in text. The reference image is the source of truth for appearance; text only describes action, expression, and injury.

When text and reference fight, the model invents a compromise. That compromise is the face swap, the costume change, the age shift. The rule is not "add more reference images." It is "stop contradicting the reference in text."

Reference sets are built per scene in a fixed order: scene → scene-specific props → characters. Empty slots are marked as text-only rather than filled with invented bindings. Creators can manually bind a script name to the right character, include or exclude props, swap to an older approved version of an asset, and route period-correct looks for time-travel or flashback scenes.

Scene art follows one discipline: empty scenes contain no people. Props are extracted per episode from the script's own wording, then merged into a show-wide catalog so nothing is lost across episodes.

Before shooting, the pipeline checks what reference art is missing and shows a list. You can skip it and shoot from text only — but quality is usually worse, which is why the professional path is to finish look dev first.

Stage 9: Scene blocking into ~10-second blocks

Episodes are not rendered as one long video. They are cut into scene blocks targeting roughly 10 seconds each, with a soft cap on scene text length and further splitting on action beats, paragraph breaks, and sentence endings when needed.

This is the equivalent of a clip list on an editing timeline. Each block has its own prompt, its own reference set, its own rendered outputs, and its own history of takes. Crowd characters and generic extras are separated from the named cast so they do not consume reference slots meant for leads.

The default deliverable is 9:16 vertical, not a landscape video cropped after the fact. Other ratios are supported, but vertical is native.

Stage 10: Engineered shot prompts

Prompts are where most teams quietly lose control. A novel-style prompt like "a dramatic confrontation in a rainy alley" asks the model to invent coverage. A shot-list prompt tells it exactly what to do.

An engineered prompt covers eight elements:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color tone
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Constraints

Complex cinematic scenes use a three-part structure: overall setup, then shot-by-shot instructions, then a constraint pack. The rules are deliberately conservative:

  • One camera move per shot; no stacking push, pull, pan, and tilt together
  • Shot numbers instead of absolute timestamps
  • A mandatory constraint pack for quality, face stability, and no watermark or logo
  • Multi-person scenes add a "no twins / no duplicates" guard
  • Non-realistic styles anchor the style explicitly
  • Actions are specific and quantified; slow continuous motion is preferred over high-dynamic bursts
  • Dialogue, sound effects, and background music use a consistent notation so they are not confused with visual instructions
  • The prompt only sees the current scene's assets; it does not pull characters or props from other scenes

Before delivery, prompts are cleaned of specific copyrighted IP references while preserving technique and aesthetic description, reducing downstream blocking risk. If a reference image is detected as a real-person photo, the workflow guides the creator to use platform-rendered art instead.

The simplest way to say it: the system is teaching the model to speak in shot-list language, not novel language.

Stage 11: Shoot, review, re-shoot

The shoot stage is where the pipeline behaves like a real set, not a render farm.

The creator opens an episode's video workspace, confirms reference art (or deliberately skips it), generates the shot prompt, reads and edits it, chooses model tier, aspect ratio, resolution, and duration, then submits. The job enters a dedicated queue separate from writing and art tasks. Re-submitting the same block while it is running is blocked to avoid duplicate charges and state confusion. Credits are reserved on submit and released on failure; stalled jobs time out and can be retried.

Crucially, rendering does not invent a new prompt. It uses the prompt you confirmed or edited, with the reference set locked at that moment. If you change the prompt and render again, that is a new take — exactly like revising shot notes and calling action again.

The professional control points map cleanly to on-set roles:

Control pointOn-set equivalent
Editing the shot promptDirector revising shot notes
Swapping character / scene / prop artChanging a look or location board
Manual character bindingFixing name mismatches and cameos
Including or excluding propsControlling visual focus in the frame
Switching styleUnifying the visual language
Choosing model tierTrading quality, cost, and speed
Reviewing historical takes per blockMulti-take selection

There is no automatic "choose the best take" engine. Final quality judgment stays with the creator and producer. The pipeline gives you multiple takes and a controlled way to re-shoot after changing something; it does not pretend to replace the director's eye.

Where Maosika fits in this pipeline

[Craftdoc note: tier 艺 — explain the industrial solution first, then state what Maosika solidifies.]

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not replace the writer, director, or producer. It solidifies the handoffs described above into an enforced sequence: intake scoring, brief lock, cast confirmation, story archive, batch writing with rule checks, shared style path, look dev for characters/scenes/props, ~10-second scene blocks, engineered shot prompts, queued rendering, and multi-take re-shoots.

The 18 digital specialists in the system mirror real crew roles — archivist, producer, writer, script supervisor, casting director, art director, costume and makeup, props, storyboard artist, director of photography, director, camera operator, editor, and VFX supervisor — and their work is visible as a streaming production log rather than hidden inside a black box.

It is honest about what it does not do:

  • It does not auto-score or auto-pick the best finished take; final review is human.
  • Reference art is not a hard gate; missing art warns you but allows text-only shooting, with the expected quality tradeoff.
  • It does not build a cross-scene automatic video continuation workflow; the unit of work is one scene block to one clip.
  • Character consistency depends on the look-dev asset chain, not on facial embedding magic; period routing is rule-based, and final look still depends on art quality and prompt discipline.

That honesty is the point. Maosika turns the repetitive, drift-prone, failure-prone parts of short drama production into a constrained pipeline, while keeping aesthetic judgment on the human side of the table.

High-quality short drama is never one long generation cut up. It is a sequence of controlled units, each reviewable, each re-doable, each anchored to assets that were approved before the camera rolled.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com