How Vertical AI Short Dramas Actually Get Made: A Production Pipeline, Not a Generate Button

Maosika Editorial | Last updated

AI short drama production fails not because the model is weak, but because teams skip the handoffs. The reliable path is a staged pipeline: brief, archive, script, look-dev, references, shot blocks, prompts, shoot, review, reshoot.

The wrong mental model is "text in, video out"

Most teams new to AI micro-dramas imagine the workflow as one prompt producing a finished show. In practice, that approach collapses at scale: characters change faces, props vanish between scenes, episode 17 forgets a setup from episode 3, and cliffhangers stop landing because the model is improvising continuity.

The working mental model is closer to a real vertical drama crew. There is a sequence of gates, each producing an artifact the next stage consumes. You can move fast through them, but you cannot skip them without paying later in reshoots.

A vertical AI short drama pipeline is a staged production chain where each phase outputs a reviewable artifact — brief, story archive, beat sheet, script, look, character sheets, scene sheets, shot blocks, shot prompts, takes, edits — and the next phase is locked from changing what was already approved.

This article walks through that pipeline as twelve handoffs, written for producers, writer-directors, and studios that need repeatable output rather than one lucky clip.

Why vertical drama is harder than generic AI video

Vertical short dramas are not just short videos shot in 9:16. They carry a specific set of production pressures that generic AI video tools ignore:

  • Episodes run roughly 1–2 minutes, sometimes extending to 3–4, and each must open with a hook and close on a cliffhanger.
  • Characters recur across tens or hundreds of episodes, so face, costume, and personality drift is immediately visible.
  • Props and injuries carry plot weight; a jade pendant, a scar, a signed contract must look like the same object across scenes.
  • Releases are serialized, so writing must keep continuity across batches rather than treating each episode as a standalone generation.
  • The default deliverable is native 9:16 vertical, not a landscape video cropped after the fact.

These constraints are why "generate a video" is the wrong unit of work. The right unit is a controlled shot block, roughly 10 seconds long, bound to approved references and an engineered shot prompt, with multiple takes to choose from.

The twelve-stage pipeline

Think of this as the call sheet for an AI crew. Each stage has a job, a deliverable, and a review gate.

#StageWhat gets producedWho decides it is ready
1Idea intake & scoringA completeness score on the core conceptProducer / showrunner
2Creative guidanceMissing dimensions filled inCreator + system prompts
3Brief lockLogline, conflict, arc, hooks, audienceCreator sign-off
4Cast & visual confirmationNamed cast with visual fieldsCreator sign-off
5Story archive buildContinuity bible: characters, ties, threads, propsSystem-maintained, creator-reviewed
6Batch script planningBeat sheet and cliffhanger per episodeWriter + producer
7Script draft + rule checkEpisode scripts passing format/rhythm gatesWriter sign-off after QC
8Look selection & look-devStyle manual, character art, scene art, prop artArt director / creator
9Shot blockingEpisodes cut into ~10s shot blocksDirector / editor mindset
10Reference mapping & shot promptsPer-shot prompt with scene → prop → character refsDirector reviews the prompt
11Shoot & queue managementRendered takes per shot blockSystem runs; creator selects takes
12Review, reshoot, handoffChosen takes, edits, notes for next batchCreator / editor final pass

The rest of this section explains what actually happens at each handoff and where teams usually cut corners.

1. Idea intake and completeness scoring

Before any writing starts, the concept needs to be graded on how complete it actually is. A vague pitch like "a CEO revenge drama" is not a brief. A usable intake covers genre, protagonist, core conflict, direction of the story, episode count, episode length, tone, hook rhythm, ending direction, and target platform / audience.

A useful pattern is a 0–100 intake score. Concepts above a high threshold can move almost directly to a locked brief; mid-score concepts get prompted for only the missing dimensions; low-score concepts go through a full guided process. The point is not bureaucracy — it is preventing the model from inventing foundational choices you never agreed to.

2. Creative guidance, not creative replacement

Guidance is the stage where missing dimensions get pulled out of the creator, not made up by the model. Good systems ask targeted questions: who the protagonist is before and after their arc, what the central contradiction is, how often a big reversal should land, what the ending shape looks like. This is the equivalent of a development meeting, not a writing pass.

The common failure here is letting the model "helpfully" fill in premise-level gaps. Those fills become invisible assumptions later, and by episode 20 the show has drifted into a different genre.

3. Brief lock: the greenlight gate

The brief is the document everyone — writers, art, prompts, editing — is legally bound to inside the production. It should contain at minimum:

  • Logline (one-sentence story)
  • Core conflict
  • Story direction
  • Ending direction
  • Beat of satisfaction and cliffhanger rhythm
  • Target platform and audience
  • Episode-by-episode synopsis
  • Notes to the writer
  • Protagonist field that explicitly includes character arc

Nothing downstream should start until the brief is locked. "We'll fix it in the script" is how you get thirty episodes with no spine.

4. Cast and visual confirmation

Before art is generated, the cast list must be complete: names are valid, visual fields are filled, and leads match the brief. This sounds trivial, but AI productions regularly fail here because a side character is never visually defined, then shows up in episode 12 with a random face that becomes canonical by accident.

This gate should enforce hard checks: cast is complete, names are clean, visual descriptions exist, leads align with the brief. If any of those fail, the pipeline does not move forward.

5. Story archive: continuity as data, not memory

A story archive is a structured, show-long record of character identities, stable traits, current states (injuries, revealed identities, relationship changes), relationships, open and resolved plot threads, episode appearance tables, batch-level plot summaries, and prop visual descriptions. It is the digital version of a writers' room continuity bible.

The archive solves a real problem: long-form context does not fit reliably in model memory. Instead of asking the model to "remember" everything, the system feeds only the relevant slice for each batch — current character states, unresolved threads, recent batch summaries, and the current batch's main line — then writes the archive back after each batch is approved.

The rule is simple: continuity comes from structure, not from hoping the model recalls episode 3.

6. Batch planning: beat sheet before dialogue

Writing should never begin with full episode prose. For each batch of episodes, the first artifact is a beat sheet: scene-by-scene beats plus the end-of-episode cliffhanger. Only after that is approved do you write the actual script pages.

Vertical drama has hard rhythm rules that belong at this planning stage, not as a polish pass:

  • Golden 3-second rule: the very first scene must open with conflict or suspense, never a slow setup or exposition dump.
  • Episode shape: opening hook (1 scene) → rising conflict → end-of-episode cliffhanger (final scene).
  • Satisfaction density: at least one small payoff per episode (a reversal, a face-slap, an identity hint, evidence secured); a larger payoff every few episodes.
  • Dialogue discipline: short lines, generally under twenty characters of spoken Chinese in the original format (kept tight in translation), no essay-like monologues, no narrator explaining the plot.

If the beat sheet does not have these, writing the script is a waste of a batch.

7. Script drafting and rule-based QC

Once the beats are locked, the draft is written episode by episode, with each episode continuing directly from the full text of the previous one. From the second batch onward, four things should be confirmed up front and treated as hard constraints: the batch's plot direction, focus characters, threads and conflicts to advance, and end-of-batch hooks. If those conflict with the archive, the freshly confirmed intent wins — that is how you keep a show responsive to audience feedback without breaking continuity.

After drafting, scripts must pass a rule-based quality check, not a vague "is this good" pass. Typical gates catch:

  • Wrong episode titles
  • Scene count mismatches
  • Missing character lines
  • Too little dialogue
  • Placeholder text like "to be continued"

Failures get sent back for an automatic rewrite with the specific error attached, capped at a small number of retries. Nothing that fails these checks should reach the director or the video stage.

8. Look selection and look-dev

Consistency starts before the first frame is rendered. A usable system ships with multiple style manuals — covering 2D, 3D, and live-action-adjacent looks such as urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink style, xianxia, 3D donghua, stop-motion clay, cyber-Chinese, and more — and each manual should include at least character art guidance (face anchors, materials, vibe, view consistency), scene and prop guidance, and video style tags.

Crucially, once a look is chosen, character art, scene art, prop art, and video prompts all share the same style path. The most common visual failure in AI dramas is "character is anime, video turned photoreal." That happens because art and video are treated as separate generators with separate style settings.

Character art is best done in two steps: first, a text pass that polishes archive descriptions into art prompts using the chosen style manual and hard gender rules (this is where a creator can override); second, the actual image generation, with reference image support, stored as reviewable version history.

Discipline notes that save hours later:

  • Scene art must be empty plates — no people in background plates.
  • Props are pulled per episode from the script's own wording, in the voice of a props master, then merged into a show-wide catalog so nothing is lost between episodes.
  • User or script descriptions override the style manual's subject-matter restrictions; the manual governs how things are drawn, not what is allowed to exist.
  • Gender is a hard rule at art generation, not a suggestion.

9. Shot blocking: episodes become ~10-second units

A finished episode is not one generation. It is a sequence of shot blocks, each targeting roughly ten seconds of screen time, with a soft cap on script length per block. Longer scenes get split further by action beats, paragraph breaks, and sentence endings. Crowd or generic characters are separated from the drawable main cast so they do not consume character reference slots.

This is the same mental model as a clip list on an editing timeline. Each block has its own prompt, its own set of references, its own rendered takes, and its own history. You can re-block an entire episode if the structure is wrong, but you should not be trying to steer a two-minute monolith.

10. Reference mapping and engineered shot prompts

This is where most "AI face swap" problems are actually solved. For each shot block, the system builds a reference table in a fixed order: scene art first, then prop art for that shot, then character look-dev art. Slots are only filled when an image exists; if something is missing, it is marked as text-only rather than silently invented.

The highest-priority rule for character consistency is worth stating plainly:

When a character has a reference image, the prompt must not re-describe that character's clothing or appearance in text. The reference image is the source of truth for how they look; text describes only action, expression, and injury state.

That single rule eliminates the majority of costume swaps and face drift, because those failures happen when text and reference fight each other and the model tries to satisfy both.

Other mapping controls that belong in a real pipeline:

  • Manual character binding, so a script reference like "the officer" is tied to the correct named character's look.
  • Manual prop include / exclude, so irrelevant props do not eat reference slots.
  • Version swapping, so you can pick a different historical image for the same character or prop.
  • Era-based routing for time-travel or flashback stories, so a character's period look is picked for period scenes and modern look for modern scenes.

The shot prompt itself should follow an engineered specification, not freeform prose. A strong prompt covers eight elements: precise subject, action detail, scene environment, lighting and color, camera movement, visual style, image quality, and constraint package. Simple scenes can be one paragraph; complex cinematic scenes use a three-part structure (overall setup → shot 1 / shot 2 / shot 3 → constraint package).

A few prompt-writing rules that materially reduce broken shots:

  • One camera movement per shot; do not stack push, pull, pan, and tilt into a single shot.
  • Refer to shots by number, not by absolute timestamps like "0–3s."
  • Always include a base constraint package: quality, facial stability, no watermark or logo. Add twin / duplicate avoidance for multi-character scenes, and style anchoring for non-photoreal looks.
  • Favor continuous, low-speed motion with quantified limb detail over high-energy explosive actions that tend to break.
  • Use clear notation for dialogue, sound effects, and score.
  • Feed only this shot's materials into the prompt; never leak another scene's references.
  • Bias parameters toward stability and control, not wild creativity.

Before prompts are handed off to rendering, a compliance pass strips specific copyrighted work or IP names while keeping technique and aesthetic descriptors, reducing downstream blocking risk. If a reference image is detected as a suspected real-person photo, the system should refuse to silently use it and instead route the creator back to the platform's own art pipeline.

11. Shoot and queue management: production, not a demo

The shooting stage is where toy tools and production tools diverge. A real pipeline treats rendering as a queued, observable production process:

  • Video tasks run on a separate queue from writing and art, with isolated concurrency slots.
  • The same shot block cannot be double-submitted while a render is in flight, preventing duplicate charges and state confusion.
  • Credits are pre-deducted and settled; failed jobs release them automatically.
  • Stuck jobs time out and are recoverable; failures are visible and retryable, not silent.
  • When an upstream vendor task ID already exists, the system polls and resumes instead of creating a second job and charging twice.
  • Output files get faststart processing and store measured runtime, not just the vendor's reported duration.

Default output specs for vertical drama:

SpecDefault
Aspect ratio9:16 vertical (other ratios supported)
DurationSmart duration, or fixed 5–15s per shot
Model tierStandard / fast / lightweight options
ResolutionPer-model allowlist, ranging from 480p to 4K
AudioGenerated by default; watermark off by default

The vertical format is the native deliverable, not a crop applied after the fact.

12. Review, reshoot, and handoff

There is no magic "auto-pick the best take" engine. What a production system can do is make reshoots cheap and legible:

  • Edit the shot prompt (equivalent to a director revising shot notes).
  • Swap character, scene, or prop references (equivalent to changing a look or a plate).
  • Rebind characters manually when the script uses an alias.
  • Toggle props in or out to control visual focus.
  • Switch the look to unify visual language.
  • Choose model tier to trade quality, cost, and speed.
  • Browse historical takes for the same shot and pick the strongest.

A new render only happens when you actually change something and submit — same prompt and same references produce a recorded take, not a reinvention. That mirrors how a real set works: you change the shot note, then you roll again.

Where the pipeline intentionally stops

Honest boundaries are part of the craft. A production system should not pretend to do things it does not do:

  • There is no automatic quality scoring or auto-reshoot engine that picks the "best" take for you; final judgment sits with the creator.
  • Reference images are not a hard block — you can skip them and render text-only, but quality is usually worse, so professional process is to lock looks before shooting.
  • Reference markers in prompts rely on the prompt spec; creators should still read the prompt and confirm references are present.
  • Shot grammar follows the platform's built-in camera specification, while the full art manual applies on the look-dev side.
  • The ~10-second shot block is an engineering heuristic, not a timecode-precise edit; longer blocks may still occur, and final trimming and assembly belong in editing.
  • There is no productized cross-shot automatic extension or video continuation; the unit of work is "one shot, with references, producing one clip."
  • Character consistency depends on the look-dev asset chain, not on facial embedding verification; era routing is rule-based, and final look still depends on art quality and on obeying the "do not re-describe referenced characters" rule.

These are not flaws to hide. They are the shape of a tool that is trying to be a production operating system, not a magic button.

What this looks like in practice

A healthy AI short drama crew does not sit around waiting for one big generation. It runs a rhythm:

  1. Lock the brief and cast before writing a page.
  2. Build and maintain the story archive so batches do not drift.
  3. Plan beats and cliffhangers before drafting dialogue.
  4. Lock the look and generate character, scene, and prop art as shared assets.
  5. Block episodes into ~10-second shot units.
  6. Map references per shot and write an engineered prompt per block.
  7. Render to a managed queue, review takes, and reshoot only what needs changing.
  8. Carry the archive and notes into the next batch.

High-volume short drama is not one long generation cut up into episodes. It is many controlled units stacked into a show. Consistency comes from assets, not luck.

About Maosika

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. Rather than offering a single "text to video" button, it structures production from idea to finished episodes in a staged chain: intake and brief lock, story archive, batch script writing with rule-based QC, look selection and look-dev, ~10-second shot blocks, reference-mapped shot prompts, queued rendering, and take-by-take reshoot control. Eighteen digital specialists mirror real crew roles — archivist, producer, writer, script supervisor, casting director, art director, DP, director, editor, VFX supervisor, and more — and expose their work as streaming production logs instead of a black box. The goal is not to remove human judgment; it is to turn the repetitive, drift-prone, failure-prone parts of the process into a controllable pipeline, while审美 decisions stay with the creator.

If you are building a vertical AI drama slate and want a process designed for serialized, repeatable output rather than one-off clips, you can explore the workflow at https://www.maosika.com.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com