How Vertical AI Short Dramas Actually Get Made: A 12-Stage Production Pipeline

Maosika Editorial | Last updated

Most AI short drama failures are not model failures—they are pipeline failures. A usable vertical drama is built stage by stage, with locked intermediate assets, not generated in one long prompt.

The real problem is not generation; it is sequence

Teams new to AI short drama production usually start the same way: write a script, throw it into a video tool, and hope a finished episode comes out the other side. What comes out instead is a familiar list of problems: the lead changes face between scenes, costumes reset every shot, props appear and disappear, episode 7 forgets a secret established in episode 2, and the cliffhanger lands flat because the pacing was never planned.

These are not separate bugs. They are symptoms of skipping production order. In a real crew, you do not shoot before the brief is locked, you do not lock cast before look development, and you do not ask the camera department to invent continuity on set. AI workflows need the same discipline—only the crew is digital, and the handoffs happen through structured assets rather than paper call sheets.

A production pipeline for vertical AI short drama is a staged sequence in which each stage produces a reviewable artifact before the next stage is allowed to begin. The artifact is the handoff.

This article walks through that pipeline in 12 stages, from intake to delivered episode, with the gates that actually prevent drift.

Why vertical short drama needs a stricter pipeline than horizontal video

Vertical micro-dramas—typically 1–2 minutes per episode, sometimes extending to 3–4—have structural properties that punish sloppy workflows harder than long-form:

  • Hook timing is brutal. The opening beat must land within the first few seconds; a slow establishing scene is a drop.
  • Cliffhangers are mandatory, not decorative. Each episode ends on an open question; if continuity broke earlier, the cliffhanger does not land.
  • Volume is high. Teams do not make one film; they make tens or hundreds of episodes, so any inconsistency multiplies.
  • The default deliverable is 9:16 from the start, not a landscape master cropped after the fact.
  • Face, costume, and prop consistency are visible on a phone screen held at arm's length. Small drift reads as obvious error.

For these reasons, a working pipeline is not a nice-to-have. It is the thing that decides whether the output looks like a show or a sequence of unrelated AI clips.

The 12 stages, in order

The stages below are presented in the order they must happen. Moving a later stage earlier is the most common cause of rework.

#StageWhat gets producedWho approves
1Idea intake & scoringA completeness score on the core conceptProducer / creator
2Guided creative developmentMissing dimensions filled inCreator
3Brief lockThe approved creative briefProducer
4Cast & visual confirmationNamed cast with visual fields lockedCreator / director
5Story archive buildThe continuity bible for the seriesWriter / producer
6Batch script writingBeat sheets, then script pages, per batchWriter / producer
7Rule-based script QCPass/fail on structural checksSystem + editor
8Style selectionOne style path shared across art and videoDirector / art lead
9Look development: characters, scenes, propsFinal approved reference artArt lead / director
10Scene blocking into ~10-second chunksPer-scene shooting units with bound referencesDirector / editor
11Engineered shot prompts & shootPrompt per shot, then rendered takesDirector
12Review, re-shoot, selectChosen takes, edits, re-records as neededCreator / producer

1. Idea intake & scoring

The first gate is not writing. It is deciding whether the idea is developed enough to write.

A useful intake captures the vertical-short-drama dimensions up front: genre, protagonist, core conflict, story direction, episode count, episode length, tone, beat and hook rhythm, ending direction, and target platform / audience. Ideas are scored on completeness. A high-completeness score can skip most of the guided development; a mid-score fills only the missing dimensions; a low score goes through full development.

The point is to prevent the writer (human or digital) from inventing fundamentals on the fly in episode 6.

2. Guided creative development

This is not a chat where the model "helps you brainstorm." It is a structured interview that fills the gaps the intake flagged. By the end of this stage, the creator has answered the questions that would otherwise be answered inconsistently later: who the lead is, what they want, what is stopping them, what the audience is supposed to feel every few episodes, and how the ending is shaped.

3. Brief lock

The brief is the first hard gate. Nothing proceeds until it is locked.

A locked brief includes the logline, core conflict, story direction, ending direction, satisfaction beats and hook rhythm, platform and audience, episode-by-episode outline, and notes for the writer. The protagonist entry must include their character arc—where they start, what breaks them, and who they become.

Think of this as the greenlight meeting. Once the room has signed off, the crew does not get to re-decide the premise on set.

4. Cast & visual confirmation

Before any art is generated, the cast list must be complete: names are valid, visual fields are filled, and the leads match the brief. This is a hard check, not a suggestion. A half-defined cast is the single largest source of face drift later, because every downstream stage starts improvising.

5. Story archive build

The story archive is a structured continuity record that travels with the entire series: character identities, stable traits, current state (injuries, revealed identities, changed alliances), relationships, open and resolved plot threads, episode-by-episode appearance tables, batch-level plot summaries, and visual descriptions of recurring props.

This is the digital equivalent of a writers' room continuity bible. Its job is to make episode 40 remember what episode 4 established. Writing does not rely on model memory; it relies on the archive.

6. Batch script writing

Scripts are written in batches of several episodes, not one enormous generation. The order inside each batch matters:

  1. A beat sheet is planned first, including the episode-end cliffhanger.
  2. Only after the beats are approved does scene-by-scene dialogue get written.
  3. When writing the next batch, the system reads the previous batch's full text, not a loose summary.
  4. From the second batch onward, the batch intent is locked first: plot direction for this batch, focus characters, threads and conflicts to advance, and the end-of-batch hook. These become hard constraints.
  5. Past the halfway point of the series, ending constraints are injected so the story does not wander.

Vertical short drama writing has hard rhythm rules that belong in the pipeline, not in the writer's mood:

  • Golden opening: the first scene must open on strong conflict or strong suspense; no slow background setup.
  • Per-episode structure: opening hook → escalating conflict → end-of-episode cliffhanger.
  • Beat density: at least one small satisfaction beat per episode (a reveal, a reversal, a piece of evidence landed, a status hint); a larger payoff every few episodes.
  • Dialogue: short lines, generally under twenty characters in the original Chinese writing convention; no lecture-style monologue and no long explanatory voiceover.

7. Rule-based script QC

After writing, scripts pass through a rule checker that catches structural failures before they reach the shoot stage: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, placeholder text like "to be continued," and similar mechanical failures. Failing scripts are sent back for rewrite with the specific error attached, up to a retry cap. Nothing half-finished is handed to production.

This is a deliberately boring step, and it is one of the highest-ROI gates in the whole pipeline. Most "AI wrote a bad script" complaints are really "no one checked the script before shooting it."

8. Style selection

Before any character or scene art is made, one style path is chosen for the whole show. A production-ready system ships with multiple style books covering 2D, 3D, and realistic directions—urban realism, period realism, mature urban romance animation, 1990s anime, Chinese ink style, xianxia, 3D donghua, stop-motion clay, cyber-Chinese, and others. Each style book includes a character sheet guide (face anchors, materials, temperament, multi-view consistency), a scene and prop guide, and video style tags.

The key rule: character art, scene art, prop art, and video prompts all consume the same style path. This prevents the common fracture where characters look like anime but the video output flips to live-action realism.

9. Look development: characters, scenes, props

This is the stage most teams skip, and the stage that does the most work.

Character art is produced in two steps. First, the archive description is polished into an art prompt using the chosen style book, with gender as a hard rule; this polished prompt can be manually overridden. Second, the final prompt is rendered, optionally with a reference image, and saved as a versioned asset. Only completed, readable art assets are consumed downstream.

Scenes are parsed from the script into structured headings—interior/exterior, location, day/night—rather than retyped into a separate form. Scene art has a discipline of its own: no people in establishing shots. Props are pulled episode by episode from the script's own wording, as a props master would, then merged into a show-wide catalog so nothing is lost between episodes.

A critical authority rule: the script and the creator's descriptions override the style book's world rules. The style book controls how things are drawn; it does not get to refuse story content.

10. Scene blocking into ~10-second chunks

A finished episode script is not shot as one video. It is cut into scene blocks targeting roughly 10 seconds each, with a soft cap on text length; longer scenes are split further on action beats, paragraph breaks, or sentence endings. Background or generic characters are separated from the drawable main cast so they do not consume character slots.

The creator sees not "Episode 3, one big video," but "Episode 3 → Scene 1, Scene 2, Scene 3…" Each scene has its own prompt, its own bound references, its own rendered output, and its own history of takes. This is the editing-room clip list, built before shooting rather than discovered after.

11. Engineered shot prompts & shoot

This is where most "prompt engineering" advice lives, but in a real pipeline the prompt is not a paragraph of vibes. It is a shot document.

A production-grade shot prompt is built from eight components:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color tone
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Constraints

Simple scenes use a single block; complex cinematic scenes use a three-part structure (overall setup → shot 1 / shot 2 / shot 3 → constraint pack). One shot gets one camera move—no stacking push, pull, pan, and tilt into a single instruction. Shots are numbered; they do not reference absolute timestamps. A mandatory constraint pack covers image quality, facial stability, and no watermarks or logos; multi-person scenes add anti-doppelgänger constraints; non-realistic styles anchor the look explicitly.

Action is written as fine-grained body movement with quantified intensity, favoring slow continuous motion over high-dynamic action that tends to break. Dialogue, sound effects, and music use explicit notation so they are not confused with visual description.

Before prompts are generated, the system checks that required scene, character, and prop references exist and surfaces a missing list. Teams can choose to proceed with text only, but they should know what they are trading away.

The single highest-leverage consistency rule at this stage is simple and worth stating on its own:

For any character with a reference image, the prompt must not re-describe clothing or appearance in text. The reference image is the authority; text describes only action, expression, and injury state.

This rule directly eliminates the classic failure where text and reference fight each other and the character changes outfit or face mid-scene.

Additional mechanics support this: manual character binding handles cases where the script uses a title ("the officer") and the archive uses a name; props can be manually included or excluded to control visual focus; historical versions of an asset can be swapped in; and time-crossing stories route to period-appropriate character art based on scene keywords.

Before delivery, prompts are cleaned of specific copyrighted work or IP names, keeping the technique and aesthetic descriptors to reduce downstream blocking risk.

12. Review, re-shoot, select

Shooting is not a one-click event. The standard path is: open the episode's video workspace, confirm references (or deliberately skip them), generate prompts, read and edit them, choose model tier / aspect ratio / resolution / duration, submit to the render queue, and then review historical takes to pick the best one.

The controllable points are the same ones a real production has:

Control pointProduction equivalent
Edit the shot promptDirector revising the shot note
Swap character / scene / prop referenceChanging a look or a location plate
Manual character bindingFixing name / title mismatches and cameos
Include / exclude propsControlling what is in frame
Switch styleUnifying the visual language
Choose model tierTrading quality, cost, and speed
Review historical takes per sceneMulti-take selection

When you change the prompt and re-submit, that is a new take—exactly like revising a shot note and rolling camera again. The system does not silently rewrite your approved prompt under you.

On the reliability side, production queues need to behave like production queues: video tasks run in their own lane separate from writing and art, in-flight scenes cannot be double-submitted, credits are pre-deducted and released on failure, stuck tasks time out and become retryable, and previously issued vendor tasks resume polling rather than being double-created. Finished videos get faststart processing and are stored with their actual measured duration rather than trusting a vendor-reported number.

Where the gates actually save you

Stepping back, the pipeline exists to enforce five things that almost never happen spontaneously in AI generation:

  1. The brief is locked before writing begins. No premise drift.
  2. The archive is the source of truth for continuity. No "the model forgot."
  3. Look development happens before shooting. No face or costume improvisation per scene.
  4. Scenes are shot as discrete blocks with bound references. No one-hour generation to carve up afterward.
  5. Takes are selected by a human after review. No pretending the model knows which take is "best."

When teams report that AI video "doesn't work for series," they almost always mean they skipped one or more of these five.

What this pipeline does not do

It is important to be explicit about the boundaries, because a production tool that lies about its limits is worse than one that states them:

  • There is no automatic scoring of finished footage and no auto-select-best-take engine. Final quality judgment sits with the creator and producer; the system provides multiple takes and the ability to revise prompts and re-shoot.
  • Reference images are not a hard gate. Missing references trigger a warning, but teams can proceed with text only—quality is usually worse, which is why a professional workflow locks looks before shooting.
  • Reference markers inside prompts are guided by convention, not forcibly re-injected; creators should still read the prompt and confirm markers are present.
  • Video shot grammar uses the built-in camera conventions; the full art manual applies to the still-image side, while video receives the style tags.
  • The ~10-second scene block is an engineering heuristic, not a timecode-precise cut. Long scenes are split further, but some blocks may still run long; final trimming and stitching happen in a later edit.
  • There is no cross-scene automatic continuation or video extension workflow as a product unit. The unit of work is "one scene, with multi-modal references, producing one clip."
  • Character consistency relies on the look-development asset chain, not on face-embedding verification. Period routing is rule-based; final look still depends on the quality of the approved art and on whether the prompt obeys the "do not re-describe appearances" rule.

These are not flaws to hide. They are the boundaries that tell a professional team where their own judgment is required.

The 9:16 question, answered once

Vertical short drama is not landscape video cropped for phones at the end. A real pipeline treats 9:16 as the default deliverable from the start: prompts are written for vertical framing, scenes are blocked for vertical composition, and the render output defaults to vertical. Teams that generate horizontal and crop after the fact consistently lose framing, facial performance, and hook timing.

A short checklist for teams evaluating a pipeline

Before you commit to a tool or a workflow for a multi-episode series, ask:

  • Is there a locked brief stage before writing, or does the system start generating immediately?
  • Is continuity stored in a structured archive, or left to model context?
  • Are scripts written from beat sheets, with cliffhangers planned before dialogue?
  • Is there a rule-based QC pass before scripts reach the shoot stage?
  • Is one style path shared across characters, scenes, props, and video?
  • Are character, scene, and prop art produced as versioned assets before shooting?
  • Are episodes cut into scene blocks before rendering?
  • Do shot prompts follow a fixed structure with explicit camera and constraint sections?
  • Can you edit prompts, swap references, and review multiple takes per scene?
  • Are video tasks handled in a production queue with failure handling and retry?

If the answer to several of these is "no," what you have is a generation tool, not a production pipeline.

How Maosika fits into this

Maosika (猫斯卡) is built around exactly this staged shape. It is an AI production operating system for vertical short dramas, not a single text-to-video button. The 12 stages above describe the order the product enforces: idea to brief, brief to archive, archive to batch-written and QC-checked scripts, scripts to shared-style look development, looks to scene blocks, blocks to engineered prompts, prompts to queued multi-take shooting, and takes to human selection.

A digital crew of 18 specialists—five covering creative and script work, thirteen covering video production—handles the handoffs between stages, with visible streaming work logs so the producer can see what each role did, rather than staring at a black box. The philosophy is not "no humans needed." It is: turn the repetitive, drift-prone, failure-prone parts of the process into a constrained pipeline, and leave taste, approval, and final judgment with the people making the show.

To explore the workflow in practice, you can start a project at maosika.com.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com