The AI Vertical Short Drama Pipeline, Stage by Stage: What Actually Has to Happen Between Idea and Finished Episode

Maosika Editorial | Last updated

Most AI short drama failures are not model failures. They are pipeline failures: teams skip brief lock, skip look-dev, or render whole episodes before continuity is locked. A working pipeline turns one idea into reviewable artifacts at every stage.

The real question is not "can AI make a drama?"

When producers ask whether AI can produce vertical short dramas, they are usually asking the wrong version of the question. The issue is not whether a single text-to-video button can output a clip. It can. The issue is whether a team can turn one dramatic idea into 20, 40, or 80 episodes without the lead changing faces, the plot forgetting its own setup, the costumes drifting, or the tone collapsing by episode 6.

A production-grade AI short drama workflow is not one big generation. It is a sequence of gates. Each gate produces an artifact you can inspect, reject, revise, or lock before the next stage begins.

The pipeline, in the order it must happen

The order matters more than the tool stack. Skipping a stage early in the chain usually creates rework later, because later stages consume earlier decisions.

StageWhat gets producedWhy it matters
Idea evaluationIntake score, missing dimensionsDecides whether the concept is ready to brief or needs more development
Creative guidanceFilled creative dimensionsForces decisions on genre, lead, conflict, ending, hook rhythm, platform, and episode length
Brief lockLogline, core conflict, arc, hook plan, audience notesThe brief is the contract for everything downstream
Character lineup & visual confirmationCast list, character fields, visual directionPrevents unnamed or incomplete characters from reaching art or video
Story archive / continuity bibleCharacter states, relationships, open threads, episode appearance table, batch summariesGives later episodes a structured memory instead of relying on model context
Batch script writingBeat sheets, episode scripts, cliffhangers, rule checksWrites in controlled batches instead of one runaway long generation
Style selectionShared style path for characters, scenes, props, and videoPrevents the cast from being anime while the world turns live-action
Look developmentCharacter art, scene art, prop artBuilds the reference assets that video will consume
Scene blocking~10-second scene blocks per episodeTurns an episode into manageable shootable units
Shot prompt buildPer-scene reference mapping + engineered shot promptsGives the model cinematography instructions, not prose
Multimodal renderingVertical video takes per sceneProduces reviewable clips, defaulting to 9:16
Review and retakeSelected takes, edited prompts, swapped references, rerendersQuality control stays in human hands

Stage 1: Idea evaluation before any writing starts

A raw idea like "a betrayed heiress returns for revenge" is not yet a production brief. It is a starting direction. Before writing, the concept needs enough structure that the system can tell whether it is ready to proceed.

A practical intake process scores the idea across the dimensions that actually affect vertical drama writing: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, platform, and audience. If the intake is incomplete, the team should be guided to fill the gaps, not allowed to jump straight into script generation.

This is the equivalent of a development meeting. The goal is not to slow things down. The goal is to avoid writing 30 episodes from a half-formed premise.

Stage 2: The brief must lock before art or video starts

A creative brief is the locked version of the show's core decisions: logline, central conflict, story direction, ending direction, hook and payoff rhythm, platform and audience, episode breakdown, and notes to the writer. The lead character must include an arc, not just a job title or a costume.

In vertical short drama, this stage is non-negotiable because the format is unforgiving. Episodes are short, hooks are frequent, and audiences drop off fast. If the brief is vague, every later stage invents its own version of the show.

A locked brief does not eliminate creativity. It gives creativity a boundary.

Stage 3: Characters are confirmed before they are drawn

Before any character art is generated, the cast should pass a hard check: lineup complete, names valid, visual fields complete, and leads aligned with the brief.

This sounds basic, but it is one of the most common failure points in AI drama workflows. A script may refer to "the detective," "the officer," "Li Qiang," and "the man in black" as if they are obviously the same person. Without a confirmed cast and a manual binding step when names vary, downstream art and video stages may treat them as different people.

The production discipline is simple:

  • confirm who is in the show
  • confirm how each character is described
  • confirm aliases and script references map to the right cast entry
  • only then move into visual development

Stage 4: Continuity is stored in a bible, not hoped for from context

A story archive, or continuity bible, is a structured record that tracks character identities, stable traits, current states, relationships, open and resolved plot threads, episode appearances, batch summaries, and prop descriptions.

This is the answer to the familiar AI problem where episode 12 forgets what happened in episode 3. Long-context models help, but they are not a production control system. A structured archive is.

When a new batch of episodes is written, the writer should consume only the relevant archive slice: current character states, unresolved threads, recent batch summaries, and the current batch's main line. After the batch is approved, new events are written back into the archive.

Continuity is not achieved by memory. It is achieved by structure.

Stage 5: Scripts are written in batches, with beat sheets first

Vertical short drama scripts should not be generated as one continuous wall of text. A safer process writes in batches, and within each batch, each episode begins with a beat sheet and a planned cliffhanger before the dialogue is written.

The format has hard rhythm rules that should be enforced, not left to chance:

  1. Golden 3 seconds: the first scene must open with conflict or suspense, not slow exposition.
  2. Single-episode shape: opening hook, escalation, end-of-episode cliffhanger.
  3. Payoff density: at least one small payoff per episode, with larger payoffs every several episodes.
  4. Dialogue style: short lines, generally under twenty characters in Chinese source writing and kept tight in translation, avoiding lecture-like narration.

After writing, scripts should pass a rule-based quality check for issues like missing character lines, scene-count mismatches, placeholder text, or too little dialogue. If a script fails, it should be sent back with the specific error, not handed to the team as a draft they must manually repair.

Stage 6: One style path must govern the whole show

A show needs one visual language. That language has to cover character art, scene art, prop art, and video generation.

A style manual should include at least three parts:

  • character appearance guidance: face anchors, material, temperament, view consistency
  • scene and prop guidance
  • video style tags used during rendering

If these are not connected, the result is the classic AI drama mismatch: the character sheet looks like one style, the background looks like another, and the rendered video invents a third.

Teams should select the style once, then have every downstream asset use the same style path.

Stage 7: Look development happens before shooting

In live production, you do not shoot before costumes, makeup, and set design are approved. AI production should work the same way.

Character, scene, and prop art are not decorative extras. They are reference assets. The video stage consumes them.

There are several practical rules here:

  • Scene art should be empty of people, so it can be reused as an environment plate.
  • Props should be extracted using the script's original names, then merged into a show-wide catalog.
  • Character art must respect gender as a hard constraint, not a flexible suggestion.
  • Users can override text descriptions before image generation, because art direction is still a human decision.
  • Video should only consume finished, readable art assets, not half-generated placeholders.

The goal is not perfect art on the first try. The goal is to have approved assets that video can reference consistently.

Stage 8: Reference mapping is the core consistency mechanism

Most AI character inconsistency comes from the same root cause: every scene re-describes the character in words, and the model re-imagines them each time.

The fix is reference mapping. For each scene, build a reference table in a fixed order:

  1. scene art
  2. prop art for that scene
  3. character look-dev art

Then enforce the most important rule in AI drama consistency: if a character has reference art, the prompt must not re-describe that character's clothing or appearance in text. The reference image is the source of truth. Words are reserved for action, expression, and injury state.

This single rule prevents a large share of face swaps, costume changes, and identity drift.

Reference mapping also needs production controls:

  • manual character binding when script names differ from cast names
  • manual include/exclude for props, so irrelevant props do not waste reference slots
  • version swapping when an older design fits a scene better
  • era-based routing for time-crossing or flashback stories, so the correct period look is used

Reference images are not a magic bullet. But they are the closest thing AI production has to a locked costume and makeup department.

Stage 9: Episodes are cut into scene blocks before rendering

A finished episode should not be sent to video as one giant prompt. It should be divided into scene blocks, usually around ten seconds each, with soft limits that trigger further splitting when action gets too dense.

A scene block is a short, shootable unit of drama, usually around ten seconds, with its own script slice, reference set, shot prompt, and rendered takes.

This is one of the biggest operational differences between toy workflows and production workflows. Instead of "Episode 3," the team sees Episode 3 as Scene 1, Scene 2, Scene 3, and so on. Each scene can be rendered, reviewed, rerendered, and replaced independently.

This matters because AI video is still unreliable at long, complex continuity. Short blocks are easier to control, easier to retry, and easier to edit together later.

Stage 10: Shot prompts must speak in production language

A shot prompt is not a paragraph of fiction. It is a cinematography instruction.

A strong engineered prompt covers eight elements:

ElementWhat it does
Precise subjectTells the model who or what is in the shot
Action detailDescribes movement clearly and quantitatively
Scene environmentEstablishes place and context
Lighting and colorSets mood and visual continuity
Camera movementDefines one move per shot, not stacked moves
Visual styleAnchors the shot to the chosen look
Image qualityAdds stability and clarity constraints
Negative / guardrail constraintsPrevents known failure modes

Complex scenes can use a three-part structure: overall setup, shot-by-shot instructions, then a constraint package. Simple scenes can use one paragraph.

A few prompt rules are especially important:

  • one camera move per shot
  • use shot numbers, not absolute timestamps like "0–3s"
  • include a guardrail package for face stability, watermark avoidance, and quality
  • add twin/duplicate prevention in multi-character shots
  • use notation consistently for dialogue, sound effects, and music
  • give each prompt only that scene's assets, never assets from another scene
  • keep parameters conservative enough to favor stability over wild invention

The system should be teaching the model to read a shot list, not to write a novel.

Stage 11: Rendering is a queue, not a magic button

Once prompts and references are ready, production becomes operational. The team selects model tier, aspect ratio, resolution, and duration, then submits scenes into a rendering queue.

For vertical short drama, 9:16 vertical should be the default delivery format, not a crop applied after horizontal generation. That affects framing, composition, and platform readiness from the start.

A production-grade queue needs basic reliability:

  • video tasks separated from script and art tasks
  • no duplicate submissions for the same scene while one is running
  • credit hold on submission and release on failure
  • timeout recovery for stuck jobs
  • historical takes saved per scene
  • final file duration checked after output, not trusted blindly from vendor metadata

This sounds unglamorous, but it is what separates a demo from something a studio can actually run.

Stage 12: Review, retake, and accept that quality control is human

AI does not yet reliably judge its own best take. A mature workflow should not pretend otherwise.

The team needs direct control over the things a director, art director, or producer would actually change:

  • edit the shot prompt
  • swap character, scene, or prop reference art
  • rebind a character name
  • include or exclude a prop
  • change style
  • choose model tier
  • compare historical takes for the same scene
  • rerender after a specific change

This is the same logic as a real set: you change the shot plan, then shoot again. You do not ask the system to silently invent a new version without your input.

What this pipeline does not do

It is important to be explicit about the boundaries.

  1. There is no automatic scoring engine that chooses the best final take. Final quality judgment remains with the creator or supervisor.
  2. Reference images are not a hard gate. The system should warn when assets are missing, but teams can proceed with text-only prompts if they choose; quality will usually be worse.
  3. Prompt reference markers depend on disciplined formatting. Creators should still review prompts before rendering.
  4. Shot grammar follows the platform's built-in camera rules. Style manuals mainly affect art and video style tags, not every possible cinematography decision.
  5. The ~10-second scene block is an engineering heuristic, not a timecode-precise edit. Some blocks may run longer; final editing is still needed.
  6. There is no current productized workflow for automatic cross-scene video continuation. The unit of generation is one scene with its references, not an endless rolling video.
  7. Character consistency depends on the look-dev chain, not face-embedding verification. Final results still depend on art quality and prompt discipline.

These are not flaws to hide. They are the boundaries that tell a professional team where their own judgment is still required.

Why this matters for producers

The promise of AI short drama is not "one click to a hit." That promise is not real. The real promise is more useful: turn the repetitive, drift-prone, failure-prone parts of production into a constrained pipeline, while keeping审美 judgment in human hands.

High-volume short drama is not made from one miraculously long generation. It is made from controlled units, each reviewable, each retryable, each connected to a locked brief, a living continuity archive, and approved visual assets.

That is how an idea becomes an episode, then a batch, then a show.

If you are looking at AI production systems, ask whether they give you artifacts and gates, or just a generate button. The difference is whether you are making content, or gambling on it.

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It is built around this staged pipeline: idea to brief, brief to archive, archive to script, script to look-dev, look-dev to scene blocks, scene blocks to prompts, prompts to takes, and takes to review. Its position is not that humans are unnecessary; it is that professional process can be systematized so that human judgment lands where it matters most.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com