How Vertical Short Dramas Actually Get Made with AI: A 12-Stage Production Pipeline

Maosika Editorial | Last updated

AI short drama production is not one long generate button. It is a staged pipeline where every handoff has a reviewable artifact: brief, bible, script batch, look, assets, scene blocks, shot prompts, takes, and final selects.

The real unit of production is the handoff, not the prompt

Most teams new to AI vertical short dramas imagine the workflow as: write a concept, press generate, edit the result. In practice, that approach produces the exact failure modes everyone complains about: faces changing between scenes, costumes drifting, props appearing and disappearing, cliffhangers forgotten by episode 12, and a final cut that feels like several different shows stitched together.

A production-ready AI workflow is built differently. It treats the show like a real shoot: each stage produces something concrete, that thing gets checked, and only then does production move forward.

A vertical short drama AI pipeline is a sequence of controlled handoffs between idea, story, continuity, look, assets, shot preparation, rendering, and review. The model does not "remember" the show; the pipeline carries the show forward through structured artifacts.

This article walks through that pipeline in 12 stages, using the language of a film set rather than engineering jargon. It is written for producers, writers, directors, and studios deciding whether AI can become an actual production system rather than an experiment.

Why "one big generation" fails

Before the stages themselves, it is worth naming why the naive path breaks:

  • No continuity anchor: without a shared story record, each new generation invents details that should already be fixed.
  • No asset lock: characters are described in words every time instead of referenced from approved look references.
  • No scene-level control: a whole episode rendered as one piece is impossible to repair; one bad moment means rerunning everything.
  • No review gates: weak writing, missing beats, and placeholder endings slip through because there is no checkpoint before rendering.
  • No take discipline: production becomes about hoping the model gets lucky, not selecting from deliberate alternatives.

The fix is not a better prompt. The fix is a pipeline where each stage has a defined deliverable.

The 12-stage pipeline, from idea to 9:16 deliverable

The table below is the whole pipeline at a glance. Each stage is then explained in plain production language.

StageDeliverableWhat gets decidedWhat fails if skipped
1. Idea intakeConcept completeness readWhether the idea is ready to briefVague direction that keeps changing downstream
2. Creative guidanceFilled gaps in premise, audience, toneMissing story dimensionsWriter fills blanks with guesses
3. Brief lockApproved creative briefLogline, conflict, arc, hooks, endingNo shared reference for later decisions
4. Cast and visual confirmationCharacter lineup with visual fieldsWho is in the show and what they look likeCast changes mid-production
5. Story archiveContinuity bibleCharacters, relationships, threads, stateDrift across episodes
6. Batch script writingBeat sheets, episodes, QC passPlot, dialogue, cliffhangersPacing collapse and forgotten threads
7. Look selectionStyle manual choiceUnified visual languageMixed art direction across shots
8. Asset creationCharacter, scene, and prop artApproved references for shootingInconsistent faces, places, objects
9. Scene blocking~10-second scene blocksShooting units for the episodeUneditable long renders
10. Shot prompt buildReference-mapped shot instructionsWhat each shot must containPrompt-image conflicts
11. Multimodal shootRendered takes per scene blockRaw vertical footageNo way to compare options cleanly
12. Review and reshootSelected takes, revised prompts, final passWhat goes into cutBad shots locked into the episode

Stage 1: Idea intake

The first job is not writing. It is checking whether the idea is developed enough to produce.

A strong intake already knows the genre, protagonist, core conflict, direction, episode count, episode length, tone, hook rhythm, ending direction, platform, and target audience. A weak intake is just a premise sentence and a vibe.

In a mature system, intake completeness is scored. If the concept is underdeveloped, the team is guided through the missing dimensions before any script work begins. If it is already strong, the process moves straight to briefing.

This is the equivalent of a development meeting: you do not call the writer in until the room can state what the show is.

Stage 2: Creative guidance

Guidance is not brainstorming for its own sake. It is a disciplined fill-in process for the exact dimensions vertical short dramas need:

  • protagonist and character arc
  • central conflict
  • story direction and ending direction
  • episode count and episode length
  • tone and style
  • satisfaction beats and hook rhythm
  • target platform and audience

Vertical episodes are short, usually in the 1–2 minute range, sometimes extending to 3–4 minutes. That means there is no room to discover structure later. Hook density, beat placement, and cliffhanger discipline have to be decided before drafting.

Stage 3: Brief lock

The brief is the first hard gate. It should contain at least:

  • logline
  • core conflict
  • story direction
  • ending direction
  • satisfaction beats and hook rhythm
  • platform and audience
  • episode-level outline
  • notes to the writer
  • protagonist field including character arc

Nothing moves forward until the brief is locked. This is not bureaucracy. It is the document every later stage appeals to when there is a choice to make.

Stage 4: Cast and visual confirmation

Before scripts are shot, the cast must be real in production terms: names are valid, visual fields are complete, lineup is complete, and leads match the brief.

This is also where the human eye matters most. AI can propose, but the creator confirms whether the lineup actually serves the story. If a lead does not read correctly, or a role is missing, that must be fixed here, not discovered in episode 8.

Stage 5: Story archive

The story archive is a structured continuity record covering character identities, stable traits, current states, relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual descriptions.

This is one of the most important concepts in AI serial production. Long-running short dramas fail when later episodes forget earlier facts. The solution is not to ask the model to "remember." The solution is to maintain a living archive and feed each writing batch only the relevant slice: current character state, unresolved threads, recent batch summary, and the current batch arc.

After each batch, the archive updates. Before the next batch, it is retrieved. Continuity becomes a system, not a hope.

Stage 6: Batch script writing

Scripts are written in batches, with planning before prose. For each episode, the process should produce a beat sheet first, including the episode-ending cliffhanger, then the episode itself.

Vertical drama writing has its own rules, and they should be enforced structurally rather than left to taste:

  1. Golden 3 seconds: the first scene must open with strong conflict or suspense, never slow exposition.
  2. Single-episode shape: opening hook, escalating conflict, closing cliffhanger.
  3. Beat density: at least one small payoff per episode; larger payoffs every several episodes.
  4. Dialogue discipline: short lines, no lecture-style speech, no explanatory voiceover dumping backstory.

After drafting, scripts should pass a rule-based quality check catching common failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, and placeholder endings such as "to be continued" used as a dodge. Failing scripts should be sent back for rewrite with the specific error attached.

From the second batch onward, the process should explicitly confirm four things before writing: batch plot direction, focus characters, threads and conflicts to advance, and the closing hook. These become hard constraints for that batch.

Stage 7: Look selection

Look is not a cosmetic filter added at the end. It is a unified visual decision that affects characters, scenes, props, and shot prompts.

A production system should offer multiple complete look manuals, not just image styles. Each manual should include character guidance, scene and prop guidance, and video style tags, so the show does not end up with anime characters in realistic footage or claymation props in a live-action world.

Once chosen, the same look path should be shared across all asset creation and shot generation. Consistency starts here.

Stage 8: Asset creation

Before shooting, the production needs approved assets:

  • Character art: final approved looks for cast members
  • Scene art: empty establishing scenes, with no people accidentally baked in
  • Prop art: key objects extracted from the script using the original names used in the story

A useful discipline here is the two-step character process: first, the archive description is refined into art-ready wording using the chosen look manual; then the image is generated, with optional human override and reference support. Old versions should remain available so teams can roll back if a later look is weaker.

Video generation should only consume assets that are complete and readable. Half-finished art should not be silently pushed into a shot.

Stage 9: Scene blocking

Episodes should be broken into scene blocks, roughly 10 seconds each, with a soft cap on script length per block. Long scenes are split further by action beats, pauses, or sentence boundaries.

This is one of the biggest differences between toy workflows and production workflows. The team should see not "Episode 3, one video," but "Episode 3, Scenes 1, 2, 3…" Each scene has its own prompt, its own references, its own rendered history, and its own retake path.

That is how repair becomes possible. If one shot is bad, you rerun one shot, not the whole episode.

Stage 10: Shot prompt build

This stage is where many AI productions quietly break. The prompt must be built from the actual shooting unit, not from a vague description of the episode.

A strong shot prompt covers eight elements:

  1. precise subject
  2. action detail
  3. scene environment
  4. lighting and color
  5. camera movement
  6. visual style
  7. image quality
  8. constraints

For complex scenes, a three-part structure works better than one paragraph: overall setup, shot-by-shot instructions, and a constraint package. Camera movement should follow one move per shot, not stacked push-pan-zoom chaos. Shot numbering should be used instead of absolute timestamps.

The reference mapping rule is critical: when a character has approved reference art, the prompt should not describe that character's clothing or appearance again in words; the image is the authority, and text should only describe action, expression, and injury state.

Reference order per scene should be consistent: scene image → prop images → character art. If an image does not exist, the slot should remain empty rather than be filled with invented material. Teams should also be able to bind script names to archive characters, include or exclude props manually, and choose alternate versions of an asset when needed.

Before prompts are finalized, there should be a completeness check for missing scene, character, and prop references. Skipping that check should be a deliberate choice, not a silent default.

Stage 11: Multimodal shoot

Shooting means submitting each blocked scene with its approved prompt and locked references. The output is vertical by default, in 9:16, with selectable model tier, resolution, duration, and audio settings.

Production discipline matters here:

  • one scene should not have duplicate jobs running at once
  • credits and billing should be handled predictably
  • failed jobs should be visible and retryable
  • completed videos should be stored as real takes with actual measured duration
  • the prompt used for a take should remain attached to that take

This is the difference between a demo and a production queue. On a real set, you do not lose track of which take used which slate.

Stage 12: Review and reshoot

The final stage is selection and repair. For each scene block, the team reviews takes, chooses the strongest one, and decides whether to:

  • keep the take
  • edit the prompt and reshoot
  • swap a character, scene, or prop reference
  • rebind a role
  • change model tier or duration
  • cut around the result in later editing

There is no automatic "best take" engine replacing human judgment here. The system provides multiple takes and a clean retake path; the creator or supervisor decides what is good enough.

What the pipeline does not do

It is important to be explicit about the limits, because honest boundaries are what make a production system trustworthy.

  • There is no automatic scoring engine that declares one take objectively best; final quality judgment remains human.
  • Reference images are not a hard forced gate; teams can proceed with text-only shots, though quality is usually weaker.
  • Prompt reference tagging relies on disciplined structure, so creators should still review prompts before shooting.
  • The ~10-second scene block is an engineering heuristic, not frame-accurate timecode editing; long blocks can still occur.
  • There is no cross-shot automatic continuation workflow as a standard product unit; the unit is one scene block to one rendered clip.
  • Character consistency depends on the asset chain, not a separate face-embedding guarantee; final look still depends on art quality and prompt discipline.

These are not flaws to hide. They are the boundaries within which a professional team can plan.

Where human judgment still sits

The point of an AI pipeline is not to remove people. It is to remove the parts of production that are repetitive, drift-prone, and hard to control manually:

  • reconstructing continuity from memory
  • rewriting the same style instructions hundreds of times
  • describing characters from scratch every scene
  • rebuilding shot structure episode after episode
  • losing track of which take used which prompt

What remains with humans is exactly what should remain with humans:

  • whether the story works
  • whether the cast reads correctly
  • whether the look fits the platform
  • whether a performance beat lands
  • whether a cliffhanger is strong enough
  • which take is the one worth keeping

That is the division of labor that scales.

A practical checkpoint list for teams

If you are evaluating or building an AI short drama workflow, use this checklist:

  1. Can the idea be blocked from production until the brief is actually complete?
  2. Is there a living story archive, or are you relying on model memory?
  3. Are scripts written in planned batches with beat sheets and rule-based QC?
  4. Is look chosen once and shared across characters, scenes, props, and video?
  5. Are approved assets created before shooting begins?
  6. Are episodes split into reviewable scene blocks rather than one long render?
  7. Do shot prompts stop describing appearance once reference art exists?
  8. Can you review, edit, and version prompts before rendering?
  9. Do you get multiple takes with attached prompts and references?
  10. Can you reshoot one scene without rerunning the whole episode?

If the answer to most of these is no, you are not running a production pipeline yet. You are running a generation experiment.

About Maosika

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not promise one-click hits or fully unattended production. Instead, it formalizes the staged workflow described above: idea evaluation, brief lock, cast confirmation, story archive, batch script writing, look development, asset creation, scene blocking, engineered shot prompts, multimodal rendering, and multi-take review. Human creators keep the审美 decisions; the system carries the repeatable structure across every episode and every scene.

If you want to dig deeper into the script side, read the pillar on vertical short drama screenwriting craft. For the most common visual failures, see the guide to AI short drama character consistency. For shot construction, see the prompt engineering standard for AI short drama camera language.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com