How Vertical Short Dramas Get Made: The 12-Stage Production Pipeline Behind AI-Generated Episodes

Maosika Editorial | Last updated

Making an AI vertical short drama is not one long video generation. It is a staged pipeline with checkpoints: lock the brief, build the continuity bible, write in batches, approve looks, split into scene blocks, shoot block by block, and pick the best takes.

The production problem is not generation—it is handoff

Most teams that try AI for vertical short dramas run into the same wall: the first episode looks promising, the third changes the lead's face, the seventh forgets a key prop, and by episode twenty the story contradicts itself. The failure is rarely a single bad render. It is a bad handoff between stages.

A film set solves this with roles: a producer locks the brief, a writer owns the script, a continuity supervisor tracks state, an art department approves looks, a director plans shots, and an editor picks takes. An AI pipeline has to do the same, even when one person is running the whole show.

This article walks through the twelve production stages that turn a vague idea into a deliverable vertical episode. The unit of thinking is not "prompt in, video out." It is "what artifact must be approved before the next stage is allowed to spend compute?"

Stage 1: Intake and idea scoring

Every project starts as a half-formed idea: a betrayed heir, a time-slip romance, a revenge plot in a luxury office. The first job is to measure how complete that idea actually is.

A good intake scores the concept across the vertical-short-drama dimensions that actually matter: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, target platform, and audience. Ideas that score high can move straight to a brief; ideas that score low need a guided development pass before any script is written.

Think of it as a development meeting. You would not greenlight a shoot where no one can state the logline or the ending direction. The same rule applies here.

Stage 2: Guided development for incomplete ideas

When the concept is too thin, the pipeline should ask the missing questions rather than invent answers. The goal is not to interrogate the creator; it is to fill the gaps that would otherwise cause drift later: who the lead is, what they want, what stands in their way, how the tone should feel, and how often the hook should land.

For vertical short dramas, episode length usually lands in the one-to-two minute range, sometimes extending to three or four. That constraint changes writing: there is no room for a slow setup, and every episode needs a reason to swipe to the next.

Stage 3: Lock the creative brief

The creative brief is the first hard gate. Until it is locked, nothing downstream should start.

A locked brief includes:

  • A one-line logline
  • The core conflict
  • Story direction
  • Ending direction
  • Payoff beats and hook rhythm
  • Target platform and audience
  • Episode-by-episode outline
  • Notes for the writer
  • A protagonist definition that includes character arc

This is the document every later stage reads from. If the brief changes, the pipeline should re-open from here rather than trying to patch episodes mid-production.

Stage 4: Cast the characters and confirm visual direction

Before any script is written, the main cast needs names, roles, relationships, and visual anchors. This is not about generating final art yet; it is about making sure the writer knows who is in the room during each scene, and the art side knows who will eventually need a look.

A cast list is production-ready only when it is complete, names are valid, visual fields are filled, and the leads match the brief. That sounds obvious, but many AI workflows let scripts run ahead of casting and then try to invent characters retroactively. That is where face inconsistency begins.

Stage 5: Build the story archive

The story archive is a structured continuity record that follows the whole series. It tracks character identities, stable traits, current state such as injuries or hidden identities, relationships, open and resolved plot threads, episode appearance tables, batch-level story notes, and visual descriptions of recurring props.

In writers' room terms, this is the continuity bible. The reason it matters for AI is simple: long-form series cannot rely on model memory. A structured archive gives each writing pass only the slice it needs: current character state, unresolved threads, recent batch summary, and the current batch's main line. After each batch, the archive is updated; before the next batch, it is reloaded.

Continuity comes from assets, not luck.

Stage 6: Batch script writing with a beat sheet first

Writing a whole season in one pass is how stories drift. The safer structure is batch writing: a few episodes at a time, with an approved plan before dialogue starts.

Each batch begins with a beat sheet. For every episode, the pipeline should lay out the scene list and the episode-end cliffhanger before writing the actual pages. That enforces the vertical-drama rhythm:

  1. A strong hook in the first scene
  2. Escalating conflict through the middle
  3. A cliffhanger in the final scene

The writing rules for vertical episodes are specific and worth treating as hard constraints:

  • The first scene must open with conflict or suspense, not exposition
  • Each episode needs at least one small payoff: a reversal, a face-slap, an identity hint, or evidence revealed
  • Every few episodes need a larger payoff beat
  • Dialogue stays short; long explanatory monologues break the format

From the second batch onward, the pipeline should confirm four things before writing: the batch's plot direction, focus characters, threads and conflicts, and end hook. Those confirmed intentions override older archive assumptions when the two conflict, because the creator's current direction wins.

Stage 7: Rule-based script quality check

A script is not done when the language sounds good. It has to pass a production check. Common failures include wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, or placeholder text such as "to be continued."

A production pipeline should catch these automatically and request a rewrite with the specific error attached. Scripts that do not pass should not be handed to the art or video stages. Fixing a script problem is cheap; fixing it after assets and renders exist is not.

Stage 8: Choose the visual style and produce look-development assets

Once scripts are approved, the visual side starts. A usable style system is more than a single aesthetic word. It should include a character sheet guide for face anchors, materials, and view consistency; a scene and prop guide; and video style tags that travel into shot prompts.

Characters, scenes, and props should all use the same style path. That is how you avoid the common AI mismatch where characters look like anime but the rendered video flips into live-action realism.

Look development happens in two steps. First, the archive description is polished into image-generation prompts using the chosen style guide and hard rules such as gender. Creators can edit this text. Second, the final prompts generate the approved character, scene, and prop sheets, with version history so earlier approved looks can be reused.

Scene sheets follow an empty-shot rule: no people in scene references. Prop sheets are pulled episode by episode from the script's actual wording, then merged into a show-wide catalog so recurring items do not get reimagined later.

Stage 9: Split each episode into scene blocks

A one-to-two minute episode should not be treated as one rendering job. It should be cut into scene blocks, usually around ten seconds each, with a soft cap on text length and further splits when action beats require it.

Each block becomes an independent production unit: its own shot prompt, its own reference set, its own render, and its own history of takes. Crowd characters or generic extras should be separated from the named cast so they do not consume reference slots meant for lead consistency.

This is the equivalent of a clip list on an editing timeline. It makes review possible: if one beat fails, you reshoot that beat, not the whole episode.

Stage 10: Bind references and build engineered shot prompts

This stage is where most AI drama workflows either win or lose consistency.

For each scene block, the pipeline should build a reference table in a fixed order: scene reference first, then prop references for that block, then character references. Slots are filled only when an approved asset exists; missing references are marked as text-only rather than invented.

The highest-priority consistency rule is simple: if a character has a reference image, the prompt must not re-describe that character's clothes or appearance in text. The image owns the look. The text only describes action, expression, and injury state. That single rule prevents a large share of face swaps and costume changes.

Manual controls belong here too:

  • Bind a script name such as "the officer" to the correct cast card
  • Include or exclude props to control visual focus
  • Swap in an older approved version of a reference
  • Route period-specific looks for time-slip or flashback scenes

The shot prompt itself should read like a shot list, not a paragraph of prose. A strong prompt covers eight elements: precise subject, action detail, scene environment, lighting and color, camera movement, visual style, image quality, and constraints. Complex scenes can be split into a general setup, numbered shots, and a constraint pack. One shot gets one camera move; absolute timestamps are replaced with shot numbers; and a safety pack covers face stability, watermark avoidance, and duplicate-person issues in multi-character scenes.

The system should teach the model to speak in production language, not novelistic language.

Stage 11: Render with model, ratio, and resolution choices

Vertical short drama is native 9:16, not a landscape video cropped after the fact. That should be the default delivery shape. Other ratios can be supported, but the pipeline should treat vertical as the primary format because composition, hook placement, and close-up staging all change in portrait.

Before rendering, the creator should be able to choose the model tier for quality, speed, and cost trade-offs; set ratio and resolution; choose smart or fixed duration within the short-block range; and decide whether audio is generated. The prompt used for rendering should be the one the creator reviewed, not a secretly rewritten version.

A production-grade queue matters here. Video tasks should run in their own lane, separate from script and image work. The same scene block should not accept parallel submissions while a render is running. Failed jobs should be visible and retryable, and timeouts should be recovered rather than left hanging. This is what turns a demo into an operational production line.

Stage 12: Review, retake, and assemble

After rendering, review happens at the block level. Each scene block keeps its own take history, so the creator can compare versions and pick the strongest one. If a block is wrong, the fix is specific: edit the shot prompt, swap a reference, rebind a character, adjust a prop, or change the model tier, then render a new take.

There is no automatic "best take" engine. Final quality judgment stays with the creator or producer. The pipeline's job is to make retakes cheap, traceable, and isolated, so one bad shot does not require rebuilding the whole episode.

Final assembly, trimming, and episode-level polish still belong in editing. The roughly ten-second block size is an engineering heuristic, not a precision timecode cut, and the current production unit is a single scene block with references, not an automatic cross-scene video continuation.

What this pipeline does and does not promise

A structured pipeline removes the failure modes that come from improvisation: forgotten plot threads, inconsistent faces, props that change shape, episodes that open too slowly, and renders based on half-approved prompts. It does not remove human judgment.

The honest boundaries are worth stating clearly:

  • There is no automatic quality-scoring engine that picks the winning take for you
  • Reference images are not a hard blockade; you can skip them, but quality usually drops
  • Reference tags still need a creator's eye during prompt review
  • Full art manuals drive look development, while video prompts use style tags and shot grammar
  • Scene blocks are approximate and may need editorial trimming
  • Cross-scene automatic video extension is not the current product unit
  • Character consistency depends on approved look assets and prompt discipline, not a magic face lock

That is the right trade-off. The goal is not to claim that AI can replace a crew with one button. The goal is to turn the repeated, drift-prone, easy-to-break parts of short-drama production into a constrained pipeline, while审美 judgment stays with the person making the show.

Where Maosika fits in

Maosika (猫斯卡) is built around this staged production model: an AI production operating system for vertical short dramas, with 18 digital specialist roles modeled after real crew positions, 17 built-in style manuals, structured story archives, batch script writing with rule checks, look-development assets, scene-block shooting, engineered shot prompts, and a retake-friendly render queue.

It does not promise a hit in one click. It enforces the order that professional production already uses: story and continuity first, then approved looks, then scene planning, then shot instructions, then multiple takes selected by a human.

If you are evaluating tools for a vertical short drama slate, ask whether the workflow gives you approvable artifacts at every stage—or whether it hides the whole chain behind one generate button. The second approach looks faster on day one. The first one is what lets you produce episode twenty without rebuilding episode one.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com