The Vertical Short Drama Assembly Line: 12 Stages From One-Line Idea to Deliverable Cut

Maosika Editorial | Last updated

A usable AI short drama is not one long generation—it is a sequence of verifiable handoffs. Each stage produces something you can read, reject, or lock before the next one starts.

The core argument

Most teams that fail at AI short dramas fail the same way: they treat the video button as the start of production. It is actually the last mile. Before you ever render a frame, you need a locked brief, a cast that looks like itself across episodes, a continuity record the writer cannot forget, and shot-level instructions that speak the language of a camera department rather than a novel.

The pipeline below is the shape a professional AI vertical short drama workflow should take. It is organized around 12 stages, each with a concrete deliverable. If a stage has no artifact you can point at, it has not happened yet.

Why "assembly line" is the right mental model

A film set does not ask the director to imagine the lead actor's face every time the camera rolls. The actor shows up in wardrobe and makeup, the script supervisor tracks continuity, and the DP works from a shot list. AI short drama production needs the same discipline—except some of the crew roles are now enforced by software rather than a call sheet.

Consistency comes from assets, not luck.

The goal is not to remove human judgment. The goal is to make sure judgment happens at the decision points that matter—story direction, character look, shot design, final take selection—and not get re-litigated in every generated clip.

Stage 1: Intake scoring

Before any writing begins, the idea is scored on completeness across the dimensions that actually determine whether a vertical short drama can be produced. A thin one-liner ("a CEO falls for a maid") is not a brief; it is a starting prompt.

A production-ready intake covers:

  • Genre and tone
  • Protagonist and their arc
  • Core conflict
  • Story direction and ending shape
  • Episode count and episode length
  • Hook and payoff rhythm
  • Platform and audience

If the intake is rich enough, you move straight to brief lock. If it is thin, you are guided through the missing dimensions. If it is nearly empty, you go through a full guided ideation flow. The point is: you do not get to write a script from a half-baked premise just because the model is willing to try.

Stage 2: Creative brief lock

The brief is the first hard gate. Until it is locked, nothing downstream proceeds.

A locked brief fixes in place:

FieldWhat it pins down
LoglineOne sentence that sells the story
Core conflictWhat is actually fighting what
Story directionWhere the plot travels, episode to episode
Ending directionRough shape of resolution, so arcs can aim somewhere
Hook / payoff rhythmHow often reversals and reveals land
Platform & audienceVertical rhythm, episode length, tonal ceiling
Episode outlinePer-episode beats, not just a theme
Writer notesConstraints the writer must obey

The protagonist entry must include a character arc. A lead without an arc is a face, not a character.

Stage 3: Cast lineup and visual confirmation

Characters are named, described, and aligned to the brief before a single script page is written. This is not a cosmetic step: later stages bind reference images to character names, so a name that drifts ("the officer", "Detective Li", "that cop") becomes a continuity error waiting to happen.

The cast is checked for:

  • Lineup completeness (no unnamed key roles)
  • Legal names
  • Visual description fields filled in
  • Lead characters matching the brief

This is the equivalent of locking cast before principal photography.

Stage 4: Building the continuity bible

The continuity bible is a structured record that travels with the show for its entire run. It tracks:

  • Character identities
  • Stable personality traits
  • Mutable current state (injuries, revealed identities, changed allegiances)
  • Relationships between characters
  • Open and resolved plot threads
  • Per-episode appearance table
  • Per-batch plot summaries
  • Prop visual descriptions

A continuity bible is what lets episode 200 remember what episode 3 set up. It replaces model memory with structured record-keeping.

When a new batch of episodes is written, the writer is fed only the relevant slice: current character states, unresolved threads, recent batch summaries, and the current batch's main line. This prevents both "amnesia" and the noise of an overstuffed context window.

Stage 5: Batched script writing

Scripts are written in batches of several episodes, not one endless generation. Each batch follows the same internal order:

  1. Beat sheet first (a scene-by-scene plan for the batch)
  2. Episode-end cliffhangers pinned before dialogue
  3. Full dialogue draft
  4. Rule-based quality check
  5. Continuity bible updated with what just happened

Starting from the second batch, four things are confirmed up front and treated as hard constraints: the batch's plot direction, focus characters, threads and conflicts in play, and the episode-end hook. If these conflict with the older bible, the freshly confirmed intent wins.

The writing rules that are enforced, not suggested

Vertical short drama has a grammar that differs from long-form. These rules belong in the production system, not in a writer's memory:

  • Golden 3 seconds: the first scene must open on conflict or suspense; no slow background setup.
  • Single-episode shape: opening hook → escalation → end-of-episode cliffhanger.
  • Payoff density: at least one small payoff per episode (a reveal, a reversal, a power move, a piece of evidence); a larger payoff every few episodes.
  • Dialogue discipline: short lines, generally under 20 characters in Chinese source pacing terms, no essay-speak, no explanatory voiceover dumping backstory.

The quality gate

After writing, each episode passes through a rule checker that catches:

  • Wrong episode titles
  • Scene count mismatches
  • Missing character headers
  • Too little dialogue
  • Placeholder text like "to be continued" used as a cop-out

Failures are sent back for rewrite with the specific error attached. A half-passing script is not delivered.

Stage 6: Look selection

Before any image is generated, a visual style is chosen from a shared look library. Each look covers:

  • Character sheet guidance (face anchors, material, temperament, multi-view consistency)
  • Scene and prop guidance
  • Video style tags

The same look path is shared by character art, scene art, prop art, and video prompts. This is what prevents the common failure mode where characters look like 2D anime but the rendered video snaps to live-action realism.

Stage 7: Character, scene, and prop look-dev

Asset creation happens in two steps:

  1. Text polish: raw character descriptions are rewritten into image-generation prompts using the selected look's character rules, with gender enforced as a hard constraint. This text is editable by hand.
  2. Image generation: the final prompt is rendered, with optional reference images, and saved as a versioned asset.

Scenes follow an empty-frame rule: scene reference images must not contain people. Props are pulled from the script using the script's own wording ("the jade pendant", "the sealed envelope") and cataloged across the whole show so nothing appears in episode 2 and vanishes in episode 9.

Crucially, video generation only consumes assets that are finished and readable. A half-rendered character sheet is never silently used as a reference.

Stage 8: Shot blocking into ~10-second chunks

A finished episode is not one big video job. It is sliced into video blocks, each targeting roughly 10 seconds of screen time, with a soft cap on body length. Long scenes are split further on action beats, paragraph breaks, or sentence endings.

The result looks like a clip list on an editing timeline: Episode 3 → Shot 1, Shot 2, Shot 3… Each shot has its own prompt, its own reference set, its own output, and its own history of takes.

Default delivery is 9:16 vertical, not a landscape crop after the fact. Other aspect ratios are supported, but vertical is the native shape.

Stage 9: Reference mapping per shot

This is the single most important mechanism for character consistency. For every shot, a reference table is built in a fixed order:

  1. Scene reference
  2. Props appearing in this shot
  3. Character look-dev images

Slots are only filled if an image exists; empty slots are marked as text-only rather than invented.

The governing rule is worth stating plainly:

When a character has a reference image, the prompt must not re-describe their clothing or appearance in text. Text describes only action, expression, and injury. The image is the source of truth for how they look.

This directly attacks the classic AI video failure: a text prompt that says "red dress" arguing with a reference image of a woman in black, producing a mid-shot costume change.

Additional handles at this stage include:

  • Manual character binding (when the script says "the captain" but the cast sheet says "Wang Li")
  • Manual prop inclusion / exclusion, so irrelevant props do not eat reference slots
  • Swapping to a different historical version of a character or scene asset
  • Era-based routing for time-travel or flashback stories, where a character has both period and modern looks

Before prompts are generated, the system checks for missing references and surfaces a list. You can skip and go text-only—but you are warned, because text-only consistency is markedly weaker.

Stage 10: Engineered shot prompts

Prompts are not freeform paragraphs. They follow a structured shot language with eight components:

ComponentRole
Precise subjectWho or what is in frame
Action detailWhat they are doing, concretely
Scene environmentWhere they are
Light and colorMood and exposure direction
Camera movementOne move per shot, no stacked pans and zooms
Visual styleTags inherited from the chosen look
Image qualityBaseline fidelity constraints
Negative constraintsStability, no watermarks, no twins in multi-person shots

Simple scenes use a single-block prompt. Complex cinematic scenes use a three-part structure: overall setup, numbered shots, and a constraint bundle. Shots are referred to by number, not by absolute timestamps. Dialogue, sound effects, and score use explicit markup so they are not confused with visual description.

The prompt writer only sees assets for the current shot—no cross-episode bleed—and treats the shot's script body as the highest-priority source. Generation parameters are kept conservative, favoring stability over spectacle.

The system is teaching the model to speak a shot list, not to write a novel.

Before delivery, prompts are cleaned of specific copyrighted IP names while retaining the stylistic and technical description, reducing downstream content-policy risk.

Stage 11: Rendering with a production queue

When you actually shoot, the workflow mirrors a real set:

  1. Open the episode's video workspace.
  2. Confirm references are present (or intentionally skipped).
  3. Generate shot prompts.
  4. Read and edit them—this is the director revising the shot list.
  5. Choose model tier, aspect ratio, resolution, and duration.
  6. Submit to the render queue.
  7. Review historical takes per shot and pick the best one.

The controls you have at this stage map cleanly to on-set roles:

ControlOn-set equivalent
Editing the shot promptDirector revising shot notes
Swapping character / scene / prop referencesChanging wardrobe or location plate
Manual character bindingFixing a name mismatch in the cast sheet
Including / excluding propsControlling what is in frame
Switching lookReigning in art direction
Choosing model tierTrading quality, cost, and speed
Reviewing takes per shotMulti-take selection

A new render only happens when you change something and resubmit—same as a new take only rolls after the director adjusts the setup. The queue itself is built for production, not demos: tasks run in isolated lanes, in-flight shots cannot be double-submitted, credits are pre-deducted and released on failure, hung jobs time out and become retryable, and successful runs can resume polling without double-charging.

Stage 12: Review, retake, and handoff

There is no automatic "best take" picker. Final quality judgment sits with the creator or supervising producer. What the system provides is:

  • Multiple takes per shot
  • The ability to change the prompt and re-render
  • The ability to swap a reference and re-render
  • Per-shot history so you can compare rather than regenerate blindly

Final assembly, trimming, and cross-shot pacing still belong in an edit. The ~10-second block size is an engineering heuristic, not a timecode-precise cut; some shots will run long, and stitching them into a finished episode is a deliberate post step.

What this pipeline deliberately does not do

Being honest about the boundaries is part of what makes the pipeline trustworthy:

  • It does not auto-score finished footage or auto-choose a winning take. Human eyes do final selection.
  • Reference images are not a hard gate; you can skip them and go text-only, at a clear quality cost.
  • Reference tokens in prompts are guided by convention, not forcibly re-injected; creators should still read the prompt and confirm references are present.
  • Shot grammar follows the built-in camera rules; the full art manual applies on the image side, not as a per-shot override.
  • The ~10-second block is heuristic, not precision editing; long scenes may still overshoot.
  • There is no cross-shot automatic video extension; the unit of work is one shot, with references, producing one clip.
  • Character consistency depends on the look-dev asset chain, not a face-embedding verification lock. Era routing is rule-based; final look still depends on asset quality and on the prompt obeying the "don't describe what the image already shows" rule.

Where Maosika fits in

Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not present itself as a single text-to-video button; it enforces the sequence above as a connected production chain, with verifiable artifacts at every handoff. The human makes the taste calls at the gates that matter—brief lock, cast look, shot prompt edits, final take selection—and the system carries those decisions consistently across episodes and shots.

If you are building or running a short drama team, the question to ask of any AI tool is not "how good does one sample look?" It is "what artifact do I get to inspect at each stage, and what happens if I reject it?" A pipeline that cannot answer that is a demo. One that can is a production system.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com