The AI Vertical Short Drama Pipeline, Step by Step: 12 Checkpoints From Idea to Finished Episode

Maosika Editorial | Last updated

AI vertical short drama production is not one long text-to-video prompt. It is a pipeline of 12 checkpoints, each with a deliverable you can inspect, edit, or reject before the next stage begins.

Why most AI short drama workflows fail

Teams that treat AI short drama as a single "generate video" button usually run into the same three problems: the story drifts after a few episodes, characters change face or outfit between scenes, and the final footage has to be re-shot so many times that the supposed speed advantage disappears.

These are not model failures. They are pipeline failures. In a real crew, you do not walk onto set without a locked brief, a cast, a continuity bible, and a shot list. The same discipline applies when the crew is digital.

A production-grade AI short drama workflow is defined by one rule: every stage produces an inspectable artifact before the next stage consumes it. Scripts, story bibles, character look sheets, scene blocks, shot prompts, and takes are all saved, versioned, and editable. The human approves the direction at the gates; the system carries that decision into every episode and every scene.

The 12-checkpoint pipeline

Below is the full pipeline from idea to finished vertical episode, organized the way a working production line actually runs.

#CheckpointWhat gets deliveredWho decides it is ready
1Idea intake & scoringA completeness score and a recommended pathProducer
2Creative guidanceMissing dimensions filled inProducer / writer
3Brief lockLogline, conflict, arc, hooks, platform, endingProducer
4Cast & visual confirmationNamed cast with aligned visual fieldsProducer / art
5Story archive setupContinuity bible: characters, threads, statesWriter / archive
6Batch script writingBeat sheets first, then dialogue, then QCWriter / supervisor
7Style selectionOne shared style path for art and footageArt director
8Look developmentCharacter, scene, and prop sheetsArt / makeup / props
9Scene blockingEpisodes cut into ~10-second scene blocksDirector / editor
10Shot prompt buildPer-scene reference map + camera promptsDirector / DP
11Multi-modal shootingVertical takes per scene, queued and loggedProducer / director
12Review & re-shootEdited prompts, swapped refs, best take selectedCreator / executive producer

This is the shape Maosika (猫斯卡) follows. It is an AI production operating system for vertical short dramas, not a single generate button; the product exists to enforce the order above, not to skip it.

Stage 1–3: From idea to locked brief

The first three checkpoints exist to answer one question: are we actually ready to write?

Idea intake scores the concept on completeness. A score high enough can go straight to brief; a mid score only fills in the missing dimensions; a low score goes through a full guided flow covering genre, protagonist, core conflict, arc, episode count, episode length, tone, hook rhythm, ending direction, and target platform and audience.

Episode length is steered toward vertical norms: roughly 1–2 minutes per episode as the main working range, extending to 3–4 minutes at most. This is not an artistic rule; it is a format rule for vertical micro-drama.

The creative brief is a hard gate. Until it is locked, no script is written. The brief contains the logline, core conflict, story direction, ending direction, hook and payoff rhythm, platform and audience, episode-by-episode outline, and notes to the writer. The protagonist entry must include a character arc.

The brief is the equivalent of a greenlit development meeting. If the brief is vague, every later stage will invent its own version of the show.

Stage 4–6: Cast, continuity, and batch writing

Once the brief is locked, the cast is assembled and visually confirmed. The gate here is strict: names must be valid, visual fields complete, and leads aligned with the brief. You do not start writing against a half-named cast.

The story archive is then initialized. This is the digital continuity bible for the whole show.

A story archive is a structured record that follows a series from episode one to the end, tracking character identities, stable traits, current states (injuries, revealed identities, changed relationships), open and resolved plot threads, episode-by-episode appearance tables, batch-level plot summaries, and prop visual descriptions.

The archive solves the classic "context amnesia" problem. Instead of asking a model to remember everything across dozens of episodes, each writing batch only consumes a slice: current character states, unresolved threads, recent batch summaries, and the current batch mainline. After each batch finishes, the archive is updated; before the next batch starts, it is reloaded. The principle is simple: do not rely on model memory; rely on a structured archive.

Scripts are written in batches of a few episodes each, with a hard order:

  1. Beat sheet for the episode, including the end-of-episode cliffhanger
  2. Full scene dialogue
  3. Rule-based quality check
  4. Archive update

Vertical short drama writing has its own embedded rules, and they are worth stating plainly:

  • Golden 3 seconds: the first scene must open with strong conflict or strong suspense; no flat setup, no long exposition
  • Single-episode structure: opening hook (1 scene) → rising conflict → end-of-episode cliffhanger (final scene)
  • Payoff density: at least one small payoff per episode (a reversal, a face-slap, an identity hint, a piece of evidence landing); a larger payoff every few episodes
  • Dialogue: short lines, generally under 20 characters in the original Chinese writing rhythm, no essay-like monologue and no narrator dumping backstory

From the second batch onward, the system locks intent before writing: the batch direction, focus characters, threads and conflicts, and end hook are marked as confirmed hard constraints. If those conflict with older archive entries, the newly confirmed creative intent wins.

A rule-based quality check catches concrete failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, and placeholder text such as "to be continued." Failing scripts are sent back for rewrite with the error attached; work that does not pass is not delivered as if it were finished.

Stage 7–8: Style and look development

Style is chosen once and shared across the whole downstream chain. Maosika ships with 17 built-in style manuals covering 2D, 3D, and realistic directions, including urban realism, period realism, mature urban romance animation, 90s Japanese anime, Chinese ink-and-brush, xianxia ancient style, 3D donghua, clay stop-motion, and cyber-Chinese styles, among others.

Each manual includes at least three parts: a character sheet guide (face anchors, material, temperament, multi-view consistency), a scene and prop guide, and video style tags. Once a style is selected, character art, scene art, prop art, and video prompts all follow the same style path. This prevents the common failure mode where characters look like anime but the footage suddenly turns photoreal.

Look development runs in two steps:

  1. Text polish: the archive description is rewritten into image-generation prompts using the chosen style manual, with gender as a hard rule; this step can be manually overridden
  2. Image generation: the final prompt is rendered, reference images are supported, and outputs are saved as reviewable historical versions

The video side only consumes art assets whose status is complete and whose files are readable. Half-finished renders are not silently bound as references.

Stage 9–10: Scene blocks and shot prompts

Episodes are not rendered as one long video. They are cut into scene blocks, with a target granularity of roughly 10 seconds per scene. Soft body text is capped around 200 characters; longer scenes are split again by action beats, paragraph breaks, and sentence endings. Extras and generic roles are separated from the drawable main cast so they do not consume character reference slots.

What the creator sees is not "Episode 3 as one big file," but Episode 3 broken into Scene 1, Scene 2, Scene 3, and so on. Each scene has its own prompt, its own reference map, its own output, and its own history of takes. This is the editing-room clip list, not a one-shot render.

Default delivery is 9:16 vertical. Other ratios are supported, but vertical is the native output shape, not a crop applied after the fact.

For each scene, the system builds a reference table in a fixed order: scene image → scene props → character look sheets. Slots are only filled when an image exists; if nothing exists, it is marked as text-only rather than inventing a binding. Users can manually bind roles (for example, the script says "the officer" and it is mapped to the character "Li Qiang"), manually include or exclude props to control visual focus, and swap in a different historical version of the same named asset.

One rule here does more for character consistency than anything else:

For any character with a reference image, the prompt must not describe clothing or appearance in text. The reference image is the source of truth for appearance; text only describes action, expression, and injury state.

This directly targets the classic AI video failure where text description fights the reference image and the character changes outfit or face mid-scene.

Shot prompts themselves follow an engineered camera language, not a prose style. The eight elements are: precise subject, action detail, scene environment, lighting and color, camera movement, visual style, image quality, and constraints. Simple scenes use one paragraph; complex cinematic scenes use a three-part structure: overall setup, shot-by-shot instructions, and a constraint pack.

A few hard prompt rules keep output stable:

  • One camera movement per shot; do not stack push, pull, pan, and tilt in one shot
  • Use shot numbers, not absolute timestamps like "0–3s"
  • Always include a fallback pack: image quality, stable faces, no watermark or logo; add twin/duplicate fallbacks for multi-person scenes; anchor style explicitly for non-realistic work
  • Favor continuous low-speed motion over explosive high-dynamic action, which breaks more often
  • Use fixed notation for dialogue {}, sound effects <>, and BGM ()
  • Only feed the current scene's assets into the prompt; never cross-contaminate with another scene

In one sentence: the system teaches the model to speak in shot-list language, not in novel prose.

Before prompts are handed off, IP and specific media titles are stripped out, leaving technique and aesthetic description, to reduce downstream copyright blocking risk. If a reference image is detected as a suspected real-person photo, the system returns explicit guidance to use the platform's own art pipeline instead of binding a real photo.

Stage 11–12: Shooting, takes, and review

The standard shooting path is straightforward: open an episode's video workspace, confirm references are complete (or deliberately skip them), generate shot prompts, read and edit them, choose model tier / ratio / resolution / duration, submit to render, wait in queue, then review historical takes and pick the best one.

The professional control points map cleanly onto real crew roles:

Control in the toolEquivalent on a real set
Editing the shot promptDirector revising the shot list
Swapping character / scene / prop referencesChanging a look sheet or location board
Manual character bindingFixing name mismatches and guest references
Including / excluding propsControlling the scene's visual focus
Switching styleUnifying the art language
Choosing model tierTrading quality, cost, and speed
Reviewing historical takes per sceneMulti-take selection

When you click generate, the system does not re-invent the prompt. It uses exactly the prompt you confirmed or edited, plus the reference map locked at that moment. Change the prompt and generate again, and that is a new take—the same as on a real set when the director revises the shot note and rolls again.

On the production reliability side, jobs run on an isolated queue separate from writing and art tasks; the same scene block cannot be submitted in parallel while a job is running, to avoid double charges and state confusion; credits are pre-deducted and released on failure; hung jobs time out and become retryable; and finished files go through faststart handling with actual measured duration stored, rather than trusting a vendor-reported length alone. This is an operational production queue, not a toy script.

Where the human still sits

It is important to be explicit about what this pipeline does not do. These boundaries are not caveats hidden in fine print; they are part of how the tool should be used.

  1. There is no automatic scoring engine that picks the best take or re-shoots by itself. Final quality judgment stays with the creator and the executive producer; the system provides multiple takes and the tools to re-shoot with revised prompts.
  2. Reference images are not a hard gate. Missing images trigger a reminder, but you can skip and go text-only; quality is usually worse, which is why a professional workflow locks look dev before shooting.
  3. Reference tokens in prompts rely on the specification, not a hidden hard-stitch step. Creators should still confirm that references are present when reviewing prompts.
  4. Video shot grammar follows the platform's built-in camera specification. The style manual injects video style tags; the full art manual is used on the image side.
  5. The ~10-second scene block is an engineering heuristic, not timecode-precise editing. Overlong scenes are re-split, but blocks can still run long; final trimming and stitching belong in a later edit pass.
  6. There is currently no productized cross-scene automatic continuation or video extension flow. The unit of work is "single scene with multi-modal references → single segment output."
  7. Character consistency depends on the look-dev asset chain, not a face-embedding verification step. Era routing uses rules to pick the right look sheet, but final appearance still depends on the quality of the art and on the prompt obeying the "do not describe appearance when a reference exists" rule.

Maosika's position is that of a scalable AI short drama operating system: it hard-bakes a professional sequence into the product, rather than claiming zero-human, zero-error output. The repeated, drift-prone, failure-prone parts become a constrained pipeline; aesthetic judgment still sits with the person in the chair.

A practical checkpoint list for your own workflow

Even if you are not using Maosika, you can apply the same discipline to any AI short drama production:

  1. Score your idea for completeness before you write; do not start from a one-sentence concept
  2. Lock a written brief; refuse to write until it is locked
  3. Build a continuity archive before episode count grows
  4. Write beat sheets before dialogue, and cliffhanger before the body
  5. Pick one visual style and share it across characters, scenes, props, and footage
  6. Render look sheets first; never send half-finished art to the video stage
  7. Cut episodes into scene blocks before shooting; do not render whole episodes as one clip
  8. Build per-scene reference maps in fixed order: scene, props, characters
  9. For characters with references, ban appearance description from the prompt
  10. Write shot prompts as shot lists, not prose
  11. Queue per scene, keep takes, and select the best one manually
  12. Treat re-shoots as prompt edits, not as random retries

High-quality short drama is never one long generation cut up after the fact. It is a stack of controllable units, each one approved before the next one starts. Consistency comes from assets, not luck.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com