The AI Vertical Short Drama Pipeline, Step by Step: From Idea Lock to Deliverable Clips

Maosika Editorial | Last updated

Most AI short drama failures are not model failures—they are skipped steps. A reliable pipeline locks the brief first, builds a continuity bible, approves looks, splits episodes into scene blocks, and films one block at a time with bound references.

Why "end-to-end" usually breaks in production

When teams say they use AI to make vertical short dramas, they often mean: type a premise, generate a script, throw it at a video model, and hope the characters stay the same across episodes. That demo works for one clip. It falls apart at episode 20, when the lead’s face drifts, props vanish, and plot threads from episode 3 are forgotten.

A production-grade pipeline is different. It treats each stage as a reviewable deliverable before the next stage starts. The goal is not to remove human judgment; it is to make the repeatable, drift-prone parts of the workflow enforceable.

Continuity is built with assets and records, not with model memory.

The 11-stage pipeline at a glance

Think of this as a real vertical drama crew, not a single generate button.

StageDeliverableWho approvesWhat goes wrong if skipped
1. Idea intakeIntake score + missing-dimension listProducer / creatorVague premise causes rewrites later
2. Creative guidanceFilled creative dimensionsCreatorWrong tone, platform, or hook rhythm
3. Brief lockOne-page creative briefCreator / showrunnerScript drifts from the original concept
4. Cast & visual confirmCharacter lineup with visual fieldsCreator / art leadIncomplete cast blocks look development
5. Continuity bibleStructured story archiveWriting leadPlot amnesia across batches
6. Batch script writingBeat sheets → script → rule checkEditor / showrunnerWeak hooks, broken format, filler lines
7. Style selectionStyle manual choiceArt leadMixed visual languages across assets
8. Look developmentCharacter / scene / prop sheetsArt lead / directorFace drift, costume drift, empty scenes
9. Scene blocking~10-second scene blocks per episodeDirector / editorLong unmanageable clips
10. Shot prompts + filmingPrompt + references → rendered takesDirectorText fights reference images
11. Review & reshootSelected takes, edits, retakesCreator / producerBad clips accepted because no take history

Stage 1–3: Intake, guidance, and brief lock

A short drama should not start filming before the brief is locked. This sounds obvious, but many AI workflows skip it because the model is willing to write immediately.

Idea intake is a completeness check, not a formality. A premise can be scored from 0–100 across the dimensions that actually matter for vertical drama: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, target platform, and audience. If the intake is strong, the system can move straight to brief; if it is partial, it only asks for the missing dimensions; if it is too thin, it runs a full guided process.

The locked brief should include at least: logline, core conflict, story direction, ending direction, satisfaction beats and hook rhythm, platform and audience, episode-by-episode outline, and notes for the writer. The protagonist entry must include character arc, not just a job title or a face claim.

This stage is the equivalent of a development meeting ending with an approved brief. After that, the writer has constraints instead of unlimited freedom.

Stage 4–5: Cast confirmation and the continuity bible

A continuity bible is a structured record that travels with the show for its entire run. It tracks character identities, stable traits, current state (injuries, revealed identity, changed status), relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual notes.

This is how a show remembers what happened in episode 3 when it is writing episode 60. The model does not "remember" the whole script. The archive feeds only the relevant slice: current character states, unresolved threads, recent batch summary, and the current batch’s main line.

Before look development, the cast must pass a hard check: lineup complete, names valid, visual fields complete, leads aligned with the brief. If a character is missing or visually undefined, the pipeline should stop, not guess.

Stage 6: Batch script writing for vertical drama

Vertical short drama scripts are not compressed TV scripts. They have their own rhythm.

The writing rules that belong in the pipeline

These rules should be enforced by structure and checks, not left as advice to the model:

  1. Golden 3 seconds: the first scene must open with strong conflict or strong suspense; no slow setup or exposition dump.
  2. Single-episode structure: opening hook → conflict escalation → end-of-episode cliffhanger.
  3. Beat density: at least one small payoff per episode (a reversal, a face slap, an identity hint, evidence obtained); a larger payoff every few episodes.
  4. Dialogue discipline: short lines, no lecture tone, no long explanatory narration.

Why batch writing matters

Writing all episodes in one shot creates drift. Writing one episode at a time loses arc control. The stable pattern is batch writing: plan a batch of episodes, write them, run checks, update the continuity bible, confirm the next batch’s intent, then continue.

Before each new batch, four things should be locked: the batch’s plot direction, focus characters, threads and conflicts, and end-of-batch hook. Those become hard constraints. If they conflict with older archive notes, the confirmed new intent wins.

Beat sheet first, then dialogue

For each episode, the beat sheet and cliffhanger should exist before the full script. This prevents "written well but going nowhere" episodes. After writing, a rule checker should catch concrete failures: wrong episode title, mismatched scene count, missing character lines, too little dialogue, or placeholder text like "to be continued." Failing drafts should be sent back with the specific error, not handed to the user as finished.

Stage 7–8: Style manual and look development

Consistency breaks when characters, scenes, props, and video are generated under different visual assumptions. A strong pipeline solves this by choosing one style path and reusing it across all asset types.

A style manual is not just a prompt like "cinematic." It should include character sheet guidance (face anchors, material, temperament, multi-view consistency), scene and prop guidance, and video style tags. The same style path feeds character art, scene art, prop art, and video prompts, so the show does not accidentally switch from anime to realism between assets.

The two-step character art process

  1. Text polish: turn the archive description into an art-ready prompt using the chosen style manual, with gender treated as a hard rule.
  2. Image generation: generate the final sheet, support reference images, and save versions as reviewable history.

Only completed, readable art assets should be used downstream. Half-finished sheets should not be bound into shots.

The reference image rule that prevents face and costume drift

This is one of the most important rules in AI video production:

When a character has a reference image bound to a scene, the prompt must not re-describe that character’s clothing or appearance in text. The reference image owns appearance; text only describes action, expression, and injury state.

If the prompt says "woman in red dress" while the reference shows a blue coat, the model has to choose, and it will choose inconsistently. The pipeline must prevent that conflict.

Scene and prop discipline

  • Scene art should be empty of people; it is an environment plate, not a poster.
  • Props should be extracted from the script using the original names the script uses, then cataloged across episodes.
  • Manual binding should be supported: if the script says "the officer," bind it to the named character in the archive.
  • Props can be manually included or excluded per scene so irrelevant items do not crowd the reference set.
  • For time-travel or flashback stories, period looks should be selected by scene context, not randomly.

Before filming, the pipeline should list any missing scene, character, or prop references. Skipping is allowed, but it should be a deliberate choice with a visible warning, not a silent fallback to lower quality.

Stage 9: Scene blocks

A vertical drama episode should not be filmed as one long generation. It should be split into scene blocks, roughly 10 seconds each, with a soft cap around 200 characters of script body per block. Longer scenes are split further by action beats, paragraph breaks, or sentence endings.

This is the equivalent of a clip list on an editing timeline. Each block has its own prompt, its own reference set, its own rendered takes, and its own version history.

Why this matters:

  • shorter generations are more stable;
  • bad blocks can be reshot without redoing the whole episode;
  • editing becomes assembly instead of rescue work;
  • reference images can be bound per scene, not guessed for a full episode.

Crowd or generic characters should be separated from the drawable main cast so they do not consume character reference slots.

Stage 10: Shot prompts and filming

A good video prompt is not a paragraph of prose. It is a shot instruction.

The eight elements of a controlled shot prompt

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color tone
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Constraints

For simple scenes, one paragraph is enough. For complex cinematic scenes, use three layers: overall setup, shot-by-shot instructions, and a constraint pack.

Prompt rules that reduce failure

  • One camera move per shot; do not stack push, pull, pan, and tilt together.
  • Use shot numbers, not absolute timestamps like "0–3s."
  • Include a baseline constraint pack: quality, facial stability, no watermark, no logo.
  • For multi-person scenes, add twin / duplicate prevention constraints.
  • Keep actions continuous and quantified; avoid explosive motion that models often break.
  • Use clear notation: dialogue in braces, sound effects in angle brackets, BGM in parentheses.
  • Feed only the current scene’s assets; do not leak other scenes’ references.
  • Keep generation settings conservative: stable and controllable beats wildly unpredictable.

The system should teach the model to speak in shot-list language, not novel language.

Before prompts are sent to video generation, protected IP names should be stripped while preserving technique and aesthetic description. If a reference image is detected as a real-person photo, the pipeline should warn clearly and route the user back to drawn assets, because binding real photos can create unstable or blocked outputs.

Stage 11: Review, reshoot, and delivery

Filming should work like a real set: the director reviews the shot, changes the instruction if needed, and shoots another take.

The production controls that matter most are:

ControlSet equivalent
Edit the video promptDirector revises shot notes
Swap character / scene / prop referencesChange costume sheet or location plate
Manually bind a characterFix name mismatch or offscreen reference
Include / exclude a propControl visual focus
Switch style manualUnify art language
Choose model tierBalance quality, speed, and cost
Review historical takes per scenePick the best take

A new take should only happen when the prompt or references change. Re-clicking without a change is just duplicate rendering. On the production side, tasks should be queued independently from script and art tasks, with no parallel submissions for the same block, pre-charge settlement on failure, timeout recovery for stuck jobs, and polling continuation when a vendor task already exists.

Default delivery is 9:16 vertical, not a horizontal video cropped after rendering. Supported controls include smart or fixed duration around 5–15 seconds, model tiers, resolutions within model allowlists, optional audio, and watermark off by default.

What this pipeline does not do

Honest limits matter. Teams evaluating tools should ask what is automated and what still belongs to human creators.

  1. There is no automatic quality score or auto-select best take engine. Final quality judgment stays with the creator and producer; the pipeline provides multiple takes and reshoot tools.
  2. Reference images are not a hard gate. Missing references trigger a warning, but text-only filming is allowed—usually with weaker consistency, so professional workflow should lock looks first.
  3. Reference markers in prompts depend on prompt discipline, not hidden magic. Creators should still review that references are present.
  4. Shot grammar follows the built-in camera spec. Style manuals mainly inject visual style tags; full art-book control lives on the asset side.
  5. The ~10-second scene block is an engineering heuristic, not frame-accurate editing. Long scenes may be split again, but final trimming and assembly still belong in editing.
  6. There is no current product workflow for automatic cross-scene video continuation. The unit is one scene block with references, rendered to one clip.
  7. Character consistency depends on the look-dev asset chain, not face-embedding verification. Period routing is rule-based, and final look quality still depends on good sheets and prompt discipline.

A better mental model for AI short drama production

The winning workflow is not "one prompt, one masterpiece." It is:

  • lock the brief;
  • build the archive;
  • approve the cast and looks;
  • split the episode into controllable blocks;
  • bind references per block;
  • write shot instructions instead of prose;
  • render multiple takes;
  • reshoot only the blocks that fail.

High-volume short drama is not produced by one long generation cut into pieces. It is produced by stacking controlled units that can be reviewed, revised, and re-rendered independently.

If you are looking for this kind of structured workflow, Maosika (猫斯卡) is built as an AI production operating system for vertical short dramas: it formalizes the sequence from idea intake to script batches, look development, scene blocks, shot prompts, and filmed takes, while keeping human approval at the creative checkpoints.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com