AI Vertical Short Drama Production Pipeline: 12 Checkpoints From Idea to Final Cut
Most AI short drama failures are not model failures—they are skipped checkpoints. A production-grade pipeline turns a vague idea into a locked brief, a continuity bible, locked looks, shot-blocked scenes, and reviewable takes, with a human decision at every gate.
Why a checkpoint model beats "generate and hope"
Vertical short drama teams often describe the same pain pattern: the first episode looks promising, episode 5 changes the lead's face, episode 12 forgets a key clue, and by episode 30 the tone has drifted into a different show. The usual response is to blame the video model. In practice, the failure usually happened much earlier—at the brief, the cast lock, the continuity record, or the reference mapping stage.
A checkpoint is not a formality. It is a point in the pipeline where the team receives a concrete artifact, reviews it, and either approves it or sends it back before money is spent on the next stage.
Continuity bible is a structured story record that tracks character identities, current states, relationships, open and resolved plot threads, episode-by-episode appearance tables, and batch-level story notes. It exists so that episode 200 still remembers what episode 3 established.
This article maps the AI vertical short drama workflow into 12 checkpoints. The order matters because later stages consume earlier artifacts. Skipping a checkpoint does not save time; it converts a cheap text revision into an expensive visual reshoot.
The 12 production checkpoints
The pipeline below is organized the way a real production room thinks: greenlight first, then story, then cast and look, then shot preparation, then shooting, then review. Each checkpoint answers one question: "What must be true before we spend the next unit of budget?"
| # | Checkpoint | Reviewable artifact | Main failure if skipped |
|---|---|---|---|
| 1 | Concept intake scoring | Intake score and missing-dimension list | Team builds on a half-formed idea |
| 2 | Guided creative development | Filled creative brief | Wrong genre, tone, ending, or hook rhythm |
| 3 | Brief lock | Approved logline, conflict, hooks, audience, episode outline | Constant rewrites downstream |
| 4 | Cast and visual confirmation | Named cast with complete visual fields | Character drift from episode 1 |
| 5 | Story archive creation | Continuity bible baseline | Forgotten clues and broken arcs |
| 6 | Batch script planning | Beat sheet and cliffhanger per episode | Scenes that do not end on a hook |
| 7 | Script draft and rule check | Passed script batch | Flat openings, weak dialogue, placeholders |
| 8 | Archive update after batch | Updated continuity bible | Next batch loses context |
| 9 | Style and look development | Character, scene, and prop art in one style | Mixed art languages across assets |
| 10 | Scene blocking | ~10-second scene blocks with references | Long uneditable clips and crossed references |
| 11 | Engineered shot prompts | Prompt per scene with reference map | Model invents costumes, faces, props |
| 12 | Shoot, review, retake | Multiple takes per scene block | No way to compare or selectively reshoot |
Checkpoint 1: Concept intake scoring
Before any writing begins, the idea should be scored for completeness. A useful intake covers genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, platform, and target audience.
A strong intake score means the team can move almost directly to brief drafting. A weak score means the idea needs guided development. The point is not gatekeeping; it is preventing the writer from inventing missing business decisions on the fly.
Checkpoint 2: Guided creative development
When the intake is incomplete, the next step is not a script—it is a structured development pass. This is where the team answers the questions that a real development meeting would answer: who the lead is, what they want, what is blocking them, how the story escalates, what the audience gets hooked on, and how the ending lands.
For vertical drama, this stage must also pin down episode economics: episodes are typically 1–2 minutes, with an outer limit around 3–4 minutes, and the first three seconds must earn the swipe-stay.
Checkpoint 3: Brief lock
The creative brief is the first hard gate. Until it is locked, no script should be written.
A locked brief includes the logline, core conflict, story direction, ending direction, satisfaction beats and hook rhythm, platform and audience, episode-by-episode outline, and notes for the writer. The protagonist section must include the character arc, not just a job title and a costume.
Think of this as the end of the greenlight meeting. After this point, changes are possible, but they are conscious revisions, not drift.
Checkpoint 4: Cast and visual confirmation
Before art generation, the cast must be complete, names must be valid, visual fields must be filled, and leads must match the brief. This is not cosmetic. If a lead's visual identity is vague at this stage, every later scene will reinterpret them.
This is also the right stage to resolve naming problems: a script may call someone "the captain," "Officer Li," and "that woman in black," but the production archive needs one canonical identity.
Checkpoint 5: Story archive creation
Before the first script batch, build the baseline continuity bible. At the start it contains character identities, stable traits, initial states, relationships, open plot threads, and the planned appearance table. As batches are written, it accumulates batch summaries and prop visual notes.
The principle is simple: do not rely on model memory. Rely on a structured archive. When a new batch starts, the writer consumes a slice of that archive—current character states, unresolved threads, recent batch summaries, and the current batch's main line—instead of dumping the entire history into context.
Checkpoint 6: Batch script planning
Scripts should be written in batches, with planning before prose. For each episode, the beat sheet and the episode-end cliffhanger are defined first. Only after that does the dialogue draft begin.
Vertical drama has hard rhythm rules that belong at this planning stage:
- The first scene must open on strong conflict or suspense; no slow background dump.
- Each episode follows hook → escalation → cliffhanger.
- Every episode needs at least one small satisfaction beat: a reveal, a reversal, an identity hint, a piece of evidence.
- A larger payoff lands every few episodes.
- Dialogue stays short, generally under twenty words per line, with no essay-like narration.
Checkpoint 7: Script draft and rule check
After drafting, scripts pass a rule check before anyone sees them as deliverables. The check catches episode title errors, scene count mismatches, missing character lines, too little dialogue, and placeholder text such as "to be continued." A failed draft is sent back with the specific error and rewritten until it passes or hits a retry ceiling.
This is one of the highest-ROI checkpoints in the whole pipeline. Text fixes are cheap; reshooting a broken scene is not.
Checkpoint 8: Archive update after batch
When a batch passes, the archive must be updated before the next batch starts. New plot threads are marked open, resolved threads are closed, character states change, and a batch summary is recorded.
From the second batch onward, the next planning pass should confirm four things before writing: the batch's plot direction, focus characters, threads and conflicts, and end-of-batch hooks. Those confirmed intentions become hard constraints. If they conflict with the old archive, the new creative decision wins—because the archive should serve the story, not freeze it.
Checkpoint 9: Style and look development
Once scripts are stable, choose a visual style and generate look-development assets. A production-ready system should offer multiple style manuals covering 2D, 3D, and realistic directions, and the chosen style must be shared across character art, scene art, prop art, and video prompts. A character drawn in anime style but shot in realistic video is a pipeline break, not an aesthetic choice.
Scene art has one discipline worth stating explicitly: establishing shots and scene plates should not contain characters. Prop extraction should follow the script's own naming, viewed through a prop master's eye, so that later scenes can reference the same object by the same name.
Checkpoint 10: Scene blocking into ~10-second blocks
Before video generation, each episode is divided into scene blocks targeted around ten seconds. Long passages are split by action beats, paragraph breaks, and sentence boundaries. The result is not "one big video per episode." It is a clip list: episode 3 becomes scene 1, scene 2, scene 3, and so on, each with its own prompt, references, outputs, and retake history.
This is where vertical delivery is decided. The default should be native 9:16 vertical, not a horizontal export cropped after the fact.
Checkpoint 11: Engineered shot prompts with reference mapping
Each scene block gets a prompt built from production materials, not invented from scratch. A strong prompt includes eight elements: precise subject, action detail, scene environment, lighting and color, camera movement, visual style, image quality, and constraints.
The reference map for each scene follows a fixed priority: scene image → scene props → character looks. If a reference exists for a character, the prompt must not re-describe that character's clothing or appearance in text. The text only describes action, expression, and injury state. That rule directly prevents the classic AI video failure where the text prompt fights the reference image and produces a costume or face swap.
Complex scenes use a three-part structure: overall setup, shot-by-shot instructions, and a constraint pack. One shot gets one camera move; do not stack push, pull, pan, and tilt into a single instruction. Use shot numbers instead of hard-coded timestamps. A constraint pack should cover face stability, image quality, no watermark, twin or duplicate prevention in multi-character scenes, and style anchoring for non-realistic looks.
Before prompts are finalized, there should be a readiness check for missing scene, character, and prop references. Missing references can be skipped deliberately, but the team should know they are skipping them.
Checkpoint 12: Shoot, review, and retake
The shooting stage should behave like a real set, not a black box. The team reviews the prompt, can edit it, can swap reference images, can manually bind a script name to a cast character, can include or exclude props, can choose model tier and resolution, and can submit the shot. Outputs land in a queue, and each scene block keeps a history of takes.
Crucially, reshooting means changing the prompt or references and generating a new take. The system should not silently rewrite the prompt between takes. That would be equivalent to a director changing the shot list without telling the crew.
Production reliability matters here too: failed jobs should be visible and retryable, in-flight scenes should not accept duplicate submissions, credits should be reserved and released cleanly, and finished videos should be stored with actual measured duration rather than trusting a vendor callback.
Where the human still decides
A checkpoint pipeline does not remove human judgment. It moves judgment to the place where it is cheapest and most powerful.
- The creator decides whether the brief is locked.
- The writer or producer decides whether a batch's cliffhangers land.
- The art lead decides whether a character look is approved.
- The director decides whether a shot prompt needs editing.
- The editor or producer chooses the best take.
The system handles the repetitive, drift-prone parts: maintaining continuity, mapping references, enforcing shot structure, running rule checks, queuing jobs, and preserving history. Taste still sits with the person making the show.
Honest boundaries
No pipeline removes the limits of current AI video production. Teams should plan around them explicitly:
- There is no automatic quality score that magically picks the best take. Final review stays with the creator.
- Reference images are not a hard gate. You can shoot without them, but quality usually drops; professional process locks looks first.
- Prompt reference markers depend on disciplined formatting, so creators should still verify them before shooting.
- The ~10-second block target is an engineering heuristic, not a timecode-accurate cut; long scenes may still need trimming.
- The current unit of work is a single scene block with references, not an automatic cross-scene video continuation.
- Character consistency depends on the look-development chain and prompt discipline, not on a magic face-lock guarantee.
Maosika (猫斯卡) is designed as an AI production operating system for vertical short dramas. It does not promise one-click hits or zero-error output. It固化 the professional sequence—story and continuity first, then locked looks, then scene blocking, then shot instructions, then multiple takes—so teams can scale production without handing creative control to a black box.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com