The AI Vertical Short Drama Pipeline: 12 Checkpoints From Idea to Finished Episode
Most AI short drama failures are not model failures — they are skipped checkpoints. A reliable pipeline turns a loose idea into an approved episode by forcing sign-offs at 12 specific gates before anything is rendered.
Why checkpoints matter more than prompts
When a team says an AI drama "fell apart," they usually mean one of a few things: the lead changed face, episode 6 forgot what episode 2 established, the tone shifted from urban realism to anime, or the ending had no relationship to the hook. These are rarely single-generation accidents. They are the result of moving forward without a checkpoint.
A checkpoint is a deliverable you can look at, reject, revise, and lock before the next stage consumes it. In live-action production, this is normal: brief, casting, wardrobe, shot list, shoot, dailies, edit. In AI workflows, teams often skip straight from idea to video and then try to fix continuity in post. That is the most expensive place to fix it.
An AI vertical short drama production pipeline is a sequence of locked intermediate deliverables, where each stage only consumes approved assets from the previous stage.
Below is a 12-checkpoint structure built for vertical short dramas in the 1–2 minute per episode range, with 9:16 as the default frame. It is not a "one-click hit" formula. It is a control structure: the repetitive, drift-prone parts become a repeatable line, while taste, story judgment, and final approval stay with the creator.
The 12 checkpoints
The table below is the whole map. Each section after it explains what to actually check.
| # | Checkpoint | What gets locked | What you should reject here |
|---|---|---|---|
| 1 | Idea intake score | Completeness of the concept | Vague loglines, undefined conflict |
| 2 | Guided brief expansion | Missing dimensions | Skipped audience, tone, or ending direction |
| 3 | Creative brief sign-off | Logline, conflict, arc, hook rhythm | "We'll figure it out later" briefs |
| 4 | Cast & visual confirmation | Named roles, protagonist arc | Incomplete cast, no visual fields |
| 5 | Story archive creation | Continuity bible baseline | No record of open threads or relationships |
| 6 | Beat sheet per batch | Scene list + cliffhanger first | Writing dialogue before structure |
| 7 | Script batch + rule check | Approved scenes per batch | Missing cast lines, placeholder endings |
| 8 | Style manual selection | One visual language for the whole show | Mixed styles across assets |
| 9 | Character / scene / prop art | Locked reference images | People in scene plates, missing props |
| 10 | Scene block split | ~10-second shootable units | Unsplit 2-minute mega-scenes |
| 11 | Engineered shot prompts | Per-block camera instructions | Novel-style prompts, stacked moves |
| 12 | Shoot, review, re-take | Final per-block video | Accepting the first take blindly |
---
1. Idea intake score
Before anyone writes a scene, the concept itself gets scored on completeness. A useful intake captures the genre, protagonist, core conflict, direction of the story, episode count, episode length, tone, hook rhythm, ending direction, platform, and target audience.
The point is not bureaucracy. The point is to route the idea correctly:
- A near-complete brief can go almost straight to sign-off.
- A partial brief needs targeted follow-ups.
- A one-sentence pitch needs a full guided development path.
Reject here: a logline like "A woman gets revenge" with no antagonist, no stakes, no platform, no tone. You cannot lock art or cast from that.
2. Guided brief expansion
This is the development meeting, compressed. The missing dimensions from the intake get filled in one by one. For vertical dramas, two dimensions are repeatedly under-specified and cause the most damage later:
- Hook rhythm: how often a reversal, reveal, or small payoff lands.
- Ending direction: whether the show ends in triumph, bittersweet closure, open sequel bait, or tragic irony.
If ending direction is left open, every batch of episodes will pull toward a different genre.
3. Creative brief sign-off
The brief is the first hard lock. It should contain at minimum: the logline, core conflict, story direction, ending direction, beat and hook rhythm, target platform and audience, episode-by-episode outline, and notes for the writer. The protagonist entry must include a character arc, not just a job title.
Treat this as a greenlight document. After this point, "I thought the show was about something else" is not a note you want to be giving at episode 40.
4. Cast & visual confirmation
Before any art is generated, the cast list must be complete: names are valid, visual fields are filled, and the leads match the brief. This is a hard gate, not a suggestion.
A common failure is generating art for a cast list that still says "Male Lead," "Boss," "Mysterious Woman." Those are not characters; they are placeholders, and every downstream asset will inherit that vagueness.
5. Story archive creation
A story archive is a structured continuity record covering identity, stable traits, current state, relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual notes. It is the production bible that survives across episodes.
The rule is simple: don't rely on model memory; rely on a structured archive. When a new batch is written, it should only consume the relevant slice — current character state, unresolved threads, recent batch summary, and the current batch's main line. After each batch, the archive is updated.
6. Beat sheet per batch
Scripts should not be written as one long wall of dialogue. For each batch of episodes, the scene list (beat sheet) and the episode-end cliffhanger come first, then the dialogue.
For vertical short dramas, the internal rhythm of a single episode is fairly strict:
- A strong hook in the very first scene — the "golden 3 seconds" rule, no slow exposition.
- Escalating conflict through the middle.
- A cliffhanger or reversal in the final scene.
Approve the skeleton before you approve the prose.
7. Script batch + rule check
Once a batch is written, it goes through a rule-based quality check before anyone sees it as a deliverable. Typical failures to catch:
- Wrong episode title or scene count mismatch.
- Missing cast headers.
- Too little dialogue, too much narration.
- Placeholder endings like "to be continued" used as a cliffhanger substitute.
A failed batch should be sent back with the specific error, not handed to a human to manually patch. The goal is that no half-finished script crosses the desk.
8. Style manual selection
One visual language must be chosen for the entire show. A useful style system includes manuals across 2D, 3D, and live-action-realistic directions — urban realism, period realism, mature urban romance animation, 90s anime, ink-wash guofeng, xianxia, 3D donghua, stop-motion clay, cyber-guofeng, and so on.
Crucially, the chosen style should drive character art, scene art, prop art, and video prompts through the same visual path. The classic disaster — "2D characters in a 3D realistic world" — happens when these are treated as unrelated choices.
9. Character / scene / prop art
This is the look-dev stage. Three disciplines are locked separately:
- Character art: final approved portraits, with gender treated as a hard rule, not a suggestion.
- Scene art: empty plates — strictly no people in them.
- Prop art: extracted from the script using the script's own naming, then cataloged across all episodes.
A two-step art process works well here: first, the archive description is polished into an art-ready prompt using the chosen style manual (with manual override allowed); then the image is generated and stored as a versioned asset. Video should only consume art that is complete and readable — never half-finished drafts.
10. Scene block split
Each episode is split into video blocks, targeting roughly 10 seconds per block. A block is the unit of shooting: one set of references, one prompt, one rendered clip, one history of takes.
This is the equivalent of the clip list on an editing timeline. You are not producing "Episode 3, one big file." You are producing Episode 3 as Block 1, Block 2, Block 3… each independently re-renderable.
Why ~10 seconds? It is an engineering heuristic, not a timecode-precise cut. Longer scenes get split further on action beats and punctuation. It is still possible to end up with a slightly long block; fine trimming and final assembly belong in a later edit.
11. Engineered shot prompts
This is where most prompt-craft advice fails creators. A production shot prompt is not a paragraph of novel prose. It is a shot list instruction, and it should contain eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color
- Camera movement
- Visual style
- Image quality
- Constraints
For simple scenes, a single-paragraph prompt works. For complex cinematic scenes, use a three-part structure: overall setup, then shot-by-shot instructions, then a constraint pack.
A few hard rules that prevent common failures:
- One camera move per shot — no "push in while panning while tilting."
- Use shot numbers, not absolute timestamps like "0–3s."
- Always include a fallback pack: quality, face stability, no watermark; for multi-person shots, add twin/duplicate prevention; for stylized shows, anchor the style explicitly.
- Keep actions low and continuous rather than high-motion and explosive.
- Mark dialogue, sound effects, and BGM with consistent symbols so they are not confused with visual instructions.
- Only feed the current block's own assets into the prompt — no cross-scene bleeding.
The system should be teaching the model to speak in shot-list language, not in novel language.
12. Shoot, review, re-take
The final stage is production, not magic. For each block:
- Confirm references are present (or deliberately skipped).
- Generate the shot prompt.
- Read and edit it — this is the director revising the shot list.
- Choose model tier, aspect ratio, resolution, and duration.
- Submit to the render queue.
- Review the take.
- If needed, change the prompt or swap a reference and shoot again.
There is no automatic "best take" picker. Final quality judgment sits with the creator. What the pipeline provides is multiple takes, deterministic re-shoots, and a queue that behaves like production infrastructure: failed jobs are visible and retryable, in-flight blocks cannot be double-submitted, and credits are reserved and released cleanly.
The reference image rule that prevents face swaps
One mechanism deserves its own section because it solves the single most common AI drama complaint: character inconsistency.
When a block has a reference image for a character, the prompt must not re-describe that character's clothing or appearance in text. The reference image is the authority; text only describes action, expression, and visible injury.
The reference stack for each block is assembled in a fixed order:
- Scene reference
- Props present in this block
- Character locked art
Slots only fill if an image exists. If there is no image, that slot is marked as text-only rather than inventing a binding. Creatures of habit will want to manually bind names ("the officer" in the script to "Li Qiang" in the archive), include or exclude props to control visual focus, or swap in an older version of an asset. For time-travel or flashback stories, era-appropriate looks can be selected based on scene keywords.
Consistency comes from assets, not from luck.
Where human judgment still lives
It is worth being explicit about what this pipeline does not do, because credible production tools have boundaries:
- There is no automatic scoring engine that picks the "best" take or re-shoots by itself. Final approval is human.
- Reference images are not a hard gate. You can skip them and go text-only — but quality is usually worse, so professional workflow should lock art first.
- The ~10-second block split is a heuristic, not a precision edit. Final trimming and stitching still belong in edit.
- The unit of production is a single block with its own references. There is no productized cross-block auto-extend workflow today.
- Character consistency depends on the art pipeline, not on a face-embedding verification layer. Era routing is rule-based; final look still depends on the quality of the locked art and on prompts obeying the "don't describe what the reference already shows" rule.
These are not flaws to hide. They are the boundary between "toy that generates a clip" and "operating system you can run a slate on."
A note on the production line
The structure described above is the workflow that Maosika (猫斯卡) has built into its AI vertical short drama production operating system: an end-to-end chain from idea intake through guided brief, archive, batched script writing, locked art, scene blocks, engineered shot prompts, and queued multi-take shooting, with a digital crew of 18 specialist roles handling the steps between human sign-offs. The product does not promise one-click hits; it enforces the order of operations that experienced production teams already follow, so that what gets repeated across dozens of episodes is the discipline, not the mistakes.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com