How AI Vertical Short Dramas Actually Get Made: A Stage-by-Stage Production Pipeline With Checkpoints
Most AI short drama failures are not model failures. They are order failures: teams render before the brief is locked, shoot before characters are designed, and extend scenes before continuity is tracked. The fix is a staged pipeline with explicit acceptance points.
If you have watched AI-generated vertical dramas, you have seen the symptoms: a lead who changes face mid-episode, props that appear and disappear, cliffhangers that forget their own setup, and episodes that feel like one long prompt instead of a scripted show. These are rarely fixed by switching models. They are fixed by changing the order in which work happens.
A production-ready AI vertical short drama pipeline is not a single "generate video" button. It is a sequence of handoffs, each producing something reviewable before the next stage starts. The stages are: idea evaluation, guided creative intake, brief lock, cast and visual confirmation, story archive setup, batch script writing, style selection, character/scene/prop look development, episode splitting into scene blocks, per-scene reference binding and shot prompts, multimodal rendering, and review with retakes.
Why order matters more than model choice
In traditional crews, you do not send actors to set before the script is approved. You do not build props before you know which scene uses them. You do not ask the editor to invent continuity that was never shot. AI workflows need the same discipline, but the temptation to skip stages is stronger because generation feels instant.
When stages are skipped, the model is forced to make uncontrolled decisions:
- If the brief is vague, the writer invents tone and ending direction episode by episode.
- If characters have no locked design, each scene re-imagines their face and outfit.
- If there is no continuity record, later batches forget earlier promises.
- If scenes are not split into units, one long prompt tries to cover too many actions and the result drifts.
- If shot prompts are written like fiction instead of shot instructions, the model invents camera moves and timing.
The principle is simple: every stage should produce a deliverable a human can inspect, reject, or approve.
Stage 1: Intake and creative brief lock
The first checkpoint is not writing. It is deciding whether the idea is ready to write.
A useful intake covers the vertical-drama dimensions that actually shape output: genre, protagonist, core conflict, story direction, episode count, episode length, tone, payoff and hook rhythm, ending direction, target platform, and audience. Episode length usually lands in the 1–2 minute range for vertical dramas, with some projects extending to 3–4 minutes.
A practical gate is to score intake completeness and route accordingly. If the idea is well-formed, the team can move quickly to a locked brief. If key dimensions are missing, those gaps should be filled before any script is drafted. If the idea is still thin, it needs a fuller development conversation.
The deliverable at this stage is a locked creative brief. It should include:
| Brief field | What it locks |
|---|---|
| Logline | The one-sentence promise of the show |
| Core conflict | What keeps episodes turning |
| Story direction | Where the plot is heading |
| Ending direction | Whether it resolves, reverses, or opens a new arc |
| Payoff and hook rhythm | How often reversals, reveals, and cliffhangers land |
| Platform and audience | Format, pacing, and expectation baseline |
| Episode breakdown | Multi-episode outline, not just a premise |
| Notes for the writer | Constraints, taboos, and required beats |
| Protagonist arc | Who the lead is at start and how they change |
This stage is complete only when the brief is locked. Writing before lock is the fastest way to create continuity debt.
Stage 2: Story archive and continuity setup
A story archive is a structured continuity record for the whole drama. It tracks character identities, stable traits, current state such as injuries or hidden identity, relationships, open and resolved plot threads, episode-by-episode appearance, batch summaries, and visual descriptions of recurring props.
This is not a nice-to-have document. It is the mechanism that prevents episode 47 from contradicting episode 5.
The archive serves two purposes:
- It gives the writing process only the slice of information needed for the current batch: relevant character states, unresolved threads, recent batch summaries, and the current main line.
- It gets updated after each batch so future writing starts from a known state rather than model memory.
The checkpoint here is simple: before writing begins, the archive exists. After each batch, the archive is updated. If a thread is opened, it is marked open. If it is resolved, it is marked resolved.
Stage 3: Batch script writing with structure rules
Vertical drama writing works best in batches rather than one giant generation. A batch covers several episodes, gets reviewed, updates the archive, and then the next batch begins.
Before writing a batch, the intended direction should be locked: what happens in this batch, which characters carry it, which threads advance, and how it ends. These become hard constraints for that batch.
Each episode should follow a tight vertical-drama structure:
- A strong hook in the first scene, usually within the first few seconds
- Escalating conflict through the middle
- A cliffhanger or forced question in the final scene
A few writing rules are especially important for vertical format:
- Open with conflict, not exposition. The first scene should not explain background. It should create a problem, shock, reveal, or power move.
- Plan before prose. For each episode, produce a beat sheet and the ending cliffhanger before writing full dialogue.
- Keep payoff frequent. Every episode should have at least one small payoff: a face-slap, reversal, identity hint, evidence gain, or power reveal. Larger payoffs land every few episodes.
- Write short lines. Dialogue should be concise and speakable. Long explanatory speeches kill vertical pacing.
- Carry continuity across episodes. The next batch should read the previous batch and inherit its consequences.
After writing, scripts should pass a rule-based quality check before delivery. Common failures include wrong episode titles, mismatched scene counts, missing character labels, too little dialogue, and placeholder text such as "to be continued." A failed script should be sent back for rewrite with the error attached.
The checkpoint: no script moves forward as a "maybe done" draft. It either passes structure checks or gets rewritten.
Stage 4: Style selection and look development
Before video generation, the show needs a locked visual language. This includes character designs, recurring locations, and recurring props.
A mature pipeline offers multiple style manuals covering 2D, 3D, and realistic directions: urban realism, period realism, mature urban romance animation, 1990s anime, Chinese brushwork, xianxia ancient style, 3D donghua, clay stop-motion, cyber-Chinese fusion, and others. The key is not the number of styles but consistency across assets: once a style is chosen, character art, scene art, prop art, and video style tags should follow the same visual path.
Look development should happen in two steps:
- Text refinement: Turn archive descriptions into art prompts using the chosen style manual and hard rules such as gender. This stage can be manually overridden by the creator.
- Image generation: Generate the actual design, optionally using reference images, and save versions to the asset library.
Character confirmation is a hard gate. Before moving on, the cast should be complete, names valid, visual fields complete, and leads aligned with the brief. A half-designed cast should not be sent to shooting.
The checkpoint: characters, recurring scenes, and recurring props have finished, readable art assets. Video should not consume half-generated placeholders.
Stage 5: Episode splitting into scene blocks
A vertical episode should not be treated as one video generation task. It should be split into scene blocks, usually around 10 seconds each. Long passages can be split further by action beats, paragraph breaks, or sentence boundaries.
A scene block is the basic production unit. It has its own script section, reference images, shot prompt, rendered output, and take history. This is closer to a clip list on an editing table than a single full-episode render.
Why this matters:
- Short blocks are easier to control.
- Failures are isolated; one bad shot does not ruin the episode.
- Retakes are cheaper because you reshoot one block, not everything.
- Reference binding is clearer because each block has a smaller cast and prop set.
The checkpoint: the episode is divided into blocks, and each block can be reviewed independently.
Stage 6: Reference mapping and shot prompts
This is where most AI consistency issues are either prevented or created.
For each scene block, build a reference table in a fixed order: scene image first, then props for that scene, then character designs. Only existing assets occupy slots. If there is no image, the slot stays text-only rather than inventing a false reference.
The most important consistency rule is this: if a character has a reference image, do not describe their clothing or appearance again in the prompt. The image is the authority. The text should describe only action, expression, and temporary state such as injury.
This rule directly prevents the common AI failure where text says one outfit and the reference shows another, causing costume swaps or face drift.
Other reference controls include:
- Manual character binding when script names differ from archive names
- Manual prop inclusion or exclusion to avoid visual clutter
- Ability to swap to another historical version of the same asset
- Era-based routing for time-travel or flashback scenes, so modern and period looks are matched correctly
Before prompts are generated, check whether required references are missing. The system should warn and offer a chance to create assets first. Teams can intentionally skip references and use text-only prompts, but quality is usually weaker, so professional workflows should finish look development before shooting.
Shot prompts should be written like production instructions, not fiction. A strong prompt includes eight elements:
| Element | Purpose |
|---|---|
| Precise subject | Who or what is in the shot |
| Action detail | What happens, with concrete motion |
| Scene environment | Where it happens |
| Lighting and color | Mood and visual tone |
| Camera movement | One clear move per shot |
| Visual style | The locked style path |
| Image quality | Stability and finish requirements |
| Constraints | What to avoid or preserve |
A few prompt rules reduce drift:
- One camera move per shot; do not stack push, pull, pan, and tilt together.
- Use shot numbers, not absolute timestamps.
- Keep actions continuous and controlled; extreme motion often breaks.
- Include a constraint package for face stability, watermark avoidance, and multi-character duplication prevention.
- Mark dialogue, sound effects, and music with clear symbols.
- Use only the current scene's references; do not pull assets from other scenes.
The checkpoint: before rendering, a human can read the prompt, see the bound references, and edit either one.
Stage 7: Rendering, review, and retakes
The shooting stage should behave like a production queue, not a toy script.
A standard path is: open the episode workspace, confirm references are present or intentionally skipped, generate shot prompts, review and edit prompts, choose model tier and output settings, submit to render, and then compare takes for each block.
Default delivery is vertical 9:16, not a horizontal video cropped afterward. Output settings usually include smart or fixed duration around 5–15 seconds, model tier selection for quality/speed tradeoffs, resolution within model support, optional audio, and watermark off by default.
Professional control points at this stage include:
| Control | Production equivalent |
|---|---|
| Editing shot prompts | Director revising shot notes |
| Replacing reference images | Changing costume/look or location board |
| Manual character binding | Fixing name mismatches or cameo references |
| Including/excluding props | Controlling visual focus |
| Switching style | Unifying the art language |
| Choosing model tier | Balancing quality, speed, and cost |
| Reviewing historical takes | Selecting the best take per setup |
A serious pipeline should also handle production reliability: failed jobs should be visible and retryable, in-progress blocks should not allow duplicate submissions that waste resources, queues should separate rendering from writing and art tasks, and successful renders should be stored with actual playable duration rather than trusting vendor metadata alone.
The checkpoint: each block is reviewed as a take. If it fails, edit the prompt or references and reshoot that block. Do not ask a later stage to fix a bad block automatically.
What this pipeline does not do
It is important to be honest about the boundaries.
There is no magic "auto-pick the best shot" engine. Final quality judgment still belongs to creators and producers. The system provides multiple takes and tools to retake, but it does not replace taste.
Reference images are not a forced blockade. You can skip them and render text-only, but results are usually less consistent. Professional teams should design first, shoot second.
Scene blocks are an engineering heuristic around 10 seconds, not frame-accurate timecode editing. Some blocks may run longer, and final episode assembly still needs editing.
Character consistency depends on the asset pipeline, not a hidden face-verification trick. If the design is weak or the prompt ignores the "do not redescribe appearance" rule, drift can still happen.
The product unit is one scene block with references going to one rendered clip. There is no mature cross-scene automatic video extension workflow; longer sequences are built from controlled blocks.
A useful acceptance checklist
Use this checklist before moving from one stage to the next:
- Brief locked: logline, conflict, ending direction, hook rhythm, platform, and protagonist arc are approved.
- Archive created: characters, relationships, threads, and recurring props are recorded.
- Batch planned: direction, key characters, threads, and ending hook are confirmed before writing.
- Scripts checked: hook, structure, dialogue, formatting, and cliffhanger pass review.
- Style chosen: one visual path applies to characters, scenes, props, and video tags.
- Assets ready: character, scene, and prop art are complete and readable.
- Episode split: each scene block is short enough to control.
- References bound: scene, props, and characters are mapped; text does not redescribe designed appearances.
- Prompts reviewed: shot language is concrete, one move per shot, constraints included.
- Takes reviewed: bad blocks are reshot; good blocks move to edit and assembly.
High-quality AI vertical dramas are not produced by one long, impressive generation. They are produced by stacking controlled units, each with a clear owner, a clear deliverable, and a clear checkpoint. The model generates. The pipeline prevents drift. The human still decides what is good enough.
Maosika (猫斯卡) is built around this staged logic: an AI production operating system for vertical short dramas that turns the repeatable, drift-prone parts of production into a reviewable pipeline while keeping creative judgment with the creator.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com