The Vertical AI Short Drama Pipeline, Step by Step: 12 Checkpoints From One Line to a Deliverable Cut
Most AI short drama teams fail not for lack of a video model, but because they skip the order of operations. The reliable path is: lock brief, build continuity, cast looks, split into short scene blocks, then shoot and re-take.
The pipeline is not one big generate button
A vertical AI short drama is not a single prompt that turns into a finished show. It is a chain of handoffs, and every handoff needs something you can look at, edit, and roll back. In real production, the expensive failures are almost never "the model is bad." They are: the brief drifted in episode 7, the lead changed face in episode 12, a prop vanished, the opening had no hook, or episode 20 forgot a thread planted in episode 3.
The working method used by teams that can actually ship is closer to a small film crew than to a chatbot. You do not start with camera moves. You start with a locked brief, a continuity record, approved looks, and a scene-by-scene shooting plan. Then you generate in small units, review, and re-take.
A production pipeline for vertical AI short drama is a sequence of checkpoints where each stage produces an inspectable artifact before the next stage is allowed to consume it.
This article walks through that pipeline in 12 checkpoints, using the language of the set rather than the language of models. The goal is not to make AI sound magical. The goal is to make the process legible enough that a writer, director, producer, or studio owner can see where quality is won and lost.
Why vertical short drama needs stricter process than horizontal video
Vertical micro-drama has a few structural constraints that punish sloppy workflow harder than long-form video:
- The first three seconds must hook. A slow establishing shot is not an option.
- Episodes are short, usually in the 1–2 minute range, sometimes extending to 3–4 minutes.
- Viewers drop off fast if faces, costumes, or relationships drift.
- Cliffhangers carry the watch-through, so continuity across episodes is part of retention, not just craft.
- 9:16 vertical is the default delivery shape, not a last-minute crop.
This means the pipeline has to enforce three things that raw generation does not naturally enforce:
- Continuity across dozens or hundreds of episodes
- Visual consistency for characters, scenes, and props
- Scene-level control so you can re-shoot one bad beat without redoing the whole episode
High-quality short drama is never one long generation chopped up. It is a stack of controllable units.
The 12-checkpoint pipeline
Below is the order that survives contact with real production. The order matters more than the tooling.
| # | Checkpoint | Main artifact | Who decides |
|---|---|---|---|
| 1 | Idea intake & evaluation | Intake score, missing-dimension list | Producer / creator |
| 2 | Creative guidance | Filled brief dimensions | Creator + system prompts |
| 3 | Brief lock | One-page creative brief | Creator sign-off |
| 4 | Cast & visual confirmation | Character lineup, core visual direction | Creator / art lead |
| 5 | Story archive build | Continuity bible | Writing team |
| 6 | Batch script planning | Beat sheets, cliffhanger plan | Writer / producer |
| 7 | Script drafting & rule check | Passed scripts | Writer / quality gate |
| 8 | Style selection & look development | Style pack, character / scene / prop art | Art lead / director |
| 9 | Episode split into scene blocks | ~10-second scene blocks | Director / editor mindset |
| 10 | Shot prompt build per scene | Prompt + reference mapping | Director / prompt review |
| 11 | Multi-modal shoot | Raw takes per scene | Queue / render pipeline |
| 12 | Review, re-take, select deliverable | Final selected takes | Creator / producer |
Each checkpoint is explained below.
1. Idea intake & evaluation
The first job is not to write. It is to find out how complete the idea actually is.
A useful intake scores the concept across the dimensions vertical drama needs: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, target platform, and audience. If the intake is rich enough, you can move straight to a brief. If it is thin, you need guided follow-up. If it is barely a premise, you need a full development conversation.
This is the equivalent of asking: "Do we actually have a show, or just a vibe?"
2. Creative guidance
At this stage, missing dimensions are filled in. The key is not to generate a script yet. The key is to force decisions that scripts depend on:
- Who is the lead, and what is their arc?
- What is the core contradiction that keeps episodes turning?
- How often do reversals and payoffs land?
- Is the ending open, closed, revenge-complete, identity-revealed, or sequel-ready?
- What platform rhythm are we writing for?
Skipping this is why many AI dramas feel generic: the model invents fundamentals episode by episode because fundamentals were never locked.
3. Brief lock
The creative brief is the first hard gate. Until it is locked, no script should be written.
A locked brief should include at least: the logline, core conflict, story direction, ending direction, payoff and hook rhythm, platform and audience, episode-by-episode outline, and notes for the writer. The lead character entry must include character arc, not just name and job.
Think of this as the development meeting ending. After this point, changes are revisions, not drift.
4. Cast & visual confirmation
Before any large batch of art or video, the cast must be confirmed. That means names are valid, visual fields are complete, lineup is complete, and leads match the brief.
This is also where you decide whether the show is 2D, 3D, realistic, anime-inspired, period stylized, or something else. This choice is not cosmetic. It determines how characters, scenes, props, and video prompts all speak the same visual language.
5. Story archive build
A story archive is a structured continuity record for the whole show, covering identity, stable traits, current state, relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual notes.
This is one of the most important pieces in the whole pipeline. Long-running short drama cannot rely on model memory. It needs a structured archive.
The archive tracks things like:
- Who knows what as of the current episode
- Which threads are still open
- What injuries, disguises, or status changes are active
- Which props must appear in later scenes
- What happened in the most recent batch of episodes
When a new batch is written, the system should consume only the relevant slice of the archive, not the entire noisy history. After the batch, the archive updates. This is how episode 80 still remembers episode 5.
6. Batch script planning
Writing should happen in batches of several episodes, not one giant dump. Before drafting each batch, four things should be locked:
- Where this batch goes in the overall arc
- Which characters carry it
- Which conflicts and threads advance
- How the batch ends on a hook
Within each episode, planning comes before prose. The beat sheet comes first, including the final cliffhanger beat. Only then does the scene-by-scene script get written.
This mirrors real writers' rooms: you do not improvise structure while writing dialogue.
7. Script drafting & rule check
Vertical short drama scripts need hard rules because the format is unforgiving. The rules are not mysterious; they are the same ones experienced short-drama writers use:
- Golden 3 seconds: the first scene must open with conflict or suspense, no slow background dump.
- Episode shape: opening hook, escalation, final cliffhanger.
- Payoff density: at least one small payoff per episode; a larger payoff every several episodes.
- Line discipline: short lines, generally under twenty characters in Chinese source writing and kept tight in translation; no essay-like monologue.
After drafting, scripts should pass a rule check that catches broken formatting, missing character lines, too little dialogue, mismatched scene counts, placeholder text like "to be continued" used as a cop-out, and episode title errors. Failed scripts should be sent back with specific feedback, not handed to the user half-finished.
8. Style selection & look development
This is where the show gets its visual identity locked into assets.
A mature system offers multiple style handbooks covering 2D, 3D, realistic, period, urban romance, anime, ink-painting, xianxia, cyber-Chinese, stop-motion clay, and other directions. Each handbook should cover at least character art guidance, scene and prop guidance, and video style tags.
The important engineering point is that character art, scene art, prop art, and video prompts should all share the same style path. Otherwise you get the classic AI break: character art is anime, but video generation defaults into live-action realism.
Character art is best done in two stages: first, turn archive descriptions into art-ready prompts using the chosen style and hard gender rules, allowing manual override; second, generate the art, optionally with reference, and store version history.
Scenes have their own discipline: scene art should be empty of people. Props should be extracted from the script using the original naming used in the show, then compiled into a show-wide catalog.
9. Episode split into scene blocks
Once art exists, each episode is split into video scene blocks. The target size is around 10 seconds per block, with a soft cap around 200 characters of script body before re-splitting by action beats, paragraph breaks, or sentence endings.
This is one of the biggest differences between toy workflows and production workflows. You do not see "Episode 3" as one opaque video. You see Episode 3 as Scene 1, Scene 2, Scene 3, and so on. Each scene has its own prompt, its own references, its own output, and its own history of takes.
The analogy is the clip bin on an editing desk. If one beat is bad, you re-shoot that beat.
10. Shot prompt build per scene
This is where the system stops writing fiction and starts speaking in shot language.
A strong shot prompt for each scene should cover eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color tone
- Camera movement
- Visual style
- Image quality
- Constraints and guardrails
For simple scenes, this can be one compact block. For more cinematic scenes, a three-part structure works better: overall setup, shot-by-shot instructions, and a constraint pack.
A few discipline rules matter a lot:
- One shot, one camera move. Do not stack push, pull, pan, and tilt into one shot.
- Use shot numbers, not hard-coded timestamps like "0–3s."
- Always include a safety pack for face stability, image quality, and no watermark or logo.
- For multi-character scenes, add guardrails against twins or duplicate bodies.
- Prefer slow, continuous motion over explosive motion that breaks the model.
- Mark dialogue, sound effects, and background music with clear notation.
- Use only this scene's assets; do not leak references from another scene.
There is also one consistency rule that deserves to be stated on its own: if a character has a reference image, the prompt should not describe that character's clothing or appearance in text. The image is the source of truth; text should describe only action, expression, and injury state.
This rule directly prevents the common AI failure where text and reference fight each other and the character changes outfit or face mid-scene.
Reference mapping for each scene should follow a fixed priority: scene image first, then prop images for that scene, then character look images. If something has no image, it should be marked as text-only rather than invented. Manual binding should be supported when the script uses a title or nickname that does not exactly match the character's archived name.
11. Multi-modal shoot
With prompts and references ready, scenes go into production. The user should be able to choose model tier, aspect ratio, resolution, and duration. Vertical 9:16 is the default. Audio can be generated; watermark should be off by default.
On the operations side, this stage needs to behave like a real render queue, not a hobby script:
- Video tasks should run in a queue separate from writing and art tasks.
- The same scene block should not allow duplicate parallel submissions while one is running.
- Credits or points should be reserved on submission and released on failure.
- Stuck tasks should time out and become retryable.
- Output files should be checked for actual playable duration rather than trusting vendor metadata blindly.
This is the difference between a demo and a system a production team can run every day.
12. Review, re-take, select deliverable
There is no magic "auto-pick the best take" engine. Human review is still the quality gate.
For each scene, the team should be able to:
- Read and edit the shot prompt
- Swap character, scene, or prop references
- Manually bind roles
- Include or exclude props
- Change style
- Choose model tier
- Review previous takes and pick the best one
If you change the prompt and generate again, that is a new take. This is exactly how live production works: you change the shot note, then you shoot again.
The final deliverable is assembled from selected takes. Final trimming, ordering, and episode-level polish still belong in editorial, especially because scene blocks are engineering heuristics around 10 seconds, not frame-accurate timecode cuts.
Where teams usually break the chain
Most production problems map to a skipped checkpoint:
| Symptom | Usually caused by |
|---|---|
| Characters change face or outfit | No locked look-dev, or text describing appearance despite reference images |
| Story contradicts itself after many episodes | No structured story archive; relying on context memory |
| Episodes feel flat or hookless | Brief never locked payoff rhythm; no beat sheet before drafting |
| One bad scene forces whole-episode regeneration | No scene-block split; episode treated as one big output |
| Style shifts between art and video | Character art and video prompts not sharing one style path |
| Props appear and disappear | No prop catalog or per-scene reference mapping |
| Slow, unpredictable delivery | No queue discipline; tasks duplicated or stuck silently |
The fix is almost never "find a better model." It is "put the missing checkpoint back in order."
What this pipeline does not do
It is important to be honest about the boundaries.
- There is no automatic scoring engine that magically chooses the best take. Final quality judgment stays with the creator or producer.
- Reference images are not a hard blockade. If assets are missing, the system can warn, but creators can still proceed with text-only results, which are usually weaker.
- Reference markers in prompts depend on disciplined prompt construction; creators should still review prompts before shooting.
- The ~10-second scene block is a useful heuristic, not a precision edit. Long scenes may still need trimming later.
- The product unit is one scene with references going to one video segment. There is no mature cross-scene automatic continuation workflow for video.
- Character consistency depends on the look-dev asset chain, not on facial embedding magic. Period-look routing helps, but final results still depend on art quality and prompt discipline.
This is exactly the point: the system turns repeatable, drift-prone, failure-prone steps into a constrained pipeline, while aesthetic judgment remains with humans.
How this maps onto Maosika
Maosika (猫斯卡) is built around this exact sequence: idea evaluation, creative guidance, brief lock, cast and visual confirmation, story archive, batch script writing with beat sheets and rule checks, style selection, look development, scene-block splitting, per-scene shot prompts with reference mapping, multi-modal shooting, and review with re-takes.
The product does not claim one-click hits. It enforces the quality-control order that short-drama production has already proven works: story and continuity first, then looks, then scenes, then shot instructions, then multiple takes.
If you are planning a vertical AI short drama slate and want a production operating system rather than a single generate button, you can explore the workflow at Maosika.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com