How AI Vertical Short Dramas Actually Get Made: A 12-Step Production Pipeline From Idea to Final Cut
AI short drama production fails not because the model is weak, but because teams skip the order of operations. The reliable path is the same one live crews use: lock the brief, build the bible, cast and costume, then shoot scene by scene.
The order matters more than the tool
Most AI vertical short drama teams start in the wrong place. They write a loose synopsis, jump straight to video generation, and then spend days fighting face swaps, costume drift, forgotten plot threads, and cliffhangers that don't land. The fix is not a better prompt. The fix is a production pipeline.
This article walks through the full pipeline in 12 steps, written the way a real micro-drama crew would describe it. Each step has a deliverable you can inspect, reject, or revise before moving on. That is the difference between a toy demo and something you can run at episode scale.
A vertical short drama is a 9:16 episodic story, usually 1–2 minutes per episode (sometimes stretching to 3–4), built around tight hooks, fast conflict escalation, and a cliffhanger at the end of almost every episode.
The 12-step pipeline, end to end
| # | Stage | Key deliverable | Who decides |
|---|---|---|---|
| 1 | Idea intake & scoring | Intake score 0–100, routed to the right depth of development | Producer + creator |
| 2 | Creative guidance | Missing dimensions filled in (genre, lead, conflict, tone, platform) | Creator |
| 3 | Brief lock | Logline, core conflict, arc, hook rhythm, ending direction, audience | Creator sign-off |
| 4 | Cast & visual confirm | Named cast with visual fields aligned to the brief | Creator sign-off |
| 5 | Story archive (continuity bible) | Characters, relationships, open/resolved threads, episode appearance table, prop descriptions | System-maintained, creator-audited |
| 6 | Batch script writing | Beat sheet first, then dialogue, then rule-based QC pass | Writer + system QC |
| 7 | Style selection | One shared style path for characters, scenes, props, and video | Creator |
| 8 | Look development | Character, scene, and prop turnarounds / key art | Art department + creator |
| 9 | Episode cut into shot blocks | ~10-second blocks, each with its own references and prompt | System + editor review |
| 10 | Shot-level prompt build | Eight-element cinematic prompts bound to local references | Director / creator edits |
| 11 | Multi-modal rendering | 9:16 vertical takes per block, queued and tracked | Queue system |
| 12 | Review, retake, select | Takes compared; re-words, swaps references, re-renders; selects best | Creator / editor final cut |
1. Idea intake & scoring
Not every idea is ready to write. A strong intake captures genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, and target platform/audience. The intake is scored 0–100:
- 80+: most guidance can be skipped, go straight to the brief.
- 40–79: only the missing dimensions are filled in.
- Below 40: a full guided development path is required.
This is the equivalent of a development meeting. If the room can't state the lead's arc and the core conflict in one breath, you don't call "action."
2. Creative guidance
Guidance is not a chatbot asking vague questions. It is a structured interview that forces decisions the script will later depend on: who the lead is, what they want, what is stopping them, how the audience is supposed to feel every 30 seconds, and what the ending is leaning toward. For vertical micro-dramas, episode length is deliberately bounded—mostly 1–2 minutes, at most 3–4—because the hook density required does not survive longer runtimes.
3. Brief lock
The creative brief is the first hard gate. Until it is locked, nothing downstream starts. A locked brief includes the logline, core conflict, story direction, ending direction, beat of payoffs and hooks, target platform and audience, episode-by-episode outline, and notes to the writer. The lead character entry must include a character arc, not just a job title and a hairstyle.
Think of this as the greenlight document. Change it later, and you are re-developing, not tweaking.
4. Cast & visual confirm
Before a single scene is written, the cast must be complete: names are valid, visual fields are filled, and leads line up with the brief. This is a hard check, not a suggestion. AI productions that skip it pay for it later when "the detective" in episode 6 is a different person than "the detective" in episode 2.
5. Story archive (continuity bible)
A story archive is a structured, episode-spanning record of who characters are, how they relate, what they currently look like, which plot threads are still open, which are resolved, who appears in which episodes, and what props look like. It is the digital equivalent of a writers' room continuity bible.
The writer does not rely on the model's memory. The writer is fed a slice of the archive: current character states, unresolved threads, recent batch summary, and the current batch's main line. After each batch, the archive is updated; before the next batch, it is re-loaded. This is how a 100-episode show still remembers the scar from episode 3.
6. Batch script writing
Scripts are written in batches of several episodes, not one giant 80-episode generation. The order inside each batch is strict:
- Beat sheet first, with the episode-end cliffhanger declared before dialogue.
- Then the full scene text.
- Then a rule-based quality check.
The writing rules embedded in the pipeline are not subtle:
- Golden 3 seconds: scene 1 must open with strong conflict or suspense; no flat setup, no lore dump.
- Episode shape: opening hook (1 scene) → rising conflict → end-of-episode cliffhanger (last scene).
- Payoff density: at least one small payoff per episode (a reveal, a reversal, a hint of identity, evidence secured); a bigger payoff every few episodes.
- Dialogue: short lines, generally under ~20 characters in Chinese originals and equivalently tight in translation; no essay-speak, no narrator explaining what the camera could show.
From the second batch onward, four things are locked before writing: the batch's plot direction, focus characters, threads and conflicts in play, and the end-of-episode hook. These are treated as hard constraints the writer cannot override, even if the archive would suggest otherwise.
The QC pass catches concrete failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, lazy placeholders like "to be continued." Failing scripts are sent back with the specific error and rewritten, up to a cap. Nothing half-finished reaches the director.
7. Style selection
One style path is chosen for the whole production. A production-ready system ships with multiple style packs covering 2D, 3D, and live-action-adjacent realism—urban realism, period realism, mature urban romance animation, 90s Japanese anime, Chinese ink-and-brush, xianxia, 3D donghua, clay stop-motion, cyberpunk-Chinese fusion, and more. Each pack includes rules for character rendering (face anchors, materials, mood, view consistency), scene/prop rules, and video style tags.
The point is that character art, scene art, prop art, and video prompts all consume the same style path. You do not get an anime character walking into a photoreal room.
8. Look development: characters, scenes, props
Look dev is a two-step process:
- Text polish: the archive description is turned into an image-generation prompt using the chosen style pack, with gender as a hard rule. The creator can override this text.
- Image generation: the final prompt is rendered, reference images are supported, and every output is stored as a reviewable version.
Scenes follow an empty-frame rule: scene art must not contain characters. Props are pulled episode by episode from the script's own wording—what the prop master on a real set would do—and merged into a show-wide catalog so nothing is lost across episodes. When the script and the style pack disagree, the script wins: the style pack controls how things are painted, not what is allowed to exist.
Before shooting starts, a readiness check lists any missing scene, character, or prop references. You can skip and go text-only, but you are explicitly warned that quality usually drops; the professional path is to finish look dev first.
9. Episode cut into shot blocks
An episode is not one long generation. It is cut into shot blocks at roughly 10 seconds each, with a soft ceiling on text length; longer scenes are split further on action beats, paragraph breaks, and sentence ends. Crowd or generic characters are separated from the named cast so they don't consume character reference slots.
The creator sees episode → block 1, block 2, block 3… Each block has its own prompt, its own references, its own output, and its own history of takes. This is the clip bin on an editing desk, not a single monolithic render.
10. Shot-level prompt build
The system builds cinematic prompts, not prose. Each prompt covers eight elements: precise subject, action detail, scene environment, lighting and color, camera movement, visual style, image quality, and constraints. Simple scenes are written as one block; complex cinematic scenes use a three-part structure (overall setup → shot 1/2/3… → constraint pack).
Rules worth stealing for any AI video workflow:
- One camera movement per shot—no push-pull-pan piled into one take.
- Shot numbers, not absolute timestamps ("shot 1," not "0–3s").
- A mandatory fallback pack: quality, face stability, no watermark or logo; twin/duplicate fallbacks for multi-character shots; style anchoring for non-realism.
- Action is written with specific limbs and quantified intensity; slow continuous motion is preferred over high-impact bursts that break.
- Dialogue in
{}, sound effects in<>, score in(). - Only this block's references are fed in—no bleed from other scenes. The block's script text is the highest-priority source.
- IP and specific title names are stripped before delivery, keeping technique and aesthetic descriptors but reducing downstream copyright blocking.
The single most important consistency rule: if a character has a reference image, the prompt must not re-describe their clothes or appearance in text. Text only describes action, expression, and injury. Appearance is owned by the reference. This is the rule that kills face-and-costume drift.
Manual controls exist where they matter: bind a script name like "the officer" to a specific cast card, include or exclude props to control visual focus, swap in a different historical version of a character for time-travel or flashback scenes, or replace a reference with an earlier version.
11. Multi-modal rendering
The render stage is a production queue, not a fire-and-forget script:
- Default delivery is 9:16 vertical, not a landscape crop after the fact. Other ratios are supported.
- Duration can be intelligent or fixed, typically in the 5–15 second range.
- Model tiers (standard / fast / lightweight) let you trade quality, cost, and speed.
- Resolution follows each model's allowed range.
- Audio can be generated; watermarking is off by default.
- Files get faststart treatment and are stored with measured runtime, not just the vendor's reported duration.
- Video tasks run in their own queue lane, isolated from script and art tasks.
- A block cannot be double-submitted while already running; credits are pre-deducted and released on failure; hung tasks time out and can be retried; existing vendor jobs are polled to completion rather than re-created.
In field terms: this is a shoot day with a call sheet, not a weekend hobby render.
12. Review, retake, select
There is no magic "auto-pick the best take" button. The creator watches the takes, compares them, and decides. The tools provided are the ones a real set has:
| Control | What it's like on set |
|---|---|
| Edit the video prompt | Director rewriting the shot note |
| Swap character / scene / prop references | Changing a costume or a location plate |
| Manual character binding | Fixing a name mismatch or a cameo reference |
| Include / exclude props | Controlling what's in frame |
| Switch style | Unifying the look of the show |
| Choose model tier | Balancing quality, cost, schedule |
| Browse historical takes per block | Multi-take selection |
When you edit the prompt and hit generate, that is a new take, exactly like changing the shot note and rolling again. The system does not silently re-invent prompts under you.
Where the pipeline is honest about its limits
No production system is credible without a list of what it does not do. The limits here are explicit:
- No automatic scoring or auto-retake engine for final footage. Final quality judgment stays with the creator and the producer; the system supplies multiple takes and the tools to re-render with changed inputs.
- Reference images are not a hard gate. Missing images trigger a warning, but you can proceed with text only—expect lower consistency, which is why professional shoots finish look dev first.
- Reference markers in prompts rely on the spec, not a hidden hard-stitch. Creators should still read the prompt and confirm references are present before rendering.
- Shot grammar uses the built-in cinematic spec; the full style pack is applied to art, while video gets the style tags. You are not getting a full cinematographer's brain in one line.
- The ~10-second block is an engineering heuristic, not a timecode-precise cut. Overlong scenes get re-split, but some blocks still run long; final trimming and stitching belong in the edit.
- There is no cross-block automatic continuation or video extension workflow right now. The unit of work is "one block, with its references, producing one clip."
- Character consistency depends on the look-dev asset chain, not a face-embedding verification step. Period/costume routing is rule-based; final look still depends on the quality of the art and on the prompt respecting the "don't describe what's in the reference" rule.
The principle behind the whole chain
The point of an AI production operating system is not to press one button and get a hit show. It is to turn the parts of production that are repetitive, drift-prone, and easy to lose control of into a constrained pipeline—while keeping every审美 decision in human hands.
High-quality vertical short dramas are never one long generation chopped up. They are stacks of controllable units, each with its own references, its own prompt, its own takes, and a human picking the winner. Consistency comes from assets, not luck. Continuity comes from a structured archive, not the model's memory.
---
*This article reflects the production pipeline built into Maosika (猫斯卡), an AI production operating system for vertical short dramas. The pipeline runs from idea intake through brief lock, continuity archive, batched script writing with rule-based QC, shared style look dev, ~10-second shot blocks, cinematic prompt building, queued rendering, and multi-take review. Maosika does not claim one-click hits or zero human error; it enforces the order of operations that real crews already know works.*
FAQ
What is the full production process for an AI vertical short drama?
It runs in 12 steps: idea intake and scoring, creative guidance, brief lock, cast and visual confirmation, story archive (continuity bible), batch script writing with beat sheets and QC, style selection, look dev for characters/scenes/props, cutting episodes into ~10-second shot blocks, building cinematic prompts bound to local references, queued rendering in 9:16 vertical, and review/retake/select.
How long should each episode of a vertical short drama be?
Most episodes are 1–2 minutes, with an upper bound around 3–4 minutes. Vertical micro-dramas rely on tight hook density, and longer runtimes usually dilute the cliffhanger rhythm.
How do you stop AI characters from changing face or clothes between scenes?
Lock the cast and visual fields first, produce reference art during look dev, bind each shot block to those references, and enforce the rule that prompts must not re-describe the clothes or appearance of a character who already has a reference image. Text should only describe action, expression, and injury.
What is a story archive in AI short drama production?
It is a structured continuity bible that tracks character identities, current states, relationships, open and resolved plot threads, episode-by-episode appearance, batch summaries, and prop descriptions. Each batch of scripts is written from a slice of this archive, then writes back to it, so long shows don't forget earlier setup.
Does AI automatically pick the best take for each shot?
No. There is no reliable automatic quality judge for final footage. The system provides multiple takes, lets you edit prompts, swap references, and re-render, but the final selection is made by the creator or producer.
Can I make an AI short drama without doing character and scene reference art first?
You can proceed with text-only prompts, and the system will warn you before you do. Consistency is usually noticeably worse without references, so the professional workflow is to finish look dev before shooting.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com