How AI Vertical Short Dramas Actually Get Made: A 12-Stage Production Pipeline From Idea to Deliverable
Most AI short drama failures are not model failures—they are pipeline failures. The fix is not a fancier video button; it is a staged production order where each stage has a reviewable artifact before the next one starts.
If you have watched AI-generated micro-dramas long enough, you already know the pattern. The first episode looks promising. By episode three, the lead's face has shifted. By episode five, the costume changed. By episode eight, the plot has forgotten its own setup. The usual diagnosis is "the model is bad." The more useful diagnosis is: there was no pipeline, only a series of generations.
This article lays out the production order that turns a one-line idea into a deliverable vertical short drama. It is written for producers, writers, and production teams who need repeatable output, not one lucky clip.
The core principle: every stage must leave an artifact
A production pipeline is not defined by how many tools it uses. It is defined by whether each stage produces something you can look at, reject, edit, or lock before moving on.
In traditional crews, those artifacts are familiar: a logline, a beat sheet, a script, a continuity bible, a character look, a location board, a shot list, a take. AI production needs the same discipline. The difference is that some of those artifacts can now be generated, checked, and versioned by the system instead of being tracked in spreadsheets and group chats.
A working AI short drama pipeline is one where:
- story direction is locked before any video is generated
- character continuity is stored as an asset, not left to model memory
- look development is shared across characters, scenes, props, and video prompts
- each episode is broken into controllable shot-sized units
- prompts are written in shot-list language, not novel language
- retakes are tracked per shot, not hidden inside one giant re-render
The 12-stage pipeline, in order
The order matters. Skipping stages is what causes the familiar drift problems.
| Stage | Artifact produced | What gets decided here | What goes wrong if skipped |
|---|---|---|---|
| 1. Idea intake | Scored intake brief | Genre, lead, conflict, tone, episode count, platform | The project starts with a half-baked premise |
| 2. Creative guidance | Filled brief gaps | Missing dimensions the writer cannot infer | The script invents assumptions nobody approved |
| 3. Locked creative brief | Approved logline & direction | Core conflict, ending direction, hook rhythm, audience | Later stages drift because there is no north star |
| 4. Cast & visual confirmation | Character roster with visual fields | Names, roles, arcs, appearance anchors | Characters are inconsistent before filming even starts |
| 5. Story archive | Continuity bible | Relationships, open threads, state changes, episode appearances | The show forgets its own plot and character facts |
| 6. Batch script planning | Beat sheet per episode | Scene list, cliffhanger, conflict escalation | Scripts become talky, hookless, or structurally flat |
| 7. Script drafting & QC | Passed script batch | Dialogue, scene action, formatting, hooks | Lazy placeholders, missing cast lines, weak cliffhangers |
| 8. Style selection | Shared look path | 2D / 3D / realistic style, color, material, rendering feel | Characters and videos end up in different visual worlds |
| 9. Look development | Character / scene / prop art | Final approved reference assets | The model improvises faces, costumes, and locations |
| 10. Shot blocking | ~10-second shot blocks | Per-shot script slice, cast, props, location | One long generation becomes impossible to control |
| 11. Engineered shot prompts | Prompt + reference map | Camera, action, lighting, style, constraints | Prompts are vague, overloaded, or fight the reference images |
| 12. Rendering & retakes | Deliverable takes per shot | Model tier, resolution, duration, selected take | Bad output cannot be fixed surgically |
Stage 1–3: from idea to locked brief
The first three stages exist to answer a simple question: what exactly are we making?
A raw idea like "a disgraced heiress returns for revenge" is not enough to write from. The system needs to know the protagonist's arc, the central contradiction, the ending direction, the hook rhythm, the target platform, and the episode length. For vertical short dramas, episode length usually lives in the 1–2 minute range, with an outer limit around 3–4 minutes.
This is also where intake completeness matters. A brief that is too incomplete should not be allowed to jump straight into scripting. In a disciplined flow:
- a near-complete brief can move forward with minimal guidance
- a partial brief gets only the missing dimensions filled in
- a thin idea goes through full creative development before writing starts
The key gate is simple: the brief must be locked before the script stage begins.
Stage 4–5: cast, visuals, and the story archive
Character inconsistency does not begin in video. It begins earlier, when the production has no stable record of who a character is.
The story archive is the structured record that travels with the show for its entire run. It tracks:
- character identity and stable traits
- changeable current state, such as injuries, disguises, or revealed identity
- relationships between characters
- open and resolved plot threads
- episode-by-episode appearance records
- batch-level plot summaries
- visual descriptions of recurring props
This is the AI-production version of a continuity bible. The rule here is straightforward: do not rely on model memory; rely on a structured archive.
Before filming, the cast must also pass a visual gate: roster complete, names valid, visual fields complete, lead characters aligned with the brief. If that gate is not passed, the project should not move into art generation.
Stage 6–7: batch writing with structure, not one long generation
Vertical short drama scripts have their own mechanics, and those mechanics should be enforced before anyone touches video.
The baseline rules are:
- Golden 3 seconds: the first scene must open with conflict or suspense, not exposition
- Single-episode structure: opening hook, escalation, final-scene cliffhanger
- Payoff density: at least one small payoff per episode; a larger payoff every few episodes
- Dialogue discipline: short lines, no lecture-style speech, no heavy narration
A strong pipeline does not ask the model to "write a good episode." It asks for a beat sheet first, then the script, then a rules check. The rules check should catch concrete failures: wrong episode titles, mismatched scene counts, missing cast lines, too little dialogue, placeholder text like "to be continued." If a draft fails, it goes back for rewrite with the error feedback attached.
Writing also happens in batches, with archive updates after each batch. That is how later episodes can inherit established facts, open threads, and character state without drifting.
Stage 8–9: look development as a shared asset system
One of the most common AI failures is visual mismatch: the character sheet is anime, but the video output turns realistic; the scene is ancient, but the costume is modern; the prop described in episode two disappears by episode six.
The fix is to choose one style path and make every downstream asset use it. A mature pipeline includes style manuals covering character rendering, scenes, props, and video style tags. Once a style is selected, character art, scene art, prop art, and video prompts all follow the same visual path.
There is also a non-obvious but critical rule for reference images:
Once a character has approved reference art, the prompt should not re-describe that character's clothing or appearance in words. The reference image is the source of truth. The text should only describe action, expression, and injury state.
That one rule prevents a huge share of face swaps and costume changes.
Scenes and props need discipline too:
- scene art should be empty plates, with no people in them
- props should be extracted from the script using the script's own names
- manual binding should be supported when the script uses a nickname or role title
- time-period routing should pick the right look for flashbacks or cross-era stories
Before shot generation, the pipeline should also check whether scene, character, and prop references are missing. It can allow a text-only path, but it should warn: skipping reference assets usually lowers consistency.
Stage 10–11: shot blocks and engineered prompts
A full episode is too big to be one controllable generation unit. The practical unit is the shot block.
A shot block is a short, filmable slice of the script, usually around 10 seconds. The script is split by action beats, paragraph breaks, and sentence boundaries, with longer passages subdivided further. The result is not "Episode 3 as one video," but "Episode 3, Shots 1 through N," each with its own prompt, reference set, and render history.
This is where prompt engineering becomes production engineering instead of poetry. A shot prompt should contain the elements a camera department actually needs:
- precise subject
- action detail
- scene environment
- lighting and color
- camera movement
- visual style
- image quality constraints
- negative constraints and stability safeguards
A few prompt rules matter a lot in practice:
- one camera move per shot; do not stack push, pull, pan, and tilt together
- use shot numbers, not hard-coded timestamps like "0–3s"
- include a stability fallback for faces, watermarks, and duplicate-person errors
- prefer slow, continuous motion over extreme action that breaks easily
- mark dialogue, sound effects, and music with consistent notation
- feed only the current shot's assets; do not leak other scenes into the prompt
In one sentence: the system should teach the model to read a shot list, not improvise a novel.
Stage 12: rendering, review, and retakes
Once prompts and references are ready, production becomes a queue, not a magic button. The team should be able to:
- edit the shot prompt like a director revising shot notes
- swap character, scene, or prop references
- manually bind ambiguous names
- include or exclude props to control visual focus
- switch style paths when needed
- choose model tier based on quality, speed, and cost
- keep multiple takes per shot and select the best one
This stage also needs production-grade reliability: queued rendering, failed-task recovery, no duplicate submissions for the same in-progress shot, and retryable failures. Without that, the pipeline is a demo, not an operating system.
Importantly, there is no automatic "choose the best take" engine. Final quality judgment still belongs to the creator or supervisor. The system's job is to make review and retake surgical, not to pretend human judgment is unnecessary.
Where teams usually break the order
Most production problems come from the same few order violations:
| Bad shortcut | Symptom |
|---|---|
| Writing video prompts before the brief is locked | Every episode goes in a different direction |
| Generating video before character art is approved | Faces and costumes change shot by shot |
| Using one long episode prompt instead of shot blocks | One bad moment forces a full re-render |
| Re-describing costumes in words after reference art exists | The model fights the image and changes the outfit |
| Letting later episodes write from loose memory | Plot threads disappear or contradict earlier episodes |
| Treating retakes as full regenerations | Teams cannot tell what changed or which version won |
The pattern is always the same: teams try to save time by removing reviewable intermediate steps, and the lost time reappears later as inconsistency.
What this pipeline does not do
Honest boundaries matter here, because AI production tools are often sold as if they eliminate craft. They do not.
A serious pipeline does not claim to:
- automatically score artistic quality or pick the perfect take
- force reference images as a hard gate when a team intentionally chooses text-only
- guarantee perfect character identity without strong approved art
- auto-extend video across shots into a finished edited episode
- replace the director, writer, cinematographer, or editor
What it can do is different: it turns the repeatable, drift-prone, failure-prone parts of production into a constrained workflow, while leaving审美 judgment—about performance, tone, casting look, pacing, and final selection—in human hands.
How Maosika fits into this order
Maosika (猫斯卡) is built as an AI production operating system for vertical short dramas rather than a single text-to-video button. It formalizes the staged order above: idea intake and guidance, locked brief, cast confirmation, story archive, batch script writing with rules checks, style selection, look development for characters/scenes/props, ~10-second shot blocks, engineered shot prompts with reference mapping, multi-modal rendering, and per-shot retakes.
Its 18 digital specialists mirror real crew roles across the creative and video-production sides, and users can see streaming work records from those roles instead of facing a black box. The product philosophy is that every stage should produce a reviewable artifact, and that human confirmation belongs at the key gates.
For teams exploring whether this kind of system fits their workflow, the test is not "can it generate a video?" The real test is whether it enforces the production order that keeps a 50-episode or 100-episode vertical drama from falling apart after the first few episodes.
If you want to go deeper into one part of the pipeline, the next most useful stops are vertical short drama script structure, AI character consistency, and shot prompt engineering for vertical drama.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com