How an AI Vertical Short Drama Actually Gets Made: The Production Pipeline From Idea to Finished Episode
An AI vertical short drama is not one big text-to-video click. It is a staged production line where each stage produces a reviewable artifact you can approve, reject, or roll back before moving on.
The core claim
Most teams that fail with AI short drama fail the same way: they treat the model like a director and ask it to "make a good episode." Professional pipelines do the opposite. They break the job into stages, force a reviewable artifact at each stage, and only move forward when the previous stage is locked.
A production-grade AI short drama pipeline is a sequence of checkpoints, not a single generation. The model executes within constraints set by humans; humans make the审美 calls at the gates.
Why "one prompt, one video" breaks vertical drama
Vertical short drama has hard structural requirements that generic video tools do not enforce:
- The first scene must hook in the first few seconds
- Each episode needs escalation and a cliffhanger
- Characters must look the same across dozens of episodes
- Props, costumes, and settings must not drift between scenes
- Lines must stay short and spoken, not narrated
- The final deliverable is 9:16 by design, not a landscape crop
When you ask a model to do all of this in one shot, you get what you would get on a real set if you skipped pre-production: inconsistent cast, forgotten plot threads, wrong costumes, and scenes that do not cut together.
The pipeline, stage by stage
Below is the staged flow a serious AI short drama production should follow. Each stage has a defined owner, a deliverable, and a gate.
| Stage | Deliverable | Gate before moving on |
|---|---|---|
| 1. Idea intake & scoring | A completeness score across core dimensions | Score determines how much guidance is needed |
| 2. Guided ideation | Filled gaps in premise, protagonist, conflict, tone, platform | No missing dimension left blank |
| 3. Creative brief lock | Logline, core conflict, arc, hook rhythm, ending direction, audience, episode outline, notes for the writer | Brief is locked; later changes require a deliberate revision |
| 4. Cast & visual confirmation | Named cast with visual fields, protagonist aligned to brief | Cast is complete and legal |
| 5. Story archive / continuity bible | Structured record of characters, relationships, open threads, episode appearances, batch summaries | Archive exists and is current |
| 6. Batch script writing | Beat sheet first, then dialogue, then rule check, then archive update | Script passes rule-based quality check |
| 7. Style selection | Chosen visual style shared across characters, scenes, props, and video | Style path is fixed for the production |
| 8. Look dev: characters / scenes / props | Approved reference art for cast, locations, key items | Only finished, readable art moves to shooting |
| 9. Scene blocking | Episode split into roughly 10-second scene blocks | Each block has its own reference set and prompt |
| 10. Shot prompt engineering | Per-block prompts with subject, action, environment, lighting, camera, style, quality, constraints | Prompt is human-reviewed before render |
| 11. Multimodal shoot | Rendered vertical clips per block | Clips land in a queue with retake history |
| 12. Review & retake | Selected takes, edited prompts, swapped references, re-renders | Final take chosen per block |
Stage 1–3: From idea to locked brief
The first gate is not art and it is not video. It is the brief.
A useful intake scores how complete the idea is across the dimensions that actually matter for vertical drama: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, and target platform / audience. If the idea is thin, the system should force a guided fill; if it is already well-formed, it should skip the filler and go straight to brief.
The creative brief is the first real lock. It is the equivalent of finishing the development meeting before anyone calls "action." A protagonist entry is not complete without a character arc. Hook rhythm and ending direction are written down, not left to the model to invent episode by episode.
Stage 4–5: Cast and the continuity bible
Character consistency does not start at render time. It starts when the cast is defined and a living story archive is built.
The story archive is a structured continuity record for the whole production. It tracks character identities, stable traits, current state (injuries, revealed identities, changed alliances), relationships, open and resolved plot threads, episode appearance tables, batch summaries, and visual descriptions of recurring props.
The archive exists because model memory is not a production tool. A 60-episode drama will eventually forget what was established in episode 3 unless the system hands the writer only the relevant slice: current character states, unresolved threads, recent batch context, and the current batch's main line. After each batch, the archive updates; before the next batch, it is reloaded.
Stage 6: Batch script writing with rules built in
Writing one giant script for an entire season is how drift happens. A production pipeline writes in batches of a few episodes, with a planning step before dialogue.
For each batch, the beat sheet comes first: what happens scene by scene, and where the cliffhanger lands at the end of the episode. Only after the beats are planned does dialogue get written.
Vertical drama writing rules that should be enforced, not suggested:
- Strong opening: the first scene must open on conflict or suspense, never on slow exposition.
- Episode shape: hook at the top, escalating middle, cliffhanger at the end.
- Payoff density: at least one small payoff per episode (a reversal, a face-slap, an identity hint, a piece of evidence); a larger payoff every few episodes.
- Line length: short spoken lines, generally under 20 words; no essay-like dialogue or long explanatory voiceover.
After writing, a rule-based check should catch concrete failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, or placeholder text like "to be continued." Failures go back for rewrite with the error attached. Half-finished scripts should not reach the director or the render queue.
From the second batch onward, the pipeline should lock intent before writing: where this batch goes, who carries it, which threads and conflicts are in focus, and how the end hook lands. Those locked intents override older archive notes when the two conflict — because the creator's current direction wins.
Stage 7–8: Style and look dev
One of the most visible AI failures is style fracture: characters drawn in one dialect, scenes painted in another, video rendered in a third. The fix is boring but effective: pick one style path early, and make characters, scenes, props, and video prompts all consume it.
A production-ready system should ship with multiple style packs covering 2D, 3D, and realistic directions — urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink style, xianxia, 3D donghua, stop-motion clay, cyber-Chinese fusion, and so on. Each pack should include guidance for character art (face anchors, material,气质, view consistency), scene and prop guidance, and video style tags.
Character art is best done in two steps: first, the system turns archive descriptions into art prompts using the chosen style's character rules, with gender treated as a hard constraint; then the image is generated, with support for reference images and version history. The user can override the prompt before generation.
A strict rule on the video side: only finished, readable art assets are used as references. Half-rendered drafts do not get bound to scenes.
Stage 9: Scene blocks, not whole episodes
A finished episode is not one render job. It is a list of scene blocks, each targeting roughly 10 seconds of screen time, with a soft cap on body text and further splitting when a block runs long.
This is the editing-room view: Episode 3 is not "a video," it is Scene 1, Scene 2, Scene 3, and so on. Each block has its own prompt, its own reference set, its own rendered takes, and its own history.
Why this matters:
- A bad take only costs one block, not the whole episode
- References stay local; Scene 5 does not accidentally inherit Scene 2's props
- Retakes are surgical: change the shot description for one block and re-render that block
- The final assembly can be cut from selected takes, not accepted as a monolith
Crowd characters and generic extras should be separated from the named, drawable cast so they do not consume reference slots meant for lead characters.
Stage 10: Shot prompts that speak production language
The prompt for a block is not a paragraph of fiction. It is a shot instruction. A well-engineered prompt covers eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color tone
- Camera movement
- Visual style
- Image quality
- Constraints
For simple scenes, one paragraph is enough. For complex cinematic scenes, a three-part structure works better: overall setup, then shot-by-shot instructions, then a constraint pack.
Production rules that reduce failure:
- One camera move per shot; do not stack push, pull, pan, and tilt into one instruction
- Label shots by shot number, not by absolute timestamps
- Always include a base constraint pack: quality, facial stability, no watermark or logo
- For multi-character shots, add anti-twinning / anti-duplicate constraints
- For non-realistic styles, anchor the style explicitly
- Favor slow, continuous motion over high-action bursts that break
- Mark dialogue, sound effects, and BGM with consistent notation
- Feed only this block's assets into this block's prompt — no cross-scene leakage
- When a character has a reference image, the prompt must not re-describe clothing or appearance in text; text only describes action, expression, and injury state
That last rule is the single highest-leverage fix for the classic AI failure where a character changes outfit or face between shots. If the reference image already defines the look, text re-describing it only creates a conflict the model has to resolve — and it resolves it randomly.
Reference mapping per scene should follow a fixed order: scene art first, then props for this scene, then character art. Slots are only filled when an image exists; the system should not invent bindings to fill empty slots. Manual override should be available: bind a script name like "the officer" to the right cast entry, include or exclude props, swap in an older approved version of an asset, and route period-correct looks for time-travel or flashback scenes.
Before prompts are generated, the pipeline should run a readiness check: are scene, character, and prop references present? If not, it should show the gap list and offer to go create the art. It should also allow a deliberate skip to text-only — with the clear understanding that quality usually drops, which is why professional workflow locks art first.
Stage 11–12: Shoot, review, retake
The shoot stage should behave like a production queue, not a toy script:
- Video tasks run in their own queue lane, separate from writing and art
- The same block cannot submit parallel jobs while one is in flight
- Credits are reserved on submit and released on failure
- Stuck jobs time out and become retryable
- Rendered clips get faststart treatment and are stored with measured duration
- Each block keeps a history of takes
Crucially, hitting "generate" does not rewrite the prompt. The system uses the prompt you approved and the reference set that was locked at the time. If you edit the prompt or swap a reference and generate again, that is a new take — exactly like revising shot notes on a real set before rolling again.
Controls a creator should expect at this stage:
| Control | What it is equivalent to on set |
|---|---|
| Edit the shot prompt | Director revising shot notes |
| Swap character / scene / prop references | Changing a look or a location board |
| Manual character binding | Fixing name mismatches and off-screen references |
| Include / exclude props | Controlling visual focus in the frame |
| Switch style | Unifying the visual language |
| Choose model tier | Trading quality, cost, and speed |
| Browse historical takes per block | Choosing the best take from multiple rolls |
There is no magic "auto-pick the best shot" engine. Final judgment sits with the creator and the producer. The pipeline's job is to make retakes cheap, traceable, and isolated.
The deliverable shape
The default output is 9:16 vertical, built vertical from the start — not a landscape video cropped after the fact. Creators should be able to choose smart duration or a fixed 5–15 second range per block, pick a model tier, and select resolution within the supported range. Audio can be generated by default; watermarks should be off by default.
Prompts should be cleaned of specific copyrighted IP names before they hit downstream renderers, keeping the cinematic and aesthetic description while reducing rights-related blocking. If a supplied reference image is detected as a suspected real-person photo, the system should refuse to silently proceed and instead guide the creator to use platform-generated art or re-shoot without the photo reference.
Where the pipeline deliberately stops
It is worth being explicit about what this kind of system does not do, because those boundaries define how to use it well:
- It does not auto-score finished videos or auto-choose the best take; human review is the final gate.
- Reference images are not a hard block — you can skip them, but text-only renders are generally weaker.
- Reference markers in prompts rely on disciplined formatting, so creators should still read the prompt before rendering.
- The shot grammar follows the platform's built-in camera rules; full style manuals live on the art side.
- The ~10-second block target is an engineering heuristic, not frame-accurate editing; long scenes may still overshoot, and final assembly belongs in editing.
- There is no cross-block automatic video extension workflow; the unit of work is "one block, one rendered clip."
- Character consistency depends on the art pipeline and disciplined prompts, not on face-embedding verification; period routing is rule-based, and final look still depends on the quality of the approved art.
This is the right tradeoff for production. The goal is not to remove humans; it is to turn the repetitive, drift-prone, failure-prone parts of the job into a constrained pipeline, while taste, story judgment, and final selection stay with the people making the show.
How Maosika fits in
Maosika (猫斯卡) is built around exactly this staged structure. It is positioned as an AI production operating system for vertical short dramas — an end-to-end line from idea to finished episode, with 18 digital specialists mapped to real crew roles, 17 built-in style packs, structured story archiving, rule-checked batch writing, reference-mapped scene blocks, engineered shot prompts, and a retake-friendly render queue. It does not promise one-click hits; it enforces the order of operations that professional short drama already relies on.
If you are building or running an AI short drama slate, the question to ask a tool is not "can it generate video?" It is "what artifacts does it force me to lock before it spends render credits, and can I review, edit, and roll back each one?" That is the difference between a demo and a production line.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com