The AI Vertical Short Drama Pipeline: A Crew-Style Workflow From Idea to Final Cut
AI short drama production breaks not when the model is weak, but when the workflow skips the order a real crew would never skip: lock the brief, build the bible, approve the cast, then shoot scene by scene.
Why most AI short dramas fail before the camera rolls
The common failure mode is not bad video. It is bad production order. Teams type a loose premise, jump straight to image or video generation, and then spend days patching continuity errors, costume swaps, forgotten props, and cliffhangers that do not land.
A real crew does not shoot first and discover the story later. It locks the brief, casts the roles, builds the continuity record, approves the look, boards the scenes, and only then calls action. AI production needs the same discipline.
A production-grade AI short drama pipeline is a staged, auditable workflow in which each stage produces a reviewable artifact before the next stage begins.
This guide walks through that pipeline in crew language, not tool hype. The goal is not "one-click viral." It is to turn the repetitive, drift-prone parts of short drama production into a constrained pipeline, while taste, story judgment, and final quality control stay with the creator.
The pipeline at a glance
Think of the process as a film set with digital roles. Each handoff has something you can inspect, reject, or revise.
| Stage | Crew equivalent | Reviewable artifact | What goes wrong if skipped |
|---|---|---|---|
| 1. Idea intake & evaluation | Development meeting | Intake score, missing dimensions | Vague premise, no audience fit |
| 2. Creative guidance | Producer notes | Filled story dimensions | Writer guesses tone and hook rhythm |
| 3. Brief lock | Greenlight meeting | Logline, conflict, ending direction, hook plan | Scope creep, contradictory notes |
| 4. Cast & visual confirmation | Casting + look development | Character lineup, approved descriptions | Characters change face episode to episode |
| 5. Story archive / continuity bible | Script department bible | Character states, relationships, open threads | Model "forgets" earlier plot beats |
| 6. Batch script writing | Writers' room | Beat sheets, episode drafts, cliffhangers | Padded scenes, weak cliffhangers |
| 7. Rule-based script QC | Script supervisor | Pass / fail with error reasons | Broken formatting, missing cast lines |
| 8. Style selection | Look of show | Style manual choice | Mixed art directions across assets |
| 9. Character / scene / prop art | Art, costume, props | Approved reference images | Text and image fight each other |
| 10. Scene blocking | Editor's clip list | ~10-second scene blocks | Long unmanageable generations |
| 11. Shot prompt engineering | Storyboard + DP notes | Per-scene shot prompts | Novel-style prompts that confuse video models |
| 12. Multi-modal shooting | Production | Takes per scene | One-shot gambling, no comparison |
| 13. Review, rewrite, reshoot | Director + editor | Selected takes, revised prompts | Teams accept first flawed output |
Stage 1–3: Intake, guidance, and locking the brief
Vertical short drama is unforgiving of vague development. A viewer decides in the first few seconds whether to stay. If the brief is soft, every later stage invents its own version of the show.
A strong intake covers the dimensions that actually shape vertical drama:
- Genre and tone
- Protagonist and character arc
- Core conflict
- Story direction
- Episode count
- Episode length, usually 1–2 minutes per episode for vertical, with a practical ceiling around 3–4 minutes
- Hook and payoff rhythm
- Ending direction
- Distribution platform and target audience
One useful mechanism is an intake completeness score from 0–100. Above 80, the idea is developed enough to move straight to brief lock. Between 40–79, the system only asks for the missing dimensions. Below 40, it runs a fuller guided development path. The point is not to gate for its own sake; it is to avoid sending an undercooked idea into scriptwriting.
The brief itself is the greenlight document. It should contain at minimum:
- A one-sentence logline
- Core conflict
- Story direction
- Ending direction
- Payoff and hook rhythm
- Platform and audience
- Episode-by-episode outline
- Notes to the writer
- Protagonist fields including character arc
The brief is locked before writing begins. In production terms, you do not rewrite the logline on set.
Stage 4–5: Cast confirmation and the continuity bible
AI does not naturally maintain continuity across dozens of episodes. It needs a structured record, not a long memory.
A continuity bible for AI short drama is a structured, episode-spanning record of character identities, stable traits, current states, relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual descriptions.
This is the digital equivalent of a script department's continuity bible. It exists so that episode 40 still knows who has a scar, which secret has been revealed, which relationship changed, and which prop was introduced in episode 6.
Before writing, the cast must pass a hard confirmation gate:
- All roles are present
- Names are valid
- Visual fields are complete
- Lead characters match the locked brief
If this gate is skipped, later stages inherit ambiguity. The video model does not know it is looking at the same person; it only knows what the prompt and references tell it in that scene.
Stage 6–7: Batch writing with beat sheets and QC
Writing vertical drama episode by episode without structure is how shows drift. A more reliable pattern is batch writing: write several episodes as a unit, archive the results, confirm direction, then continue with the next batch.
Each episode should be built in order:
- Beat sheet first
- Episode-end cliffhanger defined
- Then full scene text
This mirrors real writers' room discipline. You decide where the episode is landing before you write the dialogue that gets it there.
Vertical script rules worth hard-coding:
- Golden 3 seconds: the first scene must open with strong conflict or suspense. No slow setup, no background lecture.
- Single-episode structure: opening hook → conflict escalation → final-scene cliffhanger.
- Payoff density: at least one small payoff per episode, such as a reversal, identity hint, evidence reveal, or satisfying confrontation; a larger payoff every several episodes.
- Dialogue: short lines, generally under twenty words; no essay-style speech or explanatory narration.
Continuity is protected by a few mechanisms:
- The writer only consumes a relevant slice of the bible: current character states, unresolved threads, recent batch summaries, and current batch objectives.
- Open plot threads are tracked as open or resolved.
- Past the halfway point, ending constraints are injected so the story does not wander.
- Each new batch starts by confirming batch intent: plot direction, focus characters, key conflicts and threads, and episode-end hooks.
After drafting, scripts pass through a rule-based quality check. It should catch concrete failures, not subjective taste:
- Wrong episode titles
- Scene count mismatches
- Missing character lines
- Too little dialogue
- Placeholder text like "to be continued"
Failed drafts are sent back with error feedback for rewrite, up to a limited number of retries. Half-finished drafts should not reach production.
Stage 8–9: Style lock and reference image discipline
Consistency starts when all assets share the same style path. A show should not have anime characters appearing in live-action-looking scenes because each asset picked its own style.
A practical system uses built-in style manuals covering directions like urban live-action, period live-action, mature urban romance animation, 1990s anime, Chinese ink style, xianxia fantasy, 3D donghua, stop-motion clay, and cyber-Chinese aesthetics. Each manual should include:
- Character image rules: facial anchors, material, temperament, view consistency
- Scene and prop rules
- Video style tags
Once a style is chosen, character art, scene art, prop art, and video prompts all follow the same style path.
The most important consistency rule is simple but strict:
When a character has an approved reference image, do not describe that character's clothing or appearance again in text. The reference image is the authority. Text only describes action, expression, and injury.
This rule directly prevents the classic AI failure where prompt text says one outfit and the reference image shows another, causing costume swaps or face drift.
Reference mapping for each scene should follow a fixed priority:
- Scene reference
- Props appearing in that scene
- Character reference art
Slots are only filled when an image exists. If there is no image, that should be explicit rather than invented. Teams can also manually bind script names to cast entries, include or exclude props, swap in historical versions of an image, and route period looks for time-travel or flashback scenes.
Scene and prop discipline matters too:
- Scene headers are parsed from script structure into interior/exterior, location, and day/night.
- Establishing shots should not include people.
- Props are extracted from the script using the script's original naming, then merged into a show-wide catalog.
- User or script descriptions override style-manual content restrictions; the manual controls how things are drawn, not whether they may exist.
Before shooting, there should be a readiness check for missing scene, character, and prop references. Teams may choose to proceed with text-only prompts, but that should be a conscious trade-off, not a silent default.
Stage 10: Cut episodes into scene blocks
Trying to generate a whole episode as one long clip is one of the fastest ways to lose control. A better unit is the scene block.
A scene block is a short production unit, roughly ten seconds long, generated and reviewed independently.
In practice:
- Episodes are split into blocks targeting about 10 seconds each.
- Scene text has a soft ceiling around 200 Chinese characters; longer scenes are split by action beats, paragraph breaks, or sentence endings.
- Extras and generic roles are separated from drawable lead characters so they do not consume reference slots.
- Blocks can be re-split for an entire episode if needed.
This is the production equivalent of seeing clip bins on an editing timeline. Instead of "Episode 3," you see Episode 3 → Scene 1, Scene 2, Scene 3. Each scene has its own prompt, reference set, output history, and revision path.
Vertical format should be the default deliverable, 9:16, not a horizontal video cropped after the fact. Other aspect ratios can be supported, but vertical is native.
Stage 11: Engineering shot prompts like a storyboard, not a novel
Video models respond poorly to vague literary prompts. They respond better when prompts read like a structured shot plan.
A strong shot prompt contains eight elements:
| Element | What it does |
|---|---|
| 1. Precise subject | Who or what is in frame |
| 2. Action detail | Exact movement, quantified where possible |
| 3. Scene environment | Location and surrounding context |
| 4. Light and color | Mood, time of day, tonal palette |
| 5. Camera movement | One move per shot, no stacked camera chaos |
| 6. Visual style | Consistent look of the show |
| 7. Image quality | Stability, clarity, face consistency |
| 8. Constraints | Negative controls and anti-artifact rules |
For simple scenes, this can be one paragraph. For complex cinematic scenes, use three sections: overall setup, shot-by-shot instructions, and a constraint package.
Prompt rules that reduce failure:
- One camera move per shot.
- Use shot numbers, not absolute timestamps like "0–3s."
- Include a baseline constraint package for image quality, face stability, and no watermark or logo.
- For multi-character scenes, add twin / duplicate prevention.
- For non-realistic styles, anchor the style explicitly.
- Favor slow, continuous motion over high-energy action that breaks easily.
- Use clear notation for dialogue, sound effects, and music.
- Feed only the current scene's assets into the prompt; do not leak other scenes.
- Treat the current scene's script text as the highest-priority source.
- Keep generation parameters conservative: stable and controllable beats wildly unpredictable.
The system should teach the model to speak storyboard language, not novel language.
There is also a compliance step before prompts leave the writing stage: specific copyrighted work or IP names are stripped out, while technique and aesthetic descriptions remain. If a reference image is detected as a suspected real-person photo, the workflow should return clear guidance to use platform-drawn art instead of uploading a real likeness as reference.
Stage 12–13: Shoot takes, review, and reshoot with intent
Production should behave like a real set, not a lottery ticket.
The standard shooting path is:
- Open the episode's video workspace.
- Confirm references are complete, or consciously proceed without them.
- Generate shot prompts.
- Read and edit the prompts.
- Choose model tier, aspect ratio, resolution, and duration.
- Submit for rendering.
- Review historical takes.
- Select the best result or reshoot with revised instructions.
The professional control points map cleanly to crew roles:
| Control point | Crew equivalent |
|---|---|
| Edit video prompts | Director revising shot notes |
| Swap character / scene / prop references | Changing costume approval or location board |
| Manually bind roles | Fixing nickname / cast mismatches |
| Include or exclude props | Controlling visual focus |
| Switch style manual | Unifying show look |
| Choose model tier | Balancing quality, speed, and cost |
| Compare historical takes | Multi-take selection |
A production queue should also behave like production infrastructure:
- Videos are processed in a separate queue from scripts and art.
- The same scene cannot submit parallel runs mid-task, preventing duplicate charges and state confusion.
- Credits are pre-deducted and released on failure.
- Stuck tasks time out and can be retried.
- Existing vendor task IDs can be polled instead of creating duplicate jobs.
- Final videos are faststart-enabled and stored with measured runtime.
This is what makes the workflow operational rather than experimental.
Where human judgment remains mandatory
It is important to say what the pipeline does not do.
- There is no automatic final scoring engine that magically picks the best take. Final quality judgment remains with the creator or supervisor.
- Reference images are not a hard gate. Missing images trigger a warning, but text-only shooting is allowed; quality is usually worse, so professional teams should approve references first.
- Reference tags in prompts rely on disciplined structure, not invisible magic. Creators should still review prompts before shooting.
- The ~10-second scene block is an engineering heuristic, not a frame-accurate edit. Long scenes may need further trimming in post.
- There is no current cross-scene automatic video continuation workflow. The product unit is "single scene with references → single rendered clip."
- Character consistency depends on the asset pipeline, not hidden face-embedding verification. Final look still depends on reference art quality and prompt discipline.
That is not a weakness to hide. It is the honest boundary of a production system.
The production principle
High-quality short drama is not generated as one long miracle clip. It is built from controlled units: a locked brief, a living continuity bible, approved cast and art, scene blocks, shot-level prompts, multiple takes, and deliberate reshoots.
Maosika (猫斯卡) is built around this principle. It is an AI production operating system for vertical short dramas, not a single text-to-video button. It encodes the staged workflow above into a crew-style pipeline with 18 digital specialists covering development, writing, art, camera, and post roles, while keeping human approval at the key checkpoints.
The aim is straightforward: turn the repetitive, drift-prone, failure-prone parts of AI short drama production into a constrained pipeline, while taste and final judgment stay with the people making the show.
If you are exploring this approach for your own slate, you can see the workflow at Maosika.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com