The AI Vertical Short Drama Pipeline: From Idea to Final Clip, Stage by Stage
Most AI short drama failures are not model failures—they are pipeline failures. The fix is not a better one-click button, but enforcing the same control order a real crew uses: brief first, then continuity, then look, then shots, then takes.
Why most AI short dramas fall apart
Teams often treat AI short drama production as a single long generation job: type in a premise, wait for a video, then edit around the problems. That works for a test clip. It does not work for a 60-episode vertical drama that needs consistent characters, coherent plot threads, and a release schedule.
The recurring failures are familiar:
- Episode 4 forgets a secret set up in Episode 2.
- The lead’s face or outfit changes between scenes.
- The opening is slow because the model spent time explaining backstory.
- One scene looks like anime while the next looks photoreal.
- A prop appears, disappears, or changes shape without warning.
- The team cannot tell which version of a clip is the approved take.
These are not random accidents. They happen when a production skips the intermediate artifacts that normally keep a crew aligned.
A production-ready AI pipeline is defined by one principle: every stage produces a reviewable deliverable before the next stage starts.
The full pipeline, in order
Below is the control order used by a structured AI vertical short drama workflow. The order matters more than the tools.
| Stage | Main deliverable | What can go wrong if skipped |
|---|---|---|
| 1. Idea evaluation | Intake score + guided brief path | Vague premise causes drift later |
| 2. Creative briefing | Locked logline, conflict, hooks, ending direction | Writers and video team optimize for different goals |
| 3. Story archive | Continuity bible: characters, relationships, threads, episode appearances | Plot amnesia, personality drift, dropped clues |
| 4. Batch script writing | Beat sheets, episode drafts, cliffhangers, rule checks | Weak hooks, flat pacing, filler dialogue |
| 5. Look selection | Shared style manual across art and video | Visual style breaks between assets and clips |
| 6. Look development | Character, scene, and prop art | Face swaps, costume drift, missing locations |
| 7. Scene blocking | ~10-second scene blocks per episode | Long unmanageable clips, hard retakes |
| 8. Shot prompts | Per-scene reference mapping + shot instructions | Text and reference images fight each other |
| 9. Rendering | Queued clips with versioned takes | Silent failures, duplicate costs, no audit trail |
| 10. Review and retake | Approved takes, edited prompts, swapped references | Teams accept weak clips because reruns feel expensive |
This is not a theoretical framework. It mirrors how live-action crews already work: development lock, bible, script, art, shot list, shoot, dailies, pickups.
Stage 1: Idea intake and creative brief
The first gate is not scriptwriting. It is deciding whether the idea is specific enough to write.
A strong intake process scores the idea across the dimensions that matter for vertical drama: genre, protagonist, core conflict, story direction, episode count, episode length, tone, hook rhythm, ending direction, platform, and target audience. If the idea is incomplete, the system should ask for the missing pieces instead of guessing.
The brief is locked only when it contains the fields a writer actually needs:
- One-sentence logline
- Core conflict
- Story direction
- Ending direction
- Beat rhythm for payoffs and hooks
- Platform and audience constraints
- Episode-level outline
- Notes for the writer
- Protagonist arc
A creative brief is the written version of a greenlight meeting. If the brief is not locked, every later stage is improvising.
Stage 2: Story archive, not model memory
A large language model does not "remember" your drama across 80 episodes. It hallucinates consistency when context becomes noisy.
The production solution is a structured story archive.
A story archive is a serialized continuity record that tracks character identities, stable traits, current states, relationships, open and resolved plot threads, episode appearance tables, batch summaries, and visual prop notes.
Instead of dumping the entire project into one context window, the script stage should only receive the slice it needs:
- Current character states
- Unresolved threads
- Recent batch summaries
- The current batch’s main line
After each batch, the archive is updated. Before the next batch, the archive is loaded again. This is how long-form serialized drama avoids continuity rot.
Stage 3: Batch script writing for vertical rhythm
Vertical short drama scripts are not compressed feature films. They have strict rhythm rules.
A production-grade script pipeline enforces them as checks, not suggestions:
- Golden 3 seconds: the first scene must open with conflict or suspense, no slow exposition.
- Episode shape: opening hook → escalation → end-of-episode cliffhanger.
- Payoff density: at least one small payoff per episode; a larger payoff every few episodes.
- Dialogue: short lines, no essay-like speeches, no narrator explaining what the scene should show.
The writing order also matters. Before drafting dialogue, each episode should first produce a beat sheet and confirm the cliffhanger. After drafting, scripts pass rule-based checks for missing character lines, wrong scene counts, placeholder text, or too little dialogue. Failed drafts are sent back with explicit error feedback.
After the midpoint of the series, the system should inject ending constraints so the plot does not wander indefinitely. From the second batch onward, the next batch direction should be confirmed in four areas: plot direction, focus characters, threads and conflicts, and end hooks. Those confirmed choices become hard constraints.
Stage 4: Look development before shooting
Many teams try to generate video directly from text and then wonder why characters change face. The problem is not the model alone; it is the absence of locked visual assets.
A robust pipeline separates look development from rendering:
- Choose one style path for the whole show.
- Generate character art with consistent face anchors, costume, material, and mood.
- Generate scene art without characters in the frame.
- Extract props using the script’s own naming, then build a show-wide prop catalog.
- Use the same style path for characters, scenes, props, and video prompts.
This is the equivalent of a locked lookbook before principal photography.
The most important consistency rule is simple but strict:
When a character has approved reference art, the prompt must not re-describe that character’s clothing or appearance in text. The image owns the look. The text only describes action, expression, and injury state.
That one rule eliminates a large share of AI costume swaps and face drift.
Stage 5: Scene blocks, not one long episode video
Vertical drama episodes are usually 1–2 minutes long, sometimes extending to 3–4 minutes. Trying to generate a whole episode as one clip creates uncontrolled motion, weak continuity, and impossible retakes.
Instead, each episode should be cut into scene blocks of roughly 10 seconds each. Long passages are split further by action beats, paragraph breaks, and sentence boundaries.
A scene block is the smallest production unit:
- One block has its own script slice.
- One block has its own reference set.
- One block has its own shot prompt.
- One block has its own render queue entry.
- One block keeps its own history of takes.
This is how editing software treats clips, and it is how AI production should treat generation. If one shot fails, you rerun that shot, not the whole episode.
Stage 6: Shot prompts written like a shot list
A good AI video prompt is not a paragraph of prose. It is a shot instruction.
A strong prompt structure includes eight components:
| Component | Purpose |
|---|---|
| Subject | Who or what is in the shot |
| Action | What happens, with quantified motion |
| Environment | Where the shot takes place |
| Lighting and color | Mood, time of day, tonal palette |
| Camera movement | One move per shot, not stacked moves |
| Visual style | Anchored to the chosen look |
| Quality guardrails | Stability, face clarity, no watermark |
| Constraints | No extra characters, no style drift, no duplicate faces |
For complex scenes, use a three-part structure: overall setup, numbered shots, then constraint package. Use shot numbers instead of absolute timestamps. Keep motion continuous and moderate; high-action bursts are still a major failure point in video models.
The prompt should only see assets for that scene. It should not pull props or characters from another scene. The script text for that scene remains the highest-priority source.
Before prompts are sent, IP names should be stripped out, leaving only technique and visual language. If a reference image appears to be a real-person photograph, the workflow should flag it instead of silently feeding it into the renderer.
Stage 7: Rendering as a production queue
Rendering is where hobby tools and production tools diverge.
A production queue needs operational controls:
- Separate lanes for script, art, and video tasks
- No duplicate submissions for the same active scene block
- Pre-charge and settlement logic with automatic release on failure
- Timeout recovery for stuck jobs
- Continued polling when a vendor job already exists
- Actual measured clip duration stored after processing
- Multiple takes preserved for selection
This is not glamorous, but it is what makes a system usable by a studio. A team needs to know what failed, what can be retried, what was charged, and which take is current.
Stage 8: Review, retake, and human judgment
AI does not replace the director or editor.
At review stage, the team should be able to:
- Edit the shot prompt like revising a shot list
- Swap character, scene, or prop references
- Manually bind script names to approved characters
- Include or exclude props from a scene
- Switch model quality tiers for speed, cost, or look
- Compare historical takes of the same scene block
There is no reliable automatic "best clip" judge in current AI workflows. The final quality decision belongs to the creator. The system’s job is to make retakes cheap, traceable, and specific.
What this pipeline does not do
It is important to be clear about the limits.
- There is no automatic quality scoring or auto-select best take. Final judgment remains human.
- Reference images are not a hard gate. You can skip them and render from text, but quality usually drops; professional workflow should lock art first.
- Reference tag consistency depends on prompt discipline. Creators should still review that tags are present.
- Shot grammar is handled through the platform’s built-in camera rules. Style manuals mainly control art-side consistency.
- ~10-second blocks are heuristic, not timecode-exact editing. Final assembly still needs editing.
- There is no mature cross-scene automatic video continuation workflow yet. The unit remains one scene block to one clip.
- Character consistency depends on the art pipeline, not hidden face verification. Good look dev and prompt discipline still determine the result.
That honesty matters. The goal is not to claim AI can run a crew by itself. The goal is to turn the repetitive, drift-prone, failure-prone parts of production into a controlled pipeline while keeping审美 judgment with humans.
A practical checklist for teams
If you are evaluating or building an AI short drama workflow, use this checklist:
- Can the idea be rejected or sent back for more detail before writing starts?
- Is the brief locked before script generation?
- Is continuity stored as a structured archive, not left to model memory?
- Are scripts written in batches with beat sheets and cliffhangers first?
- Do scripts pass explicit rule checks before delivery?
- Are characters, scenes, and props art assets created before video rendering?
- Are episodes split into ~10-second scene blocks with independent takes?
- Do prompts follow shot-list structure instead of prose?
- Are text descriptions prevented from fighting approved reference art?
- Can you review, edit, rerender, and compare takes without chaos?
If the answer to several of these is no, you are not running a pipeline yet. You are running a prototype.
Where Maosika fits
Maosika (猫斯卡) is built around this exact stage order: idea evaluation, locked brief, story archive, batch script writing, look development, scene blocking, engineered shot prompts, queued rendering, and take-based retakes. Its 18 digital specialists mirror real crew roles across development, script, art, camera, and post, and every stage leaves a reviewable artifact rather than hiding decisions inside a black box.
It is not a one-click drama generator. It is an AI production operating system for vertical short dramas: the workflow enforces the sequence that professional crews already know works, while people keep control over story, look, and final selection.
If you want to explore the workflow in practice, you can start at https://www.maosika.com.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com