The Vertical Short Drama Pipeline: What to Gate at Every Stage From Idea to Final Cut
Most AI short drama failures are not model failures — they are skipped gates. The fix is to treat each stage like a real production: lock the brief, archive continuity, approve lookdev, split into field blocks, then shoot.
The problem is not generation — it is order
Teams that treat AI short drama as a single "text to video" button usually hit the same wall: characters change face mid-episode, props vanish, episode 12 forgets a secret planted in episode 3, and the opening scene wastes the golden three seconds on exposition.
The root cause is not weak models. It is missing production gates. A live-action crew would never shoot without a locked brief, cast look references, or a shot list. AI workflows need the same discipline — just implemented as structured checkpoints instead of meetings.
A vertical short drama pipeline is a sequence of gated stages where each stage produces a reviewable artifact before the next stage is allowed to start.
This article walks through those gates in order, using the language of a real set: greenlight, writers' room, makeup and wardrobe, blocking, shooting, and editorial. We will note what must be locked, what can be iterated, and where human judgment is non-negotiable.
Stage 1 — Intake and greenlight: do not shoot a half-baked idea
The first gate is idea completeness. A one-sentence pitch like "a CEO falls for a maid" is not a brief; it is a trope. Before any writing begins, the production needs concrete answers across the vertical drama dimensions:
- Genre and tone
- Protagonist and their character arc
- Core conflict
- Story direction and ending shape
- Episode count and episode length
- Hook and payoff rhythm
- Platform and audience
A useful heuristic is to score intake completeness from 0 to 100. Above 80, the idea is ready to become a locked brief. Between 40 and 79, only the missing dimensions need to be filled in. Below 40, the idea needs a full development pass. The threshold should be enforced by process, not left to the model to "decide" it has enough.
Think of this gate as the greenlight meeting. If the logline, core conflict, hook rhythm, and ending direction are not nailed down, every later stage inherits the ambiguity.
What to lock at greenlight
| Artifact | Why it matters |
|---|---|
| Logline | One-sentence north star for every future decision |
| Core conflict | Prevents the story from drifting into side plots |
| Character arc for leads | Gives the protagonist somewhere to go across episodes |
| Hook and payoff rhythm | Defines where cliffhangers and big beats land |
| Ending direction | Avoids writing into a corner at the midpoint |
| Platform and audience | Determines pacing, episode length, and tone |
Stage 2 — The writers' room: plan before you write pages
Vertical short drama writing has hard rhythm rules that do not apply to long-form. Episodes are short — typically 1 to 2 minutes, with a hard ceiling around 3 to 4 minutes — and the audience swipes away in seconds.
The four writing rules that should be non-negotiable
- Golden three seconds. The first scene must open on conflict or suspense. No slow establishing shots, no voiceover backstory dumps.
- Single-episode structure. Every episode follows: opening hook (one scene), escalating conflict, end-of-episode cliffhanger (the final scene).
- Payoff density. At least one small payoff per episode — a face slap, a reveal, an identity hint, evidence secured. A larger payoff every few episodes.
- Tight dialogue. Short lines, generally under twenty words. No essay-style speeches or narrator exposition.
Plan the beats, then write the pages
The biggest structural mistake is generating full episodes directly from the brief. A professional workflow writes a beat sheet for each episode first — scene-by-scene beats plus the end-of-episode cliffhanger — and only then expands into dialogue and action.
Writing should also happen in batches of several episodes, not one infinite stream. Between batches, four things should be re-confirmed:
- The batch's plot direction
- Which characters carry the batch
- Which open threads advance and which new ones are planted
- The episode-end hooks
These become hard constraints for the next batch. If they conflict with older notes, the freshly confirmed intent wins.
The continuity bible: do not rely on model memory
A continuity bible is a structured, living record of everything that must stay consistent across an entire series: character identities, stable traits, current state (injuries, revealed identities, changed relationships), relationship maps, open and resolved plot threads, per-episode appearance tables, batch summaries, and prop visual descriptions.
This is the single most important artifact for long-running short dramas. Models forget. A bible does not. When writing a new batch, the system should only pull the relevant slice — current character state, unresolved threads, recent batch summaries, and the current batch's main line — instead of dumping the entire history into context and hoping for the best. After each batch, the bible is updated.
Rule-based quality check
Finished episodes should pass a rule-based checker before anyone reads them as deliverable. Common failures to catch:
- Wrong episode titles
- Scene count mismatches
- Missing character headers
- Too little dialogue
- Placeholder text like "to be continued" used as a cliffhanger
A failed batch should be sent back for rewrite with the specific error attached. The point is not that AI is perfect on the first pass; the point is that known structural failures should be caught by rules, not by a tired producer reading at 2 a.m.
Stage 3 — Look development: lock the cast, wardrobe, and world before shooting
Character inconsistency — the famous "new face every scene" problem — happens because each generation re-imagines the character from text. The solution is to build visual assets first and then reference them, not describe the character's outfit in every prompt.
Pick a visual style once
A production should choose one style handbook that covers character art, scene art, props, and video style tags together. Mixing styles mid-show — anime characters in a realistic world, 3D one scene and 2D the next — breaks immersion instantly. A useful production system ships with multiple handbooks spanning 2D, 3D, and realistic directions (urban realism, period realism, 90s anime, Chinese ink, xianxia, cyberpunk, claymation, and others), but the key rule is: once chosen, character art, scene art, prop art, and video prompts all share the same style path.
Two-pass character art
Character look development works best in two passes:
- Text polish. Turn the bible's character description into a drawable prompt using the chosen style handbook, with gender treated as a hard rule. This step is human-overridable.
- Image generation. Generate the final art, optionally with a reference image, and save it as a versioned asset.
Only completed, readable art assets should be consumed by the video pipeline. Half-finished character sheets should never be silently bound to a scene.
The reference image rule that prevents face swaps
Here is the single highest-leverage rule for character consistency:
For any character with a reference image, the prompt must not describe their clothing or appearance in text. The reference image is the source of truth for how they look; text only describes action, expression, and injury state.
When text and reference fight, the model invents a compromise — which is how you get wardrobe swaps and face drift. Removing the redundant text eliminates the conflict.
Scenes and props
Scenes should be parsed from the script into structured location tags (interior/exterior, place, day/night) rather than manually re-entered. Scene art must follow an empty-plate rule: no characters in scene reference images, so the shot is not polluted by a random figure.
Props should be extracted per episode from the script's original wording — using a props-department mindset — and merged into a show-wide catalog so nothing is lost between episodes.
Pre-shoot readiness check
Before any scene is sent to generation, run a readiness check: are the scene, character, and prop references for this scene complete? If anything is missing, surface a list. Teams can choose to skip and shoot with text only, but they should do so knowingly — text-only shots are usually lower quality, and the professional path is to finish lookdev first.
Stage 4 — Blocking: split the episode into shootable field blocks
A 60-second episode is not "one video." It is a sequence of short, controllable units.
A field block is a single shootable unit of an episode, roughly 10 seconds long, with its own prompt, its own bound references, its own output, and its own history of takes.
Splitting follows simple heuristics: target around 10 seconds per block, with a soft cap around 200 characters of script body. Longer beats are split on action turns, paragraph breaks, or sentence endings. Crowd and generic characters are separated from the named cast so they do not consume reference slots.
This is the equivalent of a clip list on an editing timeline. Instead of "Episode 3," the director sees "Episode 3, Scene 1, Scene 2, Scene 3…" — each independently re-shootable.
Default delivery is vertical
Vertical format should be the default shape of production, not a crop applied after the fact. Standard output is 9:16, with configurable duration (smart or fixed 5–15 seconds), model tier (standard / fast / lightweight), resolution within model allowlists, optional audio, and watermark off by default.
Stage 5 — Shot language: engineer prompts like a shot list, not a novel
Prompts are where most teams quietly lose control. A good prompt for video generation reads like a storyboard note, not a paragraph of fiction.
The eight elements of a shot prompt
Every engineered prompt should cover:
- Precise subject
- Action detail
- Scene environment
- Lighting and color tone
- Camera movement
- Visual style
- Image quality
- Constraints
Simple scenes can be written as one block. Complex, cinematic scenes use a three-part structure: overall setup, then shot-by-shot beats, then a constraint package.
Prompt rules that reduce failure rate
- One camera move per shot. Do not stack push, pull, pan, and tilt into a single shot.
- Use shot numbers, not absolute timestamps. Write "Shot 1," "Shot 2," not "0–3s."
- Always include a fallback package. Image quality, face stability, no watermark or logo. For multi-person shots, add anti-twinning constraints. For stylized work, anchor the style explicitly.
- Describe action physically and quantitatively. Favor slow, continuous motion over high-dynamic action that tends to break.
- Use notation for audio. Dialogue in braces, sound effects in angle brackets, BGM in parentheses.
- Only feed this scene's assets. Never leak another scene's references into the prompt. The current scene's script body is the highest-priority source.
- Keep generation parameters conservative. Stability and control beat wild creativity on a production line.
Before delivery, prompts should be cleaned of specific copyrighted IP or title names — keeping the technique and aesthetic description — to reduce downstream blocking risk.
In one sentence: the system should teach the model to speak in shot-list language, not novel prose.
Stage 6 — Shooting and editorial: multi-take, review, iterate
The shoot itself should feel like a real set, not a magic button.
The standard path is: open the episode's video workspace, confirm references are ready (or intentionally skipped), generate the shot prompts, read and edit them as a director would edit storyboard notes, choose model tier / aspect ratio / resolution / duration, submit to render, then review takes and pick the best.
The professional control points
| Control point | On a real set, this is… |
|---|---|
| Editing the shot prompt | The director revising storyboard notes |
| Swapping character / scene / prop references | Changing a look or swapping a location plate |
| Manual character binding | Fixing name mismatches or off-screen references |
| Including / excluding props | Controlling what is in frame |
| Switching style handbook | Unifying the visual language |
| Choosing model tier | Trading quality, cost, and speed |
| Reviewing historical takes per scene | Shooting multiple takes and picking the best |
When you press generate, the system should use exactly the prompt you confirmed and the references that were locked at that moment — not silently rewrite them. If you change the prompt and regenerate, that is a new take, same as on a real set.
Production-grade reliability
For a tool to be usable at scale, not as a toy, it needs operational basics:
- Videos are faststart-processed and stored with measured runtime, not just the vendor's reported duration.
- Video rendering runs in an independent queue, isolated from writing and art tasks.
- The same field block cannot be submitted twice in parallel, preventing duplicate charges and state confusion.
- Credits are pre-deducted and settled; failed jobs release them automatically.
- Stuck jobs time out and become retryable with visible failure state.
- Existing vendor task IDs are reused for polling, never double-created.
These are boring details, and they are what separate a prototype from a production line.
What this pipeline does not do
Honest boundaries build trust with production teams. There are limits worth stating plainly:
- There is no automatic scoring or auto-pick of the best take. Final quality judgment sits with the creator and producer; the system gives you multiple takes and the tools to re-shoot with adjusted prompts.
- Reference images are not a hard gate. Missing images trigger a warning, but you can still shoot text-only — quality will usually suffer, which is why the professional path is to finish lookdev first.
- Reference tags in prompts rely on convention, not forced re-injection. Creators should still verify that tags are present when reviewing prompts.
- Shot grammar follows the platform's built-in camera rules. The style handbook injects video style tags; full art manuals apply on the lookdev side.
- The ~10-second field block is an engineering heuristic, not timecode-accurate editing. Overlong scenes are re-split, but blocks may still run long; final trimming and stitching happen in editorial.
- There is no productized cross-scene auto-extend workflow. The unit of work is "single field block with multi-modal references → single clip."
- Character consistency depends on the lookdev asset chain, not face-embedding verification. Era-specific costume routing uses rules, and final quality still depends on the art and on respecting the "do not describe appearance when a reference exists" rule.
Where Maosika fits
Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It does not try to replace the crew with a single generate button. Instead, it hard-encodes the gated order that professional short drama production already follows: brief locked first, then continuity bible, then lookdev, then field blocks, then engineered shot prompts, then multi-take review. Human creators make the审美 calls at each gate; the system makes sure those calls are carried consistently into every episode and every scene.
If you are building or scaling a vertical short drama slate, the question is not whether AI can generate video — it can. The question is whether your pipeline has the gates to turn that capability into repeatable, reviewable, re-shootable production. That is the gap an operating system approach is designed to close.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com