The AI Vertical Short Drama Pipeline: From Logline to Locked Cut in Industrial Order
Most AI short dramas fall apart not because the model is bad, but because the steps happen in the wrong order. Lock the brief before the script, lock the look before the shot, and treat every scene as its own unit with its own references and takes.
If you've watched enough AI-generated vertical dramas, you already know the failure modes by heart. The lead's face changes between episodes. Her jacket is red in scene 1 and blue in scene 3. A prop introduced in episode 2 vanishes by episode 6. The opening 20 seconds are backstory narration instead of a hook. Episode 8 ends on a cliffhanger that episode 9 completely forgets.
None of these are "model problems." They are pipeline order problems. A real crew doesn't walk onto set and improvise a story; a functional AI pipeline shouldn't either. This article walks through the industrial sequence that actually holds a vertical drama together — greenlight, brief, script, continuity bible, look development, scene blocking, shot prompts, shooting, and retakes — and explains why each gate exists.
Why order matters more than model choice
Vertical short dramas (typically 1–2 minutes per episode, sometimes stretching to 3–4) live or die on two things the current generation of video models is bad at on its own: consistency across dozens of episodes, and density of payoff inside a tight runtime.
A stronger model can give you a prettier single clip. It cannot, on its own, remember that your protagonist has a scar on her left wrist from episode 3, that the villain is supposed to be wearing a jade pendant in every modern-era scene, or that episode 12 needs to plant the clue that pays off in episode 15. Those are not model capabilities. They are pipeline responsibilities.
The rule of thumb is simple:
Anything you want the show to remember across episodes must be written down in a structured artifact before the shot is generated, not left to the model's context window.
This is the core difference between "playing with an AI video tool" and "running an AI short drama production line."
Stage 1 — Intake and creative brief lock
The first gate isn't writing. It's deciding whether the idea is ready to be written at all.
A useful intake scores the idea across the dimensions that actually define a vertical drama: genre, protagonist, core conflict, story arc, episode count, episode length, tone, hook and payoff rhythm, ending direction, and target platform / audience. Ideas that come in half-formed get routed through a guided fill-in process; ideas that are already complete can skip straight to the brief.
The creative brief is a locked document, not a suggestion. It contains, at minimum:
| Brief field | What it pins down |
|---|---|
| Logline | One-sentence premise |
| Core conflict | What the protagonist is fighting |
| Story arc | Where the show starts and where it ends |
| Ending direction | Rough landing zone, so mid-series writing doesn't drift |
| Payoff & hook rhythm | How often a beat lands; where cliffhangers sit |
| Platform & audience | Which viewers, which platform conventions |
| Episode-by-episode outline | Rough shape of each episode |
| Notes to the writer | Hard constraints, tonal guardrails |
| Protagonist with character arc | Who they are at start vs. end |
The brief must be locked before any script is written. This is the equivalent of finishing the development meeting before the writers' room starts. Skipping this gate is the single most common reason AI dramas meander: the model is being asked to invent both the story and the scene at the same time, with no north star.
Stage 2 — Batched script writing with a continuity bible
Vertical dramas are not written as one long document. They are written in batches of a few episodes at a time, with a structured record — a continuity bible — carried between batches.
What a continuity bible actually is
A continuity bible is a structured, show-long record that tracks:
- Character identities and stable personality traits
- Mutable current state (injuries, revealed identities, changed allegiances)
- Character relationships
- Every planted thread, marked open or resolved
- An episode-by-episode appearance table
- Per-batch plot summaries
- Visual descriptions of recurring props
The continuity bible is how episode 200 remembers what episode 3 planted. It replaces "hoping the model doesn't forget" with "reading from the same document every time."
When a new batch is written, the writer only consumes a slice of the bible: current character states, unresolved threads, recent batch summaries, and this batch's main arc. After the batch is finished, the bible is updated before the next batch starts.
The structural rules that vertical scripts must obey
Writing for 1–2 minute vertical episodes is a different craft than writing a feature. The rules are non-negotiable because the format punishes slack:
- Golden 3 seconds. The first scene must open on conflict or suspense. No slow build, no narrated backstory.
- Per-episode structure. Opening hook (1 scene) → escalating conflict → end-of-episode cliffhanger (final scene).
- Payoff density. At least one small payoff per episode (a reveal, a reversal, a hint of identity, a piece of evidence landing); a larger payoff every few episodes.
- Dialogue. Short sentences, generally under 20 Chinese characters / a tight English breath. No essay-style monologue, no explanatory voiceover dumping lore.
Before any batch's dialogue is written, a beat sheet is laid out per episode — including the exact cliffhanger the episode ends on. This is the structural spine; the dialogue is flesh on top of it. Past the halfway point of the series, ending constraints are injected to force threads toward the planned resolution rather than wandering.
Rule-based quality checks
A finished batch goes through a rule checker that catches mechanical failures before a human ever reads it: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, placeholder text like "to be continued." Failed batches are sent back for an automatic rewrite with the error feedback attached. Nothing half-finished gets handed to the production side.
Stage 3 — Look development and character lock
Once the script is stable, the show gets its face. This stage is look development, and it's where most AI productions quietly die.
A show picks one visual style from a shared style library — 2D, 3D, photorealistic, period, stylized, and so on — and that same style path is reused for character art, scene art, prop art, and video prompts. The point is to prevent the uncanny断层 where characters look like anime but the video output looks photoreal, or vice versa.
Character art is a two-step process
- Text polish. The character's bible description is rewritten into an image-generation prompt using the chosen style's character sheet rules, with gender enforced as a hard constraint. The creator can override this by hand.
- Image generation. The final prompt is rendered, optionally with a reference image, and saved as a versioned asset. Old versions are kept so a creator can roll back.
Characters are not allowed to proceed to shooting until the lineup is complete, names are valid, visual fields are filled, and the leads match the brief. This is a hard gate, not a suggestion.
Scenes and props
- Scenes are parsed from the script into structured location tags (interior/exterior, place, day/night), not typed into a separate spreadsheet by hand.
- Empty scene plates must contain no people. A scene image with a random character in it will pollute every shot that uses it.
- Props are extracted episode by episode from the script's actual wording — the way a props master would read it — and merged into a show-wide catalog so nothing is lost between episodes.
- When there's a conflict between the script's description and the style guide's default world rules, the script wins. The style guide controls *how* things are drawn, not *what* is allowed to exist.
Stage 4 — Scene blocks: the editing-bin view
A finished episode is not sent to the video model as one long prompt. It is cut into scene blocks targeting roughly 10 seconds each, with a soft cap on block length; longer beats are split along action, paragraph breaks, or sentence boundaries.
This is the single most important structural decision on the video side. Instead of "episode 3, one big video," the creator sees:
`` Episode 3 ├─ Scene 1 — prompt, references, takes ├─ Scene 2 — prompt, references, takes ├─ Scene 3 — prompt, references, takes └─ ... ``
Each block has its own prompt, its own set of reference images, its own output clip, and its own history of takes. This is how a real editing bin looks, and it's the only shape that makes retakes usable: if scene 3 is bad, you reshoot scene 3. You don't regenerate the whole episode and hope scenes 1, 2, and 4 survive.
Default delivery is 9:16 vertical, baked in from the start rather than cropped after the fact. Other aspect ratios are supported, but vertical is the native shape of the format.
Stage 5 — Reference mapping: the consistency engine
Reference mapping is the mechanism that actually prevents face-swaps and costume changes between scenes. It's worth understanding in detail because almost nothing else matters if this step is wrong.
For every scene block, a reference table is automatically assembled in a fixed order:
- Scene art
- Props appearing in this scene
- Character look-dev portraits
Slots are only filled if an image exists; empty slots are marked as text-only rather than silently invented. The creator can manually bind a script reference ("the officer") to a specific character's portrait, manually include or exclude props to control visual focus, swap in a different version of an existing asset, or swap between era-specific looks for time-travel / flashback stories.
The single highest-impact rule on the entire video side:
For any character with a reference image, the prompt is forbidden from re-describing their clothing or appearance in text. The reference image is authoritative; text only describes action, expression, and injury state.
This rule exists because the classic AI video failure — face changes, outfit swaps, accessories appearing and disappearing — happens when the text prompt and the reference image disagree. If the reference says "black suit" and the text says "blue jacket," the model has to pick, and it picks differently every shot. Stripping appearance text for referenced characters eliminates the conflict.
Before prompts are generated, a pre-flight check lists any missing scene, character, or prop references. The creator can choose to proceed with text-only — but this is a flagged decision, not a silent default, because quality is almost always worse without references.
Stage 6 — Engineered shot prompts
The prompts sent to the video model are not freeform paragraphs. They are structured shot instructions, closer to a storyboard page than a novel.
A good shot prompt covers eight elements:
| Element | Purpose |
|---|---|
| Precise subject | Who or what is in frame |
| Action detail | What they are doing, with quantified motion |
| Scene environment | Where they are |
| Lighting & color | Mood and time of day |
| Camera movement | One move per shot, never stacked |
| Visual style | Inherited from the chosen look |
| Image quality | Baseline fidelity constraints |
| Negative constraints | Stability, no watermark, no duplicates, etc. |
Complex cinematic scenes use a three-part structure — overall setup, then shot 1 / shot 2 / shot 3, then a constraint pack — rather than one dense paragraph. Shots are numbered, never timestamped ("shot 1," not "0–3s"). Motion favors slow, continuous action over high-energy bursts, which are where video models most often break. Dialogue, sound effects, and music use consistent notation so they can be routed correctly.
The system is teaching the model to speak the language of a storyboard, not the language of a novel.
Prompts are also cleaned before delivery: specific copyrighted work or IP names are stripped out (keeping the stylistic and cinematic descriptors), and prompts that would reference a real-person photo are flagged with guidance to re-create the look through the art pipeline instead.
Stage 7 — Shooting, retakes, and operational reliability
The shooting loop itself is deliberately boring, because production tools should be boring:
- Open the episode's video workspace.
- Confirm references are present (or deliberately skip).
- Generate shot prompts.
- Read and edit the prompts — this is the director rewriting the storyboard note.
- Pick model tier, aspect ratio, resolution, duration.
- Submit to the render queue.
- Review historical takes per scene and pick the best one.
The professional control points map cleanly onto real crew roles:
| Control point | Crew equivalent |
|---|---|
| Editing the shot prompt | Director revising storyboard notes |
| Swapping character / scene / prop references | Swapping wardrobe or location plate |
| Manual character binding | Fixing name mismatches, cameo references |
| Including / excluding props | Controlling visual focus in the frame |
| Switching visual style | Unifying the art department language |
| Choosing model tier | Trading quality, cost, and speed |
| Reviewing historical takes | Multi-take selection on set |
Critically, the system does not reinvent prompts between takes. A take uses the exact prompt and reference map the creator confirmed or edited. Changing the prompt and hitting generate is what creates a new take — identical in spirit to "we changed the shot note, let's go again" on a real set.
On the operations side, the boring details are what make this usable as a production tool rather than a toy: clips are faststart-processed and stored with measured durations (not just whatever the vendor reports back), video tasks run on a separate queue from script and art tasks, in-flight scenes can't be double-submitted, credits are pre-deducted and released on failure, hung jobs time out and are retryable, and existing vendor jobs are polled rather than re-created.
What this pipeline does not do
It's worth being explicit about the boundaries, because any tool that claims to eliminate human judgment is selling something:
- There is no automatic scoring or auto-pick for the best take. Final quality judgment sits with the creator and the producer; the system gives you multiple takes and the tools to reshoot, not a robotic editor.
- Reference images are not a hard gate. You can skip them and go text-only; you should expect worse results when you do, which is why the professional flow locks look-dev before shooting.
- Reference tokens in prompts are guided by convention, not force-stitched. Creators should still give the prompt a final read to confirm references are present.
- Shot grammar follows the built-in cinematic spec. The full style manual is applied on the art side; the video side inherits style tags, not the entire art book.
- The ~10-second scene block is an engineering heuristic, not a timecode-precise cut. Over-long scenes get re-split, but some blocks will still run long; fine cutting and stitching into a final episode is a downstream edit step.
- There is no cross-scene automatic video continuation workflow. The unit of work is "single scene with multi-modal references → single clip."
- Character consistency depends on the look-dev asset chain, not a face-embedding verification. Era routing is rule-based; final fidelity still depends on the quality of the character art and on the prompt respecting the "don't describe appearance when a reference exists" rule.
These are not apologies. They are the shape of a tool that is honest about where human judgment still lives.
The shape of a working AI drama production line
Put back together, the order is:
`` Idea intake → guided development → brief lock → cast lineup & visual confirmation → continuity bible → batched script writing (beat sheet → draft → rule check → bible update) → style selection → character / scene / prop look-dev → episode cut into ~10-second scene blocks → per-scene reference map + engineered shot prompts → multi-modal generation (vertical by default) → review → prompt edit / reference swap → reshoot → pick best take ``
Every stage produces a tangible, reviewable, reversible artifact: the brief, the beat sheet, the script batch, the continuity bible, the character art, the scene plate, the prop catalog, the scene block, the shot prompt, the take. You can look at each one, change it, and roll it back.
The human is not removed from this loop. The human makes the decisions at the gates — is this brief right, is this character look right, is this take the one — and the machinery makes sure those decisions are faithfully carried into every subsequent episode and every subsequent scene, instead of being forgotten the moment the context window rolls over.
That is the bet behind Maosika (猫斯卡): not that AI replaces the crew, but that the parts of production which are repetitive, drift-prone, and easy to get wrong can be turned into a constrained, repeatable industrial line, while taste and judgment stay firmly on the human side of the screen.
If you're building or running an AI vertical drama operation, the question to ask about any tool is not "how good is its best-looking clip?" — it is "in what order does it force me to work, and which of those gates does it enforce versus leave to my memory?" The answer to that question is what determines whether your show holds together at episode 30.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com