The AI Vertical Short Drama Pipeline: From Logline to Locked Cut in Industrial Order
Most AI short drama failures are not model failures—they are order failures. Teams try to render video before the brief is locked, lock looks before characters are cast, and write episode 20 without a continuity record. The fix is a strict industrial sequence.
Why "end-to-end" usually means "out of order"
When a team says they made an AI short drama "end to end," what they often mean is: they typed a premise, got a script, fed it into a video model, and patched the result in editing. That works for a demo. It does not work for a 60-episode vertical slate that has to ship on schedule, keep characters recognizable, and land cliffhangers every episode.
The recurring failure pattern is always the same:
- The logline is vague, so the script drifts episode to episode.
- Characters are described differently in every scene, so faces and outfits change.
- Props mentioned in episode 4 vanish in episode 11 because nothing tracked them.
- Shot prompts are written like fiction instead of camera instructions, so motion breaks.
- Bad takes are treated as final because there is no queue, no versioning, no re-shoot workflow.
The root cause is not that the model is weak. It is that the production order is wrong.
In live action, you do not roll camera before the brief is signed, the cast is locked, and the shot list exists. AI drama needs the same discipline—only the crew is partly digital.
The industrial sequence, step by step
This is the order professional AI vertical drama production should follow. Each stage produces a reviewable artifact. No stage should be skipped by default.
| Stage | What gets produced | Who decides | What goes wrong if skipped |
|---|---|---|---|
| 1. Creative intake & evaluation | Intake score, missing-dimension list | Producer / creator | Premise is too thin to sustain episodes |
| 2. Creative briefing | Locked brief (logline, conflict, arc, hooks, platform, audience) | Creator + producer | Script drifts; no shared north star |
| 3. Story archive setup | Continuity bible: characters, arcs, relationships, threads, props | Writer + archive keeper | Characters forget who they are by episode 20 |
| 4. Batch script writing | Beat sheets first, then pages, then rule-based QC | Writer + showrunner | Weak cliffhangers, filler episodes, format errors |
| 5. Look selection | Style manual chosen across characters, scenes, props | Director / art lead | Person looks 2D, world looks photoreal |
| 6. Look dev & cast lock | Character, scene, and prop turnarounds | Art director | Face swaps, costume swaps, scene mismatch |
| 7. Scene blocking | Episodes cut into ~10-second scene blocks | Director / editor | Long unwieldy clips that cannot be re-shot cleanly |
| 8. Shot prompt engineering | Per-scene prompts bound to reference images | Director / DP | Model improvises appearance and motion |
| 9. Rendering & multi-take | Per-scene renders, take history, queue management | Producer / post | No way to compare or recover; silent failures |
| 10. Review, re-word, re-shoot | Selected takes, edit notes, locked cuts | Creator / editor | Bad output accepted as "AI being AI" |
The rest of this article walks through why each gate exists and how to run it in practice.
Stage 1–2: Intake and locked brief
A brief is not a paragraph of vibes. It is the document the whole downstream pipeline treats as a hard constraint.
The creative brief is a locked deliverable. Before a single page of script is written, it should contain at minimum:
- A one-sentence logline
- The core conflict
- Story direction and ending shape
- Hook and payoff rhythm
- Target platform and audience
- Episode count and per-episode length
- Tone and style baseline
- The protagonist's character arc
For vertical dramas, per-episode length typically lives in the 1–2 minute range, sometimes extending to 3–4 minutes. The format is unforgiving: if the first scene does not open with conflict or a hard question, the swipe comes before the beat lands.
A useful gate is an intake completeness score. A premise that scores high enough can move straight to brief; one in the middle gets only its missing dimensions filled in; a thin premise goes through a full guided development pass. The point is to prevent the model from politely pretending a half-baked idea is ready to shoot.
Definition: A locked creative brief is a written agreement between the creator and the production system about what the show is. Nothing downstream—script, art, camera—gets to silently renegotiate it.
Stage 3: The story archive (continuity bible)
Long-form AI drama has a memory problem. Models do not reliably remember a character's scar from episode 3, a hidden identity set up in episode 8, or a prop that must reappear in episode 24. Asking them to "remember" is not a strategy.
The solution is a structured story archive.
The story archive is a series-long structured record of characters, relationships, plot threads, appearances, per-episode cast, batch summaries, and prop descriptions. It is the AI-era equivalent of a writers' room continuity bible.
It should track, at minimum:
- Character identity and stable traits
- Mutable current state (injuries, revealed identities, changed allegiances)
- Relationship map
- Open vs. resolved plot threads
- Per-episode appearance sheet
- Per-batch plot summaries
- Prop visual descriptions
The rule is simple: do not rely on model memory; rely on structured records. When writing a new batch, the writer gets a slice of the archive—current character states, unresolved threads, recent summaries, and the batch's main line—not the entire noisy history. After each batch finishes, the archive updates.
This is how a 100-episode drama stays internally consistent without the showrunner manually re-reading every previous episode.
Stage 4: Batch script writing with vertical-format rules
Vertical short drama has its own grammar. It is not feature film structure compressed; it is a format optimized for hook density and the swipe.
Non-negotiable writing rules
- Golden 3 seconds. The first scene must open with conflict or a strong hook. No slow build, no expository voiceover.
- Per-episode shape. Opening hook → conflict escalation → end-of-episode cliffhanger.
- Payoff density. At least one small payoff per episode (a reversal, a face-slap, an identity hint, evidence obtained); a larger payoff every few episodes.
- Dialogue discipline. Short lines, generally under ~20 characters in Chinese source writing and kept tight in translation; no essay-like monologues.
Structure before pages
Before writing dialogue, each episode should first produce a beat sheet and its end-of-episode cliffhanger. Only then do pages get written. This prevents the common failure where an episode rambles for 90 seconds and then remembers to end on a hook.
Additional structural guards that belong in the pipeline, not in the writer's head:
- A thread ledger tracking open and resolved plot lines
- A midpoint enforcement that injects ending constraints once the series is past its halfway mark
- Episode-to-episode continuity so each new batch inherits the full previous text
- Batch intent locking: before each new batch, confirm plot direction, focus characters, threads and conflicts, and end hooks; treat these as hard constraints
Rule-based QC before pages move on
Before a script is handed to art and camera, it should pass an automated rule check that catches:
- Wrong episode titles
- Scene count mismatches
- Missing character lines
- Too little dialogue
- Placeholder text like "to be continued"
Failures should be sent back for rewrite with the error attached. Pages that fail QC should not reach the rendering stage. The goal is not to replace editorial judgment—it is to catch the mechanical failures that turn into expensive re-renders later.
Stage 5–6: Look selection and cast lock
One of the most visible AI drama failures is the style break: a character rendered in anime style in episode 1 turns photoreal in episode 4, or the lead's costume changes between cuts because the prompt described it differently each time.
The fix is to treat look development as a dedicated stage, not an afterthought.
Pick one style manual for the whole show
A production system should maintain a library of style manuals covering 2D, 3D, and photoreal directions—urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink style, xianxia, 3D donghua, stop-motion clay, cyber-Chinese fusion, and so on. Each manual should include at least:
- Character sheet guidance (face anchors, materials, vibe, view consistency)
- Scene and prop guidance
- Video style tags
Once a style is chosen, character art, scene art, prop art, and video prompts all consume the same style path. This is what prevents "2D person in a photoreal world."
Two-stage character art
Character turnarounds should be built in two steps:
- Text polish. Archive descriptions are rewritten into image-generation prompts using the chosen style manual, with hard rules such as gender locked in. Creators can override.
- Image generation. Final prompts are rendered, optionally with reference images, and saved as versioned history.
Video should only consume finalized, readable art assets. Half-finished turnarounds should never be silently bound as references.
Cast lock is a hard gate
Before moving to production, confirm: the cast list is complete, names are valid, visual fields are filled, and leads match the brief. If any of these fail, the pipeline should stop and ask. This is the equivalent of not walking onto set before casting is done.
Stage 7: Scene blocking into ~10-second blocks
A 90-second episode should not be rendered as one 90-second clip. It should be cut into scene blocks—targeting roughly 10 seconds each—because that is the unit of control.
A scene block is a short, independently promptable, independently renderable unit of the episode, roughly 10 seconds long, with its own reference images, its own shot prompt, its own output clip, and its own take history.
Why this matters:
- If a single shot fails, you re-render that shot, not the whole episode.
- You can swap a character or scene reference for one block without touching the rest.
- You can compare multiple takes of the same beat.
- Editors work with a clip list, not one monolithic file.
Soft limits on text length per block prevent over-stuffed prompts. When a block runs long, it is split further by action beats, paragraph breaks, or sentence boundaries. Crowd characters and generic roles are separated from the named cast so they do not consume character reference slots.
The 10-second target is an engineering heuristic, not a timecode-precise cut. Some blocks will run longer. Final tightening and assembly still belong in editing.
Stage 8: Shot prompts written as camera instructions
This is where most AI dramas quietly collapse. Prompts are written like prose: "A dramatic confrontation in a rainy alley where the hero looks angry and powerful." Video models do not interpret that the way a cinematographer does.
Shot prompts should be written like a shot list. A professional prompt covers eight elements:
| Element | What it specifies |
|---|---|
| 1. Precise subject | Who is in frame, named clearly |
| 2. Action detail | What exactly they do, with quantified motion |
| 3. Scene environment | Where they are, time of day, interior/exterior |
| 4. Light & color | Lighting direction, tone, palette |
| 5. Camera movement | One move per shot—no push-pan-zoom stacking |
| 6. Visual style | Tied to the chosen style manual |
| 7. Image quality | Resolution, stability, clean output constraints |
| 8. Guardrails | No watermarks, no twins, no extra limbs, style anchor for non-photoreal |
Additional prompt discipline:
- Simple scenes use a single-block prompt; complex cinematic scenes use a three-part structure (overall setup → shot 1/2/3 → guardrail pack).
- Use shot numbers, not absolute timestamps like "0–3s."
- Prefer slow, continuous motion over high-action bursts that break.
- Mark dialogue, sound effects, and BGM with clear notation so they are not confused with visual instructions.
- Feed only the current scene's assets into the prompt—no bleed from other scenes.
- Keep generation parameters conservative: stable and controllable beats wildly creative.
The system should teach the model to speak shot-list language, not novel language.
The reference image hierarchy
Reference images are the single strongest lever for character consistency. For each scene, build a reference table in a fixed order:
- Scene reference
- Props appearing in this scene
- Character turnarounds
The rule that prevents most face/costume swaps:
When a character has a reference image, the prompt must not re-describe their clothing or appearance in text. The reference image is authoritative. Text describes only action, expression, and injury state.
Other reference mechanics that belong in the pipeline:
- Empty scenes must contain no people in their reference art.
- Props are pulled from the script using the script's own naming, then merged into a show-wide catalog.
- Manual character binding handles cases where the script says "the officer" but the archive name is "Li Qiang."
- Manual prop include/exclude controls what actually occupies reference slots.
- Historical versions of an asset can be swapped in.
- Time-period routing picks the right character look for flashbacks or cross-era stories.
Before prompts are generated, run a completeness check: are scene, character, and prop references missing? The system should surface a list and offer to go create them. Skipping to pure text should be allowed—but treated as a conscious trade-off, not the default, because quality usually drops.
Stage 9–10: Rendering, takes, and the production queue
Rendering is not a button; it is a production queue. A tool that cannot manage renders as production jobs is a toy, not a pipeline.
The standard shooting flow should look like this:
- Open the episode's video workspace.
- Confirm references are complete (or intentionally skipped).
- Generate shot prompts.
- Read and edit the prompts—this is the director revising the shot list.
- Choose model tier, aspect ratio, resolution, and duration.
- Submit to render.
- Wait in a managed queue.
- Review historical takes per scene and pick the best.
What a professional team needs control over
| Control | Equivalent on a real set |
|---|---|
| Editing shot prompts | Director revising shot notes |
| Swapping character/scene/prop references | Changing wardrobe or location boards |
| Manual character binding | Fixing name mismatches and cameos |
| Including/excluding props | Controlling visual focus in the frame |
| Switching style | Unifying the art language |
| Choosing model tier | Trading quality, cost, and speed |
| Reviewing historical takes per scene | Multi-take selection |
Critically: pressing "generate" again should produce a new take only when the prompt or references have changed, just as a new take on set means a new setup. Re-submitting identical work and double-charging for it is a pipeline bug.
Production-grade reliability
For teams running slates, not experiments, the queue must behave like production infrastructure:
- Videos are processed in independent queues separate from writing and art tasks.
- The same scene block cannot be submitted in parallel while a job is running.
- Credits are pre-deducted and released on failure.
- Stuck jobs time out and become retryable.
- Already-created vendor jobs resume polling rather than being re-created.
- Final files are faststart-enabled and stored with measured duration, not just vendor-reported duration.
This is boring infrastructure. It is also what separates "we made a short" from "we ship a slate every week."
Where the human still sits
No serious AI drama pipeline removes the human. It moves them upstream and into the judgment seats.
The machine is good at:
- Enforcing structure and order
- Tracking continuity across many episodes
- Applying the same style path across assets
- Generating first-draft shot instructions
- Running queues, versions, and retakes
- Catching mechanical errors before they become renders
The human is still required for:
- Choosing the premise worth making
- Locking the brief and the ending shape
- Approving cast and look
- Reading and rewriting shot prompts
- Selecting takes
- Deciding when a cliffhanger lands
- Final edit and assembly
There is no automatic "pick the best take" engine. There is no replace-the-showrunner button. Anyone selling one is selling a demo.
Honest boundaries
To use this pipeline well, teams should be clear about what it does not do:
- No automatic quality scoring or auto-pick of the best take. Final judgment stays with creators; the system gives you multiple takes and the tools to re-word and re-shoot.
- Reference images are not a hard gate. You can skip them and render from text alone—expect weaker consistency.
- Reference markers in prompts rely on disciplined formatting. Review prompts before rendering.
- Shot grammar follows built-in lens rules; full art manuals apply on the image side.
- The ~10-second block is a heuristic, not a precision cut. Final tightening belongs in editing.
- No cross-scene automatic video extension workflow. The unit is one scene, one clip.
- Character consistency depends on the art pipeline, not facial embedding magic. Period looks are rule-routed, and final quality still depends on good turnarounds and on not re-describing appearance in text when a reference exists.
These are not weaknesses to hide. They are the boundaries that let a team plan real schedules.
The order is the product
Most of what people call "AI video problems" are actually production order problems. A face swap is a cast-lock problem. A forgotten plot thread is an archive problem. A broken motion shot is a prompt-grammar problem. A lost render is a queue problem.
The winning pattern is not a single bigger model. It is a pipeline that enforces the sequence the industry already learned the hard way:
Lock the brief. Build the archive. Write in batches with rules. Lock the look. Block into scenes. Bind references. Write camera instructions. Render per scene. Keep takes. Let humans choose.
High-quality short drama is never one long generation cut to pieces. It is many controllable units stacked into a whole.
If you are looking for a system built around this exact order, Maosika (猫斯卡) is an AI production operating system for vertical short dramas that encodes this sequence—from idea to locked scene clips—with reviewable artifacts at every stage and human approval at the gates that matter.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com