AI Vertical Short Drama Production: 12 Checkpoints From Idea to Deliverable Cut
AI short drama production fails not because the model is weak, but because teams skip the checkpoints that real shoots enforce. Here is the 12-stop production line, in order, with what must be signed off before moving on.
Most teams new to AI vertical short dramas treat the toolchain as a single button: type a logline, wait, get a show. The teams that actually ship on a schedule treat it as a production line with mandatory handoffs. The difference is not talent; it is whether each stage produces something you can look at, reject, or lock before the next stage spends compute on top of it.
This article lays out that line in twelve checkpoints. It is written for producers, writer-directors, and small teams who need repeatable output, not one lucky viral clip.
Why a checkpoint model beats "generate and fix"
In a live-action shoot, you do not roll camera before the cast is locked, the location is scouted, and the shot list exists. AI generation does not remove those steps; it compresses them into invisible ones. If you skip them, the model invents them for you on the fly — which is why you get face swaps mid-episode, costumes that change between cuts, props that appear and vanish, and cliffhangers that forget the cliff they were hanging from.
A checkpoint is a stage with a concrete artifact. You review the artifact. You either send it back or sign off. Only then does the next stage open.
Story continuity is not held in the model's memory. It is held in a structured story file that gets updated after every writing batch and re-read before the next one begins.
The twelve checkpoints
The order matters. Moving "pick visual style" above "lock the brief" or "shoot" above "lock lookdev" is the single most common cause of wasted renders.
| # | Checkpoint | Artifact you review | What you are really deciding |
|---|---|---|---|
| 1 | Idea intake & scoring | Intake score 0–100 | Is the idea developed enough to brief, or does it need guided expansion? |
| 2 | Guided development | Filled-in creative dimensions | Genre, lead, core conflict, arc, episode count, episode length, tone, hook rhythm, ending, platform, audience |
| 3 | Brief lock | One-page creative brief | Logline, conflict, trajectory, ending, beat rhythm, platform/audience, episode synopses, notes to the writer — including the lead's character arc |
| 4 | Cast & visual confirmation | Cast roster with visual fields | Names are legal, lineup is complete, visual fields are filled, leads match the brief |
| 5 | Story bible creation | Continuity bible | Characters, stable traits, current state (injuries, revealed identities), relationships, open/resolved plot threads, per-episode appearance table, batch summaries, prop visual notes |
| 6 | Batch script planning | Beat sheet + end-of-episode cliffhanger per episode | What happens, and what the hook is, before any dialogue is written |
| 7 | Batch script writing | Episode scripts | Golden-3-second open, escalation, cliffhanger close; short lines; at least one small payoff per episode, a larger one every few episodes |
| 8 | Script rule check | Pass/fail report with auto-rewrite | Episode title errors, scene count mismatches, missing character lines, too little dialogue, placeholder text like "to be continued" |
| 9 | Style selection & lookdev | Style pack; character, scene, and prop sheets | One consistent visual language shared across art and video; empty scenes contain no people; gender is a hard rule on generation |
| 10 | Scene blocking | ~10-second scene blocks per episode | Each block has its own reference set and its own prompt; extras are separated from named cast |
| 11 | Shot prompting & shoot | Per-scene shot prompts + takes | Prompt follows the eight-element spec; references map as scene → props → characters; locked prompts generate takes you can compare |
| 12 | Review, reshoot, assemble | Selected takes per scene | Re-roll with edited prompts or swapped references; no auto-picking of "the best" — final cut stays with the creator |
Stops 1–3: from a vague idea to a locked brief
The intake score is not a vanity metric. It is a traffic router. Above a threshold, the idea is ready to go to brief; in the middle band, only the missing dimensions get filled in; below that, the idea goes through a full guided development pass. The point is to prevent two failure modes: under-developed ideas that the writer has to invent from scratch, and over-prompting boxes that force creators to answer questions they have not thought about yet.
The brief is the first real lock. Until it is signed off, nothing downstream opens. This is the same discipline as a live-action development meeting: you do not call the crew before you know what show you are making.
Stops 4–8: from brief to a script batch you can shoot from
Vertical short drama scripts have hard shape rules that are not the same as feature or TV writing:
- Golden 3 seconds. The first scene must open on conflict or a strong hook; no slow build, no expository voiceover.
- Single-episode shape. Hook scene → escalation → cliffhanger on the final scene.
- Payoff density. At least one small beat of satisfaction per episode (a reversal, a face-slap, an identity hint, evidence landing); a larger payoff every few episodes.
- Line length. Dialogue stays short; no essay-like speeches, no paragraphs of narration explaining the plot.
Before each batch of episodes is written, the beat sheet and cliffhanger are planned first; the prose comes second. After the batch, the rule check catches structural errors and triggers a rewrite with feedback, up to a capped number of attempts. Nothing that fails the check is handed over as "done."
Crucially, the writer does not rely on having "read" earlier episodes. Before each new batch, the system pulls a slice of the continuity bible — current character states, unresolved threads, recent batch summaries, the current batch's main line — and after the batch, it writes new events back into the bible. From the second batch onward, four things are pinned before writing starts: the batch's plot direction, the focal characters, the threads and conflicts in play, and the end-of-batch hook. Those pinned intents override older bible state when the two conflict.
Stops 9–10: from script to a shootable scene plan
Look development is where most AI productions quietly fall apart. A character sheet drawn in one style and a video generated in another will never match, no matter how clever the prompt. The fix is structural: pick one style manual, then have character art, scene art, prop art, and video prompts all consume the same style path.
A few lookdev rules that save enormous pain downstream:
- Empty scene plates must contain no people. If a scene reference already has a figure in it, the model will invent a second one or swap your lead in.
- Gender is a hard rule at generation time, not a suggestion.
- Props are extracted from the script in the writer's own wording, episode by episode, then merged into a show-wide catalog so nothing is forgotten between batches.
- User or script descriptions override style-manual content rules. The style manual governs how something is drawn, not whether it is allowed to exist.
Once art is locked, episodes are cut into scene blocks targeted at roughly ten seconds each, with a soft cap on body length and re-splitting rules for longer beats. You do not see "Episode 3" as one giant video; you see Episode 3 as Scene 1, Scene 2, Scene 3… each with its own reference set, its own prompt, its own take history. That is the same unit a film editor works with on the timeline.
Stop 11: the shot prompt is a shot list, not a paragraph
The prompt that goes to the video model should read like a first assistant director's shot list, not like a novelist's paragraph. A production-grade prompt covers eight elements:
- Precise subject
- Action detail
- Scene environment
- Lighting and color
- Camera movement
- Visual style
- Image quality
- Constraints
A few of the rules that prevent the most common failures:
- One camera move per shot. No "push in while panning while tilting" stacks.
- Use shot numbers, not timestamps. Write "Shot 1," "Shot 2," not "0–3 seconds."
- Always include a constraint pack for image quality, face stability, no watermark or logo; add twin/duplicate prevention for multi-character scenes; anchor style explicitly for non-realistic looks.
- Favor slow, continuous motion over high-energy bursts that break the model.
- Mark up dialogue, sound effects, and BGM with consistent symbols so they are not read as visual description.
- Only feed that scene's assets into that scene's prompt. Cross-scene contamination is how characters wear the wrong costume or appear in a set they are not in.
Reference images are mapped in a fixed priority: scene plate → props for that scene → character lookdev. If a reference exists for a character, the prompt must not re-describe that character's clothes or appearance in text; the text only describes action, expression, and visible injury. That single rule eliminates a huge share of mid-shot face and costume swaps, because it removes the fight between the text and the image.
Before prompts are generated, a readiness check flags any missing scene, character, or prop references. You can skip and shoot on text only — the system allows it — but quality is usually worse, which is why a professional pipeline locks art first.
Stop 12: review, reshoot, and the honest boundary
There is no engine that watches your takes and automatically picks the best one, or automatically re-shoots a bad take until it is good. What there is, is a production queue that treats each generated clip as a take: same scene, same locked prompt and references, multiple attempts you compare side by side. If you want something different, you change the prompt or swap a reference and generate a new take — exactly the way a director revises a shot list before rolling again.
A few operational details that matter when this is your job, not a hobby:
- Video tasks run on a separate queue from writing and art, with their own concurrency slots.
- A scene that is already rendering cannot be double-submitted, which prevents duplicate charges and state confusion.
- Credits are pre-deducted and released on failure; hung tasks time out and can be retried.
- Delivered clips get faststart processing and are stored with their measured runtime, not just the vendor's reported duration.
This is what makes the pipeline operable as a production tool rather than a toy script.
Where human judgment still sits
It is worth stating plainly what the system does not do, because pretending otherwise is how teams get burned:
- No automatic quality scoring or auto-pick of the best take. Final cut is yours.
- Reference images are not a hard gate. You can skip them; you should usually not.
- The ~10-second scene block is an engineering heuristic, not frame-accurate editing. Long scenes get re-split, but final trimming and assembly still happen in the edit.
- There is no cross-scene automatic continuation or video extension workflow. The unit of generation is one scene, with its own references, producing one clip.
- Character consistency relies on the lookdev asset chain, not on face-embedding verification. Era-specific costumes (for time-travel or flashback stories) are picked by rule, and final look still depends on the quality of the art and on the prompt respecting the "don't describe what the reference already shows" rule.
The goal of an AI production operating system is not to remove the filmmaker. It is to take the parts that are repetitive, drift-prone, and easy to lose control of — continuity across batches, reference mapping, shot-prompt structure, queue discipline — and turn them into an assembly line you can audit. Taste, casting judgment, story decisions, and the final cut still sit with a person.
If you are building out your own AI vertical short drama pipeline, the order of the twelve checkpoints above is the part worth copying first. Tools will change; the discipline of locking an artifact before you spend the next stage's budget on top of it will not.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com