How AI Vertical Short Dramas Actually Get Made: From Brief Lock to Per-Scene Delivery
Most AI short drama failures are not model failures—they are pipeline failures. The fix is not a magic generate button, but enforcing the same order real crews use: lock the brief, build continuity, approve looks, block scenes, then shoot.
The pipeline is the product, not the prompt
If you ask a model to "write a short drama and make a video," you get a demo. If you want 60, 80, or 100 episodes of vertical micro-drama that do not forget characters, swap costumes mid-scene, or lose a cliffhanger across batches, you need a production pipeline.
A production pipeline is a sequence of checkpoints where each stage produces an inspectable artifact before the next stage is allowed to start. You can see it, change it, reject it, or roll it back. That is how live-action crews avoid chaos, and it is how AI short drama teams avoid the same chaos at scale.
This article walks through that pipeline the way a producer would read a call sheet—not as a list of features, but as a sequence of decisions.
Stage 1: Intake and creative assessment
The first job is not writing. It is deciding whether the idea is ready to write.
A raw idea usually arrives as a mix of genre, a protagonist, a twist, and a vague ending. That is enough for a pitch deck; it is not enough to brief a writing room. A healthy intake process scores how complete the idea is across the dimensions that actually matter for vertical drama:
- Genre and tone
- Protagonist and character arc
- Core conflict
- Story direction and ending shape
- Episode count and episode length
- Hook and payoff rhythm
- Platform and target audience
When the intake score is high, the system can move straight to a locked brief. When it is partial, it only asks for the missing dimensions. When it is too thin, it runs a fuller guided development path. The point is not to gatekeep creativity; it is to prevent the team from building 80 episodes on a one-sentence premise.
Definition: A locked creative brief is the approved version of the logline, conflict, ending direction, hook rhythm, audience, episode breakdown, and writer notes. No script work should start before the brief is locked.
Stage 2: The continuity bible
AI does not remember your show. It remembers whatever is in the context window for that run. If you rely on model memory, you will get:
- A character who was an undercover cop in episode 3 and a real estate agent in episode 27
- A scar that appears and disappears
- A revenge motive that softens because later batches did not see earlier batches
- A prop that matters in episode 9 and is never referenced again
The solution is a structured continuity bible—sometimes called a story archive—that tracks:
| Bible section | What it records |
|---|---|
| Character identities | Names, roles, stable traits, arc direction |
| Current state | Wounds, hidden identities, alliances, status changes |
| Relationships | Who knows what, who is allied with whom, open grudges |
| Plot threads | Each thread marked open or resolved |
| Episode appearance table | Who appears in which episode |
| Batch summaries | What happened in the last writing batch |
| Prop descriptions | Visual details that must stay consistent |
The writing process should only consume the slice of the bible it needs: current character states, unresolved threads, recent batch context, and the current batch objective. After each batch, the bible is updated. Before the next batch, it is reloaded.
The principle is simple: consistency comes from structure, not luck.
Stage 3: Batched script writing
Writing 80 episodes in one shot is a bad idea for humans and a worse idea for AI. The industrial approach is batched writing: write several episodes, review them, update the archive, confirm direction, then write the next batch.
The vertical drama writing rules
Vertical short dramas have their own grammar. These are not stylistic preferences; they are format constraints:
- Golden 3 seconds: The first scene must open with conflict, shock, or a question. No slow setup, no background lecture.
- Single-episode shape: Opening hook → escalating conflict → end-of-episode cliffhanger.
- Payoff density: Every episode needs at least one small win, reveal, slap-down, identity hint, or evidence drop; larger payoffs land every few episodes.
- Dialogue discipline: Short lines. No essay speeches. No narrator explaining what the audience just saw.
Beat sheet first, then dialogue
Before writing an episode's scenes, the system should produce a beat sheet and confirm the cliffhanger. Only then does it write the actual script. This prevents the common failure where dialogue is fluent but the episode has no turn.
Batch intent lock
Starting from the second batch, the writing process should confirm four things before generating:
- Where this batch goes dramatically
- Which characters carry it
- Which threads and conflicts are active
- What the episode-end hooks should be
These become hard constraints. If they conflict with older archive notes, the newly confirmed intent wins. That is how a showrunner overrides old planning without breaking continuity.
Rule-based quality checks
After writing, scripts pass a rule checker that catches structural failures, not just grammar:
- Wrong or missing episode titles
- Scene count mismatches
- Missing character lines
- Too little dialogue
- Placeholder text such as "to be continued" used as a lazy cliffhanger
Failed scripts are sent back for rewrite with the specific error attached. A draft that fails checks should not reach the creator as if it were finished.
Stage 4: Look development and character lock
This is where most AI dramas visually fall apart: the script is fine, but every scene invents the cast again.
Look development means locking the visual language before shooting. A mature pipeline includes multiple style manuals covering 2D, 3D, anime, realistic urban, period, xianxia, cyberpunk, stop-motion, and other directions. Once a style is chosen, characters, scenes, props, and video prompts all follow the same visual path. You do not want anime characters suddenly rendered as live-action people in episode 14.
Character approval is a hard gate
Before video work begins, the cast must be complete, names valid, visual fields filled, and leads aligned with the brief. If a main character has no approved look, the pipeline should not silently proceed.
Two-step character art
A reliable character art process works in two stages:
- Text polish: Convert archive descriptions into image-generation prompts using the chosen style manual, with hard rules such as gender locked in.
- Image generation: Generate the final look, support reference images, and save versions so creators can roll back.
The video side should only consume finished, readable character art. It should never bind a half-finished render as if it were an approved costume.
Scene and prop discipline
Scenes are parsed from script structure: interior/exterior, location, day/night. Scene art should be empty plates—no people inserted into background plates, because that creates reference conflicts later.
Props are extracted episode by episode from the script's actual wording, then merged into a show-wide catalog. This prevents episode 6 from calling it a jade pendant and episode 31 from calling it the same object a bronze lock.
Stage 5: Scene blocking
A vertical drama episode is not one giant video generation. It is a sequence of short scene blocks, usually targeting around 10 seconds each.
Definition: A scene block is a single producible unit of the episode, with its own script slice, reference set, prompt, output clip, and version history.
Why this matters:
- Long generations drift: faces change, costumes change, actions collapse.
- Short blocks are re-shootable: if one beat fails, you do not regenerate the whole episode.
- Editors can work clip-by-clip, the same way live-action is cut from a bin.
When a script block is too long, it is split by action beats, paragraph breaks, or sentence boundaries. Crowd characters and generic extras are separated from named principals so they do not consume character reference slots.
The default output is 9:16 vertical, not a horizontal video cropped after the fact. That distinction matters for framing, subtitle space, and platform delivery.
Stage 6: Reference mapping per scene
This is the single most important consistency mechanism in AI video production.
For each scene block, the pipeline builds a reference table in a fixed order:
- Scene reference
- Props used in this scene
- Character look references
If an asset does not exist, the slot should say so. The system should not invent bindings or pretend a prop exists.
The rule that prevents face swaps and costume changes is blunt and effective:
When a character has an approved reference image, the prompt must not re-describe that character's clothing or appearance in text. The image owns the look. Text only describes action, expression, and injury.
That rule directly attacks the classic failure mode: prompt says "woman in red dress," reference shows a blue coat, model compromises by generating someone new every shot.
Additional controls that belong in this stage:
- Manual character binding when the script uses a title like "the officer" instead of the character's name
- Prop inclusion/exclusion so irrelevant props do not clutter the shot
- Version swapping when an older look is better for a specific scene
- Era routing for time-travel or flashback stories, where period-specific looks take priority
Before prompts are generated, the pipeline should run a completeness check and list missing references. Creators can choose to proceed with text-only, but they should be warned: text-only usually produces weaker consistency. Professional workflow is to finish looks first, then shoot.
Stage 7: Engineered shot prompts
A video prompt is not a paragraph of fiction. It is a shot instruction.
A strong prompt structure includes eight elements:
| Element | Purpose |
|---|---|
| Precise subject | Who or what is in the shot |
| Action detail | What exactly happens, with measurable motion |
| Environment | Where the scene takes place |
| Light and color | Mood, time of day, tonal palette |
| Camera movement | One move per shot, not piled-up zooms and pans |
| Visual style | The locked look path |
| Quality baseline | Stability, face clarity, clean output |
| Constraints | What must not happen |
Complex scenes use a three-part structure: overall setup, then shot-by-shot instructions, then a constraint package. Simple scenes can be one clean block.
Prompt writing rules that reduce failure:
- One camera move per shot
- Use shot numbers, not absolute timestamps like "0–3s"
- Include fallback constraints for face stability, watermark avoidance, and duplicate-person prevention in group scenes
- Favor slow, continuous motion over explosive action that models cannot hold
- Mark dialogue, sound effects, and music with clear notation
- Only feed that scene's assets into that scene's prompt—no cross-scene contamination
- Strip specific copyrighted IP names while keeping cinematic technique
In plain terms: the pipeline teaches the model to read a shot list, not write a novel.
Stage 8: Shooting, retakes, and delivery
The shooting stage should feel familiar to anyone who has worked on set.
The creator opens an episode's video workspace, confirms references are ready, generates prompts, reads and edits them, chooses model tier and output settings, then submits the shot. The result is not one final answer; it is a take. The same scene can have multiple takes, and the creator picks the best one.
What a creator should be able to control
| Control | Production equivalent |
|---|---|
| Edit the video prompt | Director revising shot notes |
| Swap character/scene/prop references | Changing a look or location plate |
| Manually bind roles | Fixing off-screen names and cameos |
| Include or exclude props | Controlling visual focus |
| Switch visual style | Unifying the show's look |
| Choose model tier | Balancing quality, speed, and cost |
| Review historical takes | Selecting the best performance |
A new take should only happen when something changes: prompt edited, reference swapped, or setting adjusted. That mirrors real production, where you do not roll camera again without changing the shot.
Production-grade queue behavior
For teams running volume, the backend matters as much as the art:
- Video tasks run in a separate queue from writing and art tasks
- The same scene block cannot submit parallel duplicate jobs
- Credits are reserved on submission and released on failure
- Stuck tasks time out and become retryable
- Finished videos are validated for actual duration, not just vendor-reported duration
- If a vendor task already exists, the system polls it instead of creating a duplicate charge
This is what turns a fun script into an operable production line.
Where human judgment still sits
It is important to say what this pipeline does not do.
There is no automatic "best clip" engine that replaces a director's eye. The system can give you multiple takes and easy retakes, but quality judgment still belongs to the creator. Reference images are not a hard block; you can skip them, though consistency usually suffers. Scene blocks are heuristic cuts around 10 seconds, not frame-accurate edits, so final assembly still belongs in editing. Character consistency depends on the quality of approved looks and disciplined prompting, not magic face locking. Cross-scene automatic video extension is not the unit of work; the unit is one scene block with its references, producing one clip.
That is not an apology. It is the correct boundary.
The goal is not to remove filmmakers from short drama. The goal is to remove the parts of the work that are repetitive, drift-prone, and operationally fragile, so creators can spend their attention on story, casting taste, hook rhythm, performance, and cut.
Where Maosika fits
Maosika (猫斯卡) is built around this exact pipeline: an AI production operating system for vertical short dramas that moves from idea to finished scenes through locked briefs, a structured continuity bible, batched script writing with rule checks, look development, per-scene reference mapping, engineered shot prompts, and multi-take shooting.
It does not promise one-click hits. It enforces the production order that experienced crews already know works: story first, then continuity, then looks, then scenes, then shots, then retakes.
If you are evaluating tools for a real short drama slate, ask whether the product gives you inspectable artifacts at each stage—or whether it hides everything behind a single generate button. The first can scale. The second usually breaks around episode twelve.
You can learn more about the system at https://www.maosika.com.
About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com