How to Make a Vertical Short Drama with AI: From Idea to Final Cut

Maosika Editorial | Last updated

Making a vertical short drama with AI works when you treat it like a real shoot: lock the brief first, build a story bible, approve cast and set looks, then shoot scene by scene in ~10-second blocks — not as one long prompt.

If you've tried AI video tools for short dramas, you already know the pattern. The first clip looks promising. By clip five the lead has a different face. By episode twelve nobody remembers the clue planted in episode three. The ending doesn't land because the model never really knew what the story was about.

This is not a model-quality problem. It's a production-order problem. Traditional crews solved it a hundred years ago with a sequence: script meeting → locked brief → continuity bible → casting and wardrobe → shot list → shoot → dailies → pickups. AI doesn't remove the need for that order. It just makes each step faster — if you respect it.

This guide walks through that order as it applies to AI vertical short dramas, in the language of a real set.

What goes wrong when you skip the order

Most failed AI short dramas fail the same way, and the failure shows up in predictable symptoms:

  • The lead changes face or outfit between scenes. This happens when each generation re-imagines the character from text instead of referencing an approved look.
  • Props appear and disappear. A jade pendant in episode 2 becomes a metal locket in episode 9 because nobody kept a prop catalog.
  • Plot threads are dropped. A mystery is set up, never resolved, because the model has no structured record of what's still open.
  • The pacing is wrong for vertical. Long establishing shots, slow dialogue, no hook in the first three seconds — the script was written like a horizontal film, not a phone-screen drama.
  • Episodes don't connect. Each batch reads like a different writer's draft because there was no handoff document between batches.

None of these is fixed by "a better prompt." They're fixed by putting the right intermediate artifact in place before you ask the model to generate anything.

The production order that actually works

Here is the sequence, end to end. Each stage produces something you can look at, reject, or approve before moving on. That's the key difference from a single "generate" button.

StageWhat you getWho approves
Idea intake & evaluationA scored brief; gaps flaggedYou
Guided creative developmentFilled-in missing dimensionsYou
Locked creative briefLogline, conflict, arc, hook rhythm, platform, endingYou
Cast lineup & visual confirmNamed characters with visual fieldsYou
Story bible / continuity fileCharacters, relationships, open threads, prop notesSystem-maintained, you review
Batch script writingBeat sheets → drafts → rule check → archiveYou, per batch
Style selectionOne shared look for art and videoYou
Look-dev (characters / sets / props)Approved reference artYou
Scene blocking~10-second scene blocks per episodeSystem, you can re-cut
Shot prompt engineeringPer-scene prompts bound to reference artYou read and edit
Multi-modal shootingVertical clips per scene, multiple takesYou select
Review & reshootRejected takes, rewritten prompts, swapped artYou

Notice that "shooting" is the second-to-last step, not the first.

Stage 1: Intake — score the idea before you write anything

Not every idea is ready to go to script. A real production has a pitch meeting where people ask: who is this for? What is the core conflict? What does the lead want, and what's in their way? How does it end?

In an AI workflow, that meeting should be formalized. The intake should score how complete the idea is across the dimensions vertical drama actually needs:

  • Genre and tone
  • Protagonist and their arc
  • Core conflict
  • Story direction
  • Episode count and episode length (vertical drama usually lands at 1–2 minutes per episode, stretching to 3–4 at most)
  • Hook and payoff rhythm
  • Ending direction
  • Target platform and audience

An idea that's under-scored gets guided through the missing pieces. One that's already complete can go straight to a locked brief. The threshold matters because vague briefs produce vague scripts, and vague scripts produce unwatchable footage.

The rule is the same as on a real set: you don't roll camera before the brief is locked.

Stage 2: Lock the creative brief

The brief is the document everyone — every later stage — is contractually bound to. For vertical short drama it should fix at least:

  • One-sentence logline
  • Core conflict
  • Story direction and ending direction
  • Satisfaction beats and cliffhanger rhythm
  • Target platform and audience
  • Episode-by-episode synopsis
  • Notes to the writer
  • Lead characters, each with a character arc

Once the brief is locked, the script stage treats it as a hard constraint. If you want to change the ending later, you re-open the brief — you don't whisper it into a prompt halfway through episode 40.

Stage 3: Build the story bible before writing batches

A story bible (sometimes called a continuity bible) is a structured record that travels with the show for its whole life. It tracks:

  • Character identities and stable traits
  • Current, changeable state (injuries, revealed identities, disguises)
  • Relationships between characters
  • Every plot thread, marked open or resolved
  • Per-episode appearance table
  • Per-batch plot summaries
  • Visual descriptions of recurring props

This is how episode 60 remembers the scar the lead got in episode 7. The model doesn't "remember." The bible does. When a new batch is written, the system pulls only the slice it needs — current character states, unresolved threads, recent batch summaries, this batch's main line — so the context stays relevant instead of drowning in noise. After each batch, the bible updates.

Continuity runs on an asset, not on the model's memory.

Stage 4: Write scripts in batches, beat sheet first

Writing 80 episodes in one shot is how you get a meandering middle and a forgotten ending. The right granularity is batches of a few episodes at a time. Within each batch:

  1. A beat sheet comes first — scene-by-scene beats plus the end-of-episode cliffhanger.
  2. Then the draft.
  3. Then a rule-based quality check.
  4. Then the batch is archived back into the bible.

Vertical short drama has its own writing rules, and they should be enforced, not suggested:

  • Golden 3 seconds. The first scene must hit a strong conflict or悬念 (hook). No slow establishing, no voiceover explaining backstory.
  • Single-episode structure. Opening hook (1 scene) → conflict escalation → end-of-episode cliffhanger (final scene).
  • Satisfaction beat density. At least one small payoff per episode (a reveal, a comeback, a clue, an identity hint); a larger payoff every few episodes.
  • Dialogue. Short lines, typically under 20 characters in Chinese or a single breath in English. No essay-speak, no paragraph-long exposition.

A rule checker should catch mechanical failures: wrong episode titles, mismatched scene counts, missing character lines, too little dialogue, lazy placeholders like "to be continued." Failing drafts get sent back for rewrite with the error attached, up to a fixed cap. Half-finished scripts shouldn't reach the director.

From the second batch onward, four things get pinned before writing starts: this batch's plot direction, focus characters, threads and conflicts, and the end-of-episode hook. Those are marked as user-confirmed hard constraints, and they override the archive if there's a conflict. That's how you keep intentional pivots from being "corrected" back by the system.

Stage 5: Pick one visual style and commit

A common AI artifact is the style drift: characters look like 2D anime, but the video comes out photoreal, or vice versa. The fix is boring: pick one style book up front, and use it for character art, scene art, prop art, and video style tags — all four pulling from the same path.

A production-grade system should offer multiple style books spanning 2D, 3D, and photoreal directions (urban modern, period realistic, mature urban romance animation, 90s anime, Chinese ink, xianxia, 3D donghua, stop-motion clay, cyberpunk-Chinese fusion, and others). Each book defines face anchors, material, mood, view-consistency rules for characters, plus scene/prop guidance and video style tags. You pick once, and the whole show speaks one visual language.

Stage 6: Lock cast, sets, and props before shooting

Before a single frame of video is generated, every recurring character, location, and prop should have approved reference art. This is look-dev, and it's non-negotiable for consistency.

Characters. The lineup has to be complete, names valid, visual fields filled, leads matching the brief. That's a hard gate — you can't proceed if it isn't.

Sets. Scenes are parsed from the script into interior/exterior, location, day/night — not hand-typed into a spreadsheet. Establishing shots (空镜, "empty frames") must not contain people; set reference art is for the space.

Props. A props specialist pulls props from the script using the script's own names, then merges them into a show-wide catalog so nothing gets lost between episodes.

Reference art is generated in two steps: a text polish that turns the archive description into a prompt suited to the chosen style (with gender as a hard rule, not a suggestion), then image generation. You can override the prompt, bring your own reference image, and swap between historical versions. The video stage only consumes art that's marked complete and readable — it never shoots against a half-finished asset.

Stage 7: Map reference art to every scene

This is the single most important mechanism for character consistency, and it's worth understanding in detail.

For each scene, a reference table is built in a fixed order: scene art → this scene's props → character look-dev art. Slots are only filled if the art exists; empty slots are marked as text-only rather than silently invented.

There is one overriding rule that should be enforced at the prompt level:

For any character with reference art, the prompt must not re-describe their clothing or appearance in text. The reference image is the authority. Text describes only action, expression, and injury state.

That rule is aimed directly at the classic AI failure: text says "red dress," reference art shows a blue one, and the model compromises by giving you a third outfit entirely. If there's a picture, the picture wins.

On top of that:

  • You can manually bind a script name ("the officer") to a specific character's look-dev ("Li Qiang").
  • You can include or exclude props per scene so irrelevant items don't eat reference slots.
  • You can swap in a different historical version of the same-named art.
  • For time-travel or flashback stories, period routing picks the right look for the scene's era based on location keywords.

Before prompts are generated, a completeness check flags any missing scene / character / prop references and offers to route you back to draw them. You're allowed to skip and go text-only — but you should know going in that text-only scenes are usually the ones that drift.

Stage 8: Cut episodes into scene blocks

An episode is not one big video. It's a list of clips on a cutting timeline. The system should split each episode into scene blocks targeted at roughly 10 seconds each, with a soft cap on body text length; longer scenes get re-split along action beats, paragraph breaks, and sentence endings. Crowd characters and generic roles are separated from the drawable main cast so they don't consume look-dev slots. You can re-split an entire episode if the automatic cut is wrong.

What you see in the interface is not "Episode 3, one file." It's "Episode 3 → Scene 1, Scene 2, Scene 3…" Each scene has its own prompt, its own reference art, its own output clip, and its own history of takes. That's how you reshoot one bad moment without regenerating the whole episode.

Default delivery is 9:16 vertical, not a horizontal crop after the fact. That matters because composition for a phone screen is different from composition for a cinema screen — framing, headroom, and close-up distance all change.

Stage 9: Engineer shot prompts like a shot list, not a novel

A video prompt for a short drama is a shot list. It should speak the language a cinematographer speaks, not the language a novelist writes. A production-grade prompt has eight elements:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color tone
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Constraints

A few of the rules that keep prompts from producing garbage:

  • One shot, one camera move. No push-pull-pan-zoom stacked into a single shot.
  • Use shot numbers, not absolute timestamps. Write "Shot 1," "Shot 2," not "0–3s."
  • Complex scenes use a three-part structure: overall setup → shot-by-shot breakdown → constraint pack.
  • Mandatory fallback pack: quality, face stability, no watermark or logo; multi-person scenes add a twin/doppelgänger fallback; non-realistic styles pin the style anchor.
  • Action principle: fine-grained limb movement with quantified intensity; favor slow, continuous motion over high-burst action that tends to break.
  • Notation conventions: dialogue in {}, sound effects in <>, score in ().
  • Only this scene's materials are fed in. No cross-scene bleeding. The scene's script body is the highest-priority source.
  • Generation parameters lean conservative, favoring stability and controllability over wild creativity.

Before delivery, prompts are cleaned of specific copyrighted work or IP names — the aesthetic and technique descriptions stay, the proper nouns go — to reduce downstream copyright blocking. If a reference image is flagged as a suspected real-person photo, the system should refuse to shoot against it and route you back to the art pipeline instead of silently generating.

The system is teaching the model to speak shot-list language, not to write fiction.

Stage 10: Shoot, review, reshoot

The shooting path itself should feel like a real set:

  1. Open an episode's video workspace.
  2. Confirm reference art is complete (or intentionally skip).
  3. Generate shot prompts.
  4. Read and edit them — this is the director revising the shot list.
  5. Pick model tier, aspect ratio, resolution, duration.
  6. Submit to the render queue.
  7. Review historical takes and pick the best one.

Every professional control point maps to something concrete:

ControlWhat it's like on a real set
Editing the shot promptDirector revising shot notes
Swapping character / set / prop artChanging wardrobe or set board
Manual character bindingHandling off-screen names and guest roles
Including / excluding propsControlling the frame's visual focus
Switching style bookUnifying the show's visual language
Choosing model tierTrading quality, cost, and speed
Historical takes per sceneShooting multiple takes and picking

When you hit generate, the worker uses exactly the prompt you confirmed and the reference mapping that was locked at that moment. Change the prompt and hit generate again, and that's a new take — same as on set when you rewrite the shot note and roll again.

For this to be usable as a production tool rather than a toy, the queue itself has to behave like production infrastructure: video tasks in a separate lane from script and art so they don't starve each other; no parallel submissions on the same scene block to avoid double charges; pre-auth and settlement of credits with automatic release on failure; zombie task reclamation; resumable polling when a vendor job ID already exists so you never double-create. Finished clips get faststart processing and are stored with their measured runtime, not just the vendor's reported duration.

This isn't a toy script. It's an operable production queue.

What this workflow still doesn't do

It's important to be honest about the boundaries. Some things are intentionally left to the human, and some are genuinely hard:

  1. There is no automatic scoring or auto-pick of the best take. Final quality judgment stays with the creator and the producer; the system gives you multiple takes and the tools to reshoot, but it doesn't declare a winner.
  2. Reference art is not a hard gate. Missing art triggers a reminder, but you can skip and go text-only. Quality usually suffers, which is why the professional path is: lock looks first, then shoot.
  3. Reference codes in prompts are guided by convention, not force-stitched. You should still read the prompt and confirm the reference codes are all present before submitting.
  4. Video shot grammar follows the built-in shot spec; the full art manual applies to the look-dev side, while video gets the style tags.
  5. The ~10-second scene block is an engineering heuristic, not a timecode-precise cut. Over-long scenes get re-split, but you may still see blocks that run long; fine cutting and stitching into a final episode still belong in a later edit.
  6. There is currently no cross-scene automatic continuation or video extension workflow. The product unit is "single scene with multi-modal references → single clip."
  7. Character consistency depends on the look-dev asset chain, not on face-embedding verification. Period routing picks art by rule, but the final look still depends on the quality of the reference art and on the prompt respecting the "don't describe appearance when there's an image" rule.

The right way to think about it: this kind of system is a scalable AI short-drama production operating system — it hard-codes the professional sequence into the product, rather than claiming human-free, zero-error output.

A checklist you can steal

Whether or not you use a dedicated system, you can apply this order to any AI short drama project:

  • [ ] Score the idea across the vertical-drama dimensions before writing.
  • [ ] Lock a written brief (logline, conflict, arc, hook rhythm, platform, ending) — and treat changes as brief re-opens.
  • [ ] Maintain a structured story bible; update it after every batch.
  • [ ] Write in batches of a few episodes; beat sheet first, draft second, rule check third.
  • [ ] Enforce golden-3-second, episode structure, payoff density, and short-dialogue rules.
  • [ ] Pick one visual style and use it for art and video alike.
  • [ ] Approve character, set, and prop reference art before shooting.
  • [ ] Map reference art per scene in order: set → props → characters.
  • [ ] Never re-describe a referenced character's clothes or face in text.
  • [ ] Cut episodes into ~10-second blocks; reshoot at the block level.
  • [ ] Write shot prompts with eight elements, one camera move per shot, shot numbers not timestamps.
  • [ ] Shoot multiple takes; pick manually.
  • [ ] Treat the render queue as production infrastructure, not a fire-and-forget button.

High-quality short drama is never one long generation hacked into shape. It's a stack of controllable units.

The sequence above — lock the story, lock the looks, block the scenes, write the shot notes, then shoot multiple takes — is what Maosika (猫斯卡) has chosen to hard-code into the product. If you want to see what that looks like running end to end, you can start a project from a one-line logline at www.maosika.com.

---

About Maosika | Maosika (猫斯卡) is an AI production operating system for vertical short dramas. It chains idea evaluation, script writing, character / set / prop look-dev, shot prompts, and shooting into a reviewable, editable, rollback-capable production line, carried out by 18 digital specialists. Official site: www.maosika.com

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com