The 12-Step AI Vertical Short Drama Pipeline: From Idea to a Finished Episode

Maosika Editorial | Last updated

A finished AI short drama is not one long generation. It is a chain of 12 handoff points, each with a reviewable artifact. The team that wins is the one that treats those artifacts like a real production, not a prompt experiment.

The real production question is not "can AI make video?"

Most teams asking about AI short drama already know the answer: yes, you can get clips out. The harder question is whether you can get episode 7, episode 24, and episode 60 to still feel like the same show — same characters, same rules, same visual language, same promise to the viewer.

That is a pipeline problem, not a model problem.

Below is a production-side breakdown of how a vertical short drama moves from one idea to a finished episode, organized as 12 stages. Each stage has a concrete deliverable, a decision point, and a failure mode you can see before it becomes a bad clip.

The 12-stage pipeline at a glance

StageDeliverableWho approvesCommon failure if skipped
1. Idea intakeScored brief completenessCreator / producerVague premise that drifts episode to episode
2. Creative guidanceFilled gaps in premise, audience, hook rhythmCreatorWriter invents missing pieces
3. Locked creative briefLogline, conflict, arc, ending direction, platform notesProducer / showrunnerNo shared north star
4. Character lineup & visual lockCast list with arcs and visual fieldsCreator / art leadCharacters change face, age, or role
5. Story archiveContinuity bible: people, ties, threads, stateSystem + writer reviewForgotten setup, resurrected plot holes
6. Batch script planningBeat sheet per episode, cliffhangersWriter / producerEpisodes wander or repeat beats
7. Script draft & rule checkPassed script with scene list and dialogueWriterFlat openings, missing hooks, placeholder lines
8. Style selectionShared look for art and videoArt lead / directorPeople are anime, world is live action
9. Look developmentCharacter, scene, and prop artArt leadReference gaps during shooting
10. Scene blocking into ~10s blocksPer-scene clip units with referencesDirector / editorOne giant uneditable video per episode
11. Engineered shot promptsShot-level instructions with referencesDirectorModel invents clothes, props, camera chaos
12. Shoot, review, reshootMultiple takes per scene, selected finalsCreator / producerFirst bad take becomes the episode

Stage 1–3: From idea to a locked brief

A short drama that starts with "a cool concept" and no brief will pay for it later. The first three stages exist to force one thing: everyone agrees on what show this is before anyone writes or draws.

Idea intake is not a formality. A premise can be scored on completeness — logline, lead, core conflict, direction, episode count, episode length, tone, hook rhythm, ending, platform, audience. If the intake is thin, guidance should fill the holes; if it is nearly complete, the team should move straight to a locked brief.

Creative guidance is the part many AI workflows skip. When a creator says "revenge drama in a wealthy family," that is not enough to write from. Guidance asks the unspoken questions: Who is wronged? What is the first proof the audience gets? How often does a reversal land? Is the ending a comeuppance, an identity reveal, or a reunion?

The locked brief is the first real contract. It should contain at minimum:

  • One-sentence logline
  • Core conflict
  • Story direction
  • Ending direction
  • Beat and hook rhythm
  • Platform and target audience
  • Episode-by-episode outline
  • Notes to the writer
  • Lead character arc

The lead character must have an arc. A protagonist who only reacts is why viewers drop off after episode 3.

Stage 4–5: Cast and continuity before pages

Character lock is not just naming people. Each lead needs identity, stable traits, current state, relationships, and visual fields. Visual lock means: before any shot is generated, the system knows who this person is and what they are supposed to look like.

A good rule used in serious AI pipelines: a character cannot move to shooting unless the lineup is complete, names are valid, visual fields are filled, and leads match the brief.

The story archive is the structured memory of the show.

A story archive is a running continuity record for the whole drama: character identities, stable traits, current state, relationships, open and resolved plot threads, episode appearance tables, batch summaries, and prop visual notes. It is the AI-era equivalent of a writer's room continuity bible.

This matters because long-form micro-dramas break in one predictable way: episode 12 forgets what episode 4 established. The fix is not to "remind the model" in a long chat. The fix is to feed the writer only the relevant slice of the archive — current character state, unresolved threads, recent batch summary, current batch arc — and write back to it after each batch.

Continuity comes from structure, not from model memory.

Stage 6–7: Batch writing that respects vertical rhythm

Vertical short drama writing is not prose writing. It is built around retention.

Before drafting, each episode should have a beat sheet and a named cliffhanger. Only after that does the scene-level draft get written. Writing in batches — a few episodes at a time — lets the team confirm direction, key characters, threads, and end-of-batch hooks before the next run.

Vertical scripts live and die by a small set of hard rules:

  1. Golden 3 seconds: the first scene must open with conflict or suspense, no slow setup.
  2. Episode shape: opening hook → escalating conflict → end-of-episode cliffhanger.
  3. Payoff density: at least one small payoff per episode (a reversal, a reveal, evidence landed, a public win); a larger payoff every several episodes.
  4. Dialogue: short lines, spoken language, no lecture-style monologue.

After drafting, scripts should pass a rule check that catches broken titles, mismatched scene counts, missing character lines, too little dialogue, and placeholder text like "to be continued." If it fails, it goes back for rewrite before anyone sees it as a deliverable.

Stage 8–9: One style, shared across art and motion

One of the most visible AI failures is style fracture: characters rendered in one look, backgrounds in another, video in a third.

The fix is to choose a style manual once and use it for everything: character art, scene art, prop art, and video style tags. A mature pipeline offers multiple style manuals covering 2D, 3D, and realistic directions — urban realism, period realism, mature urban romance animation, 90s anime, Chinese ink, xianxia, 3D donghua, stop-motion clay, cyber-Chinese, and so on — but the rule is the same: once chosen, every downstream asset uses that path.

Look development then produces three asset classes:

  • Character art — face anchors, material, temperament, view consistency
  • Scene art — empty plates, no people in them
  • Prop art — extracted from script wording, cataloged across episodes

A useful discipline: scenes are parsed from the script into interior/exterior, location, and day/night. Scene art stays empty of people. Props are pulled by the script's own wording, not renamed into generic terms.

Stage 10–11: Cut into blocks, then write shot language

This is where "AI video" stops feeling like a demo and starts behaving like a production.

Instead of generating one big video per episode, cut the episode into scene blocks targeted around 10 seconds each. Long scenes get split further by action beats and punctuation. Each block is a separate unit with its own references, prompt, output, and take history — like clips on an editing timeline.

Then, for each block, build an engineered shot prompt. A good prompt is not a paragraph of novel writing. It is shot instruction.

A strong shot prompt covers eight elements:

  1. Precise subject
  2. Action detail
  3. Scene environment
  4. Lighting and color
  5. Camera movement
  6. Visual style
  7. Image quality
  8. Constraints

For complex scenes, use a three-part structure: overall setup → shot-by-shot instructions → constraint package. One shot, one camera move. Use shot numbers, not absolute timestamps. Always include a safety package for face stability, no watermark, no logo; add twin/duplicate prevention for multi-person scenes; anchor style for non-realistic looks.

The reference mapping rule that prevents most face-and-outfit drift:

When a character has reference art, the prompt must not re-describe that character's clothing or appearance in text. The image is the authority. Text only describes action, expression, and injury.

References are loaded in a fixed order: scene → props → characters. If an image does not exist, leave the slot empty and mark it as text-only rather than inventing a binding. Manual binding handles cases where the script says "the officer" but the archive name is "Li Qiang." Props can be manually included or excluded so irrelevant items do not steal visual focus.

Before prompts go out, a readiness check should list any missing scene, character, or prop references. Teams can choose to skip and shoot text-only, but they should do so knowing consistency will usually be worse. Professional flow is: lock looks first, then shoot.

Stage 12: Shoot like a set, not a chat window

The final stage is production discipline.

A real shoot path looks like this:

  1. Open the episode's video workspace.
  2. Confirm references are complete — or intentionally skip.
  3. Generate shot prompts.
  4. Read and edit the prompts.
  5. Choose model tier, aspect ratio, resolution, duration.
  6. Submit to the render queue.
  7. Review takes per scene.
  8. Rewrite prompts or swap references, then reshoot.
  9. Pick the best take.

Vertical 9:16 should be the default output shape, not a crop applied after the fact. Duration can be intelligent or fixed in the short-form range; audio can be generated; watermarks should be off by default.

The important mindset: editing a prompt is the AI equivalent of revising the shot note; swapping a reference is changing the look; resubmitting is a new take. You do not judge a scene by its first output. You judge it after you have multiple takes and a deliberate selection.

On the operations side, production queues should behave like production queues: separate lanes for script, art, and video; no duplicate submissions for the same in-flight scene; failure states that are visible and retryable; credit hold and release; timeout recovery for stuck jobs. This is what turns a script toy into something a studio can actually run.

Where the pipeline is honest about limits

No pipeline is magic. Teams should know the boundaries up front:

  • There is no automatic "best clip" judge. Final quality calls remain with the creator and producer.
  • Reference images are not a hard gate. You can skip them, but quality usually drops.
  • Reference tags in prompts depend on disciplined formatting; creators should still review them.
  • The ~10-second scene block is an engineering heuristic, not frame-accurate editing; final trimming and assembly still belong in post.
  • The product unit is one scene with references → one clip. Cross-scene automatic video continuation is not a standard workflow.
  • Character consistency depends on the art pipeline and prompt discipline, not on a magic face lock.

These are not weaknesses to hide. They are the terms under which a team can plan capacity, assign roles, and deliver episodes on schedule.

How Maosika fits into this

Maosika (猫斯卡) is built around exactly this sequence. It is an AI production operating system for vertical short dramas, not a single text-to-video button. It hardcodes the order that professional productions already follow: brief → archive → batch script → style → look-dev → scene blocks → shot prompts → multi-take shoot.

Its 18 digital specialists mirror real crew roles — archivist, producer, writer, script supervisor, casting, art director, wardrobe, props, storyboard artist, director of photography, director, camera operator, editor, VFX supervisor — and expose streaming work logs so producers can see what each role is doing, rather than staring at a black box.

What Maosika does not do is promise one-click hits. It turns the repetitive, drift-prone, easy-to-break parts of the workflow into a constrained pipeline. Taste, story judgment, and final selection still sit with the humans making the show.

If that matches how you already think about production, you can see the full workflow at Maosika.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com