中文

AI Short-Drama Case Study: Four References to a 15-Second Fantasy Scene

Maosika Editorial | Last updated

A fantasy confrontation from Yaochi Snow (瑶池雪) shows what an AI short-drama workflow prepares before video generation. Maosika turned one scripted scene, two character references, one location image and one six-panel storyboard into a landscape clip. This case covers Episode 1, Scene 4 only, with the original inputs and output available below.

Watch the scene before judging the workflow

Yaochi Snow, Episode 1, Scene 4: Su Xue and Mo Han confront each other on a cloud-sea bridge.

This scene was produced on September 28, 2026. The production record identifies one video generation for this scene and reports no manual changes to the generated prompt. That is a result for this example, not a prediction of how many attempts another scene will need.

The request was for a 15-second, 16:9 video. The supplied file measures 15.08 seconds at 24 frames per second, with a 1280 × 736 frame. Its delivered dimensions are therefore not exactly 16:9. The clip contains Chinese dialogue.

Start with a conflict that can be seen

Su Xue protects an ice lotus linked to the mountain's protective barrier. Mo Han wants the lotus to save children afflicted by a deadly cold. On a white-stone bridge above the clouds, she raises a golden shield. He strikes it with flame, spreading cracks across the barrier.

Her warning means that taking the flower will break the mountain's protection. His reply challenges her to choose what she is willing to sacrifice. The action makes the unresolved choice visible.

The script supplied the location, the two characters, two action passages and two dialogue lines. It did not specify every camera position, reference binding or sound instruction. Those production decisions were prepared around the script.

Four reference images, four responsibilities

Su Xue's original character reference combines a facial close-up with a full-body costume view.
Su Xue's original character reference combines a facial close-up with a full-body costume view.

Su Xue's character sheet provides her face, hair, pale-blue dress and full-body appearance. The prompt assigns appearance to this image and adds posture and voice instructions in text. This avoids asking a written face description to compete with the approved image.

Mo Han's original character reference shows his face, dark costume and full-body silhouette.
Mo Han's original character reference shows his face, dark costume and full-body silhouette.

Mo Han's character sheet supplies the second identity: dark clothing, hairstyle and silhouette. Named references help distinguish them across camera angles.

The original location reference: a white-stone arch bridge above moonlit clouds, without characters.
The original location reference: a white-stone arch bridge above moonlit clouds, without characters.

The location image establishes the bridge, carved railings, moon, surrounding clouds and distant mountains. It contains no actors, so it can define the setting without introducing an extra person into the scene.

The original six-panel storyboard for this same scene, read left to right and top to bottom.
The original six-panel storyboard for this same scene, read left to right and top to bottom.

The storyboard supplies staging and composition. Its six panels move from the bridge and shield to the strike, dialogue viewpoints and final confrontation. The prompt explicitly uses those compositions without adopting the sketch's drawing style. The pencil storyboard guides a realistic scene.

For the wider reference strategy, see the AI video consistency guide.

What the 18 digital experts prepare

Maosika organises its assistance into 18 digital production roles: five for story and script work, and thirteen for production. These are software roles, not a claim that eighteen people made the film or that every role needed to act in this scene.

Story roles cover producing, writing, maintaining the story archive, checking the episode structure and recording decisions. Production roles handle scene alignment, casting, visual style, character design, locations, props, material readiness, storyboards, framing, generation submission, directing, take selection and technical monitoring. This scene needed no separate prop reference, and its single generated take did not require comparing several alternatives.

The recorded prompt has 1,621 characters, including whitespace but excluding the saved text file's final newline. It assigns the four images their jobs, describes shot order and preserves the scripted dialogue. It also specifies the shield and flame as distinct effects, camera transitions, voices, sound cues and constraints against subtitles, watermarks and duplicate characters. The complete short-drama production guide explains how these tasks fit into an episodic workflow.

What the output demonstrates—and what still needs review

Sampled frames from approximately 0, 3, 6, 9 and 12 seconds show the cloud-sea bridge, a golden dome, a hand striking the shield, bright cracks and opposing character positions. They make the intended action and visual references traceable to the output.

They do not establish perfect motion, lip synchronisation or every spoken word. Review the full clip with sound before accepting it. Check character identity across cuts, hand anatomy at contact, the shield's geography, flame colour, dialogue intelligibility and whether the last shot preserves the dramatic question. The intended dark flame appears warm red around the hand in sampled frames, illustrating why a prompt's instructions and the rendered result should be judged separately.

The original video-task record shows 7 minutes 53 seconds from processing start to completion, including preparation and delivery. The video task consumed 653.60 credits on September 28, 2026; this excludes earlier writing and reference-image preparation. It is a historical charge for this example, not a current quote or a total episode budget. Clip length is not production time. This is one scene, not a completed episode or season.

Reuse the preparation, then review your own take

The reusable lesson is to give each reference a clear responsibility, keep the dramatic action explicit and inspect the generated scene against those decisions. A successful first take is useful evidence of one production run; creative approval still belongs to the filmmaker.

Start an English-language short-drama project with your own conflict and approved reference materials.

FAQ

What exactly was generated in this case study?

One scene from Yaochi Snow: Episode 1, Scene 4, a confrontation between Su Xue and Mo Han on a white-stone bridge above the clouds. It is not a complete episode or season.

How many input reference images were used?

Four: one character sheet for each of the two characters, one location image and one six-panel storyboard. Each had a separate role in the generated prompt.

Was the video generated successfully on the first attempt?

The recorded source identifies one video generation for this scene and no manual prompt edits. This describes this example only and does not guarantee first-attempt results for other scenes.

What were the actual video dimensions and duration?

The request was 15 seconds at 16:9. The supplied file is approximately 15.08 seconds at 24 fps and 1280 × 736 pixels, which is not an exact 16:9 frame.

How long did production take and what did it cost?

The recorded video task took about 7 minutes 53 seconds from processing start to completion and consumed 653.60 credits on September 28, 2026. The charge excludes writing and reference preparation, and is not a current quote.

About Maosika — Maosika · Professional AI Video Production System. It connects briefing, scripting, look development, shot prompts and delivery into one reviewable pipeline. www.maosika.com