Wan 3.0: The Essentials

Last updated: August 24, 2026

Covers Wan 3.0.

Wan 3.0 is Alibaba's flagship video model, and it does the whole job in one pass: it renders the picture and synchronized audio together, so dialogue, ambient sound, and music arrive with the video instead of being dubbed in later. Drive it three ways: text-to-video, a first frame (with an optional matching last frame), or reference-to-video using up to ten images, five videos, and five audio clips to lock subjects, motion, and sound. Clips run 2 to 30 seconds at up to 1080P.

Open this asset in Scenario

Hero: two rival chefs trade spoken lines mid-scene, lip-synced entirely from prompt text, no reference audio uploaded. asset_VFfjCJYHcDHjvTTqTE2h3fUH


Which Model Should I Use?

Model

ID

Input

Best for

Wan 3.0

model_alibaba-wan-3-0

Text, first/last frame, or up to 10 reference images + 5 reference videos + 5 reference audio clips, plus documents and web links

Anything needing native spoken dialogue, multi-reference input combined in one call, or turning a document/webpage directly into video

Wan 2.7 T2V / I2V

model_wan-2-7-t2v / model_wan-2-7-i2v

Text or a single image

Simpler, single-purpose clips (text-only or image-only) when you do not need dialogue or combined reference inputs

Reach for Wan 3.0 whenever the brief needs a character to speak on camera, needs several kinds of reference media locked into one shot, or needs a document or webpage turned into a video brief. Use the older Wan 2.7 line for a plainer, narrower job.


How to Use the Model

How Wan 3.0 Works

Every Wan 3.0 request renders picture and audio in a single pass. The prompt field takes one plain-language description; when you attach reference images, videos, or audio, refer to them in the text as "Image 1," "Video 1," "Audio 1" (numbered per type, in the order you added them). Three ways to drive it:

  • Text-to-video: prompt only, nothing else attached.

  • First/last frame: set image for an opening frame, optionally endImage for where the clip should land.

  • Reference-to-video: referenceImages, referenceVideos, and referenceAudios lock subjects, motion, or sound from real media instead of describing them from scratch.

Writing Prompts as a Shot List

Wan 3.0 rewards prompts written like a shot list rather than one flowing sentence: numbered beats with timestamps, a camera direction per beat, then a style block, then audio as two separate cues, a music description in parentheses and a plain list of sound effects. This structure gave noticeably more coherent multi-beat action than a single-paragraph prompt in testing.

Beat one, 0 to 3 seconds: a knight plants both boots on a cracked battlefield,
grips a glowing greatsword, raises it overhead. Camera: low heroic angle.
Beat two, 3 to 6 seconds: swings down in a wide arc that shatters a goblin's
shield, sending it flying backward. Camera: whip-pan following the blade.
Beat three, 6 to 12 seconds: lowers the sword to rest point-down in the dirt,
looks up as more goblins approach. Camera: wide shot holding on the standoff.
Cinematic fantasy game trailer style, overcast stormy sky, cool grey light
with a warm blue glow from the sword.
(deep orchestral hit landing on the shield impact, fading to a tense drone)
<armor clanking, a deep whoosh on the swing, a cracking shield impact,
distant goblin screeches>

Open this asset in Scenario

Knight vs. goblin, 4:3 1080P, beat-timestamp prompt above.

Native Dialogue and Lip-Sync

Wan 3.0 will generate spoken dialogue straight from a quoted line in the prompt, no reference audio file needed, and it lip-syncs the character's mouth to the generated speech. Write the line as dialogue attributed to the character, and keep the camera on their face while they speak. Measured a spectrogram on one test clip and found concentrated speech-band energy lining up exactly with the quoted dialogue window, confirming this is genuine synthesized speech, not a music bed with captions.

...she turns to face the camera directly, holds up her wrist, and says with
a warm smile, "This thing tracked my whole run without me even thinking
about it." She taps the screen once more and gives a thumbs up.
Camera: close-up on her face and wrist, slow push-in during her line.

Open this asset in Scenario

Smartwatch testimonial, dialogue confirmed by spectrogram analysis, 16:9 1080P.

Reference-to-Video with Images

Pass up to 10 referenceImages and Wan 3.0 will build the scene around the subjects and style shown, no lengthy text description required. This is the fastest path to a consistent character or style across multiple clips.

Open this asset in Scenario

Stylized desert race, referenceImages locking two character designs and the art style

Open this asset in Scenario

Robot soldier, referenceImages driving a war-torn cinematic scene

Cinematic Range at Different Resolutions

Open this asset in Scenario

Arctic fox under the aurora, 9:16 1080P

Open this asset in Scenario

Classroom volcano demo, 4:3 480P, the cheapest resolution tier


Parameters

Every field below lives on the same call; only combine image/endImage OR the reference arrays, never both.

prompt

Required, up to 5000 characters. Describe the full shot: subject, action, camera, style, and audio cues in one string. Reference numbered inputs as "Image 1," "Video 1," "Audio 1" per type. See the beat-timestamp structure above.

image

An optional opening frame. Required if endImage is set. Cannot combine with referenceImages, referenceVideos, or referenceAudios.

endImage

An optional closing frame. Only works alongside image. Gives Wan 3.0 a start and end point to animate between.

referenceImages

Up to 10 images, 20MB each, to lock subjects or style. Used for the desert race and robot soldier examples above. Cannot combine with image/endImage.

referenceVideos

Up to 5 videos, 100MB each, 15 seconds combined, to guide motion. Not verified in this batch; test before relying on it for a production shot.

referenceAudios

Up to 5 audio clips, 15MB each, 15 seconds combined, to guide the video's sound. Not verified in this batch.

document

An optional document (up to 100MB or 50 pages) to generate a video from. Turns on enableThinking automatically. Not verified in this batch.

link

A public https:// URL, reachable without authentication, to generate a video from. Turns on enableThinking automatically. Not verified in this batch.

resolution

480P, 720P, or 1080P (default 720P). 1080P is sharpest and was used for most examples above; 480P still holds up for quick concepting, as seen in the volcano classroom example. Higher resolution costs more.

aspectRatio

adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 (default adaptive). Adaptive follows your input media if you provided any, otherwise falls back to 16:9.

duration

2 to 30 seconds, or -1 for Smart (let Wan choose). Longer clips cost more. Every example above runs 10 to 15 seconds; short clips under 10 seconds tend to read as unfinished demo reels rather than usable footage.

audio

Boolean, default true. Adds the generated soundtrack (music, ambience, and dialogue when prompted). Costs the same whether on or off, so there is no reason to turn it off.

enableThinking

Boolean, default false. Lets Wan reason about the request before generating; automatically turned on for document or link input. Used on nearly every multi-beat and dialogue example above and produced noticeably more coherent results on complex prompts than leaving it off.


Use Cases

  • Product marketing: a testimonial-style clip with a character speaking directly to camera about a product, generated with no actor, no mic, no editing pass.

  • Game trailers: multi-beat combat or reveal sequences with camera direction baked into the prompt, at up to 1080P and 30 seconds.

  • Film pre-visualization: a scene sketch with dialogue, blocking, and mood lighting to pitch a shot before a real shoot.

  • Character and style consistency: reference images carry a character design or art style across multiple clips without redescribing it each time.

  • Education: a classroom demonstration clip with narration and reactions, useful for explainer content, even at the cheapest 480P tier.

  • Nature and mood pieces: atmospheric, dialogue-free shots for background footage or social content.


Tips for Better Results

  1. Structure prompts as numbered beats with timestamps ("Beat one, 0 to 3 seconds: ...") rather than one long paragraph. This gave visibly more coherent pacing on multi-action clips.

  2. Split audio into two tagged parts: a music cue in parentheses and a plain list of sound effects in angle brackets. This reads more reliably than folding everything into one "ambient audio" clause.

  3. Give every handled object an explicit final position. A prop that is picked up, opened, or set down needs a stated end state ("sets the cap down visibly on the table") or it can vanish mid-shot. See Known Limitations.

  4. Write dialogue as a quoted line assigned to a character and keep the camera on their face while they speak. Confirmed via spectrogram that this produces real synchronized speech.

  5. Turn on enableThinking for anything with multiple beats or spoken dialogue. It was used on every complex example in this article and is required automatically for document or link input.

  6. Describe the visual quality of a genre rather than naming it ("high-contrast blue and magenta neon, rain-soaked streets" instead of "film noir style"). Naming a style directly can cause the model to render the label as garbled signage in the scene.

  7. Keep clips at 10 seconds or longer with a full, detailed prompt. Every polished example in this article runs 10 to 15 seconds; shorter clips read as rough tests rather than finished shots.


Known Limitations

  • Object permanence. Small props a character manipulates (a lid, a glove, a photograph) can vanish instead of staying on screen through the action. Workaround: state the object's visible end position explicitly in the prompt.

  • Style names can render as on-screen text. Naming a genre directly in the prompt (for example "film noir") produced a neon sign in one test clip with the phrase distorted and half-legible. Describe the visual traits instead of the label.

  • Reference video and audio inputs, plus document and link input, are exposed in the schema but were not exercised in this example batch. Test them before relying on them for production work; only text-to-video and image-reference-to-video were verified here.

  • Pinning examples on the model page is a manual step. The example gallery auto-populates from recent generations across the team and is not reliably curated through the API; a human orders and pins the final set in the Scenario UI.

  • Use Scenario CDN URLs only. Provider asset URLs can expire; always link and embed the Scenario-hosted copy.