Vidu Q4: The Essentials
Last updated: October 9, 2026
Vidu Q4 I2V turns a starting image into a moving shot. Vidu Q4 Reference2V brings separate characters, objects and environments into a new scene, with optional reference voices. Both support 3 to 16 second clips, resolutions from 540p to 4K, and synchronized audio.
Use them to direct a product reveal, a journey through impossible environments or a character performance that changes visual style. The examples below were generated on Scenario at 4K, with their original prompts and settings included.
Which Model Should I Use?
Model | Inputs | Best for |
|---|---|---|
One starting image and an optional prompt | Animating an existing composition with action, camera movement and sound | |
1 to 15 images, a required prompt and up to 3 optional voice clips | Combining separate subjects and settings, or guiding a performance with reference voices |
Choose I2V when you already have the opening composition. Choose Reference2V when you want to define the ingredients of a new scene. Reference2V provides five output aspect ratios; I2V follows its source image.
How to Use the Model
Animate a starting image with Vidu Q4 I2V
Open the I2V model, add one image and describe a complete visual sequence: what starts moving, what changes, where the camera travels and how the shot ends. Set Duration, Resolution and Audio, then generate. The two I2V examples below use Enhance Prompt off, preserving the submitted direction without automatic rewriting.
Porcelain tempest
Vidu Q4 I2V: A painted dragon emerges from a porcelain bowl, becomes a three-dimensional creature and returns to the glaze. This 16 second experiment demonstrates a complete transformation arc. The final painted motif differs from the starting image, so the return is conceptual rather than an exact loop.
16 seconds · 4K · 3840 × 2160 · Open this video in Scenario
Visual references, in input order:
Exact generation settings:
{
"duration": 16,
"resolution": "4K",
"audio": true,
"enhancePrompt": false,
"seed": 6120
}Original prompt:
The cobalt dragon painted inside the porcelain bowl in the opening image begins to move within the glaze. The 75mm macro camera creeps toward its tiny painted eye as the raised blue droplet falls back onto the ceramic with a clear musical ping. During the first four seconds the dragon swims around the bowl as living two-dimensional ink, its painted scales and cloud patterns remaining attached to the curved porcelain surface. Then its head pushes out of the glaze into three dimensions, pulling a long fluid cobalt body behind it like a ribbon of concentrated ink. The camera pulls back and lowers as the dragon spirals above the bowl, now several times larger than the object, whiskers trailing blue droplets and milky porcelain dust. Between seconds seven and eleven the dragon coils into a violent but beautiful miniature storm: its tail whips a ring of water into the air, and dozens of delicate porcelain shards orbit without the original bowl visibly disintegrating. The camera flies through the dragon's open coil for a dramatic parallax reveal, showing luminous white ceramic scales forming over its liquid-blue spine. In the last five seconds the dragon dives gracefully headfirst back into the bowl; water, shards and porcelain scales flatten into the original cobalt painted motif, and the final ripple becomes a still glossy reflection. Finish on the pristine recognizable bowl, with one tiny painted eye blinking as the last surprise. Sophisticated museum lighting, dark stone and mist, extraordinary fluid-to-ceramic transitions, no gore or combat. Porcelain chimes, brushlike ink swishes, a miniature thunder roll and deep resonant strings synchronize with the rise and return. No text, subtitles or visible technical instructions.Gold tide
Vidu Q4 I2V: A perfume bottle is framed by liquid gold, suspended droplets and metallic fish. The product remains recognizable through an ambitious fluid sequence, although cap proportions and the number of fish vary.
12 seconds · 4K · 2160 × 3840 · Open this video in Scenario
Visual references, in input order:
Exact generation settings:
{
"duration": 12,
"resolution": "4K",
"audio": true,
"enhancePrompt": false,
"seed": 6113
}Original prompt:
The amber perfume bottle with its spherical gold cap remains the exact hero product seen in the source image, standing upright on its obsidian pedestal. A 90mm macro camera begins close to the curved wave of liquid gold, following a single glistening droplet as the wave curls higher around the bottle. Over the first four seconds the wave becomes a magnificent thin golden shell surrounding the product like a breaking ocean barrel, with the camera gliding through its hollow center toward the glass. Amber light refracts through the bottle onto the wet black pedestal; its straight edges and gold sphere remain rigid while the fluid moves. At the midpoint the golden barrel bursts into a crown of suspended droplets, and the camera executes a precise half-orbit at product height. Time slows enough to reveal each droplet reflecting the bottle, then resumes as the droplets fall in a circular cascade around, never through, the glass. The final three seconds reveal two luminous golden koi formed from the remaining liquid, swimming one graceful orbit around the base before merging into a perfectly still ring. Finish in a strong centered vertical product portrait with the clean spherical cap silhouetted against warm darkness. Photorealistic fluid simulation, sculptural chiaroscuro, tactile gold and amber, premium beauty-film polish. A deep liquid swell, crystalline metallic droplets and a luxurious rising string phrase synchronize with the burst and settle into a single resonant note. No text, label or narration.Connect subjects and worlds with Vidu Q4 Reference2V
Add the reference images in the order you want to address them. In the prompt, use @1, @2 and so on, and explain the role of each reference. Then describe how those ingredients interact in the same scene. References guide the content; they do not lock the first or last frame.
The impossible express
Vidu Q4 Reference2V: Four references define the train, greenhouse, fossil canyon and storm sphere. The clip connects them through camera movement and physical openings. The final return reads as a portal exit rather than a mathematically exact nested loop.
10 seconds · 4K · 3840 × 2160 · Open this video in Scenario
Visual references, in input order:
Exact generation settings:
{
"duration": 10,
"resolution": "4K",
"aspectRatio": "16:9",
"audio": true,
"seed": 7218
}Original prompt:
A breathtaking ten-second continuous journey using the exact indigo-and-brass locomotive and single passenger carriage from @1, with the crescent ornament, amber windows and articulated wheels preserved. Begin with a low 24mm camera racing backward just ahead of its headlamp on the brass rails in the luminous orchid greenhouse from @2. Giant wet leaves sweep past in strong parallax as the train accelerates toward a circular opening between roots. At two seconds the camera passes through that opening and the same rails emerge seamlessly into the fossil-rib canyon from @3; the train follows immediately, emerald spores giving way to copper dust through the physical tunnel transition. The camera rises beside the moving train and banks toward the enormous spiral ammonite ahead, revealing the tracks entering the shell's central opening. At five seconds pass through that spiral into the storm-filled glass sphere from @4: the same train now runs along a delicate track suspended over the sphere's miniature dark ocean, with violet lightning reflecting in its indigo lacquer. At seven seconds the camera pulls backward through the curved transparent glass wall to reveal the whole brass-mounted sphere standing among orchids in the original greenhouse. Its tiny train continues moving inside. The tiny headlamp flares into the bright lamp of the original larger miniature train rushing past camera on the greenhouse rails, closing the spatial loop by second ten. Keep the vehicle moving forward, the single carriage attached and the three world transitions readable; this is a connected camera journey rather than a slideshow. Rich wet leaves, pale fossil bone, sparkling glass and precise Art Deco metalwork carry an emerald-to-copper-to-violet progression. Continuous wheel rhythm and a bright whistle pass through greenhouse drips, desert wind and contained thunder, with a triumphant orchestral accent at the final loop. No dialogue or text.Carry a performance across visual styles with Vidu Q4 Reference2V
Describe the movement that continues through each transformation, and identify the visual details that should remain recognizable. In this example, the dancer’s red costume and continuous phrase connect the changing media.
One body, four realities
Vidu Q4 Reference2V: One dancer reference anchors a sequence through live action, ink animation, a paper-flower setting and a luminous figure. The red costume and movement connect the styles; the paper section retains a partly realistic body.
10 seconds · 4K · 2160 × 3840 · Open this video in Scenario
Visual references, in input order:
Exact generation settings:
{
"duration": 10,
"resolution": "4K",
"aspectRatio": "9:16",
"audio": true,
"seed": 8217
}Original prompt:
A spectacular ten-second continuous dance transformation starring the exact shaved-head, dark-skinned dancer from @1 in the crimson ribbon-sleeved costume and bare feet. The same face, proportions, costume silhouette and red accent anchor every visual medium. Begin full-body in tactile live-action fashion cinematography on a wet black stage, using a low 35mm camera that circles steadily. The dancer launches immediately into a traveling turn and a dramatic turning leap, wide sleeves carving red arcs through the air. At two seconds a sleeve sweeps close across the lens and transforms the whole world into bold hand-painted animation on cream paper; visible ink contours and crimson brush strokes continue the exact same leap without restarting the movement. The dancer lands and enters a low sweeping turn. At four seconds the brush trail folds into an intricate stop-motion figure made of cut paper and red silk, still performing the same phrase; the paper stage opens behind the figure into a huge sculptural flower. At six seconds those paper petals dissolve into luminous red-and-gold particles, and the dancer becomes a transparent volumetric sculpture of starlight, with the original head, arms, bare feet and flowing costume clearly readable. The camera accelerates around this luminous figure as it rises into a final upright spin. At eight seconds a decisive foot stamp pulls the particles inward and restores the original photorealistic dancer and real crimson fabric. Finish on a fierce balanced full-body pose while the long sleeves settle in a strong diagonal. Every transformation is carried by moving material and camera occlusion, with continuous choreography rather than four unrelated clips. Driving foot percussion, breath, brush swishes, folded-paper clicks and a bright electronic crescendo share one uninterrupted tempo, resolving at the stamp. No words, captions, logos or duplicate bodies.Direct dialogue with Vidu Q4 Reference2V
Keep Audio enabled. Add up to three MP3 references, each 3 to 12 seconds and no larger than 50 MB. Assign each audio clip to a speaker and write the exact dialogue in the prompt. Leave time for the visible action as well as the words. After longer requests timed out, this example completed at 8 seconds while retaining 4K and all three voice inputs.
The stolen second
Vidu Q4 Reference2V: Four images and three original synthetic voice clips guide an auction-room heist. Independent transcription recovered all three requested lines in order: “Time is for sale.”, “I’ll take it.” and “Gravity sold separately.” The final levitation is more explicit than the requested time freeze. Compare the voices and mouth movement yourself before selecting a final take.
8 seconds · 4K · 3840 × 2160 · Open this video in Scenario
Visual references, in input order:
Voice references, in input order:
Auctioneer voice reference (6.08 seconds, MP3)
Woman voice reference (5.84 seconds, MP3)
Robot voice reference (5.84 seconds, MP3)
Exact generation settings:
{
"duration": 8,
"resolution": "4K",
"aspectRatio": "16:9",
"audio": true,
"seed": 9216
}Original prompt:
An audacious eight-second time-heist inside a magnificent Art Deco auction room, combining four exact visual references and three distinct speaking voices. @1 is the silver-haired, silver-bearded auctioneer in burgundy brocade with his wooden gavel. @2 is the dark-haired woman in emerald velvet holding a gold pocket watch. @3 is the brass porter robot with cyan eyes, black bow tie and silver tray. @4 is the brass-mounted glass sphere containing violet lightning, resting on a black pedestal between them. Keep their faces, clothing, materials and separate props recognizable. Assign the first audio reference to the auctioneer, the second to the woman and the third to the robot. Open in a beautifully composed 35mm moving medium-wide shot containing all three characters around the sphere, with a huge crystal chandelier above. The auctioneer raises his gavel and says clearly during the first two seconds, 'Time is for sale.' At two seconds the woman snaps her watch shut. The descending gavel and the sphere's lightning freeze visibly in mid-motion, but she continues moving. A fast semicircular camera move follows her hand as she confidently lifts the sphere from its pedestal. Between seconds two and a half and four she says, 'I'll take it.' The robot looks at the stolen sphere, lifts its tray in surprise and replies between seconds four and a half and six, 'Gravity sold separately.' Exactly one character speaks at a time, with readable mouth movement and no overlapping speech. On the robot's final word the woman clicks the watch again. The gavel strikes, the frozen lightning bursts into motion and gravity reverses: the sphere rises above her open hands, the robot's tray floats upward and hundreds of chandelier crystals lift into a brilliant spiral. During the final two seconds the camera sweeps upward through this glittering vortex to reveal all three characters below, their faces illuminated by the hovering violet storm. Keep the dramatic action centered in this one coherent room. Premium photorealistic fantasy-film lighting, gold highlights, rich emerald velvet, sharp brass detail and a decisive visual climax. Ticking stops during the freeze, speech remains clean, then a gavel crack and orchestral flourish accompany the rising crystals. No subtitles, visible words, labels or split screens.Stage a larger ensemble with Vidu Q4 Reference2V
Give every reference a distinct role and position. A large creature, a human-scale rider and a small object need different amounts of screen space. This example uses eight images and a shared activation event to organize the scene.
Worlds awaken
Vidu Q4 Reference2V: Eight visual references share one museum scene: a courier, motorcycle, knight, dragon, ceramic astronaut, jellyfish, glass moth and storm sphere. The motorcycle, violet activation ring and opening roof create a coordinated reveal. The astronaut floats below the jellyfish instead of riding its crown.
10 seconds · 4K · 3840 × 2160 · Open this video in Scenario
Visual references, in input order:
Exact generation settings:
{
"duration": 10,
"resolution": "4K",
"aspectRatio": "16:9",
"audio": true,
"seed": 9220
}Original prompt:
An extraordinary ten-second fantasy finale combining eight precise visual references in a monumental circular museum beneath a glass dome. @1 is the silver-haired courier in orange armor riding @2 the emerald-and-brass scarab motorcycle; @3 is the blue-cloaked knight with the amber lantern standing beside @4 the immense silver-white ice dragon; @5 is the tiny cobalt-patterned ceramic astronaut riding safely on the glowing bell of @6 the aurora jellyfish; @7 is the faceted dichroic glass moth hovering over @8 the brass-mounted storm sphere. Preserve these exact characters, vehicles, creatures and objects as distinct subjects with the stated relationships. Open low on a 24mm lens with the scarab motorcycle sweeping across the foreground, courier leaning naturally over its handlebars, while the lantern knight and dragon occupy the distant central aisle. The camera follows the bike's curve around the storm sphere at the center. At two seconds the knight raises the amber lantern and the moth touches the sphere with one glass wing; violet lightning races through the museum floor in an intricate luminous circle. At four seconds that circle becomes a huge column of light, lifting the jellyfish with its small ceramic passenger toward the dome as the dragon spreads its wings behind the knight. The glass roof opens in enormous mechanical petals, revealing a blazing sunset and clouds outside. The camera cranes upward between the dragon's wing and the jellyfish's luminous ribbons, keeping both the knight below and the ceramic rider above readable. During the final three seconds the motorcycle completes its arc around the sphere, the moth spirals into the light, and the camera pulls into a majestic wide tableau of all eight subjects beneath the open sky. The museum remains coherent, the characters keep their own materials, and the scale hierarchy stays clear: monumental dragon and jellyfish, human-sized courier and knight, small central glass artifacts. Premium photorealistic fantasy-film craft, warm gold sunlight, violet lightning and shimmering cyan ribbons. A coordinated soundtrack combines engine sweep, lantern chime, deep dragon breath and an orchestral crescendo, resolving as the roof opens. No dialogue, text, labels or split-screen collage.Parameters
The controls below are available on the Scenario model pages. The examples keep their exact input order, prompt and seed so you can inspect the complete setup.
images
I2V requires one starting image. Reference2V requires 1 to 15 images, addressed by upload order with @1, @2 and so on. Compare the single source in Porcelain tempest with the eight roles in Worlds awaken.
prompt
Required for Reference2V and optional for I2V, up to 20,000 characters. Describe action, transitions, camera movement and sound as one clear sequence. The impossible express connects its environments through explicit openings and camera travel.
duration
An integer from 3 to 16 seconds; default 5. Porcelain tempest uses 16 seconds, while The stolen second uses 8 seconds. Fit the number of actions and spoken words to the selected length.
resolution
540p, 720p, 1080p, 2K or 4K; default 720p. All six examples use 4K. Output dimensions vary with aspect ratio: Gold tide is 2160 × 3840, while The impossible express is 3840 × 2160.
audio
On by default. Generates audio such as dialogue, ambience and sound effects; turn it off for silent output. The stolen second includes three quoted lines, recovered in the expected order by transcription.
enhancePrompt
I2V only, on by default. Allows Vidu to rewrite the prompt. It is off in Porcelain tempest and Gold tide. These examples do not compare the quality of the on and off settings.
aspectRatio
Reference2V only: 16:9, 9:16, 1:1, 4:3 or 3:4; default 16:9. One body, four realities uses 9:16. I2V has no separate aspect-ratio control and follows its source image.
referenceAudio
Reference2V only. Supply up to three MP3 clips, each 3 to 12 seconds and up to 50 MB. The stolen second supplies three clips and assigns them to the auctioneer, woman and robot. Reference similarity needs a listening check; matching dialogue alone does not establish a voice match.
seed
Optional numeric seed. Record it with the prompt, references and other settings when revising an example. The impossible express uses seed 7218. Do not assume that a seed guarantees identical results across changes to other inputs or the model.
Use Cases
Game cinematics: animate magical artifacts, creatures and journeys through designed environments.
Product campaigns: surround a recognizable product with motion, fluid effects and a final hero view.
Film development: test a compact action sequence, camera route or transformation before building a longer scene.
Music and performance: carry a dancer or musician through expressive visual changes.
Character storytelling: combine a cast, shared props and short dialogue into a miniature narrative.
Social campaigns: use vertical framing for product and performance clips, with a readable ending.
Tips for Better Results
Build an action arc. The porcelain example has an emergence, a storm and a return. Give the scene a visible destination rather than stopping after its opening effect.
Ground the prompt in the references. Name the actual material, costume and object seen in each source. The train example names its locomotive and maps the environments separately.
Use a continuing action to connect styles. The dancer keeps moving through the changes, with the red costume acting as a recognizable detail.
Assign every reference a role. The eight-image museum example separates characters, vehicles, creatures and small objects instead of treating the inputs as an undifferentiated style board.
Keep dialogue compact. Three short lines fit into the eight-second auction clip and leave a final visual beat. Check the resulting words and delivery.
Revise duration when a complex request fails. Several 4K Reference2V requests completed after being shortened to 8 or 10 seconds. Shortening is a practical retry option, not a guarantee that a job will succeed.
Judge the output, not only the prompt. The examples show differences in object count, material transitions and positioning. Review the entire clip and listen to the audio before selecting it.
Known Limitations
References guide rather than lock identity and geometry. Fine product proportions, painted patterns and character details can change during motion. The perfume cap varies, and the dragon returns to a different painted motif.
Complex instructions may be simplified. The paper dance section retains a partly realistic body, and the museum astronaut floats below the jellyfish rather than riding it. Treat these as creative interpretations, not exact execution of every prompt clause.
The input ceiling is not a tested success guarantee. Scenario accepts up to 15 images. Three attempts at a 15-image ensemble failed with service errors; the eight-image version completed. Those failures do not establish that 15-image generation is unsupported.
Voice and lip synchronization need review. The three-audio-reference example completed and its dialogue was transcribed correctly. This confirms the words, not exact similarity to each input voice or frame-perfect mouth movement.
Generation time and completion vary. Several high-resolution Reference2V attempts timed out or returned connection errors. Preserve the setup, retry thoughtfully and simplify the timeline if needed.
Controls differ between the two entries. Reference2V exposes aspect ratio and audio references; I2V exposes prompt enhancement. Neither Scenario entry provides text-only generation, a reference-video input or an end-frame input.
Provider naming. Shengshu documents the release as Q4 Preview; the Scenario entries are named Vidu Q4 I2V and Vidu Q4 Reference2V.
These examples were checked using native file metadata, audio measurements and frames sampled throughout each clip. The dialogue example also received independent transcription. Full playback, listening and final creative approval remain part of selecting a production take.