FLUX.3 Video: The Essentials

Last updated: August 6, 2026

FLUX.3 Video is Black Forest Labs' cinematic video family, and its headline feature is native synchronized audio: every clip is generated with sound baked in from the same prompt, no separate pass. Describe a line of dialogue and it is spoken and lip synced, name the sound effects and score and they arrive in time with the motion, in multiple languages. Clips run from 5 to 20 seconds at up to 1080p.

The family covers five ways to make a video, and each one ships in two tiers: a full model for final quality and a Draft model for fast, low cost previews. The five capabilities are Text to Video (start from a prompt), Image to Video (animate a still), First and Last Frame (interpolate between two images you provide), Keyframes (hit a timed sequence of images), and Extend (continue an existing clip). Pick the capability by what you are starting from, then pick the tier by whether you are exploring or finishing.


Which Model Should I Use?

Capability

Full model

You provide

Best for

Text to Video

model_bfl-flux-3-t2v

A prompt

Building a shot from nothing: concepts, ads, establishing shots

Image to Video

model_bfl-flux-3-i2v

One image plus a prompt

Bringing a single still, character, or product shot to life

First and Last Frame

model_bfl-flux-3-first-last-frame

A start image and an end image

A controlled transition that has to land on an exact final frame

Keyframes

model_bfl-flux-3-keyframes

Two or more timed images

Choreographing a shot that must pass through specific poses at specific times

Extend

model_bfl-flux-3-extend

An existing video clip

Continuing a shot to make it longer without a visible cut

Every capability above also has a Draft twin with the same name plus -draft (for example model_bfl-flux-3-t2v-draft). The Draft models take the same inputs and the same prompt structure, they simply trade quality for speed.

The rule of thumb: explore on Draft, finish on the full model. Draft is the fast, inexpensive way to test an idea, a duration, or an aspect ratio. Once a take is working, run the same prompt on the full model for the final render.


How to Use the Models

All ten models share one prompt field and one prompting method, so learning one teaches you all of them. Write a single, detailed paragraph that moves through the shot in order: the camera move and the subject, then the action described beat by beat with rough timings, then the environment, then the audio, then the visual style, and finally any constraints (identity locks, no on screen text, realistic motion). You are folding what other tools split across many fields into one plain prose description. These cue words do not appear as text on screen, they only steer the shot.

Two habits matter most. First, describe the audio explicitly: the dialogue lines in quotes, the sound effects, and the score. That is where this family separates itself from silent video models. Second, describe motion in timed beats and vary your shots: mix durations from 8 up to 20 seconds, use different aspect ratios, and put a moving camera and real action in the frame. A static single subject at a fixed length is what makes a clip read as dated.

Full Models vs Draft Models

The most important thing to understand about this family is how the two tiers respond to prompts.

The full models reward a complete, detailed prompt. The more precisely you describe the camera, the beats, the audio, and the style, the more the full model delivers exactly that, with strong temporal coherence and clean, high resolution motion. This is where you want your finished, carefully written prompt.

The Draft models work with a less detailed prompt and are a bit more inconsistent from run to run. They are a fast, low cost preview tier, so the output is lower resolution and lower overall quality than the full models. Use Draft to check a composition or a timing quickly, then re render the winning take on the full model. Note that Draft has no resolution control (the full models expose 720p and 1080p).

The pair below is the same scene on both tiers. The full model got the long, fully specified prompt, the Draft got a trimmed version of it.

Full model, complete prompt (Text to Video, 16:9, 1080p):

Fast dolly-in with a subtle whip-pan onto a battle-worn female space marine crouched behind a scorched barricade on a burning colony street at night, embers and ash swirling through the air. She slams a fresh magazine into her rifle, snaps her head toward the camera and shouts 'Fall back, now!' just as an explosion blooms behind her and debris scatters past the lens. Handheld shake, strobing muzzle flashes reflecting off her cracked visor, sparks raining down, other soldiers rushing past in the background. Audio: her urgent shouted line, 'Fall back, now!', a concussive explosion, ringing metal debris, distant gunfire and blaring alarm klaxons. Style: gritty cinematic sci-fi, anamorphic lens flares, high-contrast orange firelight against deep shadow, fine film grain. Constraints: keep her identity consistent, lips synced to the spoken line, no on-screen text, no logos, realistic physics and motion.

Full model output. Open in Scenario: asset_ktAm5Ntq3UnGS3ziTdZvVZ6x

Draft model, trimmed prompt (Text to Video Draft, 16:9, lower resolution):

Fast dolly-in on a battle-worn female space marine crouched behind a scorched barricade on a burning colony street at night. She slams a magazine into her rifle, turns to camera and shouts 'Fall back, now!' as an explosion blooms behind her. Handheld shake, muzzle flashes, sparks, soldiers rushing past. Audio: her shouted 'Fall back, now!', a concussive explosion, gunfire, alarm klaxons. Style: gritty cinematic sci-fi, firelight, film grain.

Draft output of the same scene: faster and cheaper to iterate, lower resolution and less consistent. Open in Scenario: asset_yLW78jA6kyNzD1vA15Ubq5Tc

Image to Video (model_bfl-flux-3-i2v)

Give the model one still and a prompt, and it animates that exact frame forward. Preserve the subject's design in the prompt so identity holds through the motion. Here a single dragon still becomes a roaring, fire breathing shot with a craning camera.

The colossal red dragon on the cliff from the image suddenly unfurls its massive wings with a thunderous downbeat, rears its head back and roars, blasting a torrent of fire up into the storm sky as the camera cranes back and upward to reveal the full scale of the beast and the castle far below. Embers stream past the lens, the storm clouds churn, and the dragon's amber eyes flare. Audio: a deep guttural dragon roar, the whoosh and boom of heavy wingbeats, crackling roaring flame, rolling thunder, a low cinematic brass swell. Style: epic fantasy film, volumetric dusk light, dramatic rim light on the crimson scales, fine detail, huge sense of scale. Constraints: preserve the dragon design and cliff exactly as in the opening frame, no on-screen text, no logos, realistic motion.

Image to Video, 16:9. Open in Scenario: asset_LxYr7JDJZvw5fHbQz3bxtZA7

First and Last Frame (model_bfl-flux-3-first-last-frame)

Provide a start image and an end image, and the model fills the motion in between, landing exactly on your final frame. The key to a clean result is to write the prompt as one continuous action that resolves on that last image, not as a static scene. This 20 second, 21:9 chariot finish starts wide on the bend and settles precisely on the provided finish line frame.

An epic chariot race finish, resolving exactly on the final frame, with a fast tracking camera. Beat one (0 to 6s): the two chariots thunder neck-and-neck around the coliseum bend, dust flying, the camera tracking alongside the pounding hooves. Beat two (6 to 13s): they hurtle down the straight trading the lead, wheels rattling, the charioteers whipping the reins as the crowd surges. Beat three (13 to 20s): one chariot bursts ahead across the finish line, its driver's fist raised in triumph, the crowd on its feet roaring, settling on the final frame. Audio: thundering hooves, rattling wheels, cracking whips, a colossal roaring crowd, blaring horns, an epic orchestral surge. Style: cinematic historical epic, dusty sunlight, huge scale, motion blur. Constraints: keep the chariots, arena and crowd identical, land exactly on the provided final frame, no on-screen text, realistic motion.

First and Last Frame, 21:9, 20 seconds, resolving on the supplied end frame. Open in Scenario: asset_qpnZC47Mb1Ler9EE7NTL5QsC

Keyframes (model_bfl-flux-3-keyframes)

Keyframes lets you supply several images pinned to specific times (in frames, at 24 fps) and have the model move through all of them in order. It works best when the keyframes share a composition (derive them by editing one base image so scale and position match), and when the prompt names the beat each keyframe belongs to. This 15 second, 2:1 samurai kata passes through four poses and sheathes on the final frame.

A samurai sword kata following the four keyframes, with a slow arcing camera. Beat one (0 to 4s): the samurai holds his ready stance in the misty courtyard, then explodes into motion. Beat two (4 to 8s): he swings a powerful overhead strike, katana flashing down, cloak flaring. Beat three (8 to 12s): he flows into a spinning horizontal slash, the blade a blur, petals scattering. Beat four (12 to 15s): he settles and smoothly sheathes the katana in a calm final pose, resting on the final frame. Audio: sharp fabric snaps, the ring and whoosh of the blade, quick footfalls, a single sheath click, a taut traditional string score, wind. Style: cinematic samurai film, cool dawn mist, hard rim light, film grain. Constraints: follow the four keyframes in order, keep the samurai and courtyard consistent, land on the final frame, no on-screen text.

Keyframes, 2:1, 15 seconds, moving through four supplied poses. Open in Scenario: asset_XJh1rWZmBmZ8sbVxWiYVC6Y1

Extend (model_bfl-flux-3-extend)

Extend takes a video clip you already have and continues it, so you can grow a shot past its original length without a visible cut. Feed it a short source clip (under 15 seconds) and a prompt that carries the action forward. Here the dragon shot above is continued into a full takeoff and flight.

Continue the scene: the red dragon on the cliff with wings spread. Beat one (0 to 3s): the dragon crouches, muscles coiling, wings sweeping back. Beat two (3 to 7s): it launches off the cliff with a thunderous downstroke, the camera falling away with it as it plunges then powers upward into the storm sky, embers and debris flying. Beat three (7 to 10s): it banks hard across the burning sunset, roaring, and soars toward the distant castle as the camera tracks its flight. Audio: a deep roar, the boom and whoosh of powerful wingbeats, rushing wind, rolling thunder, an epic brass swell. Style: epic fantasy film, volumetric dusk light, dramatic scale. Constraints: preserve the dragon design, no on-screen text, realistic motion.

Extend, continuing an existing clip. Open in Scenario: asset_nteByktj6rn9Ajtr2uLjXJQN


Parameters

The controls are shared across the family, with a few inputs that change per capability.

Prompt

Required. One detailed paragraph describing camera, action in timed beats, environment, audio, and style. The full models reward the most complete prompt you can write. Draft models still work with a shorter, looser prompt, they are simply less consistent about it. Always describe the audio you want, since that is generated from this same field.

Duration

How long the clip runs, from 5 to 20 seconds. Vary it deliberately rather than defaulting everything to one length: an 8 second clip and a 20 second clip read very differently. Longer durations give the model room for a full beat by beat action, shorter ones keep a single moment tight.

Aspect Ratio

Choose the frame shape: 21:9 and 2:1 for cinematic and epic wide shots, 16:9 for standard landscape, 4:3 for a classic look, 1:1 for square, and 3:4 or 9:16 for vertical and mobile. Matching the aspect ratio to the intended platform is one of the easiest ways to make a shot feel deliberate.

Resolution

On the full models, choose 720p or 1080p. The Draft models do not expose this control and render at a lower resolution by design, which is part of why they are faster and cheaper for previews.

Generate Audio

True by default, which is the point of this family: sound is generated from your prompt in the same pass. Set it to false only when you specifically want a silent clip.

Inputs (by capability)

Text to Video needs only a prompt. Image to Video takes one source image. First and Last Frame takes a start image and an end image. Keyframes takes two or more images, each pinned to a frame index (at 24 fps, so second 5 is frame 120). Extend takes an existing video clip to continue, ideally under 15 seconds. For the image based capabilities, generate clean single subject source stills first, then bring them in.


Use Cases

  • Game marketing and cinematics: action beats with native sound design, character barks, and score, in the wide aspect ratios trailers use.

  • Advertising and product: animate a single product still with Image to Video, or storyboard a spot on Draft before committing a full render.

  • Film and previz: block out shots, then use First and Last Frame or Keyframes to hit exact staging and land on a specific final frame.

  • Social and short form: vertical 9:16 clips with spoken lines already lip synced, no separate audio step.

  • Education and explainer: narrated scenes where the voiceover and on screen action are generated together and stay in time.

  • Longer sequences: chain a shot past its native length with Extend to avoid a hard cut.


Tips for Better Results

  1. Write the full prompt for the full model. The more precisely you specify camera, beats, audio, and style, the closer the full model lands. Save your most detailed prompt for the final render.

  2. Preview on Draft, finish on full. Test compositions, durations, and aspect ratios on the Draft twin, then run the winning prompt on the full model for quality.

  3. Always describe the audio. Put dialogue in quotes and name the sound effects and score. Silence in the prompt means a weaker soundtrack.

  4. Structure motion in timed beats. Break the action into "Beat one (0 to 4s)" style segments so the model paces the shot instead of guessing.

  5. Vary duration and aspect ratio. Mix lengths from 8 up to 20 seconds and use the wide and vertical ratios. Everything at one length and one shape looks like a template.

  6. Move the camera and fill the frame. A tracking or craning camera with real action and more than one character reads as current, a locked off single subject reads as dated.

  7. For First and Last Frame and Keyframes, write continuity into the prompt. Describe one continuous action that resolves exactly on the final frame, and keep the supplied frames compositionally consistent so the motion does not morph.


Known Limitations

  • Draft quality is lower and less consistent. Draft models are a preview tier: lower resolution, more variation between runs, no resolution control. Re render the final on the full model.

  • Keep depicted action implied, not graphic. The content filter can reject explicit violence wording (for example a gun being fired). Frame tense or implied action instead, which passes and still reads well.

  • Source stills go through their own filter. When you generate the input image, some subjects and poses get blocked upstream. If an image is consistently rejected, adjust the subject or wording.

  • Keyframes want consistent frames. If the supplied images differ a lot in scale or position, the motion between them can warp. Derive them from one base image to keep composition stable.

  • No fixed seed control. These models do not expose a seed input, so an exact reproduction of a given take is not guaranteed from the same prompt. Treat each run as a fresh take.

  • Extend needs a short source. The input clip should be under roughly 15 seconds. Extend the tail of a shot rather than feeding a long clip.