Veed: The Essentials

Last updated: July 21, 2026

Veed Lipsync v2 by Veed IO re-syncs a talking video to a new audio track: feed it a clip of someone speaking plus a new voice recording, and it reshapes the mouth so the person appears to say the new words. Swap a script, recast a voice, or localize a clip into another language while keeping the original performance and framing. The Veed family also includes Veed Fabric 1.0, which animates a still image into a talking video, covered near the end.


Which Model Should I Use?

Both Veed models create talking video; pick by what you start from, an existing video clip or a single still image. This article focuses on Veed Lipsync v2; Veed Fabric 1.0 is summarized at the end.

Model

ID

Best for

Veed Lipsync v2

model_veed-lipsync-v2

Re-syncing an existing talking video to a new voice, script, or language, keeping the original framing.

Veed Fabric 1.0

model_veed-fabric-1-0

Turning a still image (photo, illustration, or character) into a talking video from audio or text.


What It Does

  • Re-syncs the mouth in an existing video to a new voice track, so the speaker appears to say new words.

  • Lets you swap a line of dialogue, recast a voice, or localize a clip into another language without reshooting.

  • Works on photoreal people and stylized characters alike, from anime faces to robots.

  • Keeps the source resolution and preserves the original performance, lighting, and framing.

  • Pairs naturally with voice tools: generate the new track with Seed Audio 1.0 or ElevenLabs Dubbing, or pull the original voice with Audio Extract.


Parameters

Veed Lipsync v2

The model takes two inputs and needs no other settings:

  • Video: the clip of a person or character talking that you want to re-sync. Their mouth is reshaped to match the new audio. This input drives the cost.

  • Audio: the voice or speech track you want the subject to appear to say. Lip movement is matched to it.

The output keeps the source resolution and runs for the shorter of the two durations, so trim the video and audio to the length you want in the result.


Examples

Widescreen talking-head clip re-synced to a new line of dialogue, with no reshoot needed · Open on Scenario

Vertical, social-ready clip re-voiced while the original framing stays intact · Open on Scenario

Square-format performance re-synced to a fresh audio track · Open on Scenario

Portrait close-up with lip movement matched to a swapped voice · Open on Scenario


Tips for Better Results

  1. Start with a clear, front-facing shot. The mouth region should be visible and unobstructed for the cleanest re-sync.

  2. Match the durations. Output runs for the shorter of the video and audio, so trim both to the same length to avoid a clip that cuts off early.

  3. Use clean speech audio. A dry voice track with little background noise or music gives the model the clearest signal to sync against.

  4. Keep the new line close in pacing to the original delivery. Audio that roughly matches the original rhythm sits more naturally on the face.

  5. Generate the new voice with a companion model. Seed Audio 1.0 or ElevenLabs Dubbing can produce the track, and Audio Extract can lift the original voice when you only want to tweak a word.

  6. For localization, keep one language per pass. Feed a single translated track per run rather than mixing languages in one clip.


Known Limitations

  • Only the mouth is re-synced. Head movement, expression, and gestures come from the original video and are not regenerated.

  • Output length is capped at the shorter of the video and audio durations, so a longer track is clipped to the video length.

  • Heavily occluded, extreme-profile, or very low-resolution mouths give the model less to work with and can reduce sync quality.

  • It re-syncs an existing performance; it does not generate a talking video from a still image or from scratch.


Also in the Veed Family: Veed Fabric 1.0

Veed Fabric 1.0 is the other Veed model on Scenario. Where Lipsync v2 works from an existing video, Fabric works from a still image: give it a photo, illustration, or character render plus an audio file or a script, and it generates a talking video, animating lip-sync, facial expressions, and head motion while preserving the source style. It outputs 480p or 720p in 16:9, 1:1, or 9:16.

When to use which: reach for Veed Lipsync v2 when you already have footage of someone speaking and want to change what they say; reach for Veed Fabric 1.0 when you only have a still image and want to bring it to life. Full details are on the Veed Fabric 1.0 model page.

Video generated from an image and an audio track with narration - Open in Scenario


Use Cases

  • Games: fix or re-record character dialogue and localize cutscenes without re-animating faces.

  • Film and video: patch a flubbed line, recast a voice, or dub a scene while keeping the original take.

  • Marketing and social: localize an ad or spokesperson clip into multiple languages from a single shoot.

  • Product and training: update the script in explainer or onboarding videos without filming again.