Sonilo V1.1 Sound Effects & Music: The Essentials

Last updated: July 31, 2026

The Sonilo V1.1 suite gives every project a soundtrack. Generate royalty-free sound effects from a text prompt or straight from a video, or score any clip with an original, commercially licensed piece of music. Everything the suite outputs is cleared for commercial use.

This family has four members. Text to Sound Effects creates a sound effect from a written description. Video to Sound Effects watches a video and returns an audio track of effects timed to the action. Video to Video Sound Effects does the same but hands back the finished video with the effects already mixed in. Video to Video Music composes a soundtrack that matches a video and mixes it back into the clip, optionally keeping the original speech.


Which Model Should I Use?

Model

ID

Input

Best for

Text to Sound Effects

model_sonilo-v1-1-text-to-sound-effects

Text prompt

A specific effect on demand: foley, UI, trailer hits, ambience

Video to Sound Effects

model_sonilo-v1-1-video-to-sound-effects

Video

Effects timed to a video, returned as a separate audio track

Video to Video Sound Effects

model_sonilo-v1-1-video-to-video-sound-effects

Video

A finished, ready-to-share clip with effects already mixed in

Video to Video Music

model_sonilo-v1-1-video-to-video-music

Video

An original soundtrack scored to a video, mixed back in

Rule of thumb: start from text when you need one precise sound, and from a video when you want audio that follows the picture. Pick the audio-only Video to Sound Effects when you will mix the track yourself, and the Video to Video variants when you want the finished clip back. Video to Video Music is the current music model in the family and replaces the earlier Video to Video model.


How to Use the Models

Text to Sound Effects

Describe the sound you want and set a length from 1 to 180 seconds. The more concrete the description, the tighter the result: name the object, the action, and how it evolves over the clip.

Heavy rain pouring on a tin roof, distant rolling thunder building then fading, occasional water drips

Text to Sound Effects, 8 seconds · Open in Scenario

It handles designed and stylized effects just as well as real-world foley.

A massive dragon roars deeply, then unleashes a powerful burst of roaring flames

Text to Sound Effects, 6 seconds · Open in Scenario

Video to Sound Effects

Feed a video and the model generates effects timed to what happens on screen, returned as a standalone audio track. Leave the prompt empty to cover every scene automatically, or add a prompt to steer the palette. Below, a prompt-steered pass on a silent clip, muxed back onto the source so you can hear the timing.

a sports car engine roaring, tires screeching and skidding, gravel spraying, wind rushing past

Generated effects muxed back onto the source clip · Open in Scenario

For frame-accurate control you can split the video into contiguous segments, each with its own description, so a single clip can move from a quiet idle to a hard rev exactly when the picture does.

Video to Video Sound Effects

The same engine as Video to Sound Effects, except it returns the finished video with the effects already mixed in (plus the separate audio track if you want it). Great for quick foley on silent AI video.

Prompt-steered forge foley on a silent clip · Open in Scenario

A prompt like an ASMR direction pushes it toward intimate, close-up detail: soft breathing, cloth, and small movements rather than big hits.

ASMR-style breathing and blanket rustle · Open in Scenario

Video to Video Music

Video to Video Music reads a video's pacing, motion, and mood and composes a soundtrack that fits, then mixes it back into the clip. Add a prompt to steer genre and instruments, or let it score straight from the footage.

Prompt: sweeping cinematic orchestral, warm strings and horns · Open in Scenario

The same footage takes a completely different character with a different prompt.

Prompt: dark retro synthwave, driving analog bass, 80s neon · Open in Scenario

Turn on Keep Speech to preserve the original speech or vocals and blend them over the new music with automatic ducking, so a talking clip keeps its voice while gaining a score.

Keep Speech on: her voice is kept over a new lo-fi beat · Open in Scenario


Parameters

The four models share a small, consistent set of controls. Only the ones relevant to a given model appear in its form.

prompt

On Text to Sound Effects the prompt is required and describes the sound. On the three video models it is optional: leave it empty for automatic coverage, or use it to steer the effects palette or the music style, as in the orchestral versus synthwave pair above.

video

Required on the three video models. The generated audio matches the video length. Any silent clip works well, since the models read the picture, not the source audio.

duration

Text to Sound Effects only. The clip length in seconds, from 1 to 180. The examples above range from 6 to 12 seconds. On Video to Video Music, an optional duration instead sets the length of the segment to score.

segments

Optional on both Video to Sound Effects models. A list of contiguous time ranges, each with its own description. The first segment must start at 0 and ranges must be back to back. Leave it empty to split into scenes automatically.

keepSpeechVocal (Keep Speech)

Video to Video Music only, false by default. When on, the original speech or vocals are isolated and blended over the generated music with ducking. When off, the original soundtrack is replaced entirely.

numSamples (Samples)

Video to Video Music only, 1 to 3. Generates that many scored variations from the same source so you can pick the best fit.

startOffset and duration

Video to Video Music only. Score just part of a video: start at Start Offset and run for Duration seconds. Start offset plus duration must stay within the video length.

audioFormat

AAC, MP3, WAV, or FLAC for the returned audio file. On Video to Video Sound Effects this controls the separate audio track; the muxed video is always AAC.


Use Cases

  • Games: generate foley, UI blips, and ability sounds from text, or drop timed effects onto captured gameplay and cutscenes.

  • Film and trailers: braaam hits and risers from text, ambient beds for a scene, or a full orchestral score matched to an edit.

  • Ads and social: add a premium soundtrack and product-reveal effects to a silent product clip, ready to post.

  • Content creation: score a vlog or talking-head video while keeping the creator's voice with Keep Speech.

  • E-commerce: give catalog and demo videos a consistent, commercially cleared audio identity at scale.


Tips for Better Results

  1. Be concrete in text prompts. Name the object, the action, and how the sound changes over time, as in the rain and dragon examples.

  2. Start with automatic mode on video. Leaving the prompt empty gives solid scene-by-scene coverage; add a prompt only to redirect the palette.

  3. Use segments for timing. When a clip has distinct beats, contiguous segments starting at 0 let each range get its own effects.

  4. Silent sources are fine. The video models read the picture, so a clean silent clip is an ideal input for both effects and music.

  5. Ask for ASMR when you want intimacy. An ASMR direction pushes Video to Video Sound Effects toward soft breathing and close detail instead of big impacts.

  6. Reach for Keep Speech on talking clips. It preserves dialogue or vocals under the new music, which is ideal for vlogs and presenters.

  7. Generate a few music variations. Set Samples to 3 and pick the take that best matches the cut.


Known Limitations

  • Keep Speech preserves speech and vocals, not generic effects. In testing, a dialogue clip kept its voice cleanly, while a source whose only audio was gunfire and helicopters had that audio replaced by the score. Use it for talking or singing clips; for other sounds, expect the music to take over.

  • Video to Sound Effects returns audio only. To share a finished clip, use Video to Video Sound Effects, or mix the returned track back onto your video.

  • Segments must be contiguous. The first must start at 0 and each range must butt up against the next; gaps or overlaps are not allowed.

  • Occasional provider stalls. A generation can rarely stall early with no progress; re-running the same request resolves it.

  • Pinned examples are set in the app. The model page example set is curated in the Scenario UI, separate from the examples embedded in this guide.