ACE-Step 1.5: The Essentials

Last updated: July 21, 2026

ACE-Step 1.5 is an open-source music suite that spans the whole life of a track, with two sides that work together. The CREATE and TRANSFORM side writes and reworks music: Text to Music builds a finished song from a prompt, Cover restyles an existing track, and Repaint rebuilds one section in place. The EDIT side works on audio you already have: Stem Extract isolates a single part, Complete Track builds a full arrangement around a bare recording, and Add Layer mixes one new instrument on top.


What It Does

ACE-Step 1.5 covers the full loop of making music with AI, from a blank page to a polished edit.

  • Text to Music builds a finished song from a one-line style prompt and an optional lyric sheet. You script the structure with tags like [Verse] and [Chorus], set BPM and key, choose from more than 50 vocal languages, and render tracks from 10 seconds up to 10 minutes.

  • Cover re-renders an existing track in a new style while keeping its structure recognizable. A single Cover strength control decides how much of the original survives, from a faithful re-recording to a loose reinterpretation that only borrows the mood.

  • Repaint rebuilds one time window of a track and leaves everything outside it untouched. You set a start and end in seconds, describe the new section, and the total duration is preserved exactly, so you can swap a chorus, rework an intro, or fix a flubbed line in place.

  • Stem Extract pulls a single instrument or vocal out of a finished mix and returns it playing on its own, chosen from 12 instrument classes.

  • Complete Track takes a bare vocal or a solo instrument and builds a full backing arrangement around it, locked to the key, tempo, and feel of your source.

  • Add Layer takes an existing track and mixes in one new instrument, optionally only within a start and end window you set.

The three generation functions each ship in two quality tiers. Quality uses the full-size checkpoint with the 4B planner and is tuned for the highest fidelity and prompt adherence. Turbo is a distilled, faster variant that costs about half as much, which makes it ideal for iterating on ideas before committing to a final Quality render. The three editing tools run on the full-quality ACE-Step 1.5 edit engine.


Which Model Should I Use?

Pick a generation model to create or transform a song, and an editing model to work on audio you already have. Within each generation function, Turbo runs at about half the cost of Quality, so a common flow is to draft on Turbo, lock in the settings, then rerun the same job on the matching Quality tier for the take you keep.

Purpose

Model

Best for

Generation

Text to Music, Quality

Final songs written from a prompt and lyrics, at the highest fidelity

Generation

Text to Music, Turbo

Fast song drafts and idea exploration at about half the cost

Generation

Cover, Quality

Final restyles of an existing track, faithful to its structure

Generation

Cover, Turbo

Quick style tests before a final cover, at about half the cost

Generation

Repaint, Quality

Final section regenerations that must match the rest of the track

Generation

Repaint, Turbo

Fast passes over a section before a final repaint, at about half the cost

Editing

Stem Extract

Isolating one part from a finished mix for remixing, sampling, karaoke, and backing tracks

Editing

Complete Track

Turning an a cappella or a riff into a fully produced song

Editing

Add Layer

Building an arrangement one instrument at a time on an existing track

Quick rule for the editing side: a full mix and you want one part out, use Stem Extract. Just a vocal or a single instrument and you want a band around it, use Complete Track. A track that is nearly there and you want to add one element, use Add Layer.


Parameters

Within each generation function, the Quality and Turbo tiers share the same controls, so the settings below apply to both tiers of that function.

Text to Music

  • Prompt: a one-line description of the sound you want, covering genre, mood, instruments, and production. Either a prompt or lyrics is required.

  • Lyrics: the words to sing, with optional section tags such as [Verse], [Chorus], and [Bridge]. Put backing vocals in (parentheses), use UPPERCASE for belted lines, and use [Instrumental] for passages without vocals.

  • Instrumental: creates a track with no vocals.

  • Vocal language: the language for the vocals, or Auto to detect it from the lyrics. More than 50 languages are supported.

  • Duration: track length in seconds, from 10 to 600. Leave it empty to let the model decide. Longer tracks cost more.

  • BPM: tempo in beats per minute, from 30 to 300.

  • Key: the musical key, such as C Major or Am.

  • Audio count: how many variations to generate, from 1 to 4.

  • Thinking: lets the model plan the track before rendering, which improves prompt adherence. On by default.

  • Seed: a number that makes the first result repeatable.

Cover

  • Source audio (required): the song you want to cover or restyle.

  • Reference audio (optional): a track whose genre, mood, and feel guide the result.

  • Style prompt: the target sound, for example acoustic folk, warm and mellow, female vocals.

  • Lyrics (optional): keep the original words or rewrite them, with section tags for structure.

  • Cover strength: how closely the result follows the original, from 0 to 1. High values stay faithful to the arrangement, 0.5 to 0.6 transforms the genre while the hook survives, and around 0.2 only borrows the mood.

  • Instrumental, Vocal language, Audio count, Thinking, and Seed work as in Text to Music.

Repaint

  • Source audio (required): the track that contains the section to regenerate.

  • Prompt: describes the music for the new section, covering genre, mood, instruments, and tempo.

  • Lyrics (optional): lyrics for the new section. To change sung lyrics in the window, start the prompt with "Repaint the selected section with new sung lyrics:", pass the full lyric sheet with the new section in place, and turn Thinking off.

  • Start and End (seconds): the region to rebuild. Set End to -1 to regenerate through to the end of the track. The total duration is always preserved.

  • Output format: the file type you get back, MP3, WAV, or FLAC.

  • Instrumental, Vocal language, Audio count, Thinking, and Seed work as in Text to Music.

Stem Extract

  • Source audio (required): the finished, mixed track you are pulling a part out of.

  • Track name: the single stem to isolate, chosen from 12 instrument classes: vocals, backing vocals, drums, bass, guitar, keyboard, percussion, strings, synth, fx, brass, or woodwinds.

  • Prompt (optional, short): hints at the character of the stem, for example an energetic female pop lead vocal.

  • Number of outputs: how many variations to generate, from 1 to 4.

  • Guidance scale: how strictly the result follows the prompt, from 1 to 15 (default 7). Moderate values are the safe starting point.

  • Seed: a number that makes the first result repeatable. Leave it empty for a fresh result each run.

Complete Track

  • Source audio (required): a partial recording, such as a bare vocal or a solo instrument.

  • Complete track classes: the set of stems to generate around your source, for example drums, bass, and guitar. Pick the instruments that fit the genre you describe in the prompt.

  • Prompt (required): sets the genre and instrumentation. Add "no added vocals" when you only want instruments.

  • Thinking: lets the model plan the arrangement first for tighter musical coherence. On by default.

  • Number of outputs, Guidance scale, and Seed work as in Stem Extract.

Add Layer

  • Source audio (required): the existing track you are adding to.

  • Track name: the single stem to add, chosen from the 12 instrument classes above.

  • Prompt (required): describes the new layer, for example a lush analog synth pad matched to key and tempo.

  • Repaint start and Repaint end (seconds): where the new layer begins and ends. Leave the end at -1 to run through to the end of the track.

  • Thinking: lets the model plan the layer first for tighter musical coherence. On by default.

  • Number of outputs, Guidance scale, and Seed work as in Stem Extract.


Examples

The clips below show the ACE-Step 1.5 family in action, first the generation functions, then the editing tools. Each is shown as an audio waveform you can play and open directly in Scenario.

Generation examples

Text to Music, Quality

A soft, emotional ballad in C Major at 72 BPM with English vocals, written from a style prompt and a lyric sheet. · Open on Scenario

Text to Music, Turbo

"Paper Boats," an indie piano ballad in C Major at 72 BPM, drafted on the Turbo tier at about half the cost of Quality. · Open on Scenario

Cover, Quality

The same song restyled with French vocals at a Cover strength of 0.7, keeping the original structure recognizable. · Open on Scenario

Repaint, Quality

A window from 84s to 116s regenerated in place as a gentle passage, while the rest of the track and its total length stay untouched. · Open on Scenario

Editing examples

Stem Extract

A finished pop song goes in and the drums come out on their own.

Open this asset in Scenario

And on an orchestral track, the string section is lifted out on its own:

Open this asset in Scenario

Complete Track

A solo a cappella vocal is joined by drums, bass, keys, and strings that the model generated around it, played together as the finished song. Style prompt used: warm cinematic soul arrangement under the vocal, gentle drums, bass, electric piano, and lush strings, no added vocals.

Open this asset in Scenario

A different genre from the same approach: a soft, whispered topline turned into an electronic track, with AI synth, bass, drums, and keys built around it.

Open this asset in Scenario

Add Layer

A ballad is shown first on its own, then again after a new synth layer was added. Layer prompt used: lush analog synth pad supporting the ballad, matched to key and tempo.

Original track

With the new synth pad layer added

Open this asset in Scenario

And on a pop track, an electric guitar layer added on top:

Original track

With the new electric guitar layer added

Open this asset in Scenario


Tips for Better Results

  • Draft on a Turbo tier and finalize on the matching Quality tier. Once the prompt, lyrics, and settings feel right, rerun the same job with the same seed on Quality for the take you keep.

  • Write the style prompt like a short brief: name the genre, mood, instruments, vocal character, and production in one line rather than a long paragraph.

  • Use structure tags in the lyrics. Marking [Verse], [Chorus], and [Bridge] gives the model a clear map and produces cleaner song forms.

  • Set BPM and key explicitly when you have a target in mind, especially if the track needs to sit alongside other material at a fixed tempo.

  • For Cover, treat Cover strength as your main dial. Start high to stay close to the original, drop to 0.5 to 0.6 to change genre while keeping the hook, and go near 0.2 to only borrow the mood.

  • For Repaint, extend the region a little past the exact bars you want to change so the new section blends smoothly at its edges. To rewrite sung words, begin the prompt with "Repaint the selected section with new sung lyrics:", supply the full lyric sheet with the new lines in place, and turn Thinking off.

  • For Complete Track, give it a source with clear melody and rhythm. A sung a cappella or a played riff works far better than a flat, sustained pad or a very short clip.

  • Generate a clean a cappella directly with a music model rather than isolating a vocal from a full mix. A natively generated vocal is cleaner and gives the model more to work with.

  • Add "no added vocals" to a Complete Track or Add Layer prompt when you only want instrumentation, so the model does not sing new parts.

  • Match the Complete track classes to the genre you describe. A country prompt with guitar, bass, drums, and piano lands better than a generic instrument set.

  • With Add Layer, use the start and end times to place a solo or a swell only in the section that needs it, rather than across the whole track.

  • Render 2 to 4 variations and audition them. Takes vary, and the best one is quick to spot.

  • Keep the Guidance scale near the default and adjust in small steps. Pushed-up values can destabilize the result.

  • Keep Thinking on for most generations. It lets the model plan first and follow the prompt more closely; turn it off mainly for the lyric-rewrite Repaint workflow above.


Known Limitations

  • Longer tracks and higher audio counts cost more, and the maximum length is 10 minutes (600 seconds).

  • Turbo trades some fidelity and prompt adherence for speed and lower cost, so fine detail can differ from a Quality render of the same job.

  • Cover, Repaint, and all three editing tools need a source track as input, so they cannot start from a blank page the way Text to Music can.

  • Repaint preserves the total duration of the track, so it cannot make a section longer or shorter, only regenerate what is inside the window.

  • Complete Track needs a musical source. A very short clip, a spoken-word line with little melody, or a sustained pad tends to produce a thin result. Feed it a clear sung or played phrase.

  • Vocal isolation quality varies. Stem Extract is strong on prominent, rhythmic parts like drums and bass, while isolating a vocal from a dense mix can be less clean, so judge the result by ear.

  • Stem Extract returns a single stem per run, and Add Layer adds a single part per run. Chain runs to build up several.

  • Vocal clarity and pronunciation vary by language and by how dense the lyrics are; very fast or crowded lines can come out less intelligible.

  • The web interface exposes a vocal language selector and audio format choices (MP3, WAV, FLAC) for the editing tools that are not currently available through the API.

  • Results are probabilistic. Reuse a seed to reproduce a first result, but expect variation across runs and across the two tiers.


Use Cases

  • Original songs and jingles for games, videos, ads, and social content, written from a prompt and a lyric sheet.

  • Instrumental beds, loops, and background scores generated vocal-free with the Instrumental option.

  • Genre restyles of a demo or reference track with Cover, from a faithful re-recording to a loose reinterpretation.

  • Localized versions of a song by rewriting the lyrics and switching the vocal language.

  • Targeted fixes and revisions with Repaint, such as reworking an intro, swapping a chorus, replacing a solo, or correcting a single line.

  • Remix and sampling by pulling a clean drum or bass stem out of a track to rework or sample it.

  • Karaoke and backing tracks by isolating or removing the vocal to sing or play over the rest.

  • Songwriting from a voice memo: record a melody or a cappella idea and let Complete Track build the band around it.

  • Scoring a vocal or narration by turning a topline or spoken line into a produced piece with Complete Track.

  • Arranging in passes with Add Layer, bringing in one instrument at a time while keeping control over the mix.

  • Fast iteration on musical ideas with the Turbo tiers before committing to a final Quality render.