MiniMax Music: The Essentials

Last updated: August 14, 2026

MiniMax Music banner

Covers MiniMax Music 3.0 (model_minimax-music-3-0), the current recommended version. Earlier releases (2.6, 2.5, 2.0) are now legacy.

MiniMax Music 3.0 writes, arranges, performs and mixes a complete song in one generation. Give it your lyrics and a style prompt for full vocal tracks, or switch to instrumental mode for score and beats from the prompt alone. Genre, mood, tempo, key, vocal timbre and instrumentation all live in the style prompt, and it sings in many languages, up to about five minutes at 44.1kHz.

Hear it first

Open this track in Scenario

A full festival pop-electronic anthem generated from a written lyric and a single style prompt (example asset_HuNPagt1q583TbtHy2AZvLWw).


Which Model Should I Use?

Model

ID

Status

Best for

MiniMax Music 3.0

model_minimax-music-3-0

Recommended

New work: fullest arrangements, strongest vocals, instrumental mode, multi-language.

MiniMax Music 2.6

model_minimax-music-2-6

Legacy

Existing 2.6 workflows you do not want to change.

MiniMax Music 2.5

model_minimax-music-2-5

Legacy

Older projects that depend on 2.5 phrasing.

MiniMax Music 2.0

model_minimax-music-2-0

Legacy

Superseded; kept for reproducibility only.

Rule of thumb: start on 3.0 for anything new. Reach for a legacy version only to reproduce a track you already made on it.


How to Use the Model

Write a Song from Your Own Lyrics

Put the words you want sung in the Lyrics field and shape the sound with the style prompt. Wrap sections in bracket tags such as [Intro], [Verse], [Pre-Chorus], [Chorus], [Bridge] and [Outro]. The tags steer the arrangement and are never sung out loud. Use a single newline between lines and a blank line between sections.

Style prompt: Classic late-night soul ballad with lo-fi warmth, around 72 BPM in D minor. Tender, intimate verses open into a softly radiant chorus, led by a smoky female vocal with restrained, breathy phrasing and gentle vibrato, supported by mellow Rhodes piano, brushed drums, warm upright bass, subtle string pads, and a vintage room-like mix with tape saturation.

Lyrics (excerpt): [Verse] Streetlights blur against the rain / I trace your name on the window pane ... [Chorus] And I still hold the light you gave / A quiet fire the years cannot save

Open this track in Scenario

The vocal sings exactly the written lyric; the bracket tags only shape the structure.

Style prompt: Modern melodic trap and dark R&B, ~140 BPM half-time feel in F minor. Confident, hungry rap verses trade with a melodic sung hook, male lead vocal with light autotune, deep bouncy 808 sub-bass, fast rolling triplet hi-hats, dark ambient piano, subtle vinyl crackle, wide moody late-night mix.

Open this track in Scenario

Rap verses and a sung [Hook] from one lyric sheet, showing how a genre shift is driven entirely by the style prompt.

Instrumental Mode

Turn on Instrumental to generate music with no vocals. The Lyrics field is ignored, so the style prompt carries the whole idea: instruments, arc, tempo and mix.

Style prompt: Epic orchestral hybrid trailer score for a fantasy adventure game, roughly 90 BPM in E minor. A quiet, mysterious intro of solo piano and airy choir pads builds through pizzicato strings, ticking clock percussion, and low taiko hits into a soaring, heroic full-orchestra climax with brass ostinatos, thunderous cinematic percussion, sweeping legato strings, and a triumphant mixed choir, then resolves to a gentle, hopeful piano outro.

Open this track in Scenario

A cinematic trailer cue built in one pass with Instrumental on (example asset_49x4eWpFoNsQW7iwGmG3AGEs).

Style prompt: Chilled lo-fi hip-hop beat for studying, ~78 BPM in C major. Warm dusty Rhodes chords, mellow jazzy guitar licks, soft boom-bap drums with brushed snares, deep round upright bass, gentle vinyl crackle and light rain ambience, a rounded low-pass warmth, relaxed and hopeful. Fully instrumental.

Open this track in Scenario

A loopable lo-fi study bed, useful as background music under video or streams.

Singing in Other Languages

Write the lyric in the target language and name the language in the style prompt. Here the whole song is performed in Spanish.

Style prompt: Upbeat Latin pop / reggaeton, ~94 BPM in B minor. Warm male lead vocal with catchy melodic hooks and background ad-libs, dembow rhythm, deep bass, bright plucky synths, acoustic guitar accents, claps and percussion, summery danceable radio mix. Spanish-language vocals.

Lyrics (excerpt): [Coro] Baila, baila, que el mundo se queda atras / Dame la mano y no mires jamas ... Tu eres mi verano, tu eres mi sol

Open this track in Scenario

Big Vocal Arrangements

Ask for lead and group or choir vocals in the style prompt and the model will stack them, with call-and-response and harmonies on the chorus.

Style prompt: Uplifting modern gospel, ~72 BPM in A-flat major. A soulful lead vocal answered by a full mass choir, gospel organ, warm piano, driving tambourine and hand claps, thick bass, building from a tender verse to a euphoric clapping climax, joyful and radiant.

Open this track in Scenario

Vintage and Lo-Fi Character

Lower the sample rate and bitrate to lean into an old, degraded aesthetic. This track was rendered at 16kHz / 32kbps for a deliberate antique-phonograph feel, so the vocal is intentionally muffled and noisy.

Style prompt: Vintage 1920s gramophone jazz, ~86 BPM swing in C major, deliberately old and lo-fi. Crackly, band-pass-filtered mono-style recording, warbling female crooner vocal with a carbon-microphone character, tinny banjo, brushed snare, tuba bass, muted cornet, heavy vinyl noise and wow-and-flutter, as if played on an antique phonograph.

Open this track in Scenario

Sample rate 16000 and bitrate 32000 (example asset_TNaYW4f8BeXo1KFXxAAKRko1). Keep both at their defaults for clean, modern fidelity.


Parameters

Two fields do the heavy lifting: the style prompt (always required) and the lyrics. The rest are simple toggles and quality settings.

prompt

Required, up to 2,000 characters. Describe genre, mood, tempo (BPM), key, vocal timbre and instrumentation as flowing prose, as in every example above. In Instrumental mode this is the only instruction the model has, so be specific about the arrangement and its arc.

lyrics

Up to 3,500 characters. Contains only the words you want sung, split into sections with bracket tags. Required unless Instrumental is on. Longer lyrics produce longer songs, up to about five minutes. See the soul, trap and Spanish examples above.

isInstrumental

true / false, default false. When true the model ignores the lyrics and produces a vocal-free track from the style prompt, as in the trailer score and lo-fi examples.

lyricsOptimizer

true / false, default false, labeled Auto lyrics in the UI. Intended to write lyrics for you from a brief. In practice on Scenario it tends to sing whatever text sits in the Lyrics field rather than rewriting it into a clean song, so for reliable results write your own lyrics and leave this off. See Known Limitations.

sampleRate

16000, 24000, 32000 or 44100 Hz, default 44100. Keep it at 44100 for full-quality music; drop it only for a deliberate lo-fi effect, as in the vintage example.

bitrate

32000, 64000, 128000 or 256000 bps, default 256000. Higher is cleaner; lower adds compression grit. Pair a low bitrate with a low sample rate for a vintage character.


Use Cases

  • Games: generate a heroic trailer cue or a loopable menu theme in Instrumental mode, tuned to your scene tempo and key.

  • Marketing: spin up a custom brand anthem or an upbeat ad bed with your own tagline sung in the hook.

  • Film and video: draft temp music and needle-drop songs in any genre before committing to a licensed track.

  • Social content: make short, catchy, on-theme songs (including other languages) as scroll-stopping audio for reels and shorts.

  • Education: turn a lesson into a memorable sing-along, like a counting and animal-sounds children track.

  • Prototyping: give a demo or pitch a finished-sounding theme song in minutes instead of commissioning a scratch track.


Tips for Better Results

  1. Put only sung words in Lyrics. Instructions, section descriptions or a brief will be sung out loud. Keep direction in the style prompt.

  2. Stack the style prompt in order: genre, tempo (BPM), key, mood and emotional arc, vocal timbre, instrument list, then mix character. The model follows that structure well.

  3. Use bracket tags to arrange. [Intro], [Verse], [Pre-Chorus], [Chorus], [Bridge] and [Outro] shape dynamics and are never vocalized.

  4. Write more lyrics for a longer song. Length scales with the lyric, so add verses and a bridge to approach the roughly five-minute ceiling.

  5. Name the language for non-English vocals and write the lyric in that language, as in the Spanish example.

  6. Ask for the vocal you want: specify male or female, a choir, group harmonies or light autotune, and the model will deliver it.

  7. Keep sample rate and bitrate at their defaults for clean music; lower them only when you deliberately want a lo-fi or vintage sound.


Known Limitations

  • Auto lyrics does not reliably rewrite a brief. With lyricsOptimizer on, the model tends to sing the text in the Lyrics field verbatim rather than turning a brief into a finished song. Workaround: write real lyrics and leave Auto lyrics off.

  • Song length is set by the lyric, not a duration control. There is no explicit length field; short lyrics yield short songs. Add sections to make a track longer.

  • Low sample rate reduces vocal clarity. 16000 or 24000 Hz muffles the voice (useful for vintage effects but not for clean vocals). Stay at 44100 when intelligibility matters.

  • Occasional small artifacts at song ends. Tracks can add a brief ad-lib or tail on the outro. Trim the export if you need a hard ending.

  • Structure follows the prompt loosely. Bracket tags and the described arc guide the arrangement but are not strict; regenerate or adjust wording if a section is missing.