Text-to-music is generative AI that turns a written description into a finished piece of music. You describe the genre, mood, tempo, and instruments — and optionally hand over lyrics — and the model composes, arranges, performs, and mixes the track in one pass. What comes back is not MIDI or a loop pack but a rendered audio file you can play, download, and drop into a video.
How it works
Music models are trained on huge libraries of audio paired with descriptions, so they learn what "warm lo-fi hip hop with a dusty Rhodes piano" actually sounds like. At generation time the model works on a compressed representation of audio rather than raw waveforms, which lets it hold a whole song's structure in mind — verse, chorus, bridge — while rendering instruments and voice together. Models that support lyrics align the vocal line to your words, so the singer sings what you wrote rather than mumbling syllables.
What it's best at
- Soundtracks for video — background music that matches the pace and mood of a clip, without digging through a stock library.
- Full songs with vocals — write the lyrics, name the style, and get a sung track with verses and a chorus.
- Instrumental beds — ambient, lo-fi, cinematic, or electronic loops for podcasts, ads, and games.
- Rapid style exploration — audition ten directions for a brand jingle in the time it used to take to brief one composer.
What a good input includes
Describe the music the way a producer would brief a session: genre and era ("90s trip-hop"), mood ("melancholic but hopeful"), tempo and energy ("slow, around 80 BPM, building to the chorus"), instrumentation ("acoustic guitar, brushed drums, warm bass"), and vocals ("female alto, breathy, close-mic'd") — or "instrumental" if you want none. When you supply lyrics, mark the structure with tags like [Verse], [Chorus], and [Bridge]; the model uses them to shape the arrangement.