共有:
Text-to-Music Model

An AI that mastered voice
now starts to make music.

ElevenLabs, known for voice synthesis and voice cloning, has shipped a new music generation model, "Music v2.5." The company says it beat its predecessor across 47,885 blind-tested pairs, and the tracks it generates can be used commercially even on the free plan. Here's what changed, and what it means in practice.

AI Navigate Editorial·2026.09.14·7 min read
VOICE — ELEVEN V3 A model that reads voice as waveform MUSIC — V2.5 A model that generates melody, instruments, and arrangement
01
Why Now

Why a voice-AI company
decided to step into music

ElevenLabs has been known for voice synthesis and cloning through "Eleven v3." Its new music model is a direct extension of that line.

ElevenLabs announced the model on its official blog under the headline "Introducing Music v2.5, our best music model yet." This marks the company's formal move into music generation, an area adjacent to the voice business it has built its name on. Text-to-music has so far been led by Suno and Udio, which grew their user bases while facing copyright lawsuits from major record labels. ElevenLabs enters this space by explicitly stating that Music v2.5 was "trained only on licensed audio."

That positioning matters. Uncertainty around rights clearance has held back commercial adoption of AI-generated music; a lower-legal-risk option could lower the barrier for marketing teams and video production shops. For existing ElevenLabs users, being able to handle voice and music under one account and one credit pool is also a real operational change.

Music v2 (previous)Music v2.5 (new)
Melodies and instrumentation felt somewhat flatRicher melodies, more depth in arrangements
Output could sound mechanicalInstruments closer to a live take
One selectable option in ElevenMusicPromoted to the default model in ElevenMusic
Commercial-use positioning was unclearDesigned with commercial use in mind

The AI that used to shape our voice into words
now reaches to fill the space between notes.


02
The Evidence

The "best model yet" claim rests on
47,885 paired comparisons

ElevenLabs backs its claim with numbers: two candidate tracks from the same prompt, judged by humans who couldn't tell which model made which.

Previous model take BLIND Which model made which take? Hidden. v2.5 take 47,885 pairs judged
FIG. Two takes are generated from the same prompt — one from the previous model, one from v2.5 — then compared blind.
01

Two takes, one prompt

The same prompt produces one track from the previous model and one from Music v2.5.

02

Compared without labels

Reviewers judged 47,885 pairs without knowing which take came from which model.

03

v2.5 wins the majority

Across the pairs, Music v2.5 was preferred the majority of the time. The gap was widest in vocal-led or acoustic-instrument-heavy genres — R&B, soul, hip hop, rock, metal, orchestral, and cinematic.

03
By The Numbers

What the comparison test shows

47,885
blind-tested pairs
7
genres with the widest gap
5/day
lossless downloads on the Free plan
400/mo
lossless downloads on the Pro plan

These figures come from ElevenLabs' official blog post and its pricing page. The exact win rate of the blind test hasn't been published, but the "preferred the majority of the time" result and the genre-level pattern are stated explicitly. Tracks made on the Free plan can be used commercially, provided you credit ElevenMusic.

04
Who It's For

Who benefits, and who won't notice much

Genres that lean on vocals or live-sounding instruments show the biggest gap over the previous model.

Vocal-led (R&B, soul, hip hop)

Where the improvement in texture is said to show most. A good fit for creators who want a track people actually listen to, not just background filler.

Live-instrument-heavy (rock, metal)

Even genres judged on raw instrument feel showed a gap. Useful as a starting point for band-style backing tracks.

Large ensembles (orchestral, cinematic)

Good for long-form scores and video soundtracks where depth of arrangement matters. If you just need a short jingle or a simple loop, the difference will be harder to notice.

For agencies and video production shops already using ElevenLabs for voice, the ability to handle voice and music under one account and one credit pool is a real advantage. For individuals who just want to try music generation once, or users already comfortable with Suno or Udio, this isn't necessarily a decisive reason to switch.


05
What To Watch

What to check, and what to do next

What to do next. Creators weighing commercial use should start with the Free plan's 5-tracks-a-day lossless download allowance and test the vocal-led and live-instrument genres directly. Existing ElevenLabs users considering a combined voice-and-music workflow should budget for integration through ElevenCreative or the API.

The risk side. The "trained only on licensed audio" claim is a clear differentiator given the copyright lawsuits Suno and Udio have faced, but it hasn't been independently verified by a third party. The 47,885-pair test is ElevenLabs' own internal evaluation, and the exact win rate has not been disclosed. The Free plan's cap of 5 lossless downloads a day also means any production at scale requires at least the Pro plan.