Model index · Music & Sound · updated 2026-08-16
Music & Sound models, priced and ranked
Every Music & Sound model we know of — 108 from 21 companies, 8 of them with a published price per second of video.
Every price on this page is per second of video. 9 priced Music & Sound models are sold in a different unit and are left out of the rankings rather than converted with a guess — they are still in the full list at the bottom, and on their own pages.
Most widely-used Music & Sound models and their cheapest credible alternative
One model per company, ordered by how prominent the company is and how many independent providers serve the model — hosts only carry what customers ask for, so that is measured demand. It is not traffic or usage data, which we do not have. A matched alternative scored within 10 points on our own probe suite; unverified means it is cheaper but we have not measured one or both sides, so no quality claim is made.
| Model | $/second of video | Quality | Hosts | Best cheaper alternative | Saving |
|---|---|---|---|---|---|
| V1.1 Text to Sound Effects Sonilo · 2026-07-21 | $0.0018 fal | not measured | 1 | nothing cheaper doing the same job | — |
| ACE Step Audio Inpaint Unattributed · 2025-05-11 | $0.0002 fal | not measured | 1 | nothing cheaper doing the same job | — |
None of these alternatives are quality-matched yet — quality is measured on request, so the savings shown are price-only. Each alternative does the same job as the model beside it; that is a hard filter, not a ranking hint.
Cheapest Music & Sound models right now
Per second of video, cheapest host per model.
| Model | Company | $/second of video | Cheapest host | What it does |
|---|---|---|---|---|
| ACE Step | Unattributed | $0.0002 | fal | music generation |
| ACE Step Prompt To Audio | Unattributed | $0.0002 | fal | music generation |
| ACE Step Audio To Audio | Unattributed | $0.0002 | fal | music generation |
| ACE Step Audio Inpaint | Unattributed | $0.0002 | fal | audio editing |
| ACE Step Audio Outpaint | Unattributed | $0.0002 | fal | audio editing |
| DiffRhythm: Lyrics to Song | Unattributed | $0.001 | fal | music generation |
| V1.1 Text to Sound Effects | Sonilo | $0.0018 | fal | sound effects |
| Sonilo V1.1 Text to Music | Sonilo | $0.0025 | fal | music generation |
* OpenRouter chooses a host for each request, and the figure above is what its default choice costs — not a fixed rate for this model. Selecting a specific host through the same account costs whatever that host charges, which is why this row can sit above or below the hosts listed beside it. Token rates are passed through from the underlying provider with no markup, so the per-1M price above is what that host charges directly — OpenRouter's own charge sits on funding instead: 5.5% (min $0.80) by Stripe or 5% by Coinbase to buy credits, 5% of the equivalent spend if you bring your own provider key. So budget about 5% above the rate shown unless you already hold credit. Fee schedule.
Newest Music & Sound models
By first-published date. In this market recency is the strongest single signal of relevance, and it is a fact we hold for every model rather than a proxy.
| Model | Company | Released | $/second of video | What it does |
|---|---|---|---|---|
| Qwen Audio 3.0 TTS (Flash) | Alibaba | 2026-07-29 | — | text to speech |
| V1.1 Text to Sound Effects | Sonilo | 2026-07-21 | $0.0018 | sound effects |
| Seed Audio 1.0 | Bytedance | 2026-06-25 | — | sound effects |
| Async Text to Speech Pro V1.0 | Async | 2026-06-18 | — | text to speech |
| Ltx 2.3 Quality (text to audio) | LTX | 2026-06-17 | — | sound effects |
| Ltx 2.3 Quality (text to audio lora) | LTX | 2026-06-17 | — | sound effects |
| Zonos2 Text to Speech | Zyphra | 2026-06-16 | — | text to speech |
| Bytedance Seed Speech Text to Speech | Bytedance | 2026-06-05 | — | text to speech |
| Sonilo V1.1 Text to Music | Sonilo | 2026-06-03 | $0.0025 | music generation |
| Stable Audio 3 (medium text to audio) | Stability AI | 2026-06-03 | — | music generation |
| Stable Audio 3 Small Music Text to Audio | Stability AI | 2026-06-03 | — | music generation |
| Stable Audio 3 Small SFX Text to Audio | Stability AI | 2026-06-03 | — | sound effects |
| Stable Audio 3 (small music base text to audio) | Stability AI | 2026-06-03 | — | music generation |
| Stable Audio 3 Medium Base Text to Audio | Stability AI | 2026-06-03 | — | music generation |
| Stable Audio 3 Medium Audio to Audio | Stability AI | 2026-06-03 | — | — |
| Stable Audio 3 (small music audio to audio) | Stability AI | 2026-06-03 | — | music generation |
| Stable Audio 3 Medium Audio Outpainting | Stability AI | 2026-06-03 | — | audio editing |
| Stable Audio 3 (small sfx audio to audio) | Stability AI | 2026-06-03 | — | — |
| Stable Audio 3 Small Music Base Audio to Audio | Stability AI | 2026-06-03 | — | music generation |
| Stable Audio 3 Medium Audio Inpainting | Stability AI | 2026-06-03 | — | audio editing |
| Stable Audio 3 Medium Base Audio to Audio | Stability AI | 2026-06-03 | — | — |
| Stable Audio 3 Small Music Base Audio Outpainting | Stability AI | 2026-06-03 | — | audio editing |
| Stable Audio 3 Medium Base Audio Inpainting | Stability AI | 2026-06-03 | — | audio editing |
| Stable Audio 3 Medium Base Audio Outpainting | Stability AI | 2026-06-03 | — | audio editing |
What Music & Sound models actually do
A category is not one market. Comparison on each model page is restricted to models doing the same job, because a price difference between two different jobs is arithmetic rather than advice.
| Job | Models |
|---|---|
| text to speech | 49 |
| music generation | 23 |
| sound effects | 18 |
| audio editing | 10 |
| unclassified | 6 |
| voice conversion | 2 |
One API for every Music & Sound model here
VernaOne fronts these providers behind a single interface, so switching model or host is a config change — with automatic fallback if quality or latency regresses.
Try VernaOne free →All 108 Music & Sound models
Everything we hold in this category, newest first, priced or not. A model with no price is a fact about what its provider publishes, not a gap we hide.
| Model | Company | Released | Price | Job |
|---|---|---|---|---|
| Qwen Audio 3.0 TTS (Flash) | Alibaba | 2026-07-29 | no published price | text to speech |
| V1.1 Text to Sound Effects | Sonilo | 2026-07-21 | $0.0018 /second of video | sound effects |
| Seed Audio 1.0 | Bytedance | 2026-06-25 | no published price | sound effects |
| Async Text to Speech Pro V1.0 | Async | 2026-06-18 | no published price | text to speech |
| Ltx 2.3 Quality (text to audio) | LTX | 2026-06-17 | $0.0024075 /image | sound effects |
| Ltx 2.3 Quality (text to audio lora) | LTX | 2026-06-17 | $0.0024075 /image | sound effects |
| Zonos2 Text to Speech | Zyphra | 2026-06-16 | no published price | text to speech |
| Bytedance Seed Speech Text to Speech | Bytedance | 2026-06-05 | no published price | text to speech |
| Sonilo V1.1 Text to Music | Sonilo | 2026-06-03 | $0.0025 /second of video | music generation |
| Stable Audio 3 (medium text to audio) | Stability AI | 2026-06-03 | no published price | music generation |
| Stable Audio 3 Small Music Text to Audio | Stability AI | 2026-06-03 | no published price | music generation |
| Stable Audio 3 Small SFX Text to Audio | Stability AI | 2026-06-03 | no published price | sound effects |
| Stable Audio 3 (small music base text to audio) | Stability AI | 2026-06-03 | no published price | music generation |
| Stable Audio 3 Medium Base Text to Audio | Stability AI | 2026-06-03 | no published price | music generation |
| Stable Audio 3 Medium Audio to Audio | Stability AI | 2026-06-03 | no published price | — |
| Stable Audio 3 (small music audio to audio) | Stability AI | 2026-06-03 | no published price | music generation |
| Stable Audio 3 Medium Audio Outpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 (small sfx audio to audio) | Stability AI | 2026-06-03 | no published price | — |
| Stable Audio 3 Small Music Base Audio to Audio | Stability AI | 2026-06-03 | no published price | music generation |
| Stable Audio 3 Medium Audio Inpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Medium Base Audio to Audio | Stability AI | 2026-06-03 | no published price | — |
| Stable Audio 3 Small Music Base Audio Outpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Medium Base Audio Inpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Medium Base Audio Outpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Small SFX Audio Outpainting | Stability AI | 2026-06-03 | no published price | sound effects |
| Stable Audio 3 Small Music Base Audio Inpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Small SFX Audio Inpainting | Stability AI | 2026-06-03 | no published price | sound effects |
| Stable Audio 3 Small SFX Base Audio to Audio | Stability AI | 2026-06-03 | no published price | sound effects |
| Stable Audio 3 Small SFX Base Audio Outpainting | Stability AI | 2026-06-03 | no published price | sound effects |
| Stable Audio 3 Small Music Audio Inpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Small Music Audio Outpainting | Stability AI | 2026-06-03 | no published price | audio editing |
| Stable Audio 3 Small SFX Base Audio Inpainting | Stability AI | 2026-06-03 | no published price | sound effects |
| Lyria 3 Pro | 2026-05-22 | no published price | music generation | |
| Mirelo SFX1.6 (text to audio) | Mirelo AI | 2026-05-18 | no published price | — |
| Mirelo SFX1.6 (extend audio) | Mirelo AI | 2026-05-18 | no published price | sound effects |
| Mirelo SFX1.6 (inpaint audio) | Mirelo AI | 2026-05-18 | no published price | — |
| Gemini 3.1 Flash Tts | 2026-04-16 | no published price | text to speech | |
| Lyria3 | 2026-04-15 | no published price | music generation | |
| Minimax Music 2.6 | Minimax | 2026-04-11 | no published price | music generation |
| Minimax Music 2.5 | Minimax | 2026-04-11 | no published price | music generation |
| Lyria 3 Pro Preview | 2026-03-30 | no published price | music generation | |
| Lyria 3 Clip Preview | 2026-03-30 | no published price | music generation | |
| Gemini TTS | 2026-03-20 | no published price | text to speech | |
| xAI Text to Speech | xAI | 2026-03-17 | no published price | text to speech |
| Inworld TTS-1.5 Max | Inworld | 2026-03-13 | no published price | text to speech |
| Tada TTS 1B | Hume AI | 2026-03-12 | no published price | text to speech |
| Personaplex (realtime) | Unattributed | 2026-02-20 | no published price | text to speech |
| Personaplex | Unattributed | 2026-02-12 | $0.001 /minute of audio | text to speech |
| MiniMax Speech 2.8 [HD] | Minimax | 2026-02-04 | no published price | text to speech |
| MiniMax Speech 2.8 [Turbo] | Minimax | 2026-02-04 | no published price | text to speech |
| Qwen 3 TTS - Text to Speech [1.7B] | Alibaba | 2026-01-26 | no published price | text to speech |
| Qwen 3 TTS - Clone Voice [1.7B] | Alibaba | 2026-01-26 | no published price | text to speech |
| Qwen 3 TTS - Voice Design [1.7B] | Alibaba | 2026-01-26 | no published price | text to speech |
| Qwen 3 TTS - Text to Speech [0.6B] | Alibaba | 2026-01-26 | no published price | text to speech |
| Qwen 3 TTS - Clone Voice [0.6B] | Alibaba | 2026-01-26 | no published price | text to speech |
| GPT Audio | OpenAI | 2026-01-19 | $4.38 /1M | text to speech |
| GPT Audio Mini | OpenAI | 2026-01-19 | $1.05 /1M | text to speech |
| ElevenLabs Voice Changer | ElevenLabs | 2026-01-14 | no published price | voice conversion |
| Elevenlabs Music | ElevenLabs | 2025-12-22 | no published price | music generation |
| Kling Video Create Voice | Kling | 2025-12-16 | no published price | text to speech |
| Maya (batch) | Unattributed | 2025-12-12 | $0.002 /minute of audio | text to speech |
| Maya (stream) | Unattributed | 2025-12-12 | $0.002 /minute of audio | text to speech |
| Maya1 | Unattributed | 2025-11-15 | $0.002 /minute of audio | text to speech |
| Minimax Music | Minimax | 2025-10-30 | no published price | music generation |
| MiniMax Speech 2.6 [HD] | Minimax | 2025-10-29 | no published price | text to speech |
| MiniMax Speech 2.6 [Turbo] | Minimax | 2025-10-29 | no published price | text to speech |
| Demucs | Meta | 2025-10-27 | no published price | text to speech |
| Index TTS 2.0 | Unattributed | 2025-10-07 | $0.002 /minute of audio | text to speech |
| Kling TTS | Kling | 2025-09-13 | no published price | text to speech |
| MiniMax (Hailuo AI) Music v1.5 | Minimax | 2025-09-11 | no published price | music generation |
| Stable Audio 2.5 (text to audio) | Stability AI | 2025-09-10 | no published price | sound effects |
| Stable Audio 2.5 (audio to audio) | Stability AI | 2025-09-10 | no published price | sound effects |
| Stable Audio 25 | Stability AI | 2025-09-10 | no published price | sound effects |
| Elevenlabs | ElevenLabs | 2025-09-09 | no published price | sound effects |
| Chatterbox (text to speech multilingual) | Resemble AI | 2025-09-04 | no published price | text to speech |
| Elevenlabs Sound Effects V2 | ElevenLabs | 2025-09-02 | no published price | sound effects |
| Elevenlabs Tts Eleven V3 | ElevenLabs | 2025-08-20 | no published price | text to speech |
| Minimax (preview speech 2.5 hd) | Minimax | 2025-08-11 | no published price | text to speech |
| Minimax (preview speech 2.5 turbo) | Minimax | 2025-08-11 | no published price | text to speech |
| MiniMax Voice Design | Minimax | 2025-07-18 | no published price | text to speech |
| Chatterboxhd (text to speech) | Resemble AI | 2025-06-02 | no published price | text to speech |
| Chatterbox (text to speech) | Resemble AI | 2025-06-01 | no published price | text to speech |
| Lyria2 | 2025-05-20 | no published price | music generation | |
| ACE Step Prompt To Audio | Unattributed | 2025-05-11 | $0.0002 /second of video | music generation |
| ACE Step Audio To Audio | Unattributed | 2025-05-11 | $0.0002 /second of video | music generation |
| ACE Step Audio Inpaint | Unattributed | 2025-05-11 | $0.0002 /second of video | audio editing |
| ACE Step Audio Outpaint | Unattributed | 2025-05-11 | $0.0002 /second of video | audio editing |
| ACE Step | Unattributed | 2025-05-08 | $0.0002 /second of video | music generation |
| MiniMax Speech-02 HD | Minimax | 2025-05-06 | no published price | text to speech |
| MiniMax Speech-02 Turbo | Minimax | 2025-05-06 | no published price | text to speech |
| MiniMax Voice Cloning | Minimax | 2025-05-06 | no published price | voice conversion |
| Sound Effects Generator | Cassette AI | 2025-04-03 | no published price | sound effects |
| music generator | Cassette AI | 2025-03-27 | no published price | music generation |
| DiffRhythm: Lyrics to Song | Unattributed | 2025-03-04 | $0.001 /second of video | music generation |
| ElevenLabs TTS Multilingual v2 | ElevenLabs | 2025-02-27 | no published price | text to speech |
| ElevenLabs TTS Turbo v2.5 | ElevenLabs | 2025-02-27 | no published price | text to speech |
| ElevenLabs Audio Isolation | ElevenLabs | 2025-02-27 | no published price | — |
| Kokoro TTS | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (British English) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (Spanish) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (Brazilian Portuguese) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (Italian) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (Japanese) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (French) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (Mandarin Chinese) | hexgrad | 2025-02-14 | no published price | text to speech |
| Kokoro TTS (Hindi) | hexgrad | 2025-02-14 | no published price | text to speech |
| MiniMax (Hailuo AI) Music | Minimax | 2024-12-17 | no published price | music generation |
| Stable Audio Open | Stability AI | 2024-01-04 | no published price | sound effects |
Other categories
Text · Image · Video · 3D · All models
Frequently asked
How many Music & Sound AI models are there?
We track 108 Music & Sound models from 21 companies, of which 8 publish a price per second of video. The list is rebuilt from provider APIs and pricing pages rather than maintained by hand, so new releases appear without us writing about them.
What is the cheapest Music & Sound model?
ACE Step from Unattributed at $0.0002 per second of video on fal. Cheapest is not automatically best — check what you get for that price, and compare against the alternatives on the model's own page.
How do you compare Music & Sound models fairly?
Only within one billing unit and one job. Prices here are all per second of video; a model billed another way is excluded from the rankings rather than converted with an assumption. Alternatives on each model page are further restricted to models that do the same job, so a generator is never presented as an alternative to an upscaler.