MusicGen
Meta
Meta's music generation model. Generate up to 1 minute of music from text descriptions.
music
β $0.0529/run
Generate music and sound from a text description.
Quick picks
3 models
Meta
Meta's music generation model. Generate up to 1 minute of music from text descriptions.
music
β $0.0529/run
Community
MAGNeT is Meta's masked, non-autoregressive audio generator. Instead of predicting tokens left to right it fills masked audio tokens in parallel over a few decoding steps, so generation is faster than autoregressive MusicGen at similar quality. This Replicate packaging exposes the text-to-music and text-to-sound variants.
metamagnetnon-autoregressive
β $0.0030/run
Community
Stability AI's Stable Audio Open generates short audio from text prompts, tuned for sound effects, drum loops, instrument riffs and production elements rather than full songs. Open weights, latent diffusion over a 44.1kHz audio autoencoder, with a configurable seconds_total up to about 47 seconds.
stability-aistable-audiosound-effects
β $0.0069/run
Google DeepMind
Generate 30-second music clips from text prompts or images with Lyria 3, Google's music generation model
music
Deactivated
Google DeepMind
Generate full-length songs up to 3 minutes from text prompts or images with Lyria 3 Pro, Google's most capable music generation model
music
Deactivated
Community
Compose a song from a prompt or a composition plan
elevenlabsmusic
Deactivated
Stability AI
Generate high-quality music and sound from text prompts
music
Deactivated
Their pages stay online, but they canβt be run at the moment.
Audio-generation models cover everything that isn't speech or transcription: music, sound effects, ambience, and full songs with vocals. Reach for one when you need royalty-free background music for a video, sound effects for a game or app, or a full song with vocals for a prototype.
Pricing models vary more than in other categories. Music-from-prompt services (Suno, Udio, Riffusion) typically bill per-generation regardless of clip length up to their built-in cap. Sound-effect generators (AudioGen, ElevenLabs Sound Effects) bill per-second of output. Open-weights models running on shared GPUs are billed by compute time; on Railwail, MusicGen and Stable Audio Open are billed this way and each model page shows the price of a typical run.
The trade-off is musicality versus controllability. Flagships like Suno V4 and Udio produce surprisingly polished songs with verses, choruses, and instrumental breaks β but they decide most of the arrangement for you. Open-weights models (MusicGen, Stable Audio Open) give you finer control over genre, BPM, key, and instrumentation, but the output is shorter and less coherent. For background music in a video, the flagships usually win on time-to-final. For sound design that has to match a specific cue, open-weights with conditioning is the way.
Watch out for vocal cloning: some music models will happily generate vocals in a specific singer's style if you prompt them, which is a copyright and platform-policy minefield. Stick to original styles or use the safety-filtered tiers.
Licensing in this category is the most heterogeneous: some providers grant full commercial rights, some restrict to personal use, and a few are still in research-preview limbo. Always read the license before shipping output in a paid channel.
Top picks above cover the song-quality flagship, the cheapest sound-effect generator, the longest-clip option, and the fastest realtime model.
Suno and Udio produce 2-4 minute songs with verses, choruses and vocals on their own platforms; these models cannot currently be run through Railwail. The open-weights models here (MusicGen, Stable Audio Open) generate instrumental music and sound.
Mostly per-generation for music platforms (a fixed price per song or per few-minute clip) and per-second of output for sound-effect generators. Open-weights models on shared GPUs bill by compute time. Check each model card for the exact rate.
Sound effects: typically 1-30 seconds. Music: 30 seconds to 4 minutes depending on the model. Some platforms allow continuation β generating an additional segment that flows from the previous one β to build longer pieces.
Open-weights models (MusicGen, Stable Audio Open) accept explicit BPM and key tags. Commercial platforms accept natural-language style prompts ('upbeat synthwave at 120 BPM, in A minor'). Fine-grained control like time signature changes still requires post-editing in a DAW.
Most paid tiers grant full commercial rights. Some free tiers and research models restrict to personal use. The model card on each detail page lists the exact license β read it before shipping output in ads, films, or apps.
WAV and MP3 are universal. Some models also ship FLAC, OGG, and stems (separate vocal/drum/bass tracks for post-mixing). Default sample rate is 44.1 or 48 kHz; high-end tiers ship 96 kHz for music production workflows.
Voice cloning in music models is policy-restricted on most platforms to avoid copyright issues. For instrument cloning or style transfer, look at conditioning-capable open-weights models or use sample-pack workflows in a DAW with AI-generated stems.
Not yet at music-track quality. Sound-effect generation can be near-realtime (1-3 seconds for a 5-second clip). Full songs typically take 30-90 seconds to render. For interactive music (game scoring, live performance), look at adaptive playback systems rather than per-call generation.
Every available model through one OpenAI-compatible API. Prepaid credits in USD, no subscription.