MuseGen

AI Music Generator Benchmark 2026: A Protocol You Can Run Yourself

MuseGen logo

MuseGen Team

10/2/2026

#AI music generator benchmark#how to test AI music generators#AI music generator export formats#can you download from Udio#Suno free tier download limit#Riffusion vs Flow Music

Most comparisons of AI music tools, including our own ranked overview, give you a verdict: which tool suits which job, scored and ordered. This page is the companion to that verdict — not a ranking but the method behind one. Three prompts written to be checkable, six rules for running them, nine things worth measuring and the free tools to measure them with, and one table already filled in: what each engine actually lets you walk away with. Run it on the shortlist you are choosing from and the numbers you get are yours, taken on the day you took them, on the tiers you are actually paying for.

A producer seen from behind in a dark studio, headphones on, one hand on the mixing desk faders and the other on a mouse, facing a monitor that shows a comparison grid of six AI music engines — ElevenLabs, Flow Music, Mureka, MuseGen, MusicGPT and Suno — against five columns: free, paid, stems, MIDI and licence. Amber squares mark what is available, slate squares what is conditional and hollow outlines what is not, and Flow Music holds the only amber square in the free column. A second screen shows a loudness meter and a tempo readout of 128.0, and a mechanical stopwatch lies on the desk

What this page gives you

Six rules, three prompts and nine measurements, written out so that you can run them against whichever engines you are actually deciding between. The one table filled in here is export and licensing, because that can be read from published terms rather than produced by testing. Nothing here is a score out of ten.

  • Three prompts, written once and never adjusted. Each names a tempo, a duration, at least one exclusion and a specific ending, because those are the four instructions generators most often quietly ignore.
  • One generation per prompt per engine, kept. The first take is the one you report. Count re-rolls separately rather than using them to improve the result.
  • Measured, not judged, wherever a measurement exists. Tempo, key, duration, sample rate, bit depth, integrated loudness and true peak all come from the file rather than from listening.
  • Two of the obvious names cannot be tested at all. Riffusion no longer exists under that name, and Udio has switched off downloads entirely. Both stories are below, because they change what you can actually do more than any quality difference does.
  • The export table is already filled in. What each engine lets you download and on which tier, read from each platform's own pricing and help pages and dated so that you can recheck it rather than trust it.

Quick facts

  • What this is: A protocol, not a results page. Six rules, three prompts, nine measurements, one working day.
  • Engines it is written for: Six: ElevenLabs Music, Google Flow Music, Mureka, MuseGen, MusicGPT and Suno.
  • Engines documented only: Four: Udio, AIVA, SOUNDRAW and Mubert — export and licensing terms recorded, nothing generated.
  • Generations to keep: 18 — three prompts × six engines, first take each, re-rolls counted but never substituted.
  • Measure with: MuseGen Detect BPM, MuseGen Key Detector, and ffprobe / ffmpeg for format and loudness.
  • Terms checked: 21 September 2026, from each platform's own pricing or help pages.
  • Not covered: Which one sounds best. That question is left where it belongs, with the person listening.

The protocol, in six rules

Write the rules down before you open the first account. That matters more than it sounds: a protocol decided halfway through is a protocol shaped by early results, and once you have seen one engine do something impressive it is very hard to keep the conditions the same for the next five. There are six rules, and they fit inside one working day — which is itself a rule, because an engine tested on Thursday may be running a model the Monday engines were not.

Protocol

  1. One account tier per engine, and write down which. Where a free tier and a paid tier produce different output, use the paid entry tier, because that is the tier most people end up on.
  2. Paste the prompts verbatim into the main text field. No genre pickers, no mood dropdowns, no style presets, even where the interface offers them. Every engine gets the same words.
  3. Keep the first take. The generation you record is the first one returned. Where an engine returns two variants per submission, take variant A and be consistent about it.
  4. Count re-rolls, never substitute them. Repeat up to four more times to record how many attempts it takes to reach a usable take, but do not let those later takes replace the first in your results.
  5. Take the highest export the tier allows. If the plan offers WAV, take WAV — then read the delivered format back off the file rather than trusting the menu.
  6. Measure the exported file. Not the in-browser player, which resamples.

Expect one deliberate asymmetry: engines differ in how many words their prompt field accepts. Where a prompt has to be truncated to fit, note the truncation point against that engine rather than rewriting the prompt shorter for everyone.

"Usable take" needs a definition or the re-roll count means nothing. The one suggested here: a take you would put in front of a client without regenerating it — no obvious artefacts, no collapsed ending, no wrong instrument carrying the melody. Have two people judge independently and record only the calls they agree on. Where they disagree, mark the take unusable; that biases the column slightly pessimistic and does so equally for every engine, which is the property you want.

List the engines alphabetically and keep the columns separate. A protocol run is raw material; weighting the columns into an overall score is a separate step, and it is the one our ranked overview takes. On this page there is a second reason to stay unranked, which is that one of the entries is ours.

Two engines that changed shape before the test could run

The obvious shortlist is the six names people actually search for, which means Suno, Udio and Riffusion alongside the newer arrivals. Two of those three cannot be put through this protocol at all, for reasons that have nothing to do with audio quality and everything to do with what you are allowed to walk away with. Check both before you spend a day on a shortlist, because a comparison written six months ago will still list them as ordinary options.

Riffusion is now Google Flow Music

Riffusion began as an open-source experiment that generated audio by treating spectrograms as images, rebranded its commercial product to Producer.ai, and was acquired by Google in February 2026. The original architecture was retired and replaced by Google's Lyria models. In April 2026 the product relaunched as Flow Music, and riffusion.com now redirects to flowmusic.app, where the word "Riffusion" does not appear anywhere on the page.

So the engine tested here is Google Flow Music, and it is entered under that name. If you are searching for Riffusion you will land on it either way, but you are not using the tool that the older reviews describe — the model underneath was replaced wholesale at acquisition, and users were told to download anything they wanted to keep before prior generations became inaccessible.

Worth knowing: at the time of checking, every Flow Music tier including the free one lists downloads in MP3, WAV and M4A, plus stem downloads. That is a more generous free tier than any other engine in this comparison offers, and it is the kind of thing that changes which tool you reach for far more than a two-BPM tempo difference.

Udio has switched off downloads

Udio settled copyright litigation with Universal Music Group in October 2025 and with Warner Music Group the following month. The settlements came with a transition period, and the transition removed the one thing a benchmark like this depends on. Udio's own help centre states it plainly:

Udio Help Centre, updated 17 February 2026

Note that downloading of audio, video, and stems has been disabled

The same article records what subscribers got in exchange: a one-time grant of 1,000 non-expiring credits, the Standard monthly limit raised from 1,200 to 2,400, and the Pro limit raised from 4,800 to 6,000.

Without a downloadable file there is nothing to run ffprobe against, nothing to measure loudness on, and no export format to report. Udio therefore appears in the export and licensing table below and cannot appear in a measured column at all — not as a bad score, but as a row that nobody can produce. Udio's companion note on the Warner arrangement confirms that the October transition terms still govern the service and that generation remains on the v1, v1.5 and v1.5 Allegro models — the platform has been feature-frozen while the next version is built.

You can still create on Udio and still share what you make through udio.com links. What you cannot do is take the audio anywhere else, which rules it out for anyone delivering to a client, a distributor, or a video timeline. Our longer explainer on Udio covers the features and pricing in more depth.

The two vacated slots go to MusicGPT and Mureka, both of which generate and both of which let you download the result — which is the minimum qualification for appearing in any column of this test.

A two-lane timeline on a dark background. The upper amber lane traces Riffusion from an open-source spectrogram experiment through the Producer.ai rebrand to the Google acquisition on 25 February 2026 and the April 2026 relaunch as Flow Music. The lower slate lane traces Udio from the Universal Music Group settlement in October 2025 and the Warner Music Group partnership on 19 November 2025 to the help-centre note that downloads are disabled, ending in a greyed-out step reading not measurable at any tier

Neither of these changes is visible on the page a search sends you to — which is why a shortlist copied from an older comparison quietly stops being true.

The three prompts, verbatim

Each prompt is built on the same six slots — genre, tempo, instrument roles, vocal, production, length and ending — laid out in our prompt formula. They are written to be checkable rather than impressive. Every one of them contains a tempo number, a duration, at least one exclusion, and an instruction about how the track ends, because those four are where the gap between what a tool claims to read and what it acts on shows up most clearly. Take them as they stand, or write three of your own on the same pattern; what matters is that they are fixed before the first generation and never touched again.

Prompt A — vocal, structure, ending

jangle pop, 104 BPM, an even backbeat. Clean electric guitar carries the melody, bass and a soft kit keep time, no synths anywhere. One female lead, close and unprocessed, one harmony line on the choruses only. Dry and close, early 2000s studio sound. 2 minutes 30, ending cold on the last downbeat.

Prompt B — tempo, exclusion, instrumental

garage house, 121 BPM. Stacked organ chords carry the harmony, rolling sub bass and a shaker keep the pulse, no vocals and no vocal samples. Wide and clean, modern club sound. 3 minutes, ending on the organ alone with the drums dropped out.

Prompt C — key, arrangement, decay

orchestral cue for a slow reveal, 72 BPM, in D minor. Solo viola carries the melody, double basses sustain underneath, a single timpani hit marks each phrase end, no drum kit and no synths. Wide hall, modern film score. 90 seconds, ending on one sustained chord that decays to silence.

Prompt C is the only one that names a key, and it is there for a reason: key is the instruction most tools treat as a suggestion, and unlike "wide hall" it can be checked against the file in about fifteen seconds.

What to measure, and with what

Nine columns: six come out of software, one out of a stopwatch, and two out of a pair of ears. Keep the last two in a table of their own rather than mixed in with the rest, so that nobody has to take your listening calls on trust in order to use your measured ones.

MeasurementHow it was takenInstrument
Tempo deviationRequested BPM minus detected BPM, in BPMMuseGen Detect BPM
Key matchRequested key against detected key and mode — Prompt C onlyMuseGen Key Detector
Duration deviationRequested length minus delivered length, in secondsffprobe
Delivered formatContainer, sample rate and bit depth read from the fileffprobe
Integrated loudnessLUFS-I across the whole fileffmpeg loudnorm, print pass
True peakHighest intersample peak, dBTPffmpeg loudnorm, print pass
Time to playableSubmit to the moment the track can be played, in secondsStopwatch
Exclusion honouredDid the named instrument or vocal stay out — yes or noTwo listeners, agreed calls only
Ending honouredDid the track end as instructed rather than fading — yes or noTwo listeners, agreed calls only

The two commands that do most of the work are short enough to keep in a note. One reads the container back off the file; the other prints loudness and true peak without writing any audio out.

Read the delivered format

ffprobe -v error -show_entries stream=codec_name,sample_rate,bits_per_raw_sample,duration -of default=noprint_wrappers=1 track.wav

Print integrated loudness and true peak

ffmpeg -i track.wav -af loudnorm=print_format=summary -f null -

Run both against the exported file, not against anything the browser played you. If bits_per_raw_sample comes back empty on a WAV, read sample_fmt instead — some encoders leave the field unset, and an empty cell is not the same finding as a 16-bit one.

ColumnWhat a result actually tells you
Tempo deviationWhether the model treats a number in the prompt as an instruction or as a mood word. A clean zero on all three prompts is a different kind of tool from one that drifts four BPM.
Key matchThe same question with less room to hide, because only one prompt names a key and the answer is binary.
Duration deviationWhether a stated length is honoured or rounded to whatever the model likes generating. Matters most if the track has to sit under something.
Delivered formatWhat you can still do with the file afterwards, and whether the download menu was telling the truth.
Integrated loudness and true peakHow much headroom survives. A heavily limited delivery has had decisions made for it that you cannot undo.
Time to playableWhat using the tool all day feels like, which no other column captures.
Exclusion and endingWhether the model reads the two instruction types that cannot be checked by software, which is exactly why they sit in a separate table.

The two shaded rows are the judged ones. Everything above them can be reproduced by anyone holding the same exported file and a copy of FFmpeg, which is the whole point: anyone can re-run it and compare notes.

Tempo deserves one warning before you write anything down. Detected tempo is not infallible — half-time and double-time readings are a known failure of every detector, ours included, and a track written at 72 BPM will sometimes read as 144. When a reading comes back at exactly double or exactly half the number you asked for, check it by hand against a metronome and record both figures, the corrected one as your result and the raw one beside it, so that the correction is visible rather than silent. Deviations that are not a clean 2:1 ratio need no adjustment; they are straightforward misses. How the ambiguity arises in the first place is covered in what BPM is.

Record the date beside every run, not just at the top of the page. These engines ship model updates without changelogs, and a column taken in March is not comparable with one taken in September even when the protocol was identical. Dating each run rather than overwriting the last is what turns a one-off afternoon into something you can watch move.

What each engine actually lets you download and sell

This is the one table you do not have to produce yourself, and it is also the one most worth having before you spend a day generating. The file you can download matters more than most comparisons admit, because it is the thing that survives after the novelty wears off. Two questions decide it: what format comes out at the tier you are on, and what the terms say you may do with it. Every cell below was read from the platform's own pricing or help pages on 21 September 2026, and the sources are listed at the foot of the page. Terms in this market change without announcement, so treat the date as part of the data.

EngineFree tier downloadPaid tier downloadStemsMIDICommercial use
ElevenLabs MusicNo download of any kind on the free tier under ElevenLabs' own commercial rights table — not just a quality shortfallLossless WAV from Creator up; Starter has neither2 and 4 stems; up to 6 on higher tiersNot offeredYes on paid, but self-serve plans and even Enterprise Music Lite exclude film, TV, radio and Studio Games; only full Enterprise Music has no exclusion
Google Flow MusicMP3, WAV and M4A, plus stem downloads — the same tick as every paid tier on Google's own comparison tableIdentical to the free tier; no download row changes across Starter, Plus or MemberYes, on every tier including FreeNot listedNo document states what a plan permits, and the pricing page never uses the word commercial. Google does say it claims no ownership of what you generate
MurekaNot statedUnlimited downloads on paid tiers; MIDI and WAV export are Premier-exclusiveUp to 12 stems on PremierPremier onlyYes from Pro upward
MuseGenNot offered — downloading is listed only on the paid plansListed as "Download music" on Pro and MaxNo stem export of generated songsNo MIDI export; separate Audio to MIDI toolYes on Pro and Max; the licence attaches at generation rather than at download, survives cancellation for works already generated, and carves out no industry. Free and pay-as-you-go stay personal
MusicGPTMP3, at 50 credits per download from a 500-credit monthly allowanceDownloads cost no credits from Plus upward; Pro and Ultra are listed as "Best Quality" without the format being namedPro and Ultra onlyNot listedPaid tiers only, with a download licence file from Plus upward
SunoNone — "no monthly song downloads"Pro: 20 downloads a month. Premier: 60None on free; 2 types on Pro; 3 on PremierPremier, through Suno StudioNone on free; yes on Pro and Premier
UdioNone — audio, video and stems all disabled since 29 October 2025None. The block covers the paid tiers and reaches back to tracks made before the change, with no published return dateDisabled with everything elseNot offeredHard to exercise at all while nothing can leave the platform
AIVA3 downloads a month, MP3 and MIDIStandard: 15 a month, MP3 and MIDI. Pro: 300 a month, all formats including WAVNot offeredYes, on every tierPro only — free and Standard leave the copyright with AIVA
SOUNDRAWNot stated — the pricing page lists paid plans onlyHigher tiers add WAV and stemsYes, on the upper tiersNot offeredYes, royalty-free, trained on an in-house catalogue
MubertMP3, 5 downloads a month, non-commercialUnlimited downloads from Creator up; commercial use only from ProRegenerate and remove stems on every tier; stem editing inside sections from ProNot offeredPro and Business only — the paid Creator tier is still non-commercial; tracks cannot be registered with Content ID systems

The bottom four rows are reference rather than test subjects. Udio cannot be exported from at all, and the other three sit outside the six names most people are choosing between; their terms are here because export and licensing are worth having in one place, and because leaving them out would make the table look like a ranking rather than a reference. Where a platform does not state a term on its public pages, the cell says "not stated" rather than guessing — and if you are relying on that cell, the answer is to ask the platform, not to assume the generous reading.

Three things in that table are worth pulling out. The first is that the most generous free tier belongs to the engine most people have not heard of under its current name: Google Flow Music gives away WAV and stems at zero, while Suno's free tier gives no downloads at all. The second is that a licence and a file are different questions — AIVA will hand a free user a MIDI file but keeps the copyright, which is the opposite trade from Suno, where paying unlocks both at once. The third is that the commercial-use column varies more than any other and gets read least: one engine publishes no document at all, one excludes film, TV and radio even from a paid enterprise tier, one has a paid tier that is still non-commercial, and one ties the right to an approved download rather than to the subscription you are paying for. Read that column before the price column.

A colour-coded grid of ten AI music engines against five columns: free tier download, paid tier download, stems, MIDI and commercial use. Amber squares mark where a file comes out, slate squares mark conditional access limited by tier, format or use, hollow outlines mark none or disabled, and dotted outlines mark terms not stated publicly. Google Flow Music is amber right across the free tier and Udio is hollow across every column

The same table read as a shape rather than as sentences. Amber is a file in your hands; slate is a file with a condition attached; hollow is nothing at all.

One recurring trap is worth naming. A menu offering WAV does not guarantee a WAV that carries more information than the MP3 beside it. If the model renders at a lower rate internally, the WAV is an upsampled container and sounds identical. That is why rule five ends with reading the format back off the file with ffprobe rather than copying it from the download menu, and it is the same distinction we make in the MP3 to WAV checklist.

▶ Watch on YouTube: "Best FREE AI Music Generators 2026" — Titled as a survey of free AI music generators and published in March 2026 — one month before Riffusion relaunched as Flow Music with the free tier that now sits at the top of the column above. A worked example of why the date belongs beside the finding rather than only at the top of the page.

What this protocol does not tell you

A published protocol invites the obvious question of what it leaves out, so here it is, before anyone has to ask. Knowing these five before you start is also what stops you over-reading your own results a week later.

  • One take is one take. These models are stochastic, and the same prompt run twice returns two different songs. A single first take per engine measures what you get on a first attempt, which is a real and useful thing to know, but it is not an average and should not be read as one.
  • Three prompts are three prompts. They were chosen to be checkable, not to be representative of everything people ask for. An engine that handles a specified tempo badly may handle a vague mood prompt beautifully.
  • Nothing here says which sounds best. There is no audio-quality score here; that judgement belongs to the person listening, and our ranked overview is where we give ours. Loudness and format get recorded instead, and they are not the same thing as sounding good. An engine that lands on your tempo is telling you it reads prompts as instructions, not that it writes better music.
  • Prices and terms move. Treat anything older than a quarter as a starting point for checking rather than as a current fact — the Riffusion and Udio entries above are what happens when nobody rechecks.
  • Model versions move faster. A result from one release says little about the next, which is why the date belongs beside every run rather than only at the top of the page. Two runs of this protocol six months apart are two data points; the same run quoted six months later is one stale one.

▶ Watch on YouTube: "Best AI Music Generator? Suno vs Udio vs Mureka vs ACE Studio" — Titled as a four-way comparison, three of whose engines are in the table above, published in June 2026 — impressions rather than a measured protocol, which is the other half of the picture.

Our own tool is in this table

There is no honest way around the conflict of interest, so the useful thing is to say exactly where it sits.

MuseGen has a row in the export table, read off our own pricing page the same way every other row was read. What that row shows: the licence attaches when a track is generated rather than when it is exported, it survives cancellation for work already made, and it carves out no industry — the three things the commercial-use column is actually for. It also shows the cells where the answer is no: no stem export of generated songs, no MIDI export. The six engines the protocol is written for are listed alphabetically, the only ordering we could think of that nobody has to take on trust. No winner is named anywhere on this page; naming one is the job of the ranked overview.

Two of the suggested instruments are also ours, which deserves a sentence rather than a shrug. Tempo is read with Detect BPM and key with the Key Detector, both of which run in the browser. They are named here because they are convenient and because you can check them — any other detector, or a metronome and a piano, will give you the same reading from the same file. Nothing in the protocol requires a MuseGen account, and if a different instrument disagrees with ours on a particular file we would like to hear about it.

If you want to run the three prompts here rather than write your own, the AI song maker takes them as written — they were composed against the six-slot template rather than against any one interface, which is also why they paste cleanly into everything else on the list.

FAQ

Why is Riffusion not one of the engines here?

Because it no longer exists as a product. Riffusion became Producer.ai, Google acquired it in February 2026, replaced the underlying model with its Lyria family, and relaunched it as Flow Music in April. The domain redirects and the old name appears nowhere on the current site. The engine is on this page under the name it actually has now. Reviews still listing Riffusion as a separate option are describing something you cannot sign up for.

Can I really not download anything from Udio?

Not at the time of writing. Udio's help centre says downloading of audio, video and stems has been disabled, a change that came with the Universal Music Group settlement of October 2025. You can still generate, still edit, and still share tracks through udio.com links that anyone can open. What you cannot do is get a file. If your work ends with handing someone an audio file, that rules Udio out until the policy changes, and it is worth rechecking rather than assuming either way.

Why does the protocol stop at six engines?

Because generating, exporting and measuring three tracks on ten platforms inside a single day is not achievable without cutting a corner, and the corner that gets cut is the "same day" rule. Six is roughly what one person can run properly between morning and evening. Four more engines are documented in the export and licensing table, which is the part of a comparison that can be read from published terms rather than produced by testing. Mixing the two and presenting all ten as equally tested is the failure this protocol is built to avoid.

Why is the first take kept instead of the best one?

Because "best of five" measures the person operating the tool as much as the tool, and because the number of attempts it takes to get something usable is itself worth reporting. Keeping the first take and counting the re-rolls separately gives you both figures without letting one contaminate the other. If you generate until something works, the re-roll column is the one that describes your experience.

Is tempo detection reliable enough to use as a score?

For this purpose, yes, with one caveat that applies to every detector ever written. Half-time and double-time confusion is real: a track at 72 BPM is genuinely also a track at 144 BPM depending on what you count as the beat, and automatic detection picks one. Re-check any reading that comes back at exactly double or exactly half the figure you asked for against a metronome, and record both the corrected number and the raw one. Deviations that are not a clean 2:1 ratio need no adjustment; they are straightforward misses.

What do I need to run this?

An account on each engine you are choosing between, the three prompts printed above, a stopwatch, and FFmpeg, which is free and runs everywhere. Tempo and key can be read in a browser. The cost is a day of your time plus whatever the entry tiers charge. What you will not get is the same audio anyone else gets, since the same prompt returns a different song every time, so what you are really comparing is the shape of the result: which engines land on a requested tempo, which honour an exclusion, and what file actually arrives.

Why measure loudness at all if you are not scoring audio quality?

Because delivered loudness tells you how much headroom you have left. A file arriving at a heavily limited level has had decisions made for it that you cannot undo, and if you plan to master the track or sit it under dialogue, that matters more than any impression of quality. It is a practical number about what you can still do with the file, not a judgement about how it sounds.

How current is the export table?

Every cell carries the date it was read, which is the honest form of that answer: terms in this market change without an announcement, so a date you can check beats a promise you cannot. Treat any cell older than a quarter as a starting point for rechecking rather than as a current fact, and the Riffusion and Udio entries above are what happens when nobody does. The protocol itself does not need updating, which is the point of writing it down before the first account is opened.

Can I run this on only two engines?

Yes, and that is the most useful version of it for most people. Nothing in the protocol depends on the number of engines; it depends on the prompts being fixed before you start and the measurements coming off the exported file rather than out of the browser player. Two engines, three prompts and one afternoon will tell you more about the choice you are actually making than a ten-way table somebody else ran on a different day.

Sources