For whom the bell tolls: the copycrab music industry
tldr: The music industry wants to kill music generation, so it bought/threatened the people who made it. Every model anyone was using was retired in one day, and the labels are now the only legitimate issuer. We can make alternatives to Suno/Udio in a way that makes lawyers eat shit. What is still missing is math and technique – to avoid using anything copyrighted in training data. Simple operation: music references come at inference time, all copy music is user supplied, and they can’t crack down on all of us.
For a while, anyone could make music. A real track, with vocals. Made on a whim. No instrument, no studio, no skill. Didn’t matter.
You wanted a funny track for your friends. Maybe you wanted vibe music making, without having to spend years on the craft first. Most of it was for fun. Some of it was genuinely good, and it is worth saying plainly: regardless of what you think about Suno and Udio, when we created something with them, we created something great.
Then the copycrabs* went to work.
(Copycrabs – parasitic entities behaving like a bucket of crabs)
What happened – a quick timeline
June 2024: the RIAA sues Suno and Udio. They don’t want money. They want training to end. No training on copyrighted recordings, ever again, plus up to $150,000 per work already learned from (The Verge).
That’s the ask: the tool stops learning anything you’d recognize.
October 2025: Udio settles with Universal. Users get 48 hours to download everything they made, then the platform becomes something else. A licensed streaming business, by design (timeline). Your library was never yours. Just a grace period.
November 2025: Warner settles and licenses, and drops out of the lawsuit. Artists opt in to having their name, image, likeness and voice used (The Verge). From this moment, every use is a licensed use. Your rights end where their catalogue begins, and anything unlicensed stops being a product and turns into evidence against you.
Nobody ever beat the service in court. It was bought and absorbed.
How does the saying go? "If you can’t beat them, join them." Seems buying was just cheaper than trying to kill them in court. Result is the same anyway.
That is what killing a service looks like – threaten and consume.
Bought means lobotomized. Udio still runs, but downloads stayed disabled since the UMG deal. You can blend styles inside. Nothing leaves. Suno stayed up on the same leash, download quotas already flagged as coming. Make it, take it, play it elsewhere?
Crippled. Users weren’t kept. They got kicked out slowly, through restrictions, until only the licensed box was left.
2026 is the same attack continued against services that were already damaged. BMG licenses in August (MBW), Believe comes on board in September (MBW), and the terms arrive with them: the download counters start (Suno), new models launch while every model anyone was actually using is retired the same day, and the distributors who had spent months refusing anything unlicensed quietly start accepting output from the new models only. What you made last year cannot be distributed, what you make today is counted, and the only door still open is the licensed one.
September 18: they sue again – not over the old models but over the new ones, built from scratch on licensed material with Warner, BMG and Believe inside. Universal’s and Sony’s claim is that it does not matter, because a model built on the outputs of a tainted model inherits the taint. Their words: v6 "is not a fresh start; it is the fruit of the same poisoned tree." They want 60,202 recordings, which at the statutory ceiling comes to a little over nine billion dollars (MBW). Even compliance does not end it. Either the catalogue comes down, or the damaged service pays to keep standing.
You think they will stop? Haha, no. They either make it irrelevant or they keep the profits coming. Either way they stay in control of the whole industry and keep extracting rent.
Why this industry fights when the others did not
The music industry is an anomaly, exactly because it is fighting. They are fighting against the new world order, where content is cheap and abundant, generated on the fly, and your ownership, your intellectual property, means nothing anymore – because anyone can repeat it.
From their point of view, sure, they do not want to lose the money: all that lobbying money, all this tyranny against people and artists big and small that they have been extorting. It is the endless attempt to take what is supposed to be a commodity now, with all the tools we have, and make it scarce. Make it one of the privileged few. Take away anyone’s capability to create and spread music.
Yeah, of course there is a lot of money involved, but that money should not be involved here at all. The cost of entry collapsed, so the profits should collapse with it, and the money should drain out. Yet they are fighting for their profits everywhere. And they get to fight because they were ready: already centralized, already organized, with years and years of practice at exactly this – the file-sharers, the lawsuits, the lobbying. The money existed, so the lawyers existed, and the playbook was already written.
Fighting isn’t winning. Luddites, unable to face a new reality, will try to keep their status quo. Our job is to give people back what was taken.
The idea
The basic principle: a module that takes whole albums as reference and imitates any artist, with the reference supplied as context at inference time. The model itself is never trained on proprietary data at all.
How it works: you hand it an album, it reads what that kind of music sounds like, and it aims there. The album does not teach it an artist, it points at what to aim for right now.
The music stays in your hands, at use time, not inside the model.
The model does few-shot on Rihanna without ever seeing her in training. You give it her album at the moment you ask. It never trained on her, it never held a single one of her songs, and it still makes something in her style.
You set how close it lands. A dial: fifty to seventy percent, in her style but never her song. Or a synth of several references at once. A likeness you control.
And never mind the gatekeepers, it is a better instrument than the flat text-only kind. A text-prompt model gives you words and hopes; this gives you control over how the music is shaped. A dense model that learned one Rihanna will always average out to the same vibe – but hand it several of her albums and you can combine them, pull from one era against another, steer between them. For working artists it means that much more: keep the vibe of your own past work, get inspired by it, try something new built on your own foundation.
The reference lives with you. Your machine, the moment you ask, then gone. One album in, the vector out, the audio deleted. Or a tiny adapter training on your PC, never leaving it. Nothing uploaded, nothing kept, nothing on any server to point at.
Where the reference comes from is your business. A recording, a torrent, a download from anywhere, something you made yourself – or no audio at all, just a representation built mathematically from a song everyone knows. What matters is that nobody can see it. The company that makes the model never touches your reference, never holds it, never knows it existed. All the blame lands on users, one by one, and there is nothing to prove against any of them: no corpus, no logs, no library. A million people each with their own album is not suable.
Everything about it is open and provable. Open, synthetic, public-domain training only. All code and pipeline published. Anyone can replicate it and check nothing proprietary sits inside.
The effect: produce any music you want from references you supply.
What exists so far
A year ago, after a quick look, this did not exist – and the look was quick because it was hard to make. Nothing open could take a reference and carry its style; the few things that tried either needed a studio’s hardware or produced mush. A year later it exists. Not openly, not for you, and not enough – but it exists, which changes the question from whether it can be done to who gets to do it.
On the open side, real numbers this time, straight off the Hub. ACE-Step 1.5 still the pick. 59K downloads in 30 days, 874 likes, highest-liked song model. MIT. Reference audio on every DiT variant. Cover, repaint, LoRA. Runs on a gaming card. Full songs, with vocals.
Magenta RealTime 2 from Google. 9.2K downloads. Steered by text, audio examples, MIDI. Weights CC-BY 4.0, code Apache-2.0, Google claims no rights in outputs. Cleanest license for reference. Catch. Instrumental only, live-streaming shape.
m-a-p/YuE2-3B new, September 2026. 26K downloads, 1020 likes. Cover mode, editable scores. Heavy. And CC BY-NC 4.0. Non-commercial. Skip for anything you want to ship.
SongGeneration v2 / LeVo 2 takes 10 seconds of prompt audio. Non-commercial too. MusicGen-Style does style reference, yeah, but instrumentals only. Weights CC-BY-NC. Code is MIT, weights aren’t. Don’t mix those up.
And the big ones that can’t take a reference. MiniMax-Music3 looks huge, 531K on the GGUF, 1411 likes. Text only. Lyrics plus description in, music out. No reference in. Stable Audio 3 family, audio-to-audio seed, not style. Base MusicGen, melody conditioning, not style. DiffRhythm 2, Muse, no reference input. HeartMuLa promises reference audio in the paper. README says TODO. Not real yet.
People run these as GGUF now, by the way. That 531K MiniMax quant, 121K Yue2 quant. audio.cpp. Not raw PyTorch. That’s local usage today.
Then the proprietary side, which proves the shape and gates the door in the same move. Suno’s Inspire takes a playlist of your own songs and generates in its vibe. Custom Models train on your own catalogue, and only your own. Both work. Both are fenced: your songs, their terms, their caps, their stamp on the way out.
What is missing is good math and good techniques for carrying a reference. Supply several songs and the model averages them – and the average of eight songs is not a song. The vibe never gets extracted coherently; it either copies one reference or smears them all into mush. Nobody has built a channel that is both wide and non-copying.
So what is needed is better thinking in this exact spot: how to hold several references at once and pull the style out of them without averaging them to death. Doable, unsolved, and close to what proprietary services already do.
Workarounds – open data, the edge from elsewhere
The dataset stays open, and the edge comes from somewhere else.
Train on what is open. The Free Music Archive is 106,574 tracks and 343 days of audio, Creative Commons, frozen before any of this existed. MTG-Jamendo adds tens of thousands more. The training pairs cost nothing – hold out one track per album and condition on the rest of it.
Then the edge arrives at use time instead of living in the corpus. You point the tool at music you like, on your own machine. That reference was never in the dataset, and nothing has to be kept afterwards: one album in, the vector out, the audio gone.
The copy is the only proof, so with nothing kept there’s nothing to find. Their forensic tool was audio fingerprinting. It worked because the corpus sat on disk, a retained library. Weights can’t be fingerprinted. No test maps a model back to its training. No library, no target.
A model needs something that tells good from bad, and a judge trained the normal way needs the very thing we refuse to keep. Three ways around that, all from outside the corpus.
First, the judge just listens. No files, no dataset, no copies kept – human ears ranking pairs of outputs, the way Google aligned a music model on 300,000 pairwise preferences. What persists is a rating, not the work, so the signal stays clean even when the thing judged is not open at all.
Second, the judge trains on slop. Generated music is the cleanest negative class in history: nobody owns it, it is unlimited, it is already catalogued, and using it triggers no rights at all. Teach the judge what bad sounds like from synthetic failures, and the human preferences only have to supply the taste.
Third, the judge chases your own model, not anyone else’s. A detector trained on someone else’s outputs will happily call your own failures human, because every model family has its own fingerprints and none of them transfer. So the negative class is your own rollouts, co-evolving: generate, judge, retrain on the winners, repeat.
The judge tracks your model’s drift instead of policing a fixed past.
If you do not care about cooperating at all, put the whole thing in a third country. The honest version of that: the judge is not a judge, it is the host. An anonymous drop survives.
They scraped the freely licensed material too. The corpus was available. They took the other stuff anyway.
The slop issue
Wait, we all hate slop! Why are you advocating for something that would increase it?
"We value human artists!"
Now the part that is your fault. Slop music industry. You let that happen. You helped them to suffocate the small uses with copyright and then left.
Walk into a cafe or a restaurant and listen. Nobody is playing the real artists. It is generated music, and you might be furious about it, and you might ask why anyone would do that when everyone prefers the real thing.
They do it because it is free, and because nobody has to be forced to pay for it, and because every alternative was made impossible. Nobody can safely put music in a room without paying the toll, because of the lobbying and the suing and the endless treating of a commodity as if it were something sacred.
More than five million songs already exist. Every one of them priced as though payment were owed. So businesses responded, and now you get bad generated music for free, and you complain about the quality. Fine – it will get better. But you are suffering exactly because you let the first party win, and if they had been stopped, you would still be hearing human music in all those places, because there is more of it than anyone could ever listen to.
Slop grew popular exactly because the copyright fuckers tried to control everything. If filler music had stayed free to play – a cafe speaker, a stream background, a small room – there would be no need for slop in half the places it now fills. Nobody needed fake music where real music was allowed. Now every filler is slop, because real music was priced out of the filler slot.
But stop and ask what slop actually is, because everyone calls it shit and almost none of it is. Every track people share with each other is still a human creation. Someone thought of it, wrote or fixed the lyrics, thought about how it should sound, listened to the result. Nobody spits out a song untouched – the majority make it for their own sake, for fun, for something interesting to share, like making memes. That is people making music.
The actual slop is the mass-produced kind, and it exists for one reason: monetization. Services that pay per stream created the incentive to generate thousands of tracks, fully automated, each earning a couple of cents, millions of them in the hope of anything at all. Blame the companies that let slop earn – Spotify and the rest built the environment where slop pays, so slop gets made. Without the per-stream cents nobody would bother cranking redundant crap at industrial scale.
Then the arithmetic. You pay five to ten dollars a month for streaming. Holding and serving that content costs somewhere in the range of one to ten of those dollars, and most of what you pay leaves as payouts concentrated on a handful of names. That is how a commodity gets repriced as a luxury, and how a few artists become absurdly unequal to everyone else.
Step back and say what the leeches actually do. By every action – the lawsuits, the buyouts, the caps, the doors – they reduce what common people can create and deny everyone the ability to become an artist. That is the product, not a side effect: fewer makers, fewer voices, and a gate in front of every one that remains.
And the whole so-called music industry is just there to collect the money. You think they care for artists themselves? No. They care whether it can produce, whether they can pump money in and receive money back. Even if it started as protecting artists, now it is about whether anything is allowed to earn.
Copyright law is the insurance that lets it continue. Without state protection it would all crumble, because nothing they sell is scarce except the permission slip. Take the protection away and the only scarcity left is the real one: physical presence. A body on a stage, in a room, on one night only. That is where artists should earn – the concert, the performance, the thing that cannot be copied. Nobody should profit from music shared as files. Files are abundant, presence is not.
The artists themselves are fine, because artists are performers. Their money is the stage, not the plastic. What dies is the rent.
We care for all of them: small people who want to just make their funny songs, wannabe artists who are finally able to produce something with AI tools, real artists who can now create something interesting or better with AI and different tools.
The labels should die, because that is what they are: a construct for extracting money, not for making music. Artists do not need them to grow, to get popular, to have ideas, to be unique. All of that happens without the machine that picks a few and starves the rest.
What we don’t care about is how the music industry keeps picking up some artists and handing them fortunes while others sit with nothing.
Go take a peek at 100 different songs. Is there really any difference nowadays? Was there any difference 10 years ago? We are in an era of abundance, and yet only the few chosen ones are able to get a lot of money out of it.
Why does such an unequal distribution of income exist, where out of a million artists only a small fraction actually makes money, just because they got popular and all that? Why isn’t everyone else able to fill that space?
There is nothing wrong with everyone getting popular. What is wrong is that they are able to profit from something that is abundant. Music is now abundant. Those companies are the ones trying to keep the scarcity of access to it, so they can continue to expand and extract their rents.
Music should be ours, not theirs
They are fine with the technology, they just want to own the only license to sell it, and everything else follows from that.
All they want is to keep extracting rent and control, and all these big fucking companies just want to hold on to the money they have been pulling out of this field for years. But they are resisting the new era, in which music is abundant and nobody should pay for it.
Creation stays free and reaching gets credentialed. Home recording did not kill music either – it killed the studio monopoly, and the industry kept the catalogue and the pipes. This lands in the same slot: unlimited making, gated amplification.
So the thing they are charging rent on is not the thing that protects them. The tools are open, the data is open, the workaround is understood, and what is still missing is a couple of unsolved problems that money was never going to solve anyway.
If they behave like a bucket of crabs – let’s kick that bucket.