All articles

You Cannot rm -rf Your Way Out of a Copyright Lawsuit

6 min read1,286 words
Audit and EvidenceEvaluation and Benchmarks
['Human Operator/Engineer', 'Computer Systems and Monitoring Equipment']. Active human-in-the-loop monitoring and control of complex digital or electronic infrastructure. Technology and Engineering.

Early in my career I trusted the delete key. Point it at the offending directory, watch the bytes vanish, close the ticket, move on. It took decades of building large systems to teach me the uncomfortable truth. In machine learning, deletion is a claim, not an event.

The reason is not mystical. A training corpus does not sit inside a model like cargo in a hold, waiting to be off-loaded when a lawyer sends a letter. It dissolves into the weights, and the statistical influence of the material stays behind in the model itself, carried along into whatever gets built from it. You can shred the source files with genuine ceremony and a signed affidavit. The influence stays.

Which is why the most instructive AI story right now is not a model launch or a leaderboard shuffle. It is a lawsuit about lineage. And it describes, with unnerving precision, how models are actually built.

The lawsuit that chases lineage, not files

Universal and Sony have brought a second lawsuit against Suno. Before I go one sentence further, the disclaimer, and it belongs in the argument rather than the footer: I am not a lawyer, and I take no position on the merits of the case. Nobody is paying me to offer one. I am a technologist who has spent his career architecting large data systems, and I am reading this dispute the way an engineer reads a postmortem.

Pascal Hetzscholdt, who analyzed the filing, calls its most novel and uncertain argument the claim that allegedly infringing model lineage can carry forward into newer models such as v6. In his summary, the suit moves beyond proving that copyrighted works entered training data. It traces the full chain from acquisition and stream-ripping through training, retention, model distillation, successor models, and downstream market harm. Among its strongest features he counts the forensic identification of works and Suno's own admissions. He also points to evidence of a functioning licensing market and a more concrete market-harm theory as additional strengths.

But one line in his analysis deserves to be laminated and taped to every AI executive's monitor. A new model, he writes, may not really be clean if knowledge, outputs, preference data, or synthetic training material derived from the earlier model are used to build it. Forget who wins for a moment. As a plain description of how modern models are made, that sentence is simply accurate, and its accuracy should worry the industry more than any single verdict.

Why the delete key never reaches the weights

When a company faces a data dispute, the instinct is a file operation. Find the offending corpus, delete it, issue a statement, treat the matter as closed. That instinct was inherited from a world where data and its effects lived in the same place. Databases work that way. Models do not.

Training does not store your data; it stores what your data did. Disputed data can nudge weights, and nudged weights can be saved as checkpoints, branched into fine-tunes, compressed into distilled students, and used to score preference data and spin up synthetic examples for successors. Deleting the source is like pulling the recipe out of the drawer after the cake has been baked, sliced, and folded into three other cakes. The recipe is gone. Everything it caused is still on the table.

The standard rebuttal is that you can retrain from scratch on clean data. In principle, sure. In practice, almost nobody starts truly from scratch, because the previous model's outputs and synthetic corpora are the cheapest training material in the building, and teams under deadline pressure reach for exactly those artifacts. That is the loop a lineage argument points at. If your v6 was raised on the exhaust of your v4, deleting the files v4 ate does not make v6 clean, and claiming otherwise is bookkeeping fraud with extra steps.

Treat every deletion as a versioned migration

So the correct mental model is not rm and done. It is a schema migration on a production database, the kind you version and rehearse, except the schema here is the statistical shape of your models. Four things have to happen before anyone gets to say the word remediated.

First, preserve a cryptographic manifest. Hash every disputed artifact before it leaves, and record what was removed, when, by whom, and under what authority. It sounds backward to keep evidence of the thing you are deleting, but a deletion you cannot prove is a deletion nobody has to believe.

Second, map the descendants. That means every checkpoint trained while the data was present, every fine-tune branched from those checkpoints, every distilled student, every synthetic dataset those models ever generated. This family tree either already exists in your MLOps records or it must be reconstructed, and reconstructing lineage after the subpoena arrives is exactly as much fun as it sounds.

Third, quarantine the affected models. Not delete, quarantine: pull them from serving paths, freeze them, label them, and above all stop them from exhaling new synthetic data into your pipeline while the review is underway. An implicated model that keeps feeding its successors is a contamination pump you are paying to run.

Fourth, rerun targeted evaluations. Probe the surviving and rebuilt models for regurgitation of the removed material and document the results in a form a skeptic could audit. Evaluation is the difference between we deleted the files and we removed the influence, and only the second claim means anything.

The part where i admit why i obsess over this

I did not arrive at this obsession through legal theory. I arrived at it by operating multi-model systems, where lineage problems multiply instead of simplifying. Coheria, the system I am building now, is a collaborative multi-model AI operating system, and working across multiple models means carrying multiple lineages at once, which forces you to take the bookkeeping seriously from day one. Any system in that position needs its decisions recorded somewhere durable, so that when someone asks what a model produced and what informed it, the answer is an evidence trail rather than a shrug. And wherever an answer can be checked deterministically, it should be, so answers get verified rather than vibed.

Routing across model families is worth recommending here too, because depending on a single vendor or gated model means one supplier's data dispute can quietly contaminate everything you do. None of this eliminates model risk, and I will not pretend it does; careful lineage bookkeeping reduces exposure, it does not perform miracles. But any team that ever has to execute the migration described above will want the manifest and the family tree already in hand rather than reconstructed under pressure. And any component that learns a customer's business over time is only trustworthy if you can prove what fed it.

Deletion is a promise, not a keystroke

The industry has developed a taste for remediation theater: a deleted bucket, a solemn statement, a promise to do better next time. That performance survives exactly until someone with discovery power asks for the lineage. The complaint Hetzscholdt summarizes traces the chain from acquisition through successor models end to end, which means at least one set of plaintiffs already knows how to ask. The next set will have a template.

If you run models in production, try the exercise this week. Pick one dataset you could imagine being disputed and answer four questions: what exactly did we have, which models touched it, which models descend from those, and what evidence would convince a skeptic that its influence is gone. If assembling those answers takes more than a day, you do not have a deletion capability. You have a keyboard.

A delete key is not a remediation plan.