All articles

Model Diversity Cut My Vendor Risk and Quietly Multiplied My Legal Homework

6 min read1,222 words
Audit and EvidenceConcentration Risk
['Individual analyst (female professional in business attire)', 'Digital network visualization system (hub-and-spoke node diagram on monitor)', 'Physical documentary evidence (tagged envelopes and documents)']. Cross-referential investigation — the analyst is correlating digital.

I spent two years building multi-model routing so that no single AI vendor could hold my product hostage. Last week I reread my own architecture diagram and realized I had not reduced my risk. I had diversified it into a shape my lawyers like even less.

The prompt for the realization was public. Sony and Warner Music sued Anthropic over songs allegedly used in AI training. Reporting on the case separately noted that Anthropic was accused of pirating lyrics by the Beatles and Taylor Swift.

I hold no position on the merits, and nobody is paying me for one. What I do hold is a routing layer that treats model families as interchangeable workers, and a lawsuit reminding me that they are not interchangeable at all. Each one arrives with a training history I did not witness and cannot inspect from the outside.

Why i route across model families at all

My route here was honest, meaning I stepped on the rake first. My early stack trusted one vendor's API the way toddlers trust staircases. Then came the deprecation notices, the silent behavior shifts, the pricing changes, and the waitlists for gated frontier access that operate with all the transparency of a nightclub door policy. Model diversity was the sane response. If one family degrades or vanishes behind an enterprise sales call, the others carry the load.

The logic still holds. I am not walking it back. Routing across families remains the only architecture I trust to survive contact with a vendor's roadmap, because roadmaps are marketing and deprecation schedules are physics. Gated access programs did not get any friendlier while I was writing this.

But diversity has a cost column I kept skimming past. Every model family I add is another organization's training pipeline entering my chain of custody. Another set of data decisions I inherit without ever seeing. I had been counting redundancy as pure upside, which is the kind of accounting that works right up until someone files a complaint.

The lawsuit that itemized the bill

The Sony and Warner suit is worth reading as an engineering document as much as a legal one. Sony Music Publishing and Warner Chappell characterized the alleged conduct as illegally torrenting, scraping, and downloading copyrighted works on a massive scale. Anthropic was sued over the alleged theft of tens of thousands of songs. The publishers pressing the case include Sony, EMI, and Warner Chappell.

That lawsuit alleged that Anthropic used thousands of works protected by the music publishers, including Ain't No Mountain High Enough. Music publishers also alleged that Anthropic's models scraped Mariah Carey's All I Want For Christmas Is You. The named defendants include not just the company but CEO Dario Amodei and co-founder Benjamin Mann.

Context makes it heavier. Anthropic had already paid authors 1.5 billion dollars after admitting to pirating more than 7 million books to train AI. The publishers argue that sum was not large enough to deter infringing conduct by a company they value at 2 trillion dollars. Whether they are right is a question for a court, and courts do not consult me.

What I noticed was smaller and more personal. I route work across multiple model families precisely so no single vendor can sink me, and every one of those families arrived with a training-data lineage that could produce its own version of this filing. Diversity multiplied my provenance surface. That was the deal all along. I just had not read the fine print I wrote myself.

Consensus cannot vote on lawfulness

Here is the uncomfortable technical part. Everything my stack does well operates at inference time. Consensus across models tells me whether outputs agree. Benchmarks tell me leaderboard positions, which mostly measures who tuned hardest for the leaderboard. Deterministic verification tells me whether a circuit simulates or a netlist synthesizes.

All of that is downstream of training, and provenance lives upstream, in decisions made inside someone else's data pipeline years before my API call ever happens. No consensus mechanism can establish that a model's training corpus was lawfully assembled. Ten models agreeing on an answer tells you nothing about where any of them learned it. A perfect benchmark score is fully compatible with a perfectly scandalous corpus.

Deterministic checks catch wrong answers; they are structurally silent on stolen ones. This is not a flaw in those tools. It is a category boundary, and I spent a long time pretending it was not there because the tools on my side of the boundary were the ones I knew how to build. The lawsuit did not teach me this so much as make the tuition public.

Provenance is a routing constraint now

So here is what changed in how I operate, offered as practice rather than prophecy. First, provenance review happens before a model family gets authorized into the routing pool, not after. That means asking what is publicly known about its training data and its litigation exposure. This is judgment work, not a checkbox exercise, and it will sometimes be wrong. It still beats discovering a model's lineage from a headline.

Second, provenance is now a routing constraint, the same class of input as latency or cost. Some workloads can tolerate lineage uncertainty and some cannot. A throwaway brainstorming pass and a customer-facing deliverable should not draw from the pool under identical rules, and the router should be the thing that knows the difference. Latency you can measure in milliseconds; lineage you measure in disclosures and filings.

Third, and this is the one most teams skip: preserve which model families participated in each decision. When a question arrives later, and the current litigation weather suggests it will, saying one of our models produced this is not an answer. You need the record before you need it. Reconstructing it afterward is fiction-writing with timestamps.

What this looks like when you actually build it

This is where I admit the whole essay is also a description of my own homework. Multi-model routing across families means no dependence on any single vendor or any gated model, which was the original point. None of that solves provenance, for exactly the reasons above. What addresses it is the boring part: keeping a record of which model families participated in each decision, so that if a family's lineage later becomes a problem, I can answer which decisions it touched without doing archaeology. That is not a miracle and I will not sell it as one.

The honest framing is that orchestration surfaced the provenance problem earlier than a single-model stack would have. When you route everything to one vendor, their lineage risk is invisible because it is total. When you route across multiple families, the boundaries show up in your own architecture diagram, staring back at you. That is uncomfortable, and I have decided uncomfortable beats blind.

I still believe model diversity is the right call. Dependency on one gated frontier vendor is a worse risk than provenance sprawl, because sprawl can at least be governed. But I have stopped describing diversity as a hedge with no premium. The premium is ongoing per-family provenance work with no end date, and the invoice arrived in the form of a complaint about tens of thousands of songs.

Read the lawsuit. Then read your routing table. Mine had more defendants in it than I expected.