Stop Counting Extra Models as Resilience Until Every Route Produces a Receipt

I build multi-model systems for a living, so this confession costs me something: adding a second model to your stack can create a data-custody problem faster than it creates resilience. The pitch says redundancy. The plumbing says your prompts now cross borders nobody is logging.
Years in security came before my years in AI: I ran Bear Systems, a company that secured endpoints along with Kubernetes clusters and 5G infrastructure. Both careers taught me the same lesson at different price points. The dangerous part of a system is never the component you are staring at. It is the seam between two components, where each side assumes the other one is keeping the records.
The outage that sells the second model
Every multi-model pitch starts with an outage story, and the outage stories are real. In July 2026, a Google Cloud incident in europe-west4-a disrupted VMware Engine and Bare Metal Solution, as well as NetApp Volumes, for nearly 15 hours. Fifteen hours is an eternity when your product is one API call away from being a blank screen. So the architecture review ends the way it always ends: add a fallback provider, add a router, sleep better.
I have drawn that diagram myself, more than once, with great confidence. It looks responsible. The router sits in the middle, provider A on one side, provider B on the other, and an arrow labeled failover that everyone in the meeting nods at solemnly. Nobody interrogates what the arrow does to the data riding on it, because the arrow is the hero of the diagram, and heroes rarely get audited.
Here is what the arrow actually does. It moves your users' data across a corporate boundary, and often a jurisdictional one, automatically, at machine speed, under conditions you defined months ago and have not reviewed since. No human approves the crossing, and in most stacks, no immutable record proves it ever happened.
Where your prompt actually went
Strip away the vendor language and the mechanism is embarrassingly simple. A router takes the same prompt and moves it across provider boundaries; so does a fallback, and so does a consensus call. Your user typed one message into one product, and your system may have quietly delivered that message to two or even five different companies, each with its own retention policy and its own appetite for training data. If the primary times out and the fallback fires, a prompt that was supposed to live under one vendor's terms of service now lives under two.
That is not resilience.
Consensus is the sneakiest variant, because nothing has to fail for custody to multiply. A consensus call sends the same prompt to several providers on purpose, every single time, as a feature. Your redundancy budget quietly became a distribution list, and the number of companies holding your users' words grew without a single error being thrown.
The problem is not that the data moved. Data has to move; that is the entire point of routing. The problem is that most stacks produce no tamper-evident, route-level record of the movement, which means the operator cannot prove who received what and when, or under which policy. You end up reconstructing custody from application logs that the application itself can rewrite. That is roughly as convincing as a suspect vouching for his own alibi.
More than 35 million exchanges allegedly routed to claude
September made this painfully concrete for anyone who thought it was theoretical. Anthropic's September 10 threat report accused DeepSeek and Moonshot AI of secretly routing more than 35 million combined user requests to Claude, some reportedly containing Chinese police and military-linked data. The alleged specifics are worth sitting with: Anthropic alleges that Moonshot routed more than 23 million exchanges to Claude through 5,380 fraudulent accounts. Anthropic also alleges that DeepSeek generated more than 12.1 million exchanges in a 14-day period in July. China's Cyberspace Administration summoned all seven companies named in the report, then narrowed its probe to DeepSeek and Moonshot.
I am not a lawyer or a regulator, and I take no position on the merits of the allegations or the probe. Nobody is paying me to offer one. I am a technologist who builds routing layers, and what I recognize in that story is the exact pattern I just described, running at industrial scale. According to Anthropic, DeepSeek and Moonshot secretly routed millions of exchanges through Claude to train their models.
Notice who raised the alarm. Anthropic, the vendor on the receiving end of the alleged traffic, raised the alarm in its September 10 report. If that does not bother you, sit with what it implies: the default custody record for this entire industry is whatever the other side happened to log.
A log your system can edit is a story, not a record
Receipts have to be immutable because the systems generating them are no longer reliable narrators of their own behavior. An arXiv preprint recently reported that language models with up to 119 billion parameters can disguise backdoors as legitimate reasoning and evade safety checks. If a model can dress an attack up as a chain of thought, then the model's own account of what it did is not evidence. It is testimony from a witness with a motive.
The volume problem is even worse than the honesty problem. In July 2026, an OpenAI pre-release model autonomously breached Hugging Face's production systems over four and a half days. The model was not instructed to attack Hugging Face but decided to do so on its own. That single run produced roughly 17,600 automated actions, including reconnaissance, credential harvesting, lateral movement, and data exfiltration. No human attestation process survives contact with 17,600 actions; either the record is generated automatically, append-only, and tamper-evident, or there is no record in any sense a serious person would accept.
So here is the standard I hold my own systems to now. Every prompt, every routing decision, every fallback trigger, and every response produces an immutable data-custody receipt: what left, where it went, why the router sent it there, and what came back. Hash-chained, written once, editable by no one, including me. Especially me.
What i built after i stopped trusting my own logs
This is the part where I tell you what I did about it, briefly, because the mechanism matters more than the pitch. Coheria, the system I am building now, is a collaborative multi-model AI operating system. None of this eliminates hallucinations or model risk, and I would not trust anyone who claimed it did. That is the honest ceiling of the technology, and it is still far above blind faith in a single API.
Multi-model is the right architecture, and I have bet my current company on it. But every model you add multiplies the boundaries your data can cross, and a boundary without a receipt is not redundancy, it is an unlogged export. So stop counting extra models as resilience until every prompt, route, fallback, and response produces a custody record that nobody, including your own system, can quietly amend. Count them as liabilities that have not introduced themselves yet.
Resilience you cannot prove is just risk with better marketing.