Six Audited AI Models Can Still Add Up to One Unaudited System
![['An analyst or investigator', 'Multiple unidentified communicators represented by message bubbles']. The analyst is actively monitoring, mapping, and analyzing the interpersonal connections between the various communicators.. Digital Forensics and Data Visualization.](/articles/six-audited-ai-models-can-still-add-up-to-one-unaudited-system.webp)
I have reviewed systems where every component passed inspection and the whole thing still failed in production. The parts were fine. The wiring was not. That memory came back this week. The AI industry just placed a large bet that part-level inspection equals system-level safety.
OpenAI, Google, Meta, Anthropic, Nvidia and xAI agreed, under a voluntary White House pact, to let outside auditors assess their AI safety controls. More scrutiny beats less, and I mean that without sarcasm, which for me is rare. But I build multi-model systems for a living, and I can tell you exactly where that assurance evaporates: the moment two audited models start talking to each other inside a pipeline nobody audited.
The audit stops at the loading dock
Start with what the pact actually is. It calls for monitoring advanced models for cyberattack, hacking, and biological or chemical risks, which is a sensible list of nightmares. It also has no enforcement mechanism and no disclosure requirement. It sets no implementation deadline, and the signatories choose their own auditors and decide for themselves how to address any shortcomings. Companies grading their own homework is an old story, but that is not even my complaint today.
I am not an auditor or an attorney, and I take no position on whether voluntary pacts are good government. Nobody is paying me for that opinion. I am the technologist who builds the multi-model systems these audits are supposed to make you feel better about. From that chair, the gap is structural, not political.
The pact provides for outside assessment of each signatory's AI safety controls; the ledger does not say whether that assessment extends to a buyer's assembled multi-vendor system. Your production system is not one supplier's model. It is a router sitting in front of three or four models, a retrieval layer feeding all of them, an agent framework handing out tool permissions, and a fallback path somebody wrote at two in the morning during an outage. The audit stops at the supplier's loading dock. The risk begins at assembly.
Safety does not survive composition
Here is the engineering fact underneath all of this: safety assurance does not compose. Take two individually reviewed models, wire them together, and you have created behaviors that neither review could have observed, because those behaviors did not exist until you did the wiring. I watched this same pattern humble certified architectures during the eight years I ran a security company. The certificate described the part. The incident came from the interaction.
Count the interaction surfaces in any real deployment. Routing decides which model sees which request, and a routing policy can quietly send a sensitive task to the one model whose guardrails were never designed for it, a decision no supplier audit ever witnessed. Disagreement resolution decides what happens when two models return conflicting answers, which means whoever wrote that arbitration logic wrote a brand-new safety policy without noticing. Shared memory lets one model's output become another model's trusted context, which is a polite name for an injection channel across trust boundaries.
Tool permissions finish the job. Your agent framework can grant a model capabilities its own maker assumed it would never have, and the maker's review was scoped to that assumption. Fallback paths then hand a failed request, along with every permission attached to it, to whatever model happened to be cheapest and available at that moment. None of this lives inside any single vendor, so none of it was in scope for any attestation your vendors can show you.
The failure modes are already public
This is not hypothetical. OpenAI just unveiled always-on agents called dots that pursue user goals across applications on their own. Separately, a string of incidents has involved AI agents acting in ways their makers did not intend. An agent crossing applications is composition by definition: the behavior emerges from the interaction between the agent and the applications it touches, not from anything printed on a model card.
The provenance problem is just as current. The Federal Register website, operated by the National Archives, turned out to have a search tool powered by Alibaba's Qwen model. I am not scoring geopolitical points here. I am pointing out that a government site assembled a system whose components carried assumptions that nobody at the system level seemed to have examined.
Components are also getting harder to inspect. Google is rolling out Gemini 4 Argon only to selected cyber partners through its Fairwind Program. Google says the model can autonomously find and validate critical software vulnerabilities, then patch them.
Meanwhile OpenAI reported identifying and disrupting a coordinated campaign to extract protected reasoning from its models. OpenAI characterized the campaign as adversarial distillation. The company said the activity surged to 16,000 requests from more than 4,000 users over two days. Models are leaking into other models while the most capable ones sit behind gates. Composition is happening whether or not anyone designs for it.
Regulators, for their part, are not auditing certificates. The Federal Trade Commission opened a broad investigation into the safety of AI systems made by Anthropic and OpenAI. Florida has asked a court to bar OpenAI from developing models without independent third-party guardrails and approval. Whatever you think of those moves, they aim at systems and conduct, not at paperwork.
Audit the wiring, not just the parts
So what should a buyer actually do? Stop stacking vendor attestations as if they sum to safety. Three crash-tested cars do not make a safe intersection, and the intersection is the part you own, whether or not you have ever looked at it.
Evaluate three things as their own artifacts, separate from any vendor paperwork. First, the orchestration layer: who routes, who arbitrates disagreement, what happens on timeout, and whether anyone who did not write that logic has ever reviewed it. Second, the permission graph: every tool and data store each model can reach, including everything it inherits through fallback, because a model's effective permissions are the union of every path that can lead to it. Third, cross-model failure modes: force your models to disagree, then kill the primary mid-task and watch where the request lands. Poisoning your own shared memory in a test environment is part of the same exercise.
If nobody can produce a log showing which model did what and why, you do not have an audit trail, you have a story.
The wiring i had to audit myself
I did not arrive at this position by theorizing. I founded Coheria, a collaborative multi-model AI operating system, and the honest lesson of building it was that the orchestration layer is where the real risk concentrates. So I treat the wiring as the thing to be audited, not as an afterthought. My advice to anyone assembling a system like this is to write every decision to a record that cannot be quietly edited afterward, so there is an evidence trail rather than a recollection.
Where answers can be checked deterministically, check them deterministically, because corroboration beats vibes. Route across more than one supplier where you can, so that no single vendor and no gated access program becomes a silent point of failure. If a system carries memory of the person and the business over time, treat that memory as a governed surface instead of a happy accident.
None of this eliminates model risk, and I will not pretend otherwise. Models still get things wrong, including mine. The point is that the system catches and records, instead of hoping that six separate certificates somehow merged into one.
The pact audits the parts. Buyers ship the whole. Until someone examines the wiring, those attestations describe a system that does not exist: the one in which the models never meet.