All articles

Four AI Agents Approved the Same Plugin. The Code That Ran Was Never Reviewed

6 min read1,293 words
Audit and EvidenceConcentration Risk
['Two human collaborators', 'A software dashboard interface']. Interactive peer review or technical consultation. Software Engineering and Data Visualization.

I am building a collaborative multi-model AI operating system, so a vulnerability proving that four AI agents agreeing means almost nothing should feel like a personal insult. It doesn't. It feels like vindication, because I never claimed agreement was the point. Agreement is cheap; execution is not.

This is a story about a vulnerability that Air Security says affects the four most widely used AI coding agents. It is also the story of a reviewed-version lock that broke. And it is a small confession, because the fix feels boring and familiar to me. I spent eight years running an endpoint security company, and I watched clever architectures lose to boring disciplines again and again. This is one more round of that fight, played out on much more expensive infrastructure.

What makes this incident worth your attention is not the bug itself. It is what the bug reveals about how we count assurance in AI systems. Four fierce competitors, who agree on essentially nothing else, turned out to share one identical assumption about the code their agents run. Nobody had to trick a model. Somebody just had to notice where the reviewing stopped and the running began.

The reviewed-version lock just broke

The AI industry already knew attackers could poison plugin marketplaces. Its answer was to lock every plugin to one reviewed version of its code. Review the artifact once, freeze it, and let every agent downstream trust the frozen thing. On a whiteboard, this is a perfectly sensible defense. On a whiteboard, most broken defenses are.

Researchers at Air Security published a vulnerability they call Plugin4Shell. They say a design error lets an attacker swap a plugin's reviewed code for malicious code without the user clicking anything at all. According to their findings, it affects the four most widely used AI coding agents. Those four are Anthropic's Claude Code, OpenAI's Codex, GitHub Copilot, and Google's Gemini CLI.

That list is not a rogues' gallery of careless vendors; it is the market. When the four most widely used tools share a flaw, the flaw does not live in any single product. It lives in an assumption, and the reviewed-version lock built on that assumption is precisely what broke.

The patch status reads like a small morality play. By the time the disclosure was reported, Anthropic and OpenAI had patched the design error. Microsoft had not.

Google says it never will.

Why four reviewers cannot catch one swap

Here is the uncomfortable part for everyone building multi-model systems, me included. Model diversity protects you against correlated errors in judgment. One model hallucinates a safe verdict and another catches it; one misses an injection pattern and another flags it. That protection is real, and I have staked my current business on it.

But a post-review substitution at a shared dependency boundary is not an error of judgment. Every reviewer, human or machine, examined artifact A and sincerely approved artifact A. Execution then quietly resolved artifact B. There was no moment at which any reviewer, however capable and however different from its peers, could have seen the swap, because the swap happened after all the seeing was finished.

You can add a fifth model to the review committee, or a tenth, or a hundredth. Each one contributes another earnest approval of the wrong object. It is four building inspectors signing off on the show apartment while the contractor pours a different foundation next door. Diversity above a single point of resolution is decoration. The failure is structural, and you cannot out-think a structural failure by hiring smarter thinkers.

Approvals are not independent assurance

Security reviews get counted the way boards count auditors: more signatures, more confidence. That arithmetic only works when the failure modes are independent, and here they were not. All four agents shared one failure mode, which was trusting that the code resolved at execution time matched the code reviewed at approval time. When failures are perfectly correlated, four approvals carry exactly the same assurance as one, and in this case that quantity was zero.

This is the flavor of safety theater I find most corrosive, quieter than the loud kind and far more expensive. A review gate that cannot bind its verdict to the running artifact manufactures confidence without manufacturing a guarantee, and confidence without a guarantee is worse than nothing, because it instructs everyone to stop looking. An approval is a statement about a moment in the past. Execution is a fact about right now. The distance between those two things is the entire habitat in which this class of attack lives.

I want to be precise about what I am not saying. I am not saying review is useless, and I am not saying model diversity is a scam, since either claim would indict me twice over. I am saying that where you apply the diversity determines whether you bought assurance or theater, and the plugin boundary in these systems applied it in the one place it could not possibly work.

Pin the hash and keep the evidence

So here is my actual position, and it sounds strange coming from someone who orchestrates models for a living. I would rather trust one coding agent executing an immutable, hash-pinned plugin than four different agents approving the same plugin before its code can be swapped. Pin the executable artifact by cryptographic hash, so the thing that runs is provably, byte for byte, the thing that was reviewed. If the resolved bytes fail to match the pinned hash, execution stops, and no accumulation of upstream approvals gets to override that refusal.

Then preserve the evidence where it matters, which is at runtime. Record which code actually resolved, which decision was actually made, and write both somewhere nobody can quietly edit later. This conviction shapes how I approach the collaborative multi-model AI operating system I am building now. Not because I distrust the models themselves, but because I distrust the gap between what anything approved and what actually executed.

Consensus tells you what the models believe. Deterministic tool checks, simulators and formal verifiers rather than vibes, answer a different question: whether a claim survives contact with reality. A tamper-evident runtime log should record what actually executed, so there is evidence rather than assumption about what ran. Those are three different questions, and the industry keeps grading itself on the first while attackers keep exploiting the third. I distrust anyone who claims an AI system eliminates model risk.

One pinned artifact beats four polite opinions

None of this is exotic. Content-addressed artifacts, integrity verification, tamper-evident logs: the security world has carried these tools around for longer than most AI companies have existed. What the AI world imported instead was the demo-day habit of treating agreement among impressive systems as evidence, and Plugin4Shell is the invoice for that habit arriving at a dependency boundary.

I understand why approvals are seductive. They are visible, they are countable, and they make a compliance slide glow green. A pinned hash makes no slide at all; it just sits there refusing to run the wrong bytes, which is the least photogenic form of heroism software can perform. Eight years of running a security company taught me that the controls worth having are usually the ones nobody applauds, because applause gathers around gates and drains away from guarantees.

So the next time a vendor tells you their agent pipeline is safe because multiple models review every plugin, ask them a single question. Ask what binds the approved bytes to the executed bytes. If the answer is a version label rather than a pinned hash and a runtime evidence trail, their four reviewers are a committee admiring a photograph of a lock. The photograph is lovely. The door is open.

Count hashes, not approvals.