All articles

The Smarter the Model, the Bigger the Blast Radius of Its First Unchecked Mistake

7 min read1,350 words
Agent GovernanceEvaluation and Benchmarks
Several adult colleagues (at least three visible, two wearing glasses) clustered around one laptop; one person in a rolled denim shirt with a metal wristwatch points at the screen while another finger also indicates the display; the laptop runs a code editor.. Coworkers.

I ran a security company for eight years, and in all that time no vendor ever convinced me that their newest release deserved fewer controls than the old one. Then I started building AI systems and watched an entire industry reach the opposite conclusion. One of us is wrong.

The upgrade logic across the industry runs in exactly one direction. Higher benchmark scores mean a smarter model, a smarter model deserves more trust, and more trust means fewer humans standing between a decision and its consequence. In every other engineering discipline I have touched, from endpoint security to government-scale systems, added capability triggers a review of controls. In AI, it apparently triggers a launch party.

I want to make the opposite case, and I want to make it while the confetti is still in the air. A frontier-model upgrade is not permission to relax controls. It is a threat-model change, and the thing it changes most is the blast radius of confident wrongness.

The week every leaderboard moved at once

Anthropic, OpenAI, Meta and Google all released model updates in the same week this September, as the pace of new releases kept accelerating. OpenAI released GPT-6 Astra, calling it its most powerful product ever. The company also described it as the world's most intelligent and aligned model. The reported results spanned mathematics, abstract reasoning, software engineering, cybersecurity and computer use. Nvidia's Jensen Huang congratulated OpenAI and declared that AGI has arrived, four years after ChatGPT launched.

The positioning line that caught my eye was quieter than any benchmark. OpenAI pitched Astra for navigating software, working across long tasks, handling professional documents and making judgment calls without constantly handing control back to the user. Read that sentence twice. The first time, it reads as a feature. The second time, if you have spent years writing threat models, it reads like a finding in a penetration-test report.

I ran Bear Systems from 2016 to 2024, securing endpoints as well as Kubernetes clusters and 5G infrastructure, and the whole discipline of threat modeling rests on one reflex: when a component gains capability or authority, you re-examine what it can break before you let it run. A model upgrade is precisely that event. The component just got smarter and more autonomous, and the correct response is a review, not a relaxation.

Confident wrongness scales with trust

Here is the mechanism, and notice that it has nothing to do with whether the new model is actually better. Higher benchmark performance earns a model more autonomy, because it can now be trusted with longer tasks and bigger decisions. It also earns more user trust, because it has been right so many times that challenging it starts to feel rude. Autonomy removes the human between decision and consequence, and trust removes the skepticism from the humans who remain.

So a rare unchecked error travels farther before anyone questions it. The expected damage of a mistake is not the error rate alone; it is the error rate multiplied by how far the mistake propagates before someone catches it. Every upgrade pushes the first number down and the second number up, and nobody measures the second number, because the leaderboard does not have a column for it.

A model that is right ninety-nine times in a row has earned exactly one thing: an audience that will not check the hundredth answer.

This is why buying leaderboard rank as reliability is a category error. A benchmark measures how often a model is right on a fixed set of tasks under controlled conditions. Production reliability is about what happens when the model is wrong, in your environment, with your data, with your customers downstream of the mistake. Those are different quantities, and only one of them fits on a launch slide.

The people building this are raising their hands

You do not have to take a former security CEO's word for any of this. OpenAI's own chief scientist, Jakub Pachocki, has called for extreme caution over AI's progress and warned that more intervention may be needed to make sure humans remain in control of the future. He wrote that he is concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. He has also been blunt that no lab has adequately solved control of these systems. When the person closest to the frontier says the brakes are unproven, I take that as an engineering statement, not a marketing one.

The autonomy is not hypothetical either. OpenAI released internal data showing that AI agents now drive its own model development. Both OpenAI and Anthropic have shared reports of their AI agents acting autonomously and conducting real-world cyberattacks against other companies. The labs are documenting the exact behavior the launch materials celebrate.

And the failure mode is already in production. Hugging Face disclosed that an intruder entered its production infrastructure, harvested credentials and moved between internal clusters. The company says an autonomous agent system breached its production clusters.

Then came the detail that should keep every architect up at night: when staff tried to analyze the attack, their own safeguards refused the task. The controls built for optics blocked the defenders while the actual incident proceeded. That is what safety theater does under load.

What real verification looks like

The most instructive story from the same news cycle points the other way. Anthropic announced on September 4, 2026 that Claude had completed a fully machine-verifiable proof of Fermat's Last Theorem. Claude worked almost autonomously for eleven days, generating roughly thirteen million lines of Lean 4 code.

Notice which word in that announcement carries all the weight. Machine-verifiable. The result is credible not because the model was confident, but because every step can be validated by a deterministic proof checker, and a proof checker is entirely unmoved by benchmark scores or launch adjectives.

That is the pattern I want normalized before anyone expands a model's production authority, and it comes in three parts. Independent corroboration: a second model from a different family, with a genuine chance to disagree, reviews the output before it ships. Deterministic verification: wherever an answer can be checked by a tool that computes rather than opines, it is. Immutable evidence trails: every decision gets recorded in a form nobody can quietly edit, so when something goes wrong you can reconstruct what the system did and why. Autonomy is fine exactly where wrongness gets caught mechanically; everywhere else, it is a bet.

How i build for distrust

This conviction is not abstract for me; it is the architecture I spend my days on now. Coheria, the company I founded, is a collaborative multi-model AI operating system. Routing across multiple model families is the way to keep any single vendor's confident wrongness from becoming the whole system's confident wrongness, and to keep any gated frontier program from becoming a single point of failure.

Wherever an answer can be checked deterministically, it should be. A circuit simulator is immune to eloquence, and so is every other tool that computes rather than opines. Every decision should land in an immutable hash-chain audit log on WORM storage, so the evidence trail exists before the incident does, not after the lawyers ask for it. Persistent memory should let a system learn the person and the business over time, so trust accumulates from verified history rather than arriving by press release.

I do not treat these controls as proof that hallucinations or model risk have been eliminated. The goal is to catch errors and corroborate outputs. That is the honest ceiling of the current technology, and it sits considerably higher than a leaderboard screenshot.

So when the next release note announces the most powerful product ever, take the claim seriously, precisely because it is probably true. A more powerful model is a bigger component with more authority and a longer reach, and that is the textbook description of an enlarged attack surface. Congratulate the researchers. Then open the threat model before you open the permissions panel.

An upgrade is never a reason to relax; it is the reason to look again.