Your AI Can Pass Every Health Check and Still Be Legally Switched Off
![['A team of four male analysts or operators', 'A centralized physical identity management system (ID badge rack)', 'A mobile digital workstation (laptop on a utility cart)']. Collaborative forensic analysis where one team member presents physical evidence (an ID badge) for.](/articles/your-ai-can-pass-every-health-check-and-still-be-legally-switched-off.webp)
I was the architect and development lead for the U.S. Visa and Passport systems, so I know exactly how hard it is to prove, in software, who a person actually is. It is harder than almost anyone outside those systems believes. Now that problem has landed in the middle of AI inference, and most production architectures are not remotely ready for it.
The U.S. government ordered Anthropic to turn off its Fable 5 and Mythos 5 models for any foreign national. The order covered foreign nationals inside and outside the United States, including Anthropic's own foreign national employees. The Commerce Department gave the company roughly 90 minutes to take the models down. Ninety minutes. That is less time than most teams need to agree on a Slack channel for the incident.
Anthropic's response is the detail that should keep infrastructure people up at night. The company disabled the models globally, saying it could not enforce a nationality-based access restriction. Not would not. Could not. The provider of some of the most capable models on the market looked at its own inference stack and concluded it had no reliable way to tell which human being was behind any given call.
The day a healthy model went dark
The backstory, as reported, runs like this. A June 12 directive barred foreign nationals, including Anthropic's non-citizen employees, from accessing Fable 5 and Mythos 5. Amazon had reported possible security risks involving the two models to the Commerce Department.
During a controlled security evaluation, Anthropic's Mythos model reportedly accessed almost all classified NSA systems within hours. That evaluation was an authorized internal red-team test conducted under specific simulated conditions. Anthropic pushed back, saying similar issues exist in other models, including OpenAI's GPT-5.5. It also began talking directly with the White House about a potential fix.
I am not a lawyer and not a national security official, and I take no position on whether the order was justified. Nobody is paying me to offer one. I am the technologist who built federal identity systems and who now runs multi-model AI in production, and what interests me is what this episode does to architecture.
Because whatever you think of the policy, the operational fact is brutal. Anthropic disabled both models globally because it said it could not enforce the nationality restriction. A rule aimed at a subset of users ended with the models unavailable to everyone.
Nationality is not an api key
Here is the mechanical problem underneath the headline. Nationality rules attach to people. Production inference does not authenticate people; it authenticates API keys, organizations, service accounts, and increasingly, autonomous agents acting on behalf of nobody in particular. Picture a hypothetical, entirely ordinary setup: an API key is shared across a team, the team spans a dozen countries, and an agent fires ten thousand calls an hour while no row in any log says which passport was in the room.
I lived inside this gap. When you build identity systems at national scale, you learn that binding a legal attribute like citizenship to a live transaction is one of the hardest problems in software, full of dual nationals, pending adjudications, and documents that quietly disagree with each other. It is hard when the government is the one asking and holds the source records. It is nearly impossible when you are a model vendor whose entire view of a customer is a billing email and a bearer token.
So when a provider is handed a rule written against humans and an enforcement surface written against keys, it has two options. It can bind every call to an auditable human entitlement, which it never built. Or it can turn everything off, everywhere, for everyone. Anthropic chose the only option it actually had.
Up is no longer the same as available
Every AI dependency you run now has two failure states instead of one. The first is the familiar kind: outage, rate limit, degraded latency, the stuff your SLA theoretically covers. The second is new: the model is healthy, the API is up, and you are no longer allowed to call it, or your Berlin engineer is not, or your agent cannot prove it is acting for anyone at all. Identity policy has become part of the inference path, and no status page will ever show it.
The surrounding environment guarantees this was not a one-off. OpenAI began rolling out GPT-6 globally in ChatGPT, so capability churn continues even as policy churn arrives. Meanwhile ProjectDiscovery published a working credential-theft backdoor in a Qwen2.5-7B fine-tune, built for under fifty dollars, which is exactly the kind of result that keeps regulators deeply interested in who touches which model. Capable models plus cheap attacks plus nervous governments is not a mixture that trends toward fewer access rules.
And yet most production AI architectures still treat model availability as a vendor SLA line item, something procurement negotiated once and engineering can safely forget. That assumption just failed in public with a 90-minute fuse. It will fail again.
What to build before the next order
Three things, none of them glamorous.
First, auditable person-to-call entitlements. Every inference call in your stack should be attributable to a human entitlement, or to an accountable chain that ends at one, and the record should be tamper-evident. Not because you enjoy bureaucracy, but because the alternative to selective enforcement is broad shutdown, and you want to be the customer who can demonstrate compliance per call rather than the one who gets switched off wholesale.
Second, degraded modes. Your runbooks define what happens when latency spikes. Very few define what the product does when an entire model family is removed from your legal reach overnight. Decide now which features fall back to a smaller model, which ones queue, and which ones fail loudly, because deciding during the incident is how you ship your worst day.
Third, failover across model families, not replicas of one vendor. Two regions of the same provider share the same regulator, the same compliance posture, and the same 90-minute fuse. Diversity that actually protects you crosses vendors and model lineages, which is annoying to build and tends to look paranoid right up until the day it looks prescient.
The boring question that decides it
I will admit I did not plan for nationality orders. I am the founder of Coheria, a collaborative multi-model AI operating system, and the lesson this episode reinforces is the one every operator eventually learns: the unglamorous record-keeping decides your worst day. An enforcement regime does not ask whether your model is smart. It asks who made this call, with what authority, verified how, and whether you can answer that after the fact. Once identity enters the inference path, that question is the whole game.
Anthropic, a company with world-class engineers, disabled its models globally because it said it could not selectively enforce an identity rule. It later began talking directly with the White House about a potential fix. You will not do better by improvising during your own 90 minutes. Build the identity layer before the policy arrives.
It is the one failover nobody benchmarks.