ScrutinEyes · 2026-07-26
Found, Not Exploited: What Anthropic's Mythos Actually Did to Classified Systems
A red-team result became a national-security incident, then a three-week policy whipsaw. The gap between how the event was described and what was actually confirmed is the whole story.
Inside a classified U.S. government red-team exercise earlier this year, an AI model was pointed at highly sensitive systems and asked to do what red teams do: find the holes. Within hours, it surfaced vulnerabilities. The model was Anthropic’s Mythos. What followed — a senator’s alarm, an anonymous confirmation, an export ban, an industry revolt, and two partial reversals inside three weeks — is a preview of nearly every hard argument the next few years of AI security will have.
What is actually confirmed
Strip the story to what a named or on-record source has stood behind, and it is narrower than the headlines.
An unnamed U.S. official told the Associated Press that Mythos “identified vulnerabilities in highly sensitive and secure U.S. government computer systems during a testing exercise.” The same official added the sentence most coverage buried: the model “identified certain vulnerabilities within hours, but that does not mean the model was able to exploit them within that time.”
Finding a vulnerability and exploiting it are different acts, separated by most of the difficulty. A scanner that flags an open port has “found” something; whether that leads anywhere is another question entirely. The confirmed claim is that Mythos was fast at discovery. It is not that Mythos walked through the front door.
The testing ran under Project Glasswing, an Anthropic initiative that brings together intelligence agencies and companies with the stated aim of hardening critical software against exactly the kind of risk advanced models pose. The model finding these holes was the purpose of the exercise, not an accident of it.
The two descriptions
The reason this story traveled is not the caveated official statement. It is a second, louder account.
Around June 11, Senator Mark Warner described the same testing in blunt terms, attributing the characterization to Gen. Joshua Rudd — the head of the NSA and U.S. Cyber Command: the tool “broke into almost all of our classified systems, not in weeks but in hours.”
“Broke into almost all of our classified systems” and “identified certain vulnerabilities, though not necessarily exploitable ones” are not the same claim. They may describe the same event. One is a secondhand characterization relayed in a political setting; the other is a firsthand-sourced statement with an explicit limit built in. Both the NSA and Anthropic declined to comment — which means the most dramatic version of events is also the least directly sourced.
For a security-literate reader, that gap is the actual news. This is how a red-team finding becomes a breach in the retelling — not through anyone lying, but through the ordinary compression that happens when a technical result crosses into a hearing room. Track which verb a source uses. “Found” and “broke into” are worlds apart, and the distance between them is where the real risk assessment lives.
The policy whipsaw
Whatever the model did, the government’s reaction was fast, broad, and then repeatedly reversed.
On or around June 12, the administration imposed an export ban barring non-U.S. nationals from accessing both Mythos 5 and its smaller sibling, Fable 5. The next day, Anthropic suspended access to both models for all users to comply — and said plainly that it “did not believe the steps taken by the government were warranted.” A company arguing that a restriction on its own most powerful products was an overreaction is not a common sight.
The counter-pressure was equally notable. More than 100 cybersecurity executives urged the administration to ease the restrictions, and their argument cut at the core premise: the Mythos models are “quite good” at finding flaws, they said, but “not uniquely good at these tasks” compared with other AI systems already in circulation. If that assessment holds, the ban’s logic weakens. Restricting one vendor’s model does little if comparable capability exists elsewhere; it mostly relocates the risk while imposing the cost on the single company that chose to run the test in the open.
The reversals came quickly. In a letter dated Friday, June 26, Commerce Secretary Howard Lutnick wrote that he had “determined that appropriate safeguards are in place to permit certain trusted partners to access the Claude Mythos 5 Model” — a limited release to select companies and agencies, described as cyber defenders and infrastructure providers. That letter pointedly did not authorize Fable. Then the reversal completed. On June 30, the White House dropped the export controls on both models — Mythos 5 and Fable 5 — outright, a lifting reported across the Washington Post, CNN, CNBC, and TechCrunch and confirmed by Anthropic.
Chatham House’s read on the sequence is worth borrowing: the volatility itself is the problem. Ban, suspend, partially permit, then lift entirely — inside three weeks — “sends confusing signals to markets and is bad for investors,” and, more importantly for security, reflects an unresolved tangle of anxieties about foreign access, uncertainty about what the models can actually do, and distrust of the partnerships needed to find out.
One boundary worth drawing
The Mythos export saga is not the only front between Anthropic and the national-security state, and it should not be read as the whole picture. A separate, larger, and still-unresolved dispute predates it. In late February 2026, after Anthropic declined a Pentagon demand for unrestricted use of Claude — the company kept guardrails against autonomous-weapons applications and mass surveillance — the administration ordered federal agencies and contractors to cease business with Anthropic and designated the company a “supply chain risk.” Anthropic sued in two courts; a federal judge blocked the business ban with a preliminary injunction in March, while a D.C. Circuit panel left the supply-chain designation standing in April. Both cases are still live. That fight is about procurement and use policy, not model capability — a different argument entirely — and anyone assessing the current state of Anthropic and the agencies has to hold both threads at once. This piece covers only the capability one.
What it actually represents
The specifics of Mythos will age. The structure of the fight will not.
First: offensive capability is now a side effect of general capability. A model good enough to be useful for defense is, by construction, good enough to find flaws — and “find” sits one short step from “exploit.” There is no version of a frontier model that is powerful for defenders and inert for attackers. That is not a policy choice anyone gets to make; it is a property of the technology.
Second: the “not uniquely good” argument is the one that will decide policy, and it is genuinely hard. If several models can do a dangerous thing, restricting one is theater. If only one can, restricting it is defensible but temporary, because capability diffuses. Regulators will keep landing on this question, and the honest answer keeps moving — which is roughly what a three-week, four-position policy scramble looks like from the outside.
Third: the incentive lesson is uncomfortable, with one caveat that keeps it honest. Anthropic ran a test that surfaced real government vulnerabilities, disclosed the capability, and was restricted for it — while quieter labs drew no comparable scrutiny. Penalizing the party that tested in the open teaches every other party not to, which is backwards from what a functioning disclosure culture needs. The caveat: as noted above, the relationship was already adversarial for unrelated reasons, so the Mythos restrictions are not a clean case of punishing disclosure. The incentive still bends the wrong way — it is just not the only force acting on it.
The takeaway for anyone defending systems is narrow and immediate: tools that find flaws in hours already exist, on both sides of the line, and the discovery half of the attack is getting cheap fast. Whether Mythos could exploit what it found is the question of the moment. Whether the next model can is the question of the year.
Sources
Anthropic’s Mythos model found vulnerabilities in classified U.S. government systems, official says — SecurityWeek: https://www.securityweek.com/anthropics-mythos-model-found-vulnerabilities-in-classified-us-government-systems-official-says/
Anthropic’s Mythos model found vulnerabilities in classified US government systems, official says — Federal News Network (Associated Press): https://federalnewsnetwork.com/artificial-intelligence/2026/06/anthropics-mythos-model-found-vulnerabilities-in-classified-us-government-systems-official-says/
US Government Allows Anthropic Limited Release of ‘Mythos’ AI Model, Saying ‘Appropriate Safeguards Are in Place’ — Slashdot (quoting Commerce Secretary Howard Lutnick’s June 26 letter): https://news.slashdot.org/story/26/06/27/0159230/us-government-allows-anthropic-limited-release-of-mythos-ai-model-saying-appropriate-safeguards-are-in-place
The US government’s latest U-turn on Anthropic’s Mythos sends mixed signals on AI governance — Chatham House: https://www.chathamhouse.org/2026/07/us-governments-latest-u-turn-anthropics-mythos-sends-mixed-signals-ai-governance
On the separate Pentagon dispute referenced in “One boundary worth drawing”:
Trump administration orders military contractors and federal agencies to cease business with Anthropic — CNN: https://www.cnn.com/2026/02/27/tech/anthropic-pentagon-deadline
Pentagon-Anthropic Dispute over Autonomous Weapon Systems — Congressional Research Service: https://www.congress.gov/crs-product/IN12669
A Timeline of the Anthropic-Pentagon Dispute — TechPolicy.Press: https://www.techpolicy.press/a-timeline-of-the-anthropic-pentagon-dispute/
Note on sourcing and confidence: The narrow, caveated claim — vulnerabilities identified within hours, exploitation not established — comes from an unnamed U.S. official speaking to the Associated Press, and is the version this piece treats as load-bearing. The broader “broke into almost all” characterization originates with Sen. Mark Warner relaying remarks he attributed to NSA/Cyber Command leadership; both the NSA and Anthropic declined to comment. The June 26 limited release is well corroborated across outlets and quotes the Commerce Secretary’s letter directly. The full lifting of controls on both models on June 30 is well corroborated across the Washington Post, CNN, CNBC, and TechCrunch. Chatham House’s framing of the three-week sequence as destabilizing volatility is one analytic read, attributed as such.
Reading this because someone's asking about your security? See exactly which rules apply to you and where you stand — check your readiness free. Five minutes, in your browser, nothing stored unless you ask. Readiness, not legal advice.