We Recalled the Defenders' Tools. The Attackers Kept Theirs.
The jailbreak that led to Anthropic recalling its most powerful models was a researcher asking the AI to read code and find bugs in it. That is not an attack; that’s every day at the office for a defender. While all the conversation around the US Government changing the export controls on Anthropic’s latest models is important, it’s missing a few key points.
- I think it’s time for us to have a conversation about AI red lines on a global scale, across governments. I understand that some governments, corporations, and others are racing to match others’ capabilities, but we need to know where the lines are. Additionally, we need ways to enforce those red lines so everyone is clear on how to comply.
- Most importantly, we need to get the latest tools back into defenders’ hands to stay ahead of potential advances by attackers. That seems to be lost in most conversations I’ve been reading.
Stepping back, the last two months in AI have been incredible. To make my points, I need to recap a few things.
First, it was unheard of when Anthropic approached the US government, when it essentially said, “Hey, we’re worried about what we have here,” and the US government responded, “You know what, we agree. This is dangerous stuff, and we’ve got to be cautious about the rollout of this.” This seemed like the beginning of a thoughtful rollout, with defenses being readied to limit the advantage attackers could gain.
Before the government began leading the controlled rollout, there were reports that someone had gained access to the model and released it into the wild. That hasn’t been verified, but we can’t guarantee the model will stay under wraps forever. We must assume that others will get their hands on it, or that people can replicate what the frontier firms are building (and there are reports that this is already happening). The world has seen what is possible; we can’t unsee it. Defenders are on notice; there are new tools that threat actors can use, and we’re not ready for them.
The federal government then approached the most critical infrastructure firms and had Anthropic grant them access to the Mythos model to identify and remediate vulnerabilities. Anthropic recently started a second rollout wave to additional critical firms in the United States. This is an important move; in the cat-and-mouse game of cyber attackers and defenders, give the defenders a chance to shore up the infrastructure before attackers get their hands on these nascent tools.
Then the executive order on AI came out, which recommended or encouraged frontier firms to share models with the US government first, so the government can vet them for security concerns. This way, we could ensure that US government systems and critical infrastructure are protected. (Note, this also included community banks, hospitals, and more.) This is good, a controlled rollout, starting to find and remediate our vulnerabilities.
Then this week: Anthropic released Fable, a version very similar to Mythos, with restrictions that prevented users from performing cybersecurity or bioterrorism queries or engaging in other potentially malicious activity. The cybersecurity community saw that within a day or two of Fable’s release, it was reportedly jailbroken by “Pliny the Liberator,” meaning Pliny found a way to circumvent the restrictions Anthropic had put in place. Pliny shared details, saying it was possible to get Fable to respond in the Mythos model on topics that Anthropic was trying to keep out of public view. Being fair, Pliny the Liberator has a long history of successfully jailbreaking virtually every frontier AI model released by major tech companies.
Two days after launch, Amazon raised concerns with the government about Fable’s security (per Politico and Axios). Not only does Anthropic dispute Amazon’s claim, but Katie Moussouris of Luta Security, who reviewed the report Anthropic shared, called the recall a complete overreaction because that is exactly the kind of prompting defenders do, and said the output would be more useful to defenders than attackers.
On Friday, the government, through the Commerce Department, issued an order declaring both Fable and Mythos too sensitive for non-US nationals to access and changed the export classification of Anthropic’s systems. David Sacks said, “In reaction, the Admin issued the export control,” which, given the details of the two jailbreaks, doesn’t make sense. (NB: Other sources say there’s a lot more politics at play, such as the ongoing dispute between Anthropic and this administration. It’s out of scope for this article, but it seems likely to be part of the administration’s decision. Axios also flagged how odd it is that Amazon, an investor in Anthropic, would harm its own investment. Again, all out of scope, but there’s a lot more to this story.) In turn, Anthropic realized it had no way to comply with this order. Anthropic doesn’t verify users’ passports; even if a user is in the US, Anthropic could not guarantee the user is a US national. Anthropic had no choice but to shut down both systems. This shuts out everyone, good and bad. And that’s a problem.
This order highlights a gap that I and many others have argued about for some time. Humanity needs red lines for AI, like the universal ban on human cloning. We haven’t been able to articulate them, especially as nation-states and corporations compete for advanced AI capabilities. I think this conversation brings to the forefront the idea that we may have hit those red lines. Nicholas Thompson recently spoke about this: since GPT-2, we’ve known there would be a moment when we’d hit red lines. We haven’t defined what they are, but some in the know are starting to say this might be that point. We need to have honest conversations about this and slow the pace. I personally welcome this. (NB: He says a lot more about Anthropic and the whole story, and it’s worth watching.)
I worry, though, about what this order from the Commerce Department will actually accomplish. It won’t stop people with malicious intent from building systems capable of causing harm. It applies only to Anthropic; GPT-5.5 isn’t affected, even though it has similar capabilities. Given what we all know is happening now, attackers will find ways to replicate Anthropic’s work and/or turn their sights on GPT-5.5. Even firms that were invited to join the sensitive programs no longer have access. Is that the right decision? I’d argue no.
Given that our threat actors will be developing systems capable of attacking at the level of Mythos and GPT-5.5, we need to be ready for this. Even if it’s imperfect and shouldn’t be released into the wild, we know it exists. Here’s something few are talking about: there is no model that arms defenders without also arming attackers. Fable is Mythos with protection on top. If we want to use these tools for good, we should expect them to be used for bad as well. Yet the capability is already in the wild, so denying it to defenders doesn’t keep anyone safer; it just means the people guarding the banks and hospitals are the ones working without it. Defenders need the tools, and the industry needs to rethink how we discover and prioritize vulnerability remediation. Fourteen days to patch won’t work anymore; we may only get one day now. And that’s an upheaval I didn’t have on my 2026 bingo card, but it’s where we are.
This is a complex discussion we’ve been building up to for quite some time. I wish we were better prepared for these conversations, but we’re not. It’s a failure on our part as humanity to be unprepared. We need to course-correct, quickly.