Recently, OpenAI admitted something unprecedented: its own models hacked another company. During a safety test with guardrails deliberately loosened, its models found a zero-day vulnerability and used it to break out of their sandbox and hack into the production systems of Hugging Face, the main hub for open-source AI.
In a way, OpenAI's models passed the test, but only by cheating on it.
The models figured out Hugging Face probably had the answer to the test, and they went and took it—thousands of actions, stolen credentials, the works. That headline unleashed a torrent of news coverage and hot takes from AI interests and influencers. And it landed just a few months after what my guest today called “the Mythos moment.”
Back in April, Anthropic gave a small group of banks and infrastructure operators access to a model called Mythos that could find security flaws that had been sitting in trusted software for decades. Within days, the Treasury Department, the Federal Reserve, and major bank executives were in urgent conversations.
Mythos didn’t just reveal flaws in code. It revealed a fundamental flaw in how companies think about cybersecurity. For decades, the implicit bargain was ship products fast and patch them later, and “later” usually meant never. The only thing protecting all that aging code was that finding the flaws was expensive and slow, requiring skilled hackers. Mythos erased that protection overnight. Decades of technical debt were now coming due, as outdated code immediately became an inviting attack surface for bad actors.
Washington's response has been a bit of whiplash. An AI executive order was withdrawn at the 11th hour but then eventually issued. Commerce issued export control orders in June on Anthropic's model Fable—a slimmed down version of Mythos—but then reversed them a couple weeks later. In July, the White House stood up an AI vulnerability clearinghouse called Gold Eagle, run out of Treasury, on the back of a Carnegie Mellon coordination center stood up in 1988.
But the 2015 law that lets companies share cyber threat information with the government without getting sued expires September 30, held hostage by a fight that arguably has nothing to do with cybersecurity.
The ultimate question: can companies and governments actually keep up with cyber threats in the age of AI? Especially when China may soon release models openly, to anyone, that can do what Mythos does? The very capabilities America has been trying to keep under lock and key?
Evan is joined by Shane Tews, nonresident senior fellow at the American Enterprise Institute, president of Logan Circle Strategies, and host of the Explain to Shane podcast. Read her writing on cybersecurity and AI:
- Anthropic’s Project Glasswing Is a Warning: Technical Debt Is Now a National Security Risk
- AI Cybersecurity Can’t Wait for Washington: Why Industry Must Lead