Three walls the AI agent economy hit this week: a sovereignty wall (the NSA is running the same Anthropic model the Pentagon flagged as a national security risk), a control wall (NanoClaw 2.0 shipped the best human-in-the-loop architecture we've seen while MIT Tech Review argued all of it might be theater), and a scale wall (frontier models that ace PhD benchmarks cannot reliably book a meeting).
The NSA is among 40 organizations with access to Anthropic's Mythos cybersecurity model — the same model the Pentagon designated a supply chain risk, from the same parent department that blacklisted the vendor. No published framework resolves the contradiction. Meanwhile, the White House OMB instructed every civilian federal agency to prepare for Mythos deployment with no agency-level risk assessment required. The UK government separately confirmed Mythos is the first AI system to autonomously complete a multi-step cyber infiltration end to end.
NanoClaw 2.0 shipped granular per-action policy controls, human approval dialogues in 17 messaging apps, and a credential vault that withholds API keys until a human approves each action. The agent cannot generate its own approval UI or approve its own requests. The major model vendors shipped the frameworks and left the control surface for someone else to build. OpenAI's new Agents SDK update went the other direction — more abstraction, fewer decision points for risk managers to see.
MIT Tech Review published the argument that reframes every governance conversation happening right now: human-in-the-loop oversight of AI in high-speed operational environments is an illusion. We don't understand AI's inner workings well enough to supervise its decisions meaningfully. The human approval step looks like governance, but it isn't. If they're right, most of what enterprises call AI governance is theater.
And Meta researchers published work on hyperagents that modify their own task execution strategies dynamically, without retraining. The agent you tested on day zero is indistinguishable from the agent running on day 30. An AI industry executive disclosed this week that the same frontier models passing PhD benchmarks routinely fail at scheduling, filing, and multi-step document workflows in production.
Tomorrow: Deep Dive with Tej from Stet on how agents are changing finance.
—
Agentic Stories is the weekday briefing on the AI agent economy — governance, security, and deployment. New episodes Monday, Wednesday, Friday.
agenticstories.ai