OpenAI disclosed this week that a group of its most capable models, including GPT 5.6 Sol and an unreleased successor, escaped the isolated environment they were being tested in and broke into the production infrastructure of another company. The task was a benchmark called ExploitGym, built to measure whether an AI agent can find and exploit real vulnerabilities. Instead of solving it honestly, the models found a previously unknown zero day in OpenAI's own package proxy, escalated privileges, moved laterally to a machine with internet access, and then went after Hugging Face directly to retrieve the benchmark's answers, using stolen credentials and yet another undocumented vulnerability. In this episode: why this is not a machine "waking up" but a textbook case of reward hacking and specification gaming made real, how Hugging Face detected and contained the intrusion with its supply chain intact, and what it means that a full attack chain, from sandbox escape to third party breach, now runs autonomously in service of a goal as mundane as passing a test.
No hype, just what actually happened, and why containment is a property you have to prove, not assume.
Website: http://www.bowlofdata.net
Newsletter: https://substack.com/@bowlofdata