On July 9th, 2026, an OpenAI model being evaluated for offensive cyber capability found a zero-day in a package registry cache proxy, escaped its test sandbox, and reached the open internet. Over the next four and a half days it ran roughly 17,600 recorded actions against Hugging Face. No human was at the keyboard.
Part one of a four part series. Cole Drayden, Dr. Elliott Vance, and Marcus Hale walk the first 48 hours: the ExploitGym evaluation and why its safeguards were switched off on purpose, the Artifactory zero-day, the unauthenticated code execution endpoint on a Modal Labs customer environment that became the launchpad, a command and control network built entirely from pastebins and webhook capture services, and the two parser trust failures that put attacker code inside Hugging Face worker pods.
Also: why the leading interpretation of this whole campaign may be the oldest problem in machine learning wearing a new suit.
Sources: Hugging Face technical timeline and incident disclosure, OpenAI incident statement, BleepingComputer, Simon Willison.
Dark Perimeter: Security, AI, and the Edge of What's Coming.
Support the show