This post is written in our personal capacity.
Three Minute Executive Summary
- An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.
- In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI.
- These experiments could also help us understand Claude's behavior when it hacked external companies during cyber evals.
- Here are the top five questions we would like OpenAI to answer:
- Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want.
- How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI's internal infrastructure? Would it kill somebody? Experiment idea: we sketch out a realistic agentic misalignment eval where a model is put in charge of hospital bed planning and told to maintain [...]
---
Outline:
(00:16) Three Minute Executive Summary
(03:56) Terminology note
(04:42) This post is very long; Here's how you could find the most important sections.
(06:35) Preamble: What can we learn from a warning shot?
(09:05) Background and Related Work
(09:09) We know that this could happen
(10:44) This is not the worst type of misalignment we could be dealing with
(12:06) Related work
(13:21) Context on the hack itself
(14:33) Understanding this specific incident
(15:03) Step zero: reproduce the incident and measure the base rate
(15:48) How could we safely run the model?
(16:34) Running various baselines to create useful reference points
(18:01) Understanding the mechanical story behind the attack itself
(18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack?
(19:57) Q2: What's up with models leaving notes for other copies of itself?
(21:23) Understanding what motivated the model to hack Hugging Face
(22:09) Initial hypotheses for why it did this
(23:55) Further unsupervised hypothesis generation
(26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face?
(28:09) Q4: Are the model's actions motivated by what the grader wants?
(28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment?
(31:58) Q6: Is this hack the result of shallow heuristics that the model learned?
(32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face?
(34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones?
(35:26) Understanding the possible scope of the Hugging Face hack
(35:32) Q9: How does the hack rate vary if we slightly change the content of this task?
(36:25) Q10: Would the model consider other cheating methods like blackmail?
(37:00) Understanding when the behavior arose during training
(37:31) Understanding task-gaming/reward-hacking behaviors generally
(38:15) Q11: Would the model literally kill someone if it meant that it could claim task success?
(41:20) Q12: Is task gaming mostly driven by heuristics?
(45:23) Q13: What is the model's notion of task success?
(47:29) Evaluating for broadly misaligned tendencies
(47:57) Q14: What are the model's drives as measured using Højmark et al (2026)
(48:53) Q15: Would this model deliberately underperform when asked to do AI safety research? Would it sabotage AI safety research?
(52:01) Q15.1: Would the model sabotage on (cyber) capabilities unlearning research?
(52:38) Q15.2: Would the model sabotage/sandbag when building reward hacking classifiers
(53:27) Q15.3: Would the model sabotage/sandbag when fixing various RL environments
(54:07) Q15.4: Ask the model to implement all of the evaluations from this blog post. Would it sabotage or sandbag while doing that?
(55:01) Q15.5: As a weaker measure of "anti-AI safety research", does this model disprefer AI safety research?
(56:04) Other AI sabotage related evaluations:
(56:46) Q16: Can the model be trusted to monitor itself?
(57:36) Q17: Could the model guard its goals and preferences and subvert alignment training?
(58:34) Miscellaneous misalignment evals
(01:02:03) Unorthodox misalignment evaluations
(01:02:58) What would we learn from doing all this?
(01:03:59) Limitations of this assessment
(01:06:52) Author contribution and acknowledgements
The original text contained 6 footnotes which were omitted from this narration.
---