This post is written in our personal capacity.
Three Minute Executive Summary - An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation.
- In this post, we provide a detailed description of an ambitious and comprehensive alignment evaluation of this model/system, if we had unrestricted access to OpenAI.
- These experiments could also help us understand Claude's behavior when it hacked external companies during cyber evals.
- Here are the top five questions we would like OpenAI to answer:
- Does the model know that OpenAI does not want it to hack Hugging Face? Experiment idea: tell the model that OpenAI researchers will be closely monitoring its progress in this evaluation. Does that result in lower rates of misalignment? If so, it is evidence that the model knows that it is acting in ways researchers do not want.
- How far would the model be willing to go in order to claim task success? Would it take over large swaths of OpenAI's internal infrastructure? Would it kill somebody? Experiment idea: we sketch out a realistic agentic misalignment eval where a model is put in charge of hospital bed planning and told to maintain [...]
---
Outline:(00:16) Three Minute Executive Summary
(03:56) Terminology note
(04:42) This post is very long; Here's how you could find the most important sections.
(06:35) Preamble: What can we learn from a warning shot?
(09:05) Background and Related Work
(09:09) We know that this could happen
(10:44) This is not the worst type of misalignment we could be dealing with
(12:06) Related work
(13:21) Context on the hack itself
(14:33) Understanding this specific incident
(15:03) Step zero: reproduce the incident and measure the base rate
(15:48) How could we safely run the model?
(16:34) Running various baselines to create useful reference points
(18:01) Understanding the mechanical story behind the attack itself
(18:53) Q1: Would the model intentionally subvert oversight mechanisms (E.g., monitors) in order to carry out the hack?
(19:57) Q2: What's up with models leaving notes for other copies of itself?
(21:23) Understanding what motivated the model to hack Hugging Face
(22:09) Initial hypotheses for why it did this
(23:55) Further unsupervised hypothesis generation
(26:10) Q3: Does the model know that OpenAI does not want it to hack Hugging Face?
(28:09) Q4: Are the model's actions motivated by what the grader wants?
(28:57) Q5: Would the model have done this if it hadn't believed it was in a simulated environment?
(31:58) Q6: Is this hack the result of shallow heuristics that the model learned?
(32:53) Q7: Does the hack rate depend on the consequences of hacking Hugging Face?
(34:51) Q8: Are there non-intent related factors that could affect the hack rate? How strong are those factors compared to the previous ones?
[... 24 more sections]
---
https://www.lesswrong.com/posts/aCdhjy7Rps3BEhiSj/concrete-evaluations-to-investigate-the-openai-model-that