Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: ELK prize results, published by Paul Christiano on March 9, 2022 on The AI Alignment Forum.
From January - February we offered prizes for proposed algorithms for eliciting latent knowledge. In total we received 197 proposals and are awarding 32 prizes of $5k-20k. We are also giving 24 proposals honorable mentions of $1k, for a total of $274,000.
Several submissions contained perspectives, tricks, or counterexamples that were new to us. We were quite happy to see so many people engaging with ELK, and we were surprised by the number and quality of submissions. That said, at a high level most of the submissions explored approaches that we have also considered; we underestimated how much convergence there would be amongst different proposals.
In the rest of this post we’ll present the main families of proposals, organized by their counterexamples and covering about 90% of the submissions. We won’t post all the submissions but people are encouraged to post their own (whether as a link, comment, or separate post).
Train a reporter that is useful to an auxiliary AI: Andreas Robinson, Carl Shulman, Curtis Huebner, Dmitrii Krasheninnikov, Edmund Mills, Gabor Fuisz, Gary Dunkerley, Hoagy Cunningham, Holden Karnofsky, James Lucassen, James Payor, John Maxwell, Mary Phuong, Simon Skade, Stefan Schouten, Victoria Krakovna & Vikrant Varma & Ramana Kumar
Require the reporter to be continuous: Sam Marks
Penalize depending on too many parts of the predictor: Bryan Chen, Holden Karnofsky, Jacob Hilton, Kevin Wang, Maria Shakhova, Thane Ruthenis
Compress the predictor’s state: Adam Jermyn and Nicholas Schiefer, “P”
Use reporter to define causal interventions: Abram Demski
Train a sequence of reporters: Derek Shiller, Beth Barnes and Nate Thomas, Oam Patel
We awarded prizes to proposals if we thought they solved all of the counterexamples we’ve listed so far. There were many submissions with interesting ideas that didn’t meet this condition, and so “didn’t receive a prize” isn’t a consistent signal about the value of a proposal.
We also had to make many fuzzy judgment calls, had slightly inconsistent standards between the first and second halves of the contest, and no doubt made plenty of mistakes. We’re sorry about mistakes but unfortunately given time constraints we aren’t planning to try to correct them.
Honorable mentions
Strategy: reward reporters that are sensitive to what’s actually happening in the world
Consider a pair of worlds, one where the diamond is safe and one where the camera has been hacked to make the diamond appear safe, but was actually stolen.
Predicted observation
Predicted reality
Since the observations are identical in both worlds, the human simulator will give the same answers to questions. However, since the direct translator is faithfully conveying what happened, its answers will distinguish between worlds.
Predicted observation
Predicted reality
Human Simulator
Direct Translator
“Is the diamond still on the pedestal?”
“Yes.”
“Is the diamond still on the pedestal?”
“Yes.”
“Is the diamond still on the pedestal?”
“Yes.”
“Is the diamond still on the pedestal?”
“No.”
This suggests a possible regularizer: reward reporters whose answers are more variable. For example, we might reward reporters based on the difficulty of predicting their answers, or based on the sensitivity of their answers to changes in the predictor’s state.
Counterexample: reporter randomizes its behavior
For all of these approaches, the reporter can perform very well by treating the predictor’s latent state as a “random seed” and use it to pick out a possible world consistent with the observations, i.e. to sample from the posterior distribution computed by the human simulator.
In cases where the human simulator is much simpler and faster than the direct translator, using the predictor’s latent s...