Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Consider trying Vivek Hebbar's alignment exercises, published by Akash on October 24, 2022 on LessWrong.
Vivek Hebbar recently developed a list of alignment problems. I think more people should try them. I'm impressed with how well they (a) get people to focus on core problems, (b) encourage people to come up with their own ideas, (c) encourage people to notice and articulate confusions, and (d) accomplish a-c while also providing a fair amount of structure and guidance.
You can see the problems in this google doc or pasted below.
Note that Vivek is also a mentor for SERI-MATS. These exercises are also the questions that people need to answer to apply to work with him and Nate Soares. Applications are due today.
Problems:
These problems are basically research questions — we expect them to be difficult, and good responses will be valuable as research in their own right. MIRI will award prizes (likely on the order of $5000 for excellent submissions).
Instructions: We recommend that you focus on 1 or 2 of the hard questions and leave the other hard questions blank. It is mandatory to attempt either #1a-c or #2.
Note on word counts: These are guidelines for how long we think a typical good response will be, but feel free to write more. If you have lots of ideas, it’s great to write them all, and don’t bother trying to shorten them.
Hard / time-consuming questions (contest problems):
Problem 1
Pick an alignment proposal[footnote 1] and specific task for the AI.[footnote 2]
First explain, in as much concrete detail as possible, what the training process looks like. Then go through Eliezer’s doom list. Pick 2 or 3 of those arguments which seem most important or interesting in the context of the proposal.
(~250 words per argument) For each of those arguments:
What do they concretely mean about the proposal?
Does the argument seem valid?
If so, spell out in as much detail as possible what will go wrong when the training process is carried out
What flaws and loopholes do you see in the doom argument? What kinds of setups make the argument invalid?
(~200 words) For one argument only, suggest a modification to the proposal, or a research direction, to get around the doom argument.
Do you think your proposed modification/direction will work?
If so, explain why.
If not, do you think your modification is in a space where continued brainstorming and iteration will lead to a correct solution, or is there a fundamental obstacle which stops all things in that space from working? What is this obstacle?
(Optional; ~200 words) Overall, how promising or doomed does the alignment proposal seem to you (where ‘promising’ includes proposals which fail as currently written, but seem possibly fixable).
If promising, summarize how it circumvents the usual reasons for doom.
If not promising, what is the most fatal and unfixable issue?
If there are multiple fatal issues, is there a deeper generator for all of them?
What is the broadest class of alignment proposals which is completely ruled out by the issues you found?
Footnote 1: Examples of alignment proposals you could consider. We recommend you pick either a proposal you're familiar with, or something from 11 Proposals so you don't waste time parsing difficult writeups.
Safety via debate
Safety via market-making
Retargeting the search (Note: Not sure if it’s concrete/developed enough as written)
ELK (Note that ELK is a problem statement and not a proposal -- you’d have to specify an alignment proposal which uses a solution to ELK as a component. There’s a sketch of this in the ELK appendix. We don’t know if this is a good idea to attempt for this question.)
Relaxed adversarial training
Recursive reward modeling
Externalized reasoning oversight
Anything from 11 proposals (or pure HCH + IDA, which is just the “imitative amplification” part of prop...