Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: AI-Plans.com - a contributable compendium, published by Iknownothing on June 25, 2023 on LessWrong.
Hello, we’re working on .
Ideas behind the site:
Right now, alignment plans are spread all over the place and it’s difficult for a layperson, or someone unfamiliar with the field to get an idea of what is the current plan for making AGI or even the models we have right now, safe. And the problems with said plans.
Having a place where all AI Alignment plans and criticism of said plans can be added and seen in an easy to read way is helpful. It offers possibilities such as; seeing the most common problems with plans, which kinds of plans have the least problems, pushing for regulations against the most poor plans etc.
Judging the quality of a plan is hard and has a lot of ways to go wrong. Judging the quality of a criticism might be less complicated and perhaps has less ways it can go wrong.
Which is why I believe this site is useful
Hasn’t this already been done? aisafety.info, aiideas.com, etc
aisafety.info is excellent, and the folks there are building a conversational agent for AI Safety, which could be incredibly useful. However it’s not very simple to use to learn about specific alignment plans yet and is more of a general purpose place to learn about ai, ai-risk, ai-safety, etc.
The purpose of ai-plans.com is to be an easy to read platform that sorts the good plans from the bad and shows the problems with each one.
I believe there was a site called ai-ideas.com or something mentioned to me? But that wasn’t working the last time I checked and there’s been no change, as far as I know.
Aims
Stage 1)
A contributable compendium with most, if not all, of the plans for alignment, and criticisms
Estimated time left for this to be done: 1 week - adding ~5 plans a day now
What’s left to do:
Functionality for responding to criticisms
The idea for this is that a plan’s author can select a criticism/criticisms that they think they have a solution to, then submit a new version of the plan, with the criticism(s) quoted in the new plan.Any criticisms they didn’t select, will be automatically added to the new version of the plan (the idea being that any unselected criticisms are ones that they don’t have solutions to, so should still apply).
Adding more plans and criticisms
A filter for spam and duplicates
Creating a template/guide on how to post plans
Improve the UX - font, colours, design, etc.
What we’re missing for this stage:
A lawyer to make sure we’re complying with GDPR and help make a cookies notice
Moderators to help filter spam and make sure plans are submitted correctly
What would be helpful, but not essential for this stage:
More copywriters to speed up the process of adding alignment plans
More red-teamers/quality testers
Stage 2)
In addition to everything from Stage 1, there is now a scoring system for criticisms and a ranking system for plans- where plans are ranked from top to bottom based on the total scores of their criticisms.
The scoring system is a very essential part of the site.
It's aim is to give the most points to the most accurate criticisms and use that to rank plans from the least total criticism points (at the top) to the most criticism points(at the bottom).
Users will be able to upvote or downvote criticisms.
Users will also have a ‘karma’ that will affect how weighted their votes on criticisms are- somewhat similar to the LessWrong and AlignmentForum system- though, we're considering having karma past a certain 'age' become spendable, rather than add to the weight of the users vote, to avoid to first vote problem of LessWrong.
Users who accumulate more points on their criticisms will have a higher karma.
Plans will have a ‘bounty’ inversely proportional to the total number of criticism points they have.
I think it’s important to get this right the...