Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how [...]
---
Outline:
(03:42) Language Models Offer Mundane Utility
(04:24) Language Models Don't Offer Mundane Utility
(07:38) Fable Disproves The Jacobian Conjecture Via Counterexample
(11:24) Claude Fable Will Remain In Max Plan Indefinitely
(13:39) Huh, Upgrades
(14:42) On Your Marks
(19:48) Deepfaketown and Botpocalypse Soon
(20:42) Fun With Media Generation
(20:51) Cyber Lack of Security
(22:07) They Took Our Jobs
(22:56) Get Involved
(24:47) Introducing
(25:46) In Other AI News
(28:02) More on Kimi K3
(33:08) Show Me the Money
(33:55) Quiet Speculations
(37:35) Potential Trouble At UK AISI
(39:29) Pick Up The Phone
(40:30) OpenAI Has Some Alignment Problems
(46:48) The Quest for Sane Regulations
(52:02) Chip City
(53:10) The Week in Audio
(53:27) People Just Say Things
(56:42) Rhetorical Innovation
(58:34) The Rome Declaration
(01:04:02) Aligning a Smarter Than Human Intelligence is Difficult
(01:07:52) Anthropic Surveys Things It Calls Misalignment
(01:13:33) Cooperative Alignment
(01:17:54) Other People Are Not As Worried About AI Killing Everyone
(01:19:35) The Lighter Side
---