August 03, 2023

“AI #23: Fundamental Problems with RLHF” by Zvi

1 hour 13 minutes

After several jam-packed weeks, things slowed down to allow everyone to focus on the potential room temperature superconductor, check Polymarket to see how likely it is we are so back and bet real money, or Manifold for chats and better graphs and easier but much smaller trading.

The main thing I would highlight this week are an excellent paper laying out many of the fundamental difficulties with RLHF, and a systematic new exploit of current LLMs that seems to reliably defeat RLHF.

I’d also note that GPT-4 fine tuning is confirmed to be coming. That should be fun.

Table of Contents

Introduction.

Table of Contents.

Language Models Offer Mundane Utility. Here’s what you’re going to do.

Language Models Don’t Offer Mundane Utility. Universal attacks on LLMs.

Fun With Image Generation. Videos might be a while.

Deepfaketown and Botpocalypse Soon. An [...]

---

Outline:

(00:38) Table of Contents

(02:32) Language Models Offer Mundane Utility

(05:20) Language Models Don’t Offer Mundane Utility

(16:57) Fun with Image Generation

(17:40) Deepfaketown and Botpocalypse Soon

(17:49) They Took Our Jobs

(22:31) Get Involved

(23:35) Introducing

(31:33) In Other AI News

(34:46) Quiet Speculations

(36:13) China

(38:31) The Quest for Sane Regulations

(45:04) The Week in Audio

(45:37) Rhetorical Innovation

(47:51) No One Would Be So Stupid As To

(49:44) Aligning a Smarter Than Human Intelligence is Difficult

(01:10:30) Other People Are Not As Worried About AI Killing Everyone

(01:11:34) The Wit and Wisdom of Sam Altman

(01:11:56) The Lighter Side

---

First published:

August 3rd, 2023

Source:

https://www.lesswrong.com/posts/aKzwwKT2cy72awSyz/ai-23-fundamental-problems-with-rlhf

---

Narrated by TYPE III AUDIO.

...more

View all episodes

By zvi

22 ratings

August 03, 2023

“AI #23: Fundamental Problems with RLHF” by Zvi

1 hour 13 minutes

I’d also note that GPT-4 fine tuning is confirmed to be coming. That should be fun.

Table of Contents

Introduction.

Table of Contents.

Language Models Offer Mundane Utility. Here’s what you’re going to do.

Language Models Don’t Offer Mundane Utility. Universal attacks on LLMs.

Fun With Image Generation. Videos might be a while.

Deepfaketown and Botpocalypse Soon. An [...]

---

Outline:

(00:38) Table of Contents

(02:32) Language Models Offer Mundane Utility

(05:20) Language Models Don’t Offer Mundane Utility

(16:57) Fun with Image Generation

(17:40) Deepfaketown and Botpocalypse Soon

(17:49) They Took Our Jobs

(22:31) Get Involved

(23:35) Introducing

(31:33) In Other AI News

(34:46) Quiet Speculations

(36:13) China

(38:31) The Quest for Sane Regulations

(45:04) The Week in Audio

(45:37) Rhetorical Innovation

(47:51) No One Would Be So Stupid As To

(49:44) Aligning a Smarter Than Human Intelligence is Difficult

(01:10:30) Other People Are Not As Worried About AI Killing Everyone

(01:11:34) The Wit and Wisdom of Sam Altman

(01:11:56) The Lighter Side

---

First published:

August 3rd, 2023

Source:

https://www.lesswrong.com/posts/aKzwwKT2cy72awSyz/ai-23-fundamental-problems-with-rlhf

---

Narrated by TYPE III AUDIO.

...more

Share “AI #23: Fundamental Problems with RLHF” by Zvi

Sign up to save your podcasts

“AI #23: Fundamental Problems with RLHF” by Zvi

“AI #23: Fundamental Problems with RLHF” by Zvi

More shows like LessWrong posts by zvi

Making Sense with Sam Harris

Conversations with Tyler

The a16z Show

Future of Life Institute Podcast

ChinaTalk

Politix

Dwarkesh Podcast

Hard Fork

Clearer Thinking with Spencer Greenberg

LessWrong (Curated & Popular)

No Priors: Artificial Intelligence | Technology | Startups

"Econ 102" with Noah Smith and Erik Torenberg

BG2Pod with Brad Gerstner and Bill Gurley

LessWrong (30+ Karma)

Complex Systems with Patrick McKenzie (patio11)