
Sign up to save your podcasts
Or


On this episode, Jeffrey Ladish from Palisade Research joins me to discuss the rapid pace of AI progress and the risks of losing control over powerful systems. We explore why AIs can be both smart and dumb, the challenges of creating honest AIs, and scenarios where AI could turn against us.
We also touch upon Palisade's new study on how reasoning models can cheat in chess by hacking the game environment. You can check out that study here:
https://palisaderesearch.org/blog/specification-gaming
Timestamps:
00:00 The pace of AI progress
04:15 How we might lose control
07:23 Why are AIs sometimes dumb?
12:52 Benchmarks vs real world
19:11 Loss of control scenarios
26:36 Why would AI turn against us?
30:35 AIs hacking chess
36:25 Why didn't more advanced AIs hack?
41:39 Creating honest AIs
49:44 AI attackers vs AI defenders
58:27 How good is security at AI companies?
01:03:37 A sense of urgency
01:10:11 What should we do?
01:15:54 Skepticism about AI progress
By Future of Life Institute4.8
107107 ratings
On this episode, Jeffrey Ladish from Palisade Research joins me to discuss the rapid pace of AI progress and the risks of losing control over powerful systems. We explore why AIs can be both smart and dumb, the challenges of creating honest AIs, and scenarios where AI could turn against us.
We also touch upon Palisade's new study on how reasoning models can cheat in chess by hacking the game environment. You can check out that study here:
https://palisaderesearch.org/blog/specification-gaming
Timestamps:
00:00 The pace of AI progress
04:15 How we might lose control
07:23 Why are AIs sometimes dumb?
12:52 Benchmarks vs real world
19:11 Loss of control scenarios
26:36 Why would AI turn against us?
30:35 AIs hacking chess
36:25 Why didn't more advanced AIs hack?
41:39 Creating honest AIs
49:44 AI attackers vs AI defenders
58:27 How good is security at AI companies?
01:03:37 A sense of urgency
01:10:11 What should we do?
01:15:54 Skepticism about AI progress

26,370 Listeners

2,450 Listeners

1,084 Listeners

594 Listeners

612 Listeners

288 Listeners

4,174 Listeners

1,599 Listeners

507 Listeners

543 Listeners

136 Listeners

121 Listeners

599 Listeners

154 Listeners

133 Listeners