LessWrong posts by zvi

“o3, Oh My” by Zvi


Listen Later

OpenAI presented o3 on the Friday before Thanksgiving, at the tail end of the 12 Days of Shipmas.

I was very much expecting the announcement to be something like a price drop. What better way to say ‘Merry Christmas,’ no?

They disagreed. Instead, we got this (here's the announcement, in which Sam Altman says ‘they thought it would be fun’ to go from one frontier model to their next frontier model, yeah, that's what I’m feeling, fun):

Greg Brockman (President of OpenAI): o3, our latest reasoning model, is a breakthrough, with a step function improvement on our most challenging benchmarks. We are starting safety testing and red teaming now.

Nat McAleese (OpenAI): o3 represents substantial progress in general-domain reasoning with reinforcement learning—excited that we were able to announce some results today! Here is a summary of what we shared about o3 in the livestream.

---

Outline:

(03:48) GPQA Has Fallen

(04:21) Codeforces Has Fallen

(05:32) Arc Has Kinda of Fallen But For Now Only Kinda

(09:27) They Trained on the Train Set

(15:26) AIME Has Fallen

(15:58) Frontier of Frontier Math Shifting Rapidly

(19:09) FrontierMath 4: We're Going To Need a Bigger Benchmark

(23:10) What is o3 Under the Hood?

(25:17) Not So Fast!

(28:38) Deep Thought

(30:03) Our Price Cheap

(36:32) Has Software Engineering Fallen?

(37:42) Don't Quit Your Day Job

(40:48) Master of Your Domain

(43:21) Safety Third

(47:56) The Safety Testing Program

(48:58) Safety testing in the reasoning era

(51:01) How to apply

(53:07) What Could Possibly Go Wrong?

(56:36) What Could Possibly Go Right?

(57:06) Send in the Skeptic

(59:25) This is Almost Certainly Not AGI

(01:02:57) Does This Mean the Future is Open Models?

(01:07:17) Not Priced In

(01:08:39) Our Media is Failing Us

(01:14:56) Not Covered Here: Deliberative Alignment

(01:15:08) The Lighter Side

---

First published:

December 30th, 2024

Source:

https://www.lesswrong.com/posts/QHtd2ZQqnPAcknDiQ/o3-oh-my

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more
View all episodesView all episodes
Download on the App Store

LessWrong posts by zviBy zvi

  • 5
  • 5
  • 5
  • 5
  • 5

5

2 ratings


More shows like LessWrong posts by zvi

View all
Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,386 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,419 Listeners

a16z Podcast by Andreessen Horowitz

a16z Podcast

1,087 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

107 Listeners

ChinaTalk by Jordan Schneider

ChinaTalk

288 Listeners

Politix by Politix

Politix

93 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

75 Listeners

Hard Fork by The New York Times

Hard Fork

5,470 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

130 Listeners

LessWrong (Curated & Popular) by LessWrong

LessWrong (Curated & Popular)

13 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

130 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

153 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

504 Listeners

LessWrong (30+ Karma) by LessWrong

LessWrong (30+ Karma)

0 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

133 Listeners