LessWrong (30+ Karma)

“o3, Oh My” by Zvi


Listen Later

OpenAI presented o3 on the Friday before Thanksgiving, at the tail end of the 12 Days of Shipmas.

I was very much expecting the announcement to be something like a price drop. What better way to say ‘Merry Christmas,’ no?

They disagreed. Instead, we got this (here's the announcement, in which Sam Altman says ‘they thought it would be fun’ to go from one frontier model to their next frontier model, yeah, that's what I’m feeling, fun):

Greg Brockman (President of OpenAI): o3, our latest reasoning model, is a breakthrough, with a step function improvement on our most challenging benchmarks. We are starting safety testing and red teaming now.

Nat McAleese (OpenAI): o3 represents substantial progress in general-domain reasoning with reinforcement learning—excited that we were able to announce some results today! Here is a summary of what we shared about o3 in the livestream.

---

Outline:

(03:48) GPQA Has Fallen

(04:21) Codeforces Has Fallen

(05:32) Arc Has Kinda of Fallen But For Now Only Kinda

(09:27) They Trained on the Train Set

(15:26) AIME Has Fallen

(15:58) Frontier of Frontier Math Shifting Rapidly

(19:09) FrontierMath 4: We're Going To Need a Bigger Benchmark

(23:10) What is o3 Under the Hood?

(25:17) Not So Fast!

(28:38) Deep Thought

(30:03) Our Price Cheap

(36:32) Has Software Engineering Fallen?

(37:42) Don't Quit Your Day Job

(40:48) Master of Your Domain

(43:21) Safety Third

(47:56) The Safety Testing Program

(48:58) Safety testing in the reasoning era

(51:01) How to apply

(53:07) What Could Possibly Go Wrong?

(56:36) What Could Possibly Go Right?

(57:06) Send in the Skeptic

(59:25) This is Almost Certainly Not AGI

(01:02:57) Does This Mean the Future is Open Models?

(01:07:17) Not Priced In

(01:08:39) Our Media is Failing Us

(01:14:56) Not Covered Here: Deliberative Alignment

(01:15:08) The Lighter Side

The original text contained 22 images which were described by AI.

---

First published:

December 30th, 2024

Source:

https://www.lesswrong.com/posts/QHtd2ZQqnPAcknDiQ/o3-oh-my

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (30+ Karma)By LessWrong


More shows like LessWrong (30+ Karma)

View all
Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,346 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,388 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

7,994 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,133 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

87 Listeners

Your Undivided Attention by Tristan Harris and Aza Raskin, The Center for Humane Technology

Your Undivided Attention

1,444 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

8,913 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

87 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

373 Listeners

Hard Fork by The New York Times

Hard Fork

5,417 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,281 Listeners

Moonshots with Peter Diamandis by PHD Ventures

Moonshots with Peter Diamandis

465 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

122 Listeners

Latent Space: The AI Engineer Podcast by swyx + Alessio

Latent Space: The AI Engineer Podcast

76 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

450 Listeners