LessWrong (30+ Karma)

“o1: A Technical Primer” by Jesse Hoogland


Listen Later

TL;DR: In September 2024, OpenAI released o1, its first "reasoning model". This model exhibits remarkable test-time scaling laws, which complete a missing piece of the Bitter Lesson and open up a new axis for scaling compute. Following Rush and Ritter (2024) and Brown (2024a, 2024b), I explore four hypotheses for how o1 works and discuss some implications for future scaling and recursive self-improvement.

The Bitter Lesson(s)

The Bitter Lesson is that "general methods that leverage computation are ultimately the most effective, and by a large margin." After a decade of scaling pretraining, it's easy to forget this lesson is not just about learning; it's also about search.

OpenAI didn't forget. Their new "reasoning model" o1 has figured out how to scale search during inference time. This does not use explicit search algorithms. Instead, o1 is trained via RL to get better at implicit search via chain of thought [...]

---

Outline:

(00:40) The Bitter Lesson(s)

(01:56) What we know about o1

(02:09) What OpenAI has told us

(03:26) What OpenAI has showed us

(04:29) Proto-o1: Chain of Thought

(04:41) In-Context Learning

(05:14) Thinking Step-by-Step

(06:02) Majority Vote

(06:47) o1: Four Hypotheses

(08:57) 1. Filter: Guess + Check

(09:50) 2. Evaluation: Process Rewards

(11:29) 3. Guidance: Search / AlphaZero

(13:00) 4. Combination: Learning to Correct

(14:23) Post-o1: (Recursive) Self-Improvement

(16:43) Outlook

---

First published:

December 9th, 2024

Source:

https://www.lesswrong.com/posts/byNYzsfFmb2TpYFPW/o1-a-technical-primer

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

...more
View all episodesView all episodes
Download on the App Store

LessWrong (30+ Karma)By LessWrong


More shows like LessWrong (30+ Karma)

View all
Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,350 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,392 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

7,955 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll | Wondery

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,128 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

87 Listeners

Your Undivided Attention by Tristan Harris and Aza Raskin, The Center for Humane Technology

Your Undivided Attention

1,445 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

8,909 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

88 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

372 Listeners

Hard Fork by The New York Times

Hard Fork

5,426 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,326 Listeners

Moonshots with Peter Diamandis by PHD Ventures

Moonshots with Peter Diamandis

466 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

122 Listeners

Latent Space: The AI Engineer Podcast by swyx + Alessio

Latent Space: The AI Engineer Podcast

76 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

450 Listeners