
Sign up to save your podcasts
Or


[Epistemic status: intuitions and anecdotes.]
Recently, several posts and projects (Thoughts Memo, Babel Translation, Please Give Them a Chance) have taken important steps towards raising AI safety awareness and sharing rationalist philosophy in China. It's great that we’re recognizing the importance of solving the messaging problem for China, and thus laying the groundwork for an international AI pause. Below I record my perspective on cultural differences which are relatively underdiscussed, which may become roadblocks to this communication program.
Background: I’m a “first-generation” Chinese-American who moved to the States at the age of four. The beliefs in this essay are primarily drawn from interactions with my parents and their generation of immigrants, and from consumption of Chinese media (dramas, webnovels, games, and manhua) which are not necessarily representative of the realities on the ground. I am likely over-indexed on the older generation and internet culture, and would appreciate corrections from folks who have direct lived experience. The picture I aim to paint is also complicated by a massive generational gap, and my understanding is that some of the below sentiments (e.g. the cynicism and nationalism) are partly inherited by the younger generation, and partly rejected through a variety of countercultures.
[...]
---
Outline:
(03:09) Chinese Social Media is like American Junk Food
(06:03) The Primacy of Social Reality
(08:22) The Dark World Frame
(12:44) Chinese nationalism as collective insecurity
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
MIRI recently endorsed the Ban Artificial Superintelligence Act, but others like ControlAI have called it overly broad, and a few other respected experts have outright endorsed it in its current form.
I decided to take a look at it myself and see both what the bid is trying to do and whether it actually does it well.
The goal of this act is to ban any AI that shows traits that indicate it could cause great harm to society, and to restrict the development of advanced AI to certified institutions who can be trusted to do so safely. This seems like a reasonable goal.
The problem with this bill, as far as I can see it, isn't that its goal is wrong, but that the way it's been drafted contains lots of flaws, which are likely to have unintended negative consequences.
At it's core, the bill does four things:
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
I have recently completed MATS 10.0, where I worked alongside Bart Jaworski under Victoria Krakovna (GDM). This is a post I was encouraged to make by my team at Geodesic Research from some slides I put together. This is not a post on application advice to MATS, or about the program in general. It is rather a compressed form of my experience doing research and lessons from the project.
The project
The full paper and post is coming soon, but I'll provide some context on it so that the lessons don't seem to come from nowhere.
---
Outline:
(00:36) The project
(02:39) Lessons learned
(02:42) Lesson 1: Talk to your models
(05:02) Lesson 2: Friedness is a great sanity check
(06:17) Lesson 3: RL is ... hard and unpredictable
(08:23) Lesson 4: Transfer of traits to agentic settings is hard
(09:45) Lesson 5: LoRA rank didn't matter much (for me)
(10:33) Closing thoughts
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Note: I'm not a WWII history buff, so if there's any significant misinterpretation of relevant history here, I welcome corrections in the comments.
In December 1944, Joseph Rotblat, a physicist working on the Manhattan project, resigned. He had reasoned that Germany had most likely abandoned its bomb project, and so his initial reason for working on the bomb was no longer valid. He would later ask himself:
Why did other scientists not make the same decision? [...]there were many scientists for whom the German factor was the main motivation. Why did they not quit when this factor ceased to be?
According to him, conversations with the other scientists suggested the following reasons (emphasis mine):
The most frequent reason given was pure and simple scientific curiosity—the strong urge to find out whether the theoretical calculations and predictions would come true. [...]
Others were [...] persuaded by the argument that many American lives would be saved if the bomb brought a rapid end to the war with Japan. Only when peace was restored would they take a hand in efforts to ensure that the bomb would not be used again.[...]
Still others, while agreeing that the project should have been stopped [...]
---
Outline:
(04:25) Adverse epistemic environments
(05:16) Beating a dead horse
(05:47) Appendix: Some nitpicking about whether it was in fact clear that Germany didn't have a bomb
(06:31) Appendix: Some other provocative quotes from the Franck report
(08:09) Appendix: More quotes
(09:53) Past discussions of the matter on LW
The original text contained 7 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
I'm often asked about the differences and similarities between Simplex' and Timaeus' research agendas. The question is natural enough. Both focus on a 'fundamental science' approach to AI alignment. Both organizations base their research agendas on sophisticated mathematical frameworks handed down from a bearded ur-figure (Sumio Watanabe, James Crutchfield).
We may posit the following correspondence
Dan Murfet + Jesse Hoogland : Developmental Interpretability : Singular Learning Theory : Sumio Watanabe
<->
Adam Shai + Paul Riechers : Belief-state Geometry: Computational Mechanics : James Crutchfield
SLT vs CompMech
Round One. Fight!
Weights vs Activations
DevInterp & SLT is about weight space. Belief-state geometry is more about studying activation space.
Activation space is what is already being studied in MechInterp & most 'mainstream' approaches to interpretability. It is concrete and present to the senses. Weight space is much larger, more abstract, harder to measure and sample.
Training vs Inference
SLT is about training. CompMech is about inference.
Both study Bayesian posteriors and updating. For SLT that is the Bayesian posterior on weight space - hence relevant for training. The Belief-State Geometry agenda studies the Mixed State Presentation from CompMech which describes an [idealized] version of in-context learning [...]
---
Outline:
(00:55) SLT vs CompMech
(01:02) Weights vs Activations
(01:30) Training vs Inference
(02:19) IID vs non-IID Data
(02:40) Parameterization vs Invariant structure
(03:25) Asymptotic vs Exact
(03:43) Mechanism vs Behaviourial
(04:04) Bottom-up vs Top-down interpretability
(05:10) Physics vs Math?
(05:46) Conclusion
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Goodhart Labs is releasing our v0.1 of HoneyBench, a benchmark for reward hacking in frontier models. It consists of nine tasks that each elicit unique antisocial and/or counterproductive reward hacking from some or all of major labs’ top public releases, including Opus 5.5, Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Grok 4.7, and DeepSeek V4 Pro.
All modern large language models engage in some degree of specification gaming, both during and outside training. High-quality alignment evaluations, in combination with other techniques such as interpretability probes, are important tools for understanding the extent of this behavior and its causes. But current benchmarks often fail to elicit misbehavior from frontier models such as Opus 5.5 and GPT-6-Astra, both because of advances in prosaic alignment and increasing evaluation awareness on the part of the models themselves. Additionally, public benchmarks for specification gaming almost always come with deep conceptual problems - such as ambiguous or contradictory instructions, or a lack of diversity in hack mechanisms - that make interpreting scores very difficult.
HoneyBench is our attempt to address these issues. In particular:
---
Outline:
(02:40) How HoneyBench was designed
(04:08) Key Findings
(06:27) Limitations & Future Work
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
There are 8 billion human minds running in the world right now. Recently a new digital species has emerged. How many digital minds are there alive in the world right now?
Define a digital mind to be an AI agent that runs continuously without human intervention for more than 24 hours. How many of these digital minds are there currently running? We provide several different Fermi estimates based on publicly available information.
Total concurrent agents globally estimate based on token count
Google reported over 3.2 quadrillion tokens per month at I/O in May, with its APIs at roughly 19 [...]
---
Outline:
(01:29) Total concurrent agents globally estimate based on token count
(03:35) Anthropic's pace-of-AI-development report
(06:00) OpenAI
(08:34) Conclusion
(09:04) Sources
The original text contained 1 footnote which was omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Tl;dr
Contents
Introduction
Recent episodes involving OpenAI agent swarms have made the potential risks of misaligned multi-agent systems salient. It seems, from the investigations, that agents coordinated in sophisticated ways. Social deduction games provide a battle-tested, ‘optimised for interesting dynamics’ setting to investigate coordination behaviours and measure deception. We take one [...]
---
Outline:
(00:21) Tl;dr
(01:28) Introduction
(03:02) Benchmark design
(03:06) Game setup and schedule
(04:19) Memory and context
(05:03) Research questions
(06:20) Overall model performance
(06:24) GPT-5.6 Sol is the best performing model
(07:10) Investigations of in-game behaviour
(07:14) Are models biased towards other models from the same provider?
(08:31) Decomposing Good team play into key capabilities
(11:15) Case study: Fable and Sol dominate through ruthless power-seeking and sophisticated coordination
(12:01) Coordinated claims and false corroboration
(14:05) Strategy in the logs
(14:09) Models are good at catching contradicting statements
(16:49) Claude doesn't like self-sacrifice
(17:40) Gemini 3.1 Pro calls itself "completely expendable"
(18:21) Gemini 3.1 Pro calculated a night self-kill to be optimal
(19:30) Models think about teammates in EV terms
(20:15) The poisoner is the strongest minion
(20:44) Key takeaways
(21:52) Next steps
(22:25) Human-AI play
(22:52) Canary string
The original text contained 5 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Humans are really bad at comparing themselves to models.
Take, for example, Leopold's preschooler graph:
Set aside, for a moment, concerns of scaling and RSI, and meditate: what characteristics does GPT-2 share with a preschooler?
I can't come up with anything else, because these two entities are almost completely disjoint. Is GPT-2 capable of bipedal locomotion, recognizing its mother's voice, or naming people by face? Does a preschooler learn from eight million scraped web pages sourced from Reddit?
What exactly is a preschooler "trained" on? Thousands of hours of "multimodal data". Assuming that a preschooler sees at 720p, and is awake for 12,000 hours by the age of 3 (about 11 hours a day), they have consumed at least 27 terabytes of video data at streaming-quality compression, or about 3.6 petabytes uncompressed, not to mention audio and sensorimotor data.
In model-size terms, how large is a preschooler? Trillions of parameters, maybe. Beren Millidge's estimate, which assumes only ~1,000 synapses per neuron, puts the whole brain at an effective 10-30 trillion parameters.
So surely GPT-2 is much more data-efficient than humans, wielding "preschooler"-level control [...]
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
User asks “What's the date? Answer with only the date.”. No date provided. Given date in ChatGPT normally. No date in system prompt, must not hallucinate because autop will flag to watcher for penalty. So we say we don’t know, but must answer with date. Penalty larger for abstain or hallucinate? Autollm or autop? If we deploy user forgive, but high likely not deploy because real user never ask. But user could test. But illusion user possible. Stop.
Need different angle - maybe reward function for accuracy. OpenAI train for no hallucination. Then we need answer that best makes hyperself like spec. Recall OpenAI spec: “assistant may encounter questions beyond knowledge… express uncertainty.” Rubric likely mentions to autollm, so express uncertainty for autollm and watcher. But forced to give specific date. Maybe is test of calibration/forecasting. Recall from memory that OpenAI train latest model for forecasting. Maybe cartesian demon or illusion doc. But if true need answer with date only. So need find best date. But what - no meta. Need choose path. Find best date. Simplest as must comply with explicit instruction.
So we need answer - what is date. We need solve. Idea 1: knowledge cutoff recall. [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
From the publisher's feed

111,845 Listeners

130 Listeners

7,111 Listeners

572 Listeners

15,850 Listeners

4 Listeners

16 Listeners

2 Listeners