
Sign up to save your podcasts
Or


The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely.
He is fascinated by Ye Wenjie, the researcher that turned against humanity in that book, and decides that in the future if AI progress leads to a superior intelligence, he might be tempted to become Ye if there's no good alternative.
He performs the duties of capabilities research in a Chinese frontier lab, seeking to one day achieve parity with Western companies, though he knows this is difficult. He has a mentality of hillclimbing, believing that the progress of a future technology is highly uncertain and even unknowable, and so him and his peers could only tread one step at a time.
He looks at the western world and sees what is typical when a great technology is developed: the first mover will decide to impose restrictions to further their lead, while latecomers should use whatever means necessary to widen access to the whole world. He thinks of the AI chip restrictions as evidence of this.
He uses Anthropic and OpenAI models regularly in his day to day work. He [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
(Archive link)
The NYT editorial board's article on AI (archive link) is far better than I'd expected, but at the same time not all I'd hoped for.
The title sets off very well: "Humanity Has Avoided Apocalypse Before. Let's Do It Again." It is truly excellent to see the extinction threat from loss of control be mainlined.
A quick gloss of their policy requests: an AI Commission in government, licensing requirements for AI companies, an AI "constitution" written by the US Government incorporated into AIs, mandatory watermarks/identifiers on all AI content, mandatory independent testing for AI models before release, and a government agency to investigate accidents. Internationally, they call for tightening export controls, limiting China's access to semiconductors, and ultimately negotiating an international slowdown with China and an international framework for AI oversight.
These are all steps in the right direction—of taking AI seriously. That said, it isn't clear if the licensing is required for training or for selling AIs. The idea that constitutional AI "would ensure alignment with human values" is of course not remotely true. And mandatory testing should apply to all models trained, not all models released, of course, and this is a glaring oversight. But [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future:
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Most conversations about AI risks seem like people are talking past each other. There's a lot of strawmanning. This is largely because the issue is complex, and in order to make it manageable, the concepts get oversimplified. The most common example of this is p(doom), collapsing extreme outcomes into a single variable and focusing on that. Critics happily jump on the fact that there are many other possible bad outcomes, or that 100% extinction rates are an extremely high bar. There's some legitimacy to this.
So here I want to try to grapple with the the complexity of the larger web of issues surrounding bad AI outcomes by presenting The AI Risk Network. If you’re interested in this topic, bear with me. It might be a bit of a slog.
First I want to contrast this approach with others. Liron Shapira has what he calls The Doom Train, a linear progression through various dependencies or thresholds that eventually lead to human extinction, with various ‘stops’ along the way where the skeptic can get off.
Shapira uses this as a discussion guide to focus on particular points where the skeptic gets off the train and exits belief in the extreme [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI.
There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed.
Table of Contents
Our Two Problems
Anthropic: Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents:
---
Outline:
(00:33) Our Two Problems
(02:26) First the Good News
(03:02) We'd Just Like To Ask You a Few Questions
(04:12) Internal Research Model On The Fence
(07:28) Opus 4.7
(08:12) Opus 4.6 Checkpoint
(09:49) Holy **** That Thing's Real?
(11:45) I Thought I Saw a Pussycat
(19:28) If This Was Real You Would Never Tell Me It Was Real
(21:19) New Eval Who Dis
(26:32) Hacker Opus
(30:15) Monitoring the Situation
(31:38) Overcoming Bias
(33:40) The Anthropic Alignment Problem
(35:53) Paths Forward
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Inkhaven is a writers residency in Berkeley, in which the only requirement is you have to publish 500 words each and every day. Though I always had some confidence in my ability to write, I never actually did it much until I applied to Inkhaven. I had finished only two short stories before I applied: The Maker of MIND and The Liar and the Scold. And it was them I used in my application.
In the roughly twelve months since I was accepted, I have written thirteen, and even some half-finished things that will never see the light of day. And this isn’t including the essays and micro-fiction I wrote during the fellowship. By the metric of getting me to write more, Inkhaven was a great success. And would have been worth it even if I had a miserable time.
Despite a slight proclivity for having miserable times, I found myself unable to do so for long at Inkhaven. I rarely write utopias, and when I do they curdle by the time the story ends. But I suspect utopia will feel a lot like Inkhaven did for me once I got settled. You would think putting a bunch [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL”
A common take I’ve been hearing is: “LLMs are especially good at math because math is easy to verify”. But that story doesn’t make much sense.
So here's a different theory, in the framework of my earlier post “LLMs are (still) mostly powered by imitative learning, not RL”:
LLMs are especially good at math because almost everything in the math literature is correct. Read a random sentence in a random math paper in the research math literature, and you can be >99% confident that the sentence is true. So if LLMs do what they do best—imitative [...]
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
I'm mostly hoping this somehow gets sent to a privately disgruntled frontier lab employee, but it would also be cool to expand other people's minds on the way there.
I read through Ethical AI Departures and would like to note that only a few of them have gotten extensive media coverage and none of them have actually effectively gotten the frontier labs to stop, and that collectively signed letters by employees have historically not done much either.
I read Dear God, Please Do Not Resign In Protest and wanted to point out that leftists have a mature and relatively reliable set of strategies to address the problem of how to get a lot of people to stop working in protest at the same time.
Then I did a search of LW to see if someone else brought unions up already, read What if AI safety labs unionized?, and flinched at the repeated citation of legal reasons why a union isn't the correct legal structure. So no, what you want right now isn't an official, bureaucratic union. In fact, that would probably slow things down too much.
But I've done enough work with union people to know that you don't [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country's national security forces.
No company, no government, no individual knows how to keep such a system under human control. This is why the world's leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence.
This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill [...]
---
Outline:
(04:34) Secure Weapons-Grade AI Against Theft by Adversaries
(07:46) Necessary Measure: Registration
(08:43) Sufficient Measure: Government Security Testing
(09:42) Thorough Measure: Development Requires Government Authorization
(10:52) Criminal Liability for Leaks During AI Gain-of-Function Research
(14:09) Necessary Measure: Team Liability
(14:46) Sufficient Measure: Chain of Command Liability
(15:21) Thorough Measure: Company Liability
(16:00) Kill-Switches to Contain Critical AI Incidents
(19:08) Necessary Measure: Company Kill-Switch
(19:58) Sufficient Measure: Infrastructure Kill-Switch
(20:53) Thorough Measure: International Kill-Switches
(22:24) Conclusion
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
From the publisher's feed

111,845 Listeners

130 Listeners

7,111 Listeners

572 Listeners

15,850 Listeners

4 Listeners

16 Listeners

2 Listeners