March 20, 2024

“New report: Safety Cases for AI” by joshc

1 minute

Here's the Tweet thread:
https://twitter.com/joshua_clymer/status/1770467746951868889

The idea for this paper occurred to me when I saw Buck Shlegeris' MATS stream on "Safety Cases for AI." How would one justify the safety of advanced AI systems? This question is fundamental. It informs how RSPs should be designed and what technical research is useful to pursue.

For a long time, researchers have (implicitly or explicitly) discussed ways to justify that AI systems are safe, but much of this content is scattered across different posts and papers, is not as concrete as I'd like, or does not clearly state the assumptions being made.

I hope this report provides a helpful birds-eye view of safety arguments and moves the AI safety conversation forward by identifying their key assumptions.

Thanks to my coauthors: Nick Gabrieli, David Krueger, and Thomas Larsen -- and to everyone who gave feedback: Henry Sleight, Ashwin Acharya, Ryan Greenblatt, Stephen [...]

---

First published:

March 20th, 2024

Source:

https://www.lesswrong.com/posts/HrtyZm2zPBtAmZFEs/new-report-safety-cases-for-ai

---

Narrated by TYPE III AUDIO.

...more

View all episodes

By LessWrong