May 17, 2024

“Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems” by Joar Skalse

5 minutes

Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.

I want to draw attention to a new paper, written by myself, David "davidad" Dalrymple, Yoshua Bengio, Stuart Russell, Max Tegmark, Sanjit Seshia, Steve Omohundro, Christian Szegedy, Ben Goldhaber, Nora Ammann, Alessandro Abate, Joe Halpern, Clark Barrett, Ding Zhao, Tan Zhi-Xuan, Jeannette Wing, and Joshua Tenenbaum.

In this paper we introduce the concept of "guaranteed safe (GS) AI", which is a broad research strategy for obtaining safe AI systems with provable quantitative safety guarantees. Moreover, with a sufficient push, this strategy could plausibly be implemented on a moderately short time scale. The key components of GS AI are:

A formal safety specification that mathematically describes what effects or behaviors are considered safe or acceptable.
A world model that provides a mathematical description of the environment of the AI system.
A verifier that provides a formal proof [...]

---

First published:

May 17th, 2024

Source:

https://www.lesswrong.com/posts/LkECxpbjvSifPfjnb/towards-guaranteed-safe-ai-a-framework-for-ensuring-robust-1

---

Narrated by TYPE III AUDIO.

...more

View all episodes

By LessWrong