Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is Truthful AI, published by Owen Cotton-Barratt on October 20, 2021 on The Effective Altruism Forum.
This post contains the abstract and executive summary of a new 96-page paper (subtitle: Developing and governing AI that does not lie) from authors at the Future of Humanity Institute, Global Priorities Institute, and OpenAI. Also posted on LessWrong.
I believe that over the next several years this might be an important field on longtermist grounds for two disjoint reasons:
Maybe there's a window of opportunity to set standards for behaviour of linguistic AI systems, and higher standards could improve global epistemics and coordination.
AI truthfulness is more concrete and approachable a target than "avoiding AI risk", but has several structural similarities, and I think we'd learn a lot (on both the technical and governance sides) from tackling it.
Abstract
In many contexts, lying – the use of verbal falsehoods to deceive – is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the question of how we should limit the harm caused by AI “lies” (i.e. falsehoods that are actively selected for). Human truthfulness is governed by social norms and by laws (against defamation, perjury, and fraud). Differences between AI and humans present an opportunity to have more precise standards of truthfulness for AI, and to have these standards rise over time. This could provide significant benefits to public epistemics and the economy, and mitigate risks of worst-case AI futures.
Establishing norms or laws of AI truthfulness will require significant work to:
identify clear truthfulness standards;
create institutions that can judge adherence to those standards; and
develop AI systems that are robustly truthful.
Our initial proposals for these areas include:
a standard of avoiding “negligent falsehoods” (a generalisation of lies that is easier to assess);
institutions to evaluate AI systems before and after real-world deployment;
explicitly training AI systems to be truthful via curated datasets and human interaction.
A concerning possibility is that evaluation mechanisms for eventual truthfulness standards could be captured by political interests, leading to harmful censorship and propaganda. Avoiding this might take careful attention. And since the scale of AI speech acts might grow dramatically over the coming decades, early truthfulness standards might be particularly important because of the precedents they set.
Executive Summary & Overview
The threat of automated, scalable, personalised lying
Today, lying is a human problem. AI-produced text or speech is relatively rare, and is not trusted to reliably convey crucial information. In today’s world, the idea of AI systems lying does not seem like a major concern.
Over the coming years and decades, however, we expect linguistically competent AI systems to be used much more widely. These would be the successors of language models like GPT-3 or T5, and of deployed systems like Siri or Alexa, and they could become an important part of the economy and the epistemic ecosystem. Such AI systems will choose, from among the many coherent statements they might make, those that fit relevant selection criteria — for example, an AI selling products to humans might make statements judged likely to lead to a sale. If truth is not a valued criterion, sophisticated AI could use a lot of selection power to choose statements that further their own ends while being very damaging to others (without necessarily having any intention to deceive – see Diagram 1). This is alarming because AI untruths could potentially scale, with one system telling personalised lies to millions of people.
Aiming for robustly beneficial standards
Widespread and damaging A...