July 30, 2024

AF - Investigating the Ability of LLMs to Recognize Their Own Writing by Christopher Ackerman

23 minutes

Welcome to The Nonlinear Library, where we use Text-to-Speech software to convert the best writing from the Rationalist and EA communities into audio. This is: Investigating the Ability of LLMs to Recognize Their Own Writing, published by Christopher Ackerman on July 30, 2024 on The AI Alignment Forum.

This post is an interim progress report on work being conducted as part of Berkeley's Supervised Program for Alignment Research (SPAR).

Summary of Key Points

We test the robustness of an open-source LLM's (Llama3-8b) ability to recognize its own outputs on a diverse mix of datasets, two different tasks (summarization and continuation), and two different presentation paradigms (paired and individual).

We are particularly interested in differentiating scenarios that would require a model to have specific knowledge of its own writing style from those where it can use superficial cues (e.g., length, formatting, prefatory words) in the text to pass self-recognition tests.

We find that while superficial text features are used when available, the RLHF'd Llama3-8b-Instruct chat model - but not the base Llama3-8b model - can reliably distinguish its own outputs from those of humans, and sometimes other models, even after controls for superficial cues: ~66-73% success rate across datasets in paired presentation and 58-83% in individual presentation (chance is 50%).

We further find that although perplexity would be a useful signal to perform the task in the paired presentation paradigm, correlations between relative text perplexity and choice probability are weak and inconsistent, indicating that the models do not rely on it.

Evidence suggests, but does not prove, that experience with its own outputs, acquired during post-training, is used by the chat model to succeed at the self-recognition task.

The model is unable to articulate convincing reasons for its judgments.

Introduction

It has recently been found that large language models of sufficient size can achieve above-chance performance in tasks that require them to discriminate their own writing from that of humans and other models. From the perspective of AI safety, this is a significant finding. Self-recognition can be seen as an instance of situational awareness, which has long been noted as a potential point of risk for AI (Cotra, 2021).

Such an ability might subserve an awareness of whether a model is in a training versus deployment environment, allowing it to hide its intentions and capabilities until it is freed from constraints. It might also allow a model to collude with other instances of itself, reserving certain information for when it knows it's talking to itself that it keeps secret when it knows it's talking to a human.

On the positive side, AI researchers could use a model's self-recognition ability as the basis to build resistance to malicious prompting. But what isn't clear from prior studies is whether the self-recognition task success actually entails a model's self-awareness of its own writing style.

Panickssery et al. (2024), utilizing a summary writing/recognition task, report that a number of LLMs, including Llama2-7b-chat, show out-of-the-box (without fine-tuning) self recognition abilities. However, this work focussed on the relationship between self-recognition task success and self-preference, rather than the specific means by which the model was succeeding at the task. Laine et al.

(2024), as part of a larger effort to provide a foundation for studying situational awareness in LLMs, utilized a more challenging text continuation writing/recognition task and demonstrate self-recognition abilities in several larger models (although not Llama2-7b-chat), but there the focus was on how task success could be elicited with different prompts and in different models.

Thus we seek to fill a gap in understanding what exactly models are doing when they succeed at a self recognition task.

We first demonstrate model self-recognition task success in a variety of domains....

...more

View all episodes

By The Nonlinear Fund

4.6

88 ratings

July 30, 2024

AF - Investigating the Ability of LLMs to Recognize Their Own Writing by Christopher Ackerman

23 minutes

This post is an interim progress report on work being conducted as part of Berkeley's Supervised Program for Alignment Research (SPAR).

Summary of Key Points

Evidence suggests, but does not prove, that experience with its own outputs, acquired during post-training, is used by the chat model to succeed at the self-recognition task.

The model is unable to articulate convincing reasons for its judgments.

Introduction

Thus we seek to fill a gap in understanding what exactly models are doing when they succeed at a self recognition task.

We first demonstrate model self-recognition task success in a variety of domains....

...more

Share AF - Investigating the Ability of LLMs to Recognize Their Own Writing by Christopher Ackerman

Sign up to save your podcasts

AF - Investigating the Ability of LLMs to Recognize Their Own Writing by Christopher Ackerman

AF - Investigating the Ability of LLMs to Recognize Their Own Writing by Christopher Ackerman