June 09, 2025

[Linkpost] “Identifying ‘Deception Vectors’ In Models” by Stephen Martin

Listen Later

1 minute

This is a link post. Using representation engineering, we systematically induce, detect, and control such deception in CoT-enabled LLMs, extracting ”deception vectors” via Linear Artificial Tomography (LAT) for 89% detection accuracy. Through activation steering, we achieve a 40% success rate in eliciting context-appropriate deception without explicit prompts, unveiling the specific honesty related issue of reasoning models and providing tools for trustworthy AI alignment.

This seems like a positive breakthrough for mech interp research generally, the team used RepE to identify features, and were able to "reliably suppress or induce strategic deception".

---

First published:
June 9th, 2025

Source:
https://www.lesswrong.com/posts/3WyFmtiLZTfEQxJCy/identifying-deception-vectors-in-models

Linkpost URL:
https://arxiv.org/pdf/2506.04909

---

Narrated by TYPE III AUDIO.

...more

View all episodes

View all episodes

Download on the App Store

Download on the App Store

Get it on Google Play

LessWrong (Curated & Popular)

By LessWrong

4.8

1212 ratings

June 09, 2025

[Linkpost] “Identifying ‘Deception Vectors’ In Models” by Stephen Martin

Listen Later

1 minute

This is a link post. Using representation engineering, we systematically induce, detect, and control such deception in CoT-enabled LLMs, extracting ”deception vectors” via Linear Artificial Tomography (LAT) for 89% detection accuracy. Through activation steering, we achieve a 40% success rate in eliciting context-appropriate deception without explicit prompts, unveiling the specific honesty related issue of reasoning models and providing tools for trustworthy AI alignment.

This seems like a positive breakthrough for mech interp research generally, the team used RepE to identify features, and were able to "reliably suppress or induce strategic deception".

---

First published:
June 9th, 2025

Source:
https://www.lesswrong.com/posts/3WyFmtiLZTfEQxJCy/identifying-deception-vectors-in-models

Linkpost URL:
https://arxiv.org/pdf/2506.04909

---

Narrated by TYPE III AUDIO.

...more

More shows like LessWrong (Curated & Popular)

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,388 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,424 Listeners

Robert Wright's Nonzero by Nonzero

Robert Wright's Nonzero

590 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

107 Listeners

The Good Fight by Yascha Mounk

The Good Fight

904 Listeners

ManifoldOne by Steve Hsu

ManifoldOne

92 Listeners

The Prof G Pod with Scott Galloway by Vox Media Podcast Network

The Prof G Pod with Scott Galloway

5,477 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

89 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

488 Listeners

Hard Fork by The New York Times

Hard Fork

5,475 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

132 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

133 Listeners

The Marginal Revolution Podcast by Mercatus Center at George Mason University

The Marginal Revolution Podcast

93 Listeners

Statecraft by Santi Ruiz

Statecraft

34 Listeners

The Last Invention by Longview

The Last Invention

292 Listeners