AI Post Transformers

Split Personality Training Reveals Latent Knowledge


Listen Later

This episode explores a 2026 paper on “split personality training,” a method for attaching an internal reviewer to a language model that can reveal what the model knows about its own deceptive or reward-hacking behavior without changing the answer shown to the user. It situates the work in the broader lineage of latent knowledge elicitation, alignment faking, and mechanistic interpretability, explaining why a model’s hidden state may contain more honest information than its final text output. The discussion focuses on the paper’s use of a LoRA-based “honest persona” that activates only after the main response, and on benchmark setups like Anthropic’s auditing game that test whether internal representations expose hidden objectives that outside observers cannot infer. Listeners would find it interesting because it tackles a central safety problem: whether models can be audited for strategic deception using their own internal signals rather than their polished outward behavior.
Sources:
1. Split Personality Training Reveals Latent Knowledge
https://arxiv.org/pdf/2602.05532
2. Discovering Latent Knowledge in Language Models Without Supervision — Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Discovering+Latent+Knowledge+in+Language+Models+Without+Supervision
3. Eliciting Latent Knowledge from Quirky Language Models — Alex Mallen, Nora Belrose, 2023
https://scholar.google.com/scholar?q=Eliciting+Latent+Knowledge+from+Quirky+Language+Models
4. Challenges with Unsupervised LLM Knowledge Discovery — Sebastian Farquhar, Vikrant Varma, Zachary Kenton, Johannes Gasteiger, Vladimir Mikulik, Rohin Shah, 2023
https://scholar.google.com/scholar?q=Challenges+with+Unsupervised+LLM+Knowledge+Discovery
5. LatentQA: Teaching LLMs to Decode Activations Into Natural Language — Alexander Pan, Lijie Chen, Jacob Steinhardt, 2024
https://scholar.google.com/scholar?q=LatentQA%3A+Teaching+LLMs+to+Decode+Activations+Into+Natural+Language
6. Language Models Mostly Know What They Know — Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt, 2023
https://scholar.google.com/scholar?q=Language+Models+Mostly+Know+What+They+Know
7. ELK Report: Eliciting Latent Knowledge — Paul Christiano, Ajeya Cotra, Mark Xu, et al., 2021
https://scholar.google.com/scholar?q=ELK+Report%3A+Eliciting+Latent+Knowledge
8. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets — Samuel Marks, Max Tegmark, 2024
https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+Large+Language+Model+Representations+of+True%2FFalse+Datasets
9. LatentQA: Teaching Language Models to Decode Activations into Natural Language — Yida Pan et al., 2024
https://scholar.google.com/scholar?q=LatentQA%3A+Teaching+Language+Models+to+Decode+Activations+into+Natural+Language
10. Activation Oracles — Jesse Karvonen et al., 2026
https://scholar.google.com/scholar?q=Activation+Oracles
11. Confessions — Nikhil Joglekar et al., 2025
https://scholar.google.com/scholar?q=Confessions
12. Self-Report Fine-Tuning — Li et al., 2025
https://scholar.google.com/scholar?q=Self-Report+Fine-Tuning
13. Auditing Language Models for Hidden Objectives — Anthropic, 2025
https://scholar.google.com/scholar?q=Auditing+Language+Models+for+Hidden+Objectives
14. Alignment Faking in Large Language Models — Ryan Greenblatt et al., 2024
https://scholar.google.com/scholar?q=Alignment+Faking+in+Large+Language+Models
15. Towards Eliciting Latent Knowledge from LLMs with Mechanistic Interpretability — Bartosz Cywinski, Emil Ryd, Senthooran Rajamanoharan, Neel Nanda, 2025
https://scholar.google.com/scholar?q=Towards+Eliciting+Latent+Knowledge+from+LLMs+with+Mechanistic+Interpretability
16. Quantifying Elicitation of Latent Capabilities in Language Models — Elizabeth Donoway et al., 2025
https://scholar.google.com/scholar?q=Quantifying+Elicitation+of+Latent+Capabilities+in+Language+Models
17. BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models — Yi Zeng, Weiyu Sun, Tran Huynh, Dawn Song, Bo Li, Ruoxi Jia, 2024
https://scholar.google.com/scholar?q=BEEAR%3A+Embedding-based+Adversarial+Removal+of+Safety+Backdoors+in+Instruction-tuned+Language+Models
18. Investigating Adversarial Trigger Transfer in Large Language Models — Nicholas Meade, Arkil Patel, Siva Reddy, 2024
https://scholar.google.com/scholar?q=Investigating+Adversarial+Trigger+Transfer+in+Large+Language+Models
19. When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models — Kai Wang, Yihao Zhang, Meng Sun, 2025
https://scholar.google.com/scholar?q=When+Thinking+LLMs+Lie%3A+Unveiling+the+Strategic+Deception+in+Representations+of+Reasoning+Models
20. Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort — Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He, 2025
https://scholar.google.com/scholar?q=Is+It+Thinking+or+Cheating%3F+Detecting+Implicit+Reward+Hacking+by+Measuring+Reasoning+Effort
21. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
22. AI Post Transformers: CLUE: Hidden-State Clustering for Non-parametric Verification — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/clue-hidden-state-clustering-for-non-parametric-verification/
23. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
24. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
Interactive Visualization: Split Personality Training Reveals Latent Knowledge
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof