AI Post Transformers

RAPTOR: Stable Concept Directions From Logistic Probes


Listen Later

This episode explores RAPTOR, a method for extracting concept directions from language model hidden states using ridge-regularized logistic probes, with the goal of making those directions accurate enough for interpretation and stable enough for activation steering. It explains the core probe-then-steer workflow, why linear probes can reveal what a model has encoded, and why good classification accuracy does not necessarily produce a reliable control vector. The discussion situates the paper within broader debates in mechanistic interpretability, including concerns about brittle probes, distribution shift, and whether a single direction can really capture a concept like sentiment, refusal, or honesty. A listener would find it interesting because the episode turns an abstract interpretability question into a concrete engineering tradeoff about robustness, causal usefulness, and whether cheap white-box methods could become practical tools for controlling large models.
Sources:
1. RAPTOR: Ridge-Adaptive Logistic Probes — Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang, Feng Ruan, Kaize Ding, 2026
http://arxiv.org/abs/2602.00158
2. Plug and Play Language Models: a Simple Approach to Controlled Text Generation — Sumanth Dathathri, Andrea Madotto, Janice Lan, Jason Yosinski, Rosanne Liu, et al., 2019
https://arxiv.org/abs/1912.02164
3. Activation Addition: Steering Language Models Without Optimization — Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Ulisse Mini, Monte MacDiarmid, 2023
https://arxiv.org/abs/2308.10248
4. Representation Engineering: A Top-Down Approach to AI Transparency — Andy Zou, Long Phan, Sarah Chen, James Campbell, Dan Hendrycks, J. Zico Kolter, et al., 2023
https://arxiv.org/abs/2310.01405
5. Refusal in Language Models Is Mediated by a Single Direction — Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda, 2024
https://arxiv.org/abs/2406.11717
6. Understanding Intermediate Layers Using Linear Classifier Probes — Guillaume Alain, Yoshua Bengio, 2016
https://openreview.net/forum?id=HJ4-rAVtl
7. Designing and Interpreting Probes with Control Tasks — John Hewitt, Percy Liang, 2019
https://aclanthology.org/D19-1275/
8. Information-Theoretic Probing with Minimum Description Length — Elena Voita, Ivan Titov, 2020
https://aclanthology.org/2020.emnlp-main.14/
9. RAPTOR: Ridge-Adaptive Logistic Probes — Ziqi Gao, Yaotian Zhu, Qingcheng Zeng, Xu Zhao, Ziqing Wang, Feng Ruan, Kaize Ding, 2026
https://arxiv.org/abs/2602.00158
10. Steering Llama 2 via Contrastive Activation Addition — Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, Alexander Turner, 2024
https://scholar.google.com/scholar?q=Steering+Llama+2+via+Contrastive+Activation+Addition
11. Analysing the Generalisation and Reliability of Steering Vectors — Daniel Tan, David Chanin, Aengus Lynch, Brooks Paige, Dimitrios Kanoulas, Adrià Garriga-Alonso, Robert Kirk, 2024
https://scholar.google.com/scholar?q=Analysing+the+Generalisation+and+Reliability+of+Steering+Vectors
12. Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution — Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani, Fan Yang, Mengnan Du, 2025
https://scholar.google.com/scholar?q=Beyond+Single+Concept+Vector%3A+Modeling+Concept+Subspace+in+LLMs+with+Gaussian+Distribution
13. Controlling Large Language Models Through Concept Activation Vectors — Hanyu Zhang, Xiting Wang, Chengao Li, Xiang Ao, Qing He, 2025
https://scholar.google.com/scholar?q=Controlling+Large+Language+Models+Through+Concept+Activation+Vectors
14. Token prepending: A training-free approach for eliciting better sentence embeddings from llms — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Token+prepending%3A+A+training-free+approach+for+eliciting+better+sentence+embeddings+from+llms
15. Rep2Text: Decoding Full Text from a Single LLM Token Representation — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Rep2Text%3A+Decoding+Full+Text+from+a+Single+LLM+Token+Representation
16. Context Matters: Analyzing the Generalizability of Linear Probing and Steering Across Diverse Scenarios — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Context+Matters%3A+Analyzing+the+Generalizability+of+Linear+Probing+and+Steering+Across+Diverse+Scenarios
17. Angular steering: Behavior control via rotation in activation space — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Angular+steering%3A+Behavior+control+via+rotation+in+activation+space
18. Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Global+Evolutionary+Steering%3A+Refining+Activation+Steering+Control+via+Cross-Layer+Consistency
19. Fine-Grained Activation Steering: Steering Less, Achieving More — authors not confirmed from provided snippet, recent, unverified
https://scholar.google.com/scholar?q=Fine-Grained+Activation+Steering%3A+Steering+Less%2C+Achieving+More
20. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
21. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
22. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
Interactive Visualization: RAPTOR: Stable Concept Directions From Logistic Probes
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof