This episode explores CONTXT, a training-free method for correcting distribution shift by adding a single precomputed "context vector" directly into a model's internal activations—no fine-tuning, no paired prompts, and no gradient updates required. The discussion traces the paper's neuroscience grounding in dual-process theory, where the hippocampus rapidly encodes context and hands it to the prefrontal cortex to amplify relevant features and suppress irrelevant ones, and examines how faithfully that analogy maps onto a simple additive vector operation. It also clarifies the distinction between domain generalization (no access to target data at all) and test-time adaptation (unlabeled target data available at inference), situating CONTXT within existing activation-steering approaches like those requiring token-level paired prompts. Listeners get a walkthrough of the core math—h plus alpha times an index vector, extendable to multiple stacked contexts for simultaneous edits like adjusting tone while removing sarcasm—before the hosts turn to concrete demonstrations, including a striking out-of-distribution image classification example. The conversation is notable for its skepticism: one host pushes back hard on whether a two-region brain theory can really license a one-line vector subtraction, making this as much a critique of steering-paper rigor as an explainer of the method itself.
Sources:
1. Context is All You Need — Jean Erik Delanois, Shruti Joshi, Ryan Golden, Teresa Nick, Maxim Bazhenov, 2026
http://arxiv.org/abs/2604.04364
2. In Search of Lost Domain Generalization — Ishaan Gulrajani, David Lopez-Paz, 2020
https://scholar.google.com/scholar?q=In+Search+of+Lost+Domain+Generalization
3. Deep CORAL: Correlation Alignment for Deep Domain Adaptation — Baochen Sun, Kate Saenko, 2016
https://scholar.google.com/scholar?q=Deep+CORAL%3A+Correlation+Alignment+for+Deep+Domain+Adaptation
4. Domain-Adversarial Training of Neural Networks — Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, Victor Lempitsky, 2016
https://scholar.google.com/scholar?q=Domain-Adversarial+Training+of+Neural+Networks
5. Deeper, Broader and Artier Domain Generalization (the PACS dataset) — Da Li, Yongxin Yang, Yi-Zhe Song, Timothy M. Hospedales, 2017
https://scholar.google.com/scholar?q=Deeper%2C+Broader+and+Artier+Domain+Generalization+%28the+PACS+dataset%29
6. Tent: Fully Test-Time Adaptation by Entropy Minimization — Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, Trevor Darrell, 2021
https://scholar.google.com/scholar?q=Tent%3A+Fully+Test-Time+Adaptation+by+Entropy+Minimization
7. Improving Robustness Against Common Corruptions by Covariate Shift Adaptation — Steffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann, Wieland Brendel, Matthias Bethge, 2020
https://scholar.google.com/scholar?q=Improving+Robustness+Against+Common+Corruptions+by+Covariate+Shift+Adaptation
8. Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation — Jian Liang, Dapeng Hu, Jiashi Feng, 2020
https://scholar.google.com/scholar?q=Do+We+Really+Need+to+Access+the+Source+Data%3F+Source+Hypothesis+Transfer+for+Unsupervised+Domain+Adaptation
9. MEMO: Test Time Robustness via Adaptation and Augmentation — Marvin Zhang, Sergey Levine, Chelsea Finn, 2022
https://scholar.google.com/scholar?q=MEMO%3A+Test+Time+Robustness+via+Adaptation+and+Augmentation
10. Steering Llama 2 via Contrastive Activation Addition — Panickssery, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., Turner, A. M., 2023
https://scholar.google.com/scholar?q=Steering+Llama+2+via+Contrastive+Activation+Addition
11. Representation Engineering: A Top-Down Approach to AI Transparency — Zou, A., Gao, L., Greenblatt, R., et al., 2023
https://scholar.google.com/scholar?q=Representation+Engineering%3A+A+Top-Down+Approach+to+AI+Transparency
12. Extracting Latent Steering Vectors from Pretrained Language Models — Subramani, N., Suresh, N., Peters, M. E., 2022
https://scholar.google.com/scholar?q=Extracting+Latent+Steering+Vectors+from+Pretrained+Language+Models
13. Distributionally Robust Neural Networks for Group Shifts (GroupDRO) — Sagawa, S., Koh, P. W., Hashimoto, T. B., Liang, P., 2020
https://scholar.google.com/scholar?q=Distributionally+Robust+Neural+Networks+for+Group+Shifts+%28GroupDRO%29
Interactive Visualization: Context is All You Need: Fixing OOD Drift Without Retraining