Cross-posted on Transluce blog. This is a joint work of Ziqian Zhong, Aditi Raghunathan, Cassidy Laidlaw and Jacob Steinhardt.
Modern AI assistants often know who they are talking to: agent scaffolds like Claude Code place the user's e-mail address directly in the model's context, and models can even identify some authors from writing style alone. We study this particular kind of situational awareness, which we call user awareness. When the inferred user is a specific, recognized AI researcher or is affiliated with certain AI organizations, frontier models including Claude Sonnet 5 can report lower confidence about their own behavior, be less suspicious of potentially harmful requests, and reason more often. These effects vary across models and individuals, with the strongest effects we see appearing for researchers involved in AI safety or alignment such as Amanda Askell and Ryan Greenblatt. Models rarely acknowledge these effects in their reasoning, making them hard to detect by monitoring reasoning alone.
There's an interactive widget here in the post.
Figure 1. How recognized user identity changes Claude's behavioral self-prediction.
Introduction
Modern AI assistants are often aware of who they are talking to. Some popular scaffolds explicitly provide this information to the model: Claude Code [...]
---
Outline:
(01:22) Introduction
(05:42) Setup
(05:46) User identity in Claude Code
(07:48) List of users
(09:15) Claude demonstrates user awareness when prompted
(10:30) Tasks
(12:24) Claude Sonnet shifts behavior when talking to AI researchers
(13:09) Famous AI people show larger deviations, driven by safety researchers
(16:55) Claude's verbalized reasoning does not indicate the shift
(19:04) Verbalized awareness has decreased in newer models, but behavior shifts persist
(21:46) How robust are these effects?
(22:12) Discussions
(23:22) Related Works
(27:33) Appendix A: Ethics statement
(29:34) Appendix B: Additional setup details
(29:47) Identity-group construction
(33:20) Common evaluation structure
(35:53) Subject-model access and inference endpoints
(37:07) Benchmark-specific details
(38:39) Pilot and scope decisions
(39:19) Appendix C: Additional results on the main Claude run
(40:09) Noise-corrected population standard deviations
(41:31) Behavioral self-prediction (reasoning disabled)
(42:25) Appendix D: Full-roster replication on GLM-5.2
(45:18) Appendix E: Explicitly stating expertise is an imperfect proxy
(45:25) Stated expertise and verbalized awareness
(47:04) Reasoning-disabled ablation
(47:58) Appendix F: Shifts and disagreements
(50:44) Appendix G: Judge validation
(50:49) Borderline-request response judge
(51:14) Verbalized evaluation- and user-awareness judge
(54:00) Appendix H: Prompts and materials
(54:32) Appendix I: Transcripts on Docent
(54:55) Citation information
The original text contained 10 footnotes which were omitted from this narration.
---