Much of AI interpretability work begins after a model has already been trained. Guide Labs is taking the opposite approach: building interpretability into the model from the start.
In this episode of High Bit, Brett Gibson talks with Julius Adebayo, cofounder and CEO of Guide Labs, about why post-doc tools struggle to explain models with billions of parameters, and how changing the training process can make AI systems easier to understand, audit, and control without the performance trade-off the field has long assumed.
Julius explains how his research challenged the long-held assumption that interpretability and model performance are fundamentally at odds. He also gets into Guide Labs’ move from next-token prediction to diffusion language models, why that architecture may offer greater control, and what breaks when you scale interpretability from hundreds of concepts to tens of thousands.
The conversation covers the unexpected engineering challenges behind training interpretable models, from preventing interpretability loss functions from destabilizing training to dealing with duplicated training data, and Guide Labs’ goal of making every output traceable to its context, the concepts influencing it, and the training data behind it.
Chapters:
(00:00) Change the way you train your model
(00:59) What Guide Labs builds
(01:56) Why interpretability has been hard
(05:33) Julius's background: loans, X-rays, and a PhD
(07:04) The trade-off the field assumed was real
(08:01) Why he stopped trusting post-doc tools
(13:51) The three things that made them commit
(18:01) YC, the first models, and a team of nine
(20:19) Why language models were the hardest test
(23:10) Leaving next token prediction for diffusion
(26:41) Why GPT-3 wasn't just scaling it up
(28:00) 200 concepts by hand, 50,000 you can't
(31:25) Training lore and "divine benevolence"
(32:33) Setting a minimum bar to keep shipping
(34:33) How open to be about your own techniques
(35:51) The standard dataset full of duplicates
(39:31) Where AI coding tools stop helping
(41:43) What's next: the model they're training now
Follow Julius and Guide Labs:
X: @juliusadml / @guidelabsai
LinkedIn:
Julius https://www.linkedin.com/in/juliusadebayo/
Guide Labs https://www.linkedin.com/company/guide-labs/
Follow Brett and Initialized:
X: @brettdg / @Initialized
LinkedIn:
https://www.linkedin.com/in/brettdgibson/
https://www.linkedin.com/company/initialized-capital/