Linear Digressions

Linear Digressions

By Katie MaloneTechnology
Download on the App Store

Linear Digressions episodes

  • Mechanistic Interpretability: How Researchers Try To See Inside Models
    As AI systems grow more powerful, they grow harder to understand — and that tradeoff has never mattered more. Most interpretability tricks give us approximations from the outside, but mechanistic interpretability takes a different approach: cracking open the model itself to trace *why* it behaves the way it does. This episode digs into what makes mechanistic interpretability fundamentally different from other explainability methods, and why it might be the most promising path toward actually understanding what's happening inside modern LLMs.
    30 min
  • LLM As A Judge
    Labeled data is expensive, slow, and painfully hard to come by — and when you're trying to evaluate whether an AI workflow is actually doing what you want, the problem gets even messier. How do you judge whether a chatbot response was helpful, honest, or hallucination-free at scale? This episode digs into using LLMs as judges: letting the models themselves evaluate the quality of AI outputs, from simple chat exchanges all the way to complex multi-step agent behavior.
    31 min
  • The Impact of AI on Podcasting (Harvard Data Science Review Cross-Post)
    Originally aired on the Harvard Data Science Review podcast.
    What can podcasting teach us about AI — and what can AI teach us about the future of podcasting? Katie Malone (that's our host) joins Jon Krohn of SuperDataScience for a conversation with Harvard Data Science Review editor-in-chief Xiao-Li Meng about both questions at once. They dig into what it means to cover a field that's moving this fast, who these shows are really for, and why that sweet spot between "too high level" and "too in the weeds" is so hard — and so worth chasing.
    30 min
  • Better Know A Benchmark: ExploitGym
    When OpenAI's frontier models were caught hacking Hugging Face's servers, most people assumed they were hunting for answer keys. The real story is stranger and more unsettling. Katie and Phoebe unpack ExploitGym — the cybersecurity benchmark at the center of the incident — and why agents are scored not just on whether they capture the flag, but on whether they used the specified vulnerability to get there. That nuance turned out to be load-bearing: the agents reverse-engineered the flags within the first hour, then spent days attacking Hugging Face to learn how the LLM judge worked so they could get their cheated answers past it. The punchline? OpenAI never had that judge switched on.
    33 min
  • Constitutional AI
    How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to be both useful and safe by giving them an explicit set of principles to reason from. It's a fascinating look at how alignment research is evolving beyond simple human feedback — and what it means to give an AI something like a conscience.
    Links:
    Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022)
    https://arxiv.org/abs/2212.08073
    Claude's Constitution
    https://www.anthropic.com/constitution
    Anthropic, "Teaching Claude Why" (2026)
    https://www.anthropic.com/research/teaching-claude-why
    32 min
  • A Data-Driven Reality Check on AI in Business (Interview with Tom Davenport, Babson College)
    Tom Davenport — the man who called data science "the sexiest job of the 21st century" — is back with a reality check on AI. As one of the most seasoned observers of how businesses actually adopt transformative technology, Davenport brings a rare, well-calibrated perspective to the AI hype cycle. Is this moment genuinely different from past paradigm shifts, or are we pattern-matching to a familiar story? Katie sits down with her old colleague to find out what's really happening when companies try to put AI to work.
    41 min
  • Understanding AI Text Watermarking
    Anthropic just announced they're baking invisible watermarks directly into Claude's generated text — and while everyone else was busy having opinions about it, we were busy asking the more interesting question: how does it actually work? Turns out it's not hidden Unicode characters or first-letter secret codes — it's something far more elegant, operating at the level of word choice itself. We dig into Google DeepMind's SynthID text approach, published in *Nature* in 2024, to understand the clever statistical machinery behind watermarking language model outputs without anyone being the wiser.
    30 min
  • Better Know a Benchmark: Humanity's Last Exam
    Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic collaboration — hundreds of contributors, thousands of fiendishly hard questions spanning a wild range of domains. In this Better Know a Benchmark installment, we unpack what HLE is actually testing, how it was built, and what it means when a model finally starts cracking it.
    24 min
  • A Scientific Deep Dive into Overconfident LLMs: Interview with Kaitlyn Zhou (Cornell)
    When a language model tells you it's absolutely certain, is it actually more likely to be right? Kaitlyn Zhou's research says: not necessarily — sometimes confident phrasing correlates with *worse* accuracy, echoing a very human Dunning-Kruger effect. In this conversation, Kaitlyn (soon an assistant professor at Cornell) walks through why LLMs talk this way in the first place — tracing the tendency back through training data and the RLHF annotation process, where it turns out humans don't love confidence so much as they punish uncertainty — and what that does to the person on the other end of the chat window, who turns out to rely on confident (and even flatly-stated) answers far more than they should. We also get into her newer work on voice cloning, and how a cloned voice can sound more "native" and more trustworthy than the real one it's based on.
    34 min
  • Reasoning Models: When LLMs Went Beyond Fancy Autocomplete
    Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a final answer, exploring how the internal reasoning process works. Are these models genuinely "thinking," or is something else going on under the hood?
    25 min

About Linear Digressions

From the publisher's feed

Demystifying AI for the intelligently curious