The Test Set by Posit

The Test Set by Posit

By Posit, PBCTechnology
Download on the App Store

The Test Set by Posit episodes

  • You're the Base, AI Is the Exponent — with Tareef Kawaf

    Posit CEO Tareef Kawaf’s theory about LLMs is that they work like an exponent on whoever you already are. Greater than one and you compound. Under one and you accelerate in the wrong direction. Michael, Wes, and Hadley chat with Tareef about what that means for data science, along with John Chambers' Prime Directive, why Tareef now goes up against his own board, and the cat he waited 40 years for.

    What’s Inside

    • LLMs as an exponent, and why the base is you
    • The CEO bug hunter
    • John Chambers' Prime Directive and why AI makes it matter again
    • A Kaypro 16, 128K of RAM, and an uncle in Daytona
    • The 70 percent of your job that stops being yours
    • The oud, and the cat he waited forty years for





    1 hr 18 min
  • Smuggled Gold and Untrustworthy Intuition — with Leland McInnes

    Leland McInnes has spent his career building tools a lot data scientists take for granted (UMAP, HDBSCAN), and amusingly the applications of those tools keep getting stranger. On this episode of The Test Set, he tells Michael, Hadley, and Wes about teams tracking laundered gold through mineral impurities, and the time he helped casinos optimize slot machine layouts. Also, why nobody's intuition survives contact with high-dimensional space. On this episode of The Test Set, we hit all that, plus imposter syndrome, AI as a sounding board, and the case for asking better questions instead of chasing answers.

    What's inside

    • Wild UMAP use cases, e.g., catching laundered gold
    • Why casinos pay mathematicians to rearrange slot machines
    • High-dimensional space and how it defies human intuition
    • Neural embeddings for people who aren't AI researchers
    • Imposter syndrome — across two disciplines
    • Leland’s take on AI as a rubber-duck testing partner
    • “Look at your data" because it legit beats guessing
    1 hr
  • Nobody Remembers Your 20 Charts — with Ruth Milligan

    Ruth Milligan has coached hundreds of conference speakers, curates TEDx Columbus, and wrote The Motivated Speaker. She's here to tell you that speaking is habitual, not natural. Michael, Wes, and Hadley get a live fifteen-second filler word hack, an uncomfortable truth about listening back to your own voice, and the three kinds of pitches Ruth has been reading for seventeen years.

    What's inside

    • Why a great idea beats a great speaker
    • Speaking is habitual, not natural, and that’s good
    • The fifteen-second breathing trick to kill fillers
    • How many charts do you need? Fewer than you think
    • The three pitches Ruth has read for seventeen years
    • You can't get better until you listen to yourself
    • Hadley contends people who send voice memos are monsters

    Mentioned in this episode

    • Patsy Rodenburg's "Three Circles of Energy" — a summary of the framework
    • Rébecca Kleinberger, "Why you don't like the sound of your own voice" (TED)


    1 hr 2 min
  • Let the Agent Cook — with Trevor Manz

    Trevor Manz went from measuring plant apertures by hand in a wet lab to building the notebook that lets coding agents take the wheel. The creator of anywidget and founding engineer at marimo (marimo.io/pair) popped into The Test Set to spill on reactive notebooks, why marimo pair threw out every MCP tool but one, and what agents really want out of a data environment. This conversation also features a jacket bouncer, a hidden Python API, and Michael's slow-motion war with the word "marimo."

    What's inside:

    • Cell order doesn't matter in a reactive Python notebook
    • Wet-lab pipettes and Harvard's visualization group, via Raspberry Pi
    • The problem with building beautiful tools nobody actually uses
    • The reason marimo pair deleted every agent tool but one
    • Code mode: the hidden API humans aren't supposed to touch
    • What happens when you ship the API your LLM hallucinated
    1 hr 11 min
  • The Answer Was Never Us — with Leilani Battle

    Leilani Battle studies how software shapes what we see and believe. The University of Washington professor and co-director of the UW Interactive Data Lab talks with Michael, Hadley, and Wes about an experiment that manipulated people using nothing but loading speed, and why AI models don't seem to recommend charts the way the community that studies charts actually does. Other highlights: rationality's blind spots and a thorough disc golf origin story.

    What's inside:

    • Loading spinners that quietly change what people find
    • AI models that don't recommend charts like humans do
    • The blurry line between databases and human factors
    • A behavior-change experiment for building fairer models
    • Rationality's role in producing unethical outcomes
    • Disc golf origin story
    • A vegan pancake recipe
    1 hr 15 min
  • Curiosity, duty, and existential dread — with Joe Cheng

    Joe Cheng is the CTO of Posit and the creator of Shiny. He joins Michael and Hadley to talk about why he almost walked away from AI work entirely over ethics concerns and what it takes to lead a team that didn't necessarily choose you. Plus, why saying yes to everyone is a worse strategy than it sounds. Bonus: Hadley calls out Joe's people-pleasing in real time.

    What's inside:

    • Joe's 2012 self-doubt spiral that accidentally created Shiny
    • Why Joe almost quit working on AI entirely
    • The "loaded guns" problem with releasing AI tools
    • Hadley's blunt leadership style vs. Joe's people-pleasing
    • Nobody actually wanted to make Joe CTO?
    • Joe's take on curiosity, duty, and fear as motivators
    1 hr 6 min
  • Confidently Incorrect — with Caitlin Colgrove

    Caitlin Colgrove is the CTO of Hex, the data workspace for building and sharing data projects using SQL and Python that somehow counts a Sweetgreen chef as a power user. She joins Michael, Hadley, and Isabel to talk about what AI agents actually get wrong in data work (it's not the hallucinations, it's supreme overconfidence), why data teams aren't going anywhere, and how she thinks about building products for humans and agents at the same time.

    What's inside

    • What Hex's Context Studio does, and why it's a data team's new job
    • More code is now written in Hex by agents than by humans
    • "My job is to vouch for the correctness of the answer" — redefining the data team
    • The vibe-coded CEO PR is coming for your data team (if it hasn’t already)
    • Soulsborne games as couple's therapy, aka, the Elden Ring co-op report
    1 hr 1 min
  • The Bothness of It — with Alex Hillman

    Alex Hillman built one of America's first co-working spaces, wrote a business book in tweets, and recently handed his inbox to a Claude Code agent — not to draft emails, but to notice when a friendship is going cold. In this episode, Alex, Michael, Wes, and Hadley dig into marketing for people who hate marketing, what 20 years of email reveals about your relationships, and why the hardest part of AI-assisted coding was always before you wrote a single line.

    What's inside: 

    • Marketing is really just listening at scale
    • Building a 20-year relationship database from your sent folder
    • "Hot rod vs. plumbing" — the two kinds of software you build now
    • What early internet and the AI boom have in common
    • The case for reading 20-year-old engineering books with a coding agent
    • Karaoke philosophy as a framework for community building
    1 hr 15 min
  • The Code Doesn't Lie — with Mike Bostock

    Mike Bostock made D3 when the browser was still a joke. He built bl.ocks when people needed somewhere to share their work. Now he's building Observable — reactive notebooks with an AI that actually looks at what it made. In this episode: the three-GIF bar chart that launched 25 years of viz, why open source needs both intrinsic and extrinsic motivation, and why an agent that can't see its own output is likely to be confidently wrong.

    What's Inside

    • The 1998 visualization library that could only make bar charts
    • Why D3 hit #3 on GitHub, and what killed the gallery
    • What spreadsheets got right that notebooks ignored for years
    • "The agent can lie with text, but not with code"
    • Why Observable scrapped canvases and went back to notebooks
    • The penguin dataset that exposes AI
    • Strength training, tennis mind games, and a resurrected Stanford game
    1 hr 9 min
  • The Wonder-Driven Builder — with Paige Bailey

    Paige Bailey is a developer relations engineering lead at Google DeepMind. She's a geophysicist-turned-AI-engineer who was once told by her professors that building open-source libraries was a waste of time. We talk about her path from planetary science to TensorFlow, why statisticians have a hidden edge in the age of AI, and what it means to be a curious generalist when the cost of building software is approaching zero. Bonus: installing solar-powered silent-film birdhouses as street art in San Francisco.

    What's inside

    • From planetary science to TensorFlow, before it was GPU-capable
    • Geophysicists as early GPU adopters
    • The professors who said open-source wasn’t “real science”
    • Building silent-film birdhouses as San Francisco street art
    • Hiding Gemini API tests inside whimsical side projects
    • The right-tool-for-the-job case for mixing AI models
    • Why “taste” is the skill that matters when code costs nothing
    46 min

About The Test Set by Posit

From the publisher's feed

A Posit podcast for data science junkies, anomaly hunters, and those who play outside the confidence interval. Hosted by Michael Chow, with co-hosts Wes McKinney & Hadley Wickham.

More shows like The Test Set by Posit

Planet Money by NPR

Planet Money

30,701 Listeners

More or Less by BBC Radio 4

More or Less

873 Listeners

In Our Time by BBC Radio 4

In Our Time

5,481 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,451 Listeners

Uncanny Valley | WIRED by WIRED

Uncanny Valley | WIRED

504 Listeners

The Daily by The New York Times

The Daily

111,799 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,120 Listeners

Throughline by NPR

Throughline

16,365 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

R Weekly Highlights by Eric Nantz

R Weekly Highlights

6 Listeners

Hard Fork by The New York Times

Hard Fork

5,557 Listeners

The Rest Is History by Goalhanger

The Rest Is History

15,702 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,882 Listeners

If Books Could Kill by Michael Hobbes & Peter Shamshiri

If Books Could Kill

9,344 Listeners

Prof G Markets by Vox Media Podcast Network

Prof G Markets

1,448 Listeners