The Test Set by Posit

The Test Set by Posit

By Posit, PBCTechnology
Download on the App Store

The Test Set by Posit episodes

  • Widgets Are Lego Bricks (and Other Things People Are Sleeping On) — with Vincent Warmerdam

    Vincent Warmerdam has been the first full-time hire at a startup, a spacey punster who accidentally got himself a job, a bartender at an Amsterdam comedy theater, and a Dutch bike tour guide — and he'll tell you all of it was career development. Now doing DevRel at Marimo, Vincent makes the case for reactive notebooks, Lego-brick widgets, and why "number go up" is not a data science strategy. Also: chickens die. The model doesn't know. This matters more than you think.

    What's inside

    • How a spacey pun accidentally launched Vincent's career
    • Why Marimo's constraints make it better for LLMs, not just humans
    • The gorilla hiding in your dataset — and why the model missed it
    • Vibe coding vs. notebooks: three cells at a time as a discipline
    • Widgets as Lego bricks: reusable, composable, criminally underused
    • Cognitive debt, confirmation bias, and sycophantic data science
    • Why natural intelligence is still, actually, a pretty good idea
    1 hr 16 min
  • Everything's a Fad (Including This Podcast) — with Benn Stancil

    Benn Stancil built Mode Analytics, spent a decade in the data trenches, and now writes some of the sharpest, funniest essays in the data world. On The Test Set, he talks about the cultural shift from Nate Silver to Rick Rubin why AI might kill the analytics dashboard, and what happens when a thousand startups all build the same thing. Plus: boy bands as a model for collaboration, and why the best creative work starts with cheating.

    What's inside: 

    • Why the modern data stack was basically big data 2.0
    • The cultural flip from Nate Silver to Rick Rubin
    • Gas Town, tar pits, and the AI startup zero-sum game
    • Software is becoming content, and that changes things
    • Benn's creative process: Lorde lyrics, Codenames, and cheating
    • The boy band as a model for small-team collaboration
    • BI is (mostly) dead, and vibes might replace SQL
    1 hr 36 min
  • Deeply Unsexy: SQL's Redemption Arc — with Tristan Handy

    dbt Labs CEO Tristan Handy drops into The Test Set to map the fault lines between the data science world and the enterprise data world — and explain why analytics engineers are basically pissed-off data analysts who decided to organize the bookshelf. We get into SQL's glow-up, the community magic of dbt Slack, what AI agents mean for data warehouses, and why everyone's building iOS apps with Claude now.

    What's inside:

    • What analytics engineers *actually* do
    • SQL's journey from deeply unsexy to indispensable
    • How dbt turned source control into a source of truth
    • Building a tech community without the RTFM energy
    • AI agents on your data lake: permissions get personal
    • Will LLMs kill the open-source package ecosystem?
    • Edible gardening, welding dreams, and digital dysphoria
    1 hr 6 min
  • Your VP Is Doing a Rogue Analysis in Cursor Right Now — with Nell Thomas

    Nell Thomas has spent two decades in data — from equity research to the DNC to Facebook to leading a 400-person data org at Shopify. She walks Michael and Wes through the modern data stack role by role, gets honest about what AI is and isn't changing about data work, and admits the semantic layer has been her greatest leadership failure. Plus: Sneakers gets the respect it deserves.

    Episode Notes
    What does it actually look like to run data infrastructure for millions of merchants while the entire industry reinvents itself in real time? Nell Thomas (VP of Data, Shopify) talks vibe-coded dashboards, political campaign data scarcity, blameless postmortems, and why no one should be locking in on an AI strategy just yet. Recorded live in Times Square.

    What’s Inside

    • Mapping the modern data stack, role by role
    • Why data quality is still the #1 problem
    • What "good scrutiny" looks like on a data team
    • Vibe coded dashboards and the trust problem
    • Shopify's MCP for their data warehouse
    • The throwaway tech problem in political campaigns
    • Why the semantic layer is so damn hard
    • Sneakers!


    1 hr 23 min
  • Sleeping Rats and Sociopathic Agents — with Phillip Cloud

    Phillip Cloud has been shaping the Python data ecosystem since the early pandas days — and he has *opinions*. Now a principal engineer at NVIDIA leading the Ibis project, Phillip talks about how he stumbled into open source via an eye movement lab, why he prefers his coding agents cold and emotionless, and what happens when you ask an LLM for woodworking trig. Plus: terminal user interfaces, the file hierarchy standard hot take nobody asked for, and the pineapple-on-pizza hill he's willing to die on.

    Episode Notes

    Phillip Cloud (NVIDIA, Ibis project) joins Michael Chow, Wes McKinney, and Hadley Wickham to talk about his path from eye movement labs to pandas core team, why developer productivity tools have quietly gotten amazing, his brutally honest take on coding agents, and what it would actually take to impress him. Also: VisiData love, NixOS evangelism, and yard work as therapy.

    What’s Inside

    • From eye movement labs and MATLAB to pandas core team
    • Column multi-indexes: the feature nobody likes but Phillip needed
    • Why your command line tool better have a sweet TUI
    • VisiData: the terminal data tool you're sleeping on
    • Are we writing code for humans or for agents now?
    • Phillip's AI skepticism journey: Cursor, Claude Code, and frustration
    • The Numba CUDA test suite port that would finally impress him
    57 min
  • More productive but a lot less fun — with Charlie Marsh

    Charlie Marsh built Ruff, uv, and Ty — the tools that mass-fixed Python's worst pain points. Now he's grappling with what happens when agents start writing most of the code. In this episode, Charlie gets real about his team trusting his PRs less, the gnarly middle of coding with agents, and whether Python is even the right language for an agentic future. It's honest, a wee existential, and deeply relatable if you ship code for a living.

    Episode Notes
    Charlie Marsh is the founder and CEO of Astral — the company behind Ruff, uv, and Ty. He sits down with Michael and Wes to talk about what it's actually like building with coding agents every day, why his team's code review dynamics completely changed, and big open questions about code quality, open source community, and Python's future nobody has answers to yet.

    What’s Inside

    • How Ruff convinced mass adoption when switching tools is painful
    • Why "just uv run it" became the killer feature
    • Hiring outside Python's ecosystem to build tools for it
    • His team said "we trust your PRs less now"
    • Engineers screen-sharing their actual agent workflows at Astral
    • The "Lisp Curse" reborn: cheap code breaking open sourceIs 
    • Python the wrong language for an agentic world?
    1 hr 36 min
  • Alenka Frim: What yoga teaches us about discipline and collaboration in data science

    Alenka Frim went from teaching yoga full-time to becoming a committer and PMC Member on Apache Arrow. In this episode, Alenka joins The Test Set hosts to talk about how Arrow grew from spec to critical infrastructure, and why she started contributing to a project she had never even used. She reflects on imposter syndrome, the discipline of showing up (on the mat and in GitHub), and how agents are changing what it means to write code. Plus: managing 4,000 open issues without losing your mind.


    Episode Notes

    Alenka's path into Arrow is unconventional: Sshe wasn't looking for a job, she wasn't using the tool, and she'd spent the previous five years focusing on mind-body fitness. But open source felt like the right place to learn, have fun, and figure things out, so she jumped in. What followed was a journey from her first R bindings to becoming a PMC member on one of the most critical pieces of data infrastructure in the world.


    What’s Inside

    • Alenka's journey, from yogi to Arrow committer
    • Signs of a healthy open source community: people, dialogue, and turnover
    • Arrow as critical infrastructure: DuckDB, Polars, Pandas, and the spec that unifies them
    • Managing 4,000 open issues without losing your mindImposter syndrome in open source
    • What the yoga mat teaches you about discipline and collaborationAI and the future of programming: 100x more software or 10x better software?
    1 hr 2 min
  • Emily Riederer: Column selectors, data quality, and learning in public

    Emily Riederer writes Python with an R accent, and we’re all comfortable with it. In this episode, Emily reflects on her journey through R, Python, and SQL — from lessons learned in averaging default values (oops, we're not all rich!) to discovering that column selectors are way cooler than they sound. She weighs in on the delicate art of learning in public, why frustration often makes the best teacher, and how to find your niche by solving the boring problems. Oh, Oh, and the crew casually drops that she's keynoting posit::conf 2026!

    Episode Notes
    Emily’s had a wild ride through modeling, data engineering, machine learning, and back again, and she knows a thing or three about the evolution of SQL tooling (from nightmare multi-page scripts to the dbt renaissance). She reveals how building internal packages became her gateway to making work enjoyable. Plus: the surprising Stata origins of column selectors, the eternal struggle of naming packages across R and Python, and why watching people code teaches you more than any tutorial ever could. The conversation gets real about imposter syndrome and the magic of tacit knowledge.

    What’s Inside

    • Why real-world data is chaos, not truthThe path from modeling to data engineering (and back)
    • What a data pipeline really is (extract, load, transform) and why organization matters
    • How dbt changed the SQL game Learning by watching: Tacit knowledge and coding over the shoulder Imposter syndrome and learning in public 
    • Building internal tools to escape busyworkposit::conf 2025 keynote preview
    59 min
  • Rebecca Barter: Persistent learning, tool building, and ‘Will code even exist?’

    Rebecca Barter, senior data scientist at Arine and adjunct assistant professor at the University of Utah, refuses to work on things she doesn’t care about. Lucky for us, she cares about a lot, most of all impact. In this episode, Rebecca joins The Test Set to talk about learning fast, building better tools, and staying motivated and adaptable.

    She shares how moving between R, Python, SQL, and dashboards reshaped how she thinks about expertise. Plus a reflection on her recent posit::conf talk, “AI: Hype, Help, or Hindrance.”

    Episode Notes

    Rebecca digs into what it’s really like to work with AI every day and why humans still rule, especially in exploratory data analysis. She explains how tool building can be the fastest way out of busywork and how teaching beginners sharpened her ability to communicate clearly. The conversation circles a bigger question too: As AI keeps improving, are we headed toward a future where code looks completely different … or maybe disappears altogether?

    What’s Inside

    • Why motivation matters even more than productivity
    • Escaping busywork by building better tools
    • From R to Python to dashboards: Learning fast as a survival skill
    • Reality check on AI in the IDE
    • Why exploratory analysis still needs human intuition
    • The 80/20 of coding: Automate the boring, protect the judgment
    • Teaching beginners and earning trust
    • The uncertain future of code
    57 min
  • Marco Gorelli: Narwhals, ecosystem glue, and the value of boring work

    You’ve probably used Narwhals without realizing it. It’s the compatibility layer helping apps and libraries like Plotly play nice with Pandas, Polars, Arrow, and more — while keeping computation native instead of converting everything to Pandas. In this episode, Marco Gorelli explains how his weekend experiment turned into essential ecosystem infrastructure and why data types, not APIs, are where interoperability gets tricky. Plus what it takes to build trust and community around an open-source project.


    Episode Notes
    Marco shares the Narwhals origin story (including the meme-powered name), the hard edge cases that live in data types and null semantics, and why he’s cautious about using AI for code generation when correctness hinges on tiny details. We also jam on proactive “GitHub surfing,” conference talks as trust-building exercises, celebrating contributors, and how early commit messages capture the genuine excitement of building something new.


    What’s Inside

    • Narwhals 101: You’ve probably used it (even if you didn’t know it)
    • The real interoperability traps: data types, null semantics, and “looks-the-same” operations
    • Why expression systems won, and how they shaped Marco’s approach — with nods to Ibis, Polars, and Pandas
    • Open source as social work: proactive outreach, trust-building, and a Discord-powered community
    • Extending Narwhals to new engines, starting with the Daft plugin
    52 min

About The Test Set by Posit

From the publisher's feed

A Posit podcast for data science junkies, anomaly hunters, and those who play outside the confidence interval. Hosted by Michael Chow, with co-hosts Wes McKinney & Hadley Wickham.

More shows like The Test Set by Posit

Planet Money by NPR

Planet Money

30,701 Listeners

More or Less by BBC Radio 4

More or Less

873 Listeners

In Our Time by BBC Radio 4

In Our Time

5,481 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,451 Listeners

Uncanny Valley | WIRED by WIRED

Uncanny Valley | WIRED

504 Listeners

The Daily by The New York Times

The Daily

111,799 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,120 Listeners

Throughline by NPR

Throughline

16,365 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

R Weekly Highlights by Eric Nantz

R Weekly Highlights

6 Listeners

Hard Fork by The New York Times

Hard Fork

5,557 Listeners

The Rest Is History by Goalhanger

The Rest Is History

15,702 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,882 Listeners

If Books Could Kill by Michael Hobbes & Peter Shamshiri

If Books Could Kill

9,344 Listeners

Prof G Markets by Vox Media Podcast Network

Prof G Markets

1,448 Listeners