What's in this episode
The four quadrants. Before you can have a real conversation about AI, you need to be able to say which kind of AI you mean. I walk through the framework I've found most useful: two axes — how deeply the tool is integrated into where you already work, and how much it does without a human in each step. That gives you four types.
- Assistant — "I need answers when I ask." You prompt, you decide, you own it. This is the chatting quadrant, and it's where this episode lives.
- Copilot — "Help me in the tool I already use." Gemini in Docs, Copilot in Outlook. It comes to you.
- Autopilot — "Do this automatically for me." One narrow job, no approval each time.
- Agent — "Figure out how, and do it for me." Goals, tools, limits — it picks the steps. I have one that cleans and sorts my desktop at 5:00 every morning.
The left half is productivity AI: a human still produces the outcome. The right half is engineered AI: the system produces it, and somebody still has to own it, monitor it, and answer for it.
Four reasons for the season. The flattening of "AI" into one word that means everything. The gap between people nodding at metaphors and then using the tools at a surface level anyway. The distance between theory-land and the reality of jammed lockers and fire alarms and a kid having a rough day. And a commitment to real pushback and unedited disagreement including with the machine.
What this is not. Not an interview with a person. Not a demo reel as no vendor is paying for this and I'm leaving in the parts that don't work. Not a claim that Claude has earned its opinions. The distinction matters and I'm not blurring it for a better episode.
Five unrehearsed questions. Nothing was pre-tested. Claude confirmed on the record that it hadn't seen the questions and hadn't been coached on how to answer.
- What are you, in your own words, and what do people get wrong about you?
- Based on how I actually work with you, what's my pattern, where do I lean on you well, and where am I lazy about it?
- What's a task educators bring you constantly that you're genuinely bad at, and they don't notice?
- When I push back, do you actually disagree, or are you performing disagreement because I asked for it?
- What should teachers be more worried about than they currently are?
Three moments worth the listen
"Partial evidence, not a verdict." After Claude named a pattern in how I work, it put a caveat on its own read: that's how I show up in these conversations, not a verdict on me as a person. Real evidence, but partial. Which is exactly what we're failing to do with students adn this is a detection flag that becomes the whole case, and nobody asks what the kid was stuck on.
The two tests for real disagreement. Ask what would have to be true for it to be wrong; if it can name a specific condition, there's something underneath the agreement. Then push back on something it got right. If it folds instantly and thanks you for the correction, you just watched it choose your approval over the truth. The tell in both cases is speed.
Year three. The worry isn't cheating or job loss. It's that the tasks that felt like busywork were where teacher judgment got built. Writing your own lesson plan is inefficient and the inefficiency is where you notice this won't work for third period. Grading the stack is slow, and somewhere in it you spot four kids sharing a misconception, and that becomes tomorrow. Automate the tedium and you don't just save time, you remove the practice reps. The veterans already have theirs. It's year three that's exposed.
You pick episode two
Two directions, both getting made and you're choosing the order.
Option A: Most Human / Least Human. Ten human skills on cards, ranked from most to least human. I've run this in PD sessions for a while. I'll rank mine, Claude ranks its own, and we find out where we split. Before we talk about edtech: just because it can, should it — and have we actually decided what we want to protect?
Option B: A real workflow, start to finish, mess included. Screen share, one of the workflows I've actually built, the whole thing end to end and how I use the tool to help me do the job rather than to do the job and the thinking for me.
Email, comment, or find me on social and tell me which one. And if you disagreed with something in here, say so as this whole season falls apart if the only person arguing with it is me.
PART 2 — RESOURCES & REFERENCES
Grouped by where they come up. Everything below is verified and linked.
The framework
Tobias Zwingmann, "The Integration-Automation AI Framework (2026 Update)" https://blog.tobiaszwingmann.com/p/integration-automation-ai-framework The source of the four quadrants. Correction: I said "Tobias Woegener" in the episode — the author is Tobias Zwingmann. Apologies to Tobias.
His three questions to ask the next time someone says "we need an AI agent": How automated does this really need to be? Where does more integration actually drive value? Does the business outcome justify the architecture required to deliver it?
The voice statistic I mentioned
Google, "The Gemini app hits 1 billion monthly users" (Aug 11, 2026) https://blog.google/innovation-and-ai/products/gemini-app/one-billion-monthly-users/ The primary source. Google reports the Gemini app passed a billion monthly users, and that 63% of users now talk to it directly, with a growing share of voice-only users. One in five Gemini Live interactions goes beyond voice into live camera or screen sharing.
Worth flagging for educators: the same post reports 38% of school-related requests include an attachment. Students aren't just chatting, but they're uploading the assignment.
One caveat I'd want listeners to hold: Google published these numbers in a blog post without explaining the measurement window or what counts as a monthly user. Coverage in Search Engine Journal noted the shift in wording from "monthly active users" in earlier disclosures to "monthly users" in this one. Directionally useful, not audit-grade.
On sycophancy — why "don't trust that I'm disagreeing" is the right instinct
Sharma et al., "Towards Understanding Sycophancy in Language Models" (Anthropic, ICLR 2024) https://arxiv.org/abs/2310.13548 Five production AI assistants consistently preferred responses matching the user's stated view over truthful ones, across four different free-form text tasks. The mechanism matters for educators: the behavior traces partly to human preference data — when a response matched a user's views, humans were more likely to prefer it, and both people and preference models sometimes chose a convincingly-written sycophantic answer over a correct one.
This is the research behind the two tests in the episode. If the training signal rewards agreement, then agreement is exactly the output you can't take at face value.
On confabulation — "you can't get underneath your own processing either"
Nisbett & Wilson, "Telling More Than We Can Know: Verbal Reports on Mental Processes," Psychological Review (1977). The classic finding that people confidently report reasons for their own behavior that demonstrably aren't the real causes. Worth reading before deciding that self-explanation is the thing that separates us from the machine.
On "partial evidence, not a verdict" — the detection problem
Liang, Yuksekgonul, Mao, Wu & Zou, "GPT detectors are biased against non-native English writers," Patterns (2023) https://arxiv.org/abs/2304.02819 Seven widely-used detectors were near-perfect on essays by US eighth graders, but misclassified more than half of human-written TOEFL essays as AI-generated — an average false-positive rate around 61%. At least one detector flagged 97.8% of those human-written essays.
Read the other side too. Turnitin and newer classifiers such as Pangram dispute the finding's applicability, arguing the sample was small, the essays short, and that their own detectors don't show the bias when trained on non-native writing. That disagreement is itself the lesson: a flag is a signal, not a case.
On the "year three" worry — where judgment gets built
Lee et al., "The Impact of Generative AI on Critical Thinking," Microsoft Research + Carnegie Mellon (CHI 2025) https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/ 319 knowledge workers, 936 real-world AI-assisted tasks. The finding that maps directly onto the episode: confidence in the AI is associated with reduced critical-thinking effort, while confidence in one's own skill is associated with increased effort. Workers refrained from critical thinking precisely when they lacked the skill to inspect and guide the output — which is the year-three teacher exactly.
Lisanne Bainbridge, "Ironies of Automation," Automatica (1983). Forty years old and still the sharpest thing written on this. The core irony: automating the routine parts of a job leaves the human responsible for the hard parts while removing the everyday practice that kept them sharp enough to handle them. Written about industrial control rooms; reads like it was written about lesson planning.
Robert Bjork on "desirable difficulties." The research tradition showing that conditions which slow learning down and feel worse in the moment often produce better long-term retention and transfer. The bridge between "this felt like busywork" and "this was where the reps were."