The biggest bet in tech history is being placed against a workload that is busy becoming free. In 2026 the industry will spend more than a trillion dollars building cloud AI compute — the first trillion-dollar capex year ever — while the cost of running a model at any fixed capability falls roughly tenfold a year and small, efficient models grow good enough to do most real work on the laptop or phone already in your hand. Compute isn't consolidating into the cloud; it's splitting in two: a thin, genuinely expensive frontier for the hardest reasoning, and a commodity tier — classification, extraction, summarisation, the thousand tool-calls inside an agent — sliding toward free and increasingly running on hardware you already own.
This is a builder's map of that split: how far the price floor has fallen (LLMflation), what a 3-to-4-billion-parameter model can now actually do, the local-plus-cloud stacks already shipping from Perplexity, Microsoft, Google and Apple, why the real obstacle isn't the model but the harness, and a concrete rule for what to run on your own machine versus a cloud API — a split that cuts inference bills sixty to eighty percent today. The sharpest read on the whole trend belongs to the company everyone says lost the AI race: Apple, which spends more than ten times less than any single rival, refuses to build a frontier model, and is wiring efficient AI into a billion devices because it thinks owning the trusted edge beats owning the biggest model.
I'm Dan. AI moves too fast to keep up with, so I built my own stack of AI tools to research, analyse, verify and illustrate the questions I can't stop thinking about — mostly to learn it myself, and I share what I find. AI-assisted, fact-checked, worth a second look.
Follow Dan's AI Intel in your podcast app so new episodes find you.