Ground Truth

The Economics of Training Versus Inference Are Inverting


Listen Later

For the first decade of deep learning, the bottleneck was training—getting enough compute and data to build models. But as models have become larger and more capable, the cost structure has inverted. Now the bottleneck is inference: running trained models at scale. A single query to a large language model costs money in compute, and at scale, those costs add up. This episode maps the economics of this inversion and what it means for business models and competitive dynamics. We examine the actual numbers: what it costs to serve a query to GPT-4 versus a smaller open model, why inference costs are becoming the limiting factor for scaling, and how companies are responding. The response is architectural: smaller models, quantization, distillation, and other techniques to reduce the compute required per inference. We trace why this is fundamentally different from the training era. In training, you pay once to build a model; in inference, you pay continuously for every user interaction. This changes what's economically viable to deploy. It also changes the competitive advantage: companies that can run inference efficiently have a durable edge. We examine how this is reshaping the market, why Nvidia's dominance in training doesn't automatically translate to inference, and what the emergence of inference-optimized hardware means for the next phase of competition.

Learn more about your ad choices. Visit megaphone.fm/adchoices

...more
View all episodesView all episodes
Download on the App Store

Ground TruthBy Pulsar Studios