Episode Summary
Artificial intelligence is advancing at an extraordinary pace. Every week, a new model claims to outperform another on reasoning, coding, mathematics, or scientific knowledge. New leaderboards emerge, benchmark scores climb, and headlines celebrate the latest breakthrough. But beneath all the excitement lies a deceptively simple question: how do we know any of these claims are truly comparable?
In this episode, we explore A Common Ruler, a thought-provoking paper by Thasmika Gokal that argues the AI industry has reached a point where it needs universal standards for measuring intelligence. Just as engineering, science, aviation, and global commerce depend on common units of measurement, AI may require a shared evaluation framework that enables meaningful comparison across models, organisations, and nations.
Rather than focusing on building bigger or faster models, this episode examines the foundations of trust. What happens when every AI company creates its own benchmarks? Can governments confidently regulate systems measured using different standards? Can enterprises make informed investment decisions when every model is evaluated using a different ruler? And can the public place confidence in claims of intelligence if there is no universally accepted way to verify them?
Through real-world analogies—including the Olympic Games, Formula One, international engineering standards, and the evolution of the metric system—we explore why shared measurement has historically accelerated innovation instead of restricting it. Competition thrives when everyone agrees on the rules, and AI may be approaching the same inflection point.
The discussion also considers how a universal evaluation framework could coexist with proprietary enterprise assessments. Organisations will always need private evaluations tailored to their own objectives, industries, and risk profiles. However, those internal measures answer a different question from universally recognised standards. One determines whether an AI system is fit for a specific purpose; the other establishes whether its capabilities can be compared fairly across the broader ecosystem.
As artificial intelligence becomes increasingly embedded in healthcare, engineering, finance, education, government, and critical infrastructure, consistent measurement may prove to be just as important as technological capability itself. Standards create confidence, confidence enables adoption, and adoption ultimately determines whether transformative technologies fulfil their potential.
This episode explores why evaluation is no longer simply a technical exercise—it is becoming the foundation of AI governance, public trust, and responsible innovation. More importantly, it asks whether the next great breakthrough in artificial intelligence will not be another model, but rather a universally accepted way of measuring intelligence itself.
If AI is to become trusted infrastructure for society, perhaps the first step is agreeing on a common ruler.