BLUF:
- To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big leap/small increment’). We need a good y-axis: an interval scale of AI capability which means +1 unit always represents the same degree of ‘how much better’, in the same way +1 degree Celsius is always the same amount of ‘how much hotter’.
- Yet there is no good y-axis for AI capability. All our measures are of something related-to but clearly not identical-with it, thus ‘true’ AI capability can be a funhouse-mirror reflection of whatever was measured. Specifically:
- Benchmark score: One small step in benchmark score can be a giant leap in capability, or the opposite, or whatever else. (My 6/10 vs. your 4/10 ≠ I’m 50% better at maths than you).
- Elo et al: Can give a real y-axis in terms of winning chances, but doesn’t translate outside of beating others. (Going from 50% to 73% to 88% chance to get a higher score than you on a maths test ≠ gaining 0 → 1 → 2 units of maths ability over you)
- Epoch Capabilities Index: Analogous to IQ, so [...]
---
Outline:
(03:05) Introduction
(04:26) Both a poor reflection and a dark glass
(07:51) Human benchmarking also has a y-axis problem
(12:27) A metrological elegy
(12:49) The base case: benchmarks (cf. exams)
(13:30) Elo et al.
(16:25) (And maybe not quite 'game ability ≡ winning games', after all?)
(18:28) ECI (cf. IQ)
(20:56) Intervals Rarely True
(25:01) Measure endogeneity
(31:12) Forking IRT
(34:45) (Dimensions of being, and beating, a bat)
(39:40) Prediction (cf. chronometry)
(41:58) Time horizons
(44:34) Human "capability" is also exponential in time horizon
(47:38) Intuitive/interpretative prelude
(52:56) Time horizons and ECI share an axis kink
(57:13) Perhaps money, as a measure, stinks the least
(01:01:07) Finale: AI as normal epistemics
(01:06:48) Acknowledgements
The original text contained 33 footnotes which were omitted from this narration.
---