“Until we know we are wrong, being wrong feels exactly like being right”
― Kathryn Schulz, Being Wrong: Adventures in the Margin of Error
Two things today:
- Some recent media appearances and interviews
- On the benchmarking games at Google's Gemini
Let's do this.
1. Media Appearances and Interviews
While I don't always remember to mention these. here are three recent apperances and interviews :
- GQG (video discussion with analysts / portfolio managers)
- CBC radio (audio only)
- Search Engine (audio only, and just some excerpts)
2. On Benchmark Games, Kimi, and Gemini
Declining returns to higher training costs among large language models is a favorite topic of mine. The topic is largely misunderstood, and that became even clearer with this week's announcement of Google's new Gemini model.
As a reminder, here is the striking Gemini benchmarking table that went everywhere on announcement day. It seems to be kicking ass, to use the technical term.

I'm going to argue four things about these results:
- Even taken at face value, most people wouldn't notice the above differences in the real world.
- These benchmarks are worryingly cherry-picked.
- On a more balanced set of benchmarks, Gemini is 3, at best, on par with the above models from its competitors.
- The sub-linear improvement of large language models at super-linear cost improvements remains the dominant feature.
Most People Wouldn't Notice
The above table shows relatively small gains on tests where all leading models already cluster tightly. As a rule of thumb in a non-deterministic domain, most people don't notice gains of less than 50%.
These gaps, as a result, do not translate into different behavior for typical users. Minor shifts on saturated tasks do not change how a model reasons, follows instructions, writes code, or handles multi-step problems. When people interact with these systems, prompt phrasing, conversation history, and other sources of randomness matter more than small gaps on polluted benchmarks.
Cherry-Picked Benchmarks
The Gemini table omits nearly every high-signal test used to measure real reasoning depth and general capability. These include the harder math and logic suites, multi-step reasoning sets, and contamination-resistant exams used in Epoch's recent PCA analyses¹. These tests load most strongly on the first principal component of model capability, yet almost none of them appear in Google’s comparison.