Skip to content
4 min read Newsletters

On Benchmark Games, Gemini, and Declining Returns to Scale

Benchmark gains are tiny, tests are curated, real comparisons erase Gemini’s edge, and escalating training costs keep driving diminishing returns across models.

On Benchmark Games, Gemini, and Declining Returns to Scale
Photo by Nguyen Dang Hoang Nhu / Unsplash
“Until we know we are wrong, being wrong feels exactly like being right”
― Kathryn Schulz, Being Wrong: Adventures in the Margin of Error

Two things today:

  1. Some recent media appearances and interviews
  2. On the benchmarking games at Google's Gemini

Let's do this.


1. Media Appearances and Interviews

While I don't always remember to mention these. here are three recent apperances and interviews :

2. On Benchmark Games, Kimi, and Gemini

Declining returns to higher training costs among large language models is a favorite topic of mine. The topic is largely misunderstood, and that became even clearer with this week's announcement of Google's new Gemini model.

As a reminder, here is the striking Gemini benchmarking table that went everywhere on announcement day. It seems to be kicking ass, to use the technical term.

I'm going to argue four things about these results:

  1. Even taken at face value, most people wouldn't notice the above differences in the real world.
  2. These benchmarks are worryingly cherry-picked.
  3. On a more balanced set of benchmarks, Gemini is 3, at best, on par with the above models from its competitors.
  4. The sub-linear improvement of large language models at super-linear cost improvements remains the dominant feature.

Most People Wouldn't Notice

The above table shows relatively small gains on tests where all leading models already cluster tightly. As a rule of thumb in a non-deterministic domain, most people don't notice gains of less than 50%.

These gaps, as a result, do not translate into different behavior for typical users. Minor shifts on saturated tasks do not change how a model reasons, follows instructions, writes code, or handles multi-step problems. When people interact with these systems, prompt phrasing, conversation history, and other sources of randomness matter more than small gaps on polluted benchmarks.

Cherry-Picked Benchmarks

The Gemini table omits nearly every high-signal test used to measure real reasoning depth and general capability. These include the harder math and logic suites, multi-step reasoning sets, and contamination-resistant exams used in Epoch's recent PCA analyses¹. These tests load most strongly on the first principal component of model capability, yet almost none of them appear in Google’s comparison.