Some more thoughts on Google's Gemini 3, given continuing comments about it, questions, misunderstandings, etc. By way of preamble, Google has made huge progress, having reached parity with AI model competitors after years of it wandering in the wilds of its own corporate confusion.
Related reading:
- Google’s Gemini 3 Means AI’s “Resource Grab” Phase Is On (Bridgewater)
The Gist
Google’s Gemini 3 isn’t proof that scaling is healthy, contrary to claims, most recently from Bridgewater. It is, instead, evidence of how far the industry has drifted into diminishing returns. The model reflects brute-force computation delivering only marginal gains—an expensive way to stand still that is being framed as progress.
Key Facts
- Gemini 3’s improvement is small relative to the apparent compute spend.
If Google used materially more pre-training compute—2–3× that of GPT-4o, according to some claims, possibly more—the benchmark deltas look narrow per unit compute. - Nothing about Gemini 3 suggests a new data regime.
No new corpus, no evidence of higher-quality tokens, no training-pipeline breakthrough. Just vastly more compute applied to a more stable TPU stack. - This is late-stage scaling.
Capability improves, but only through massively increasing FLOPs. The marginal return per FLOP is declining quickly, not improving. - Other recent gains in the industry have come from post-training, not scaling.
o1/o3, Claude 3.5→4.x: all technique-driven improvements, not size-driven. Gemini 3 is a clean test of raw scaling—and it shows that the curve is flattening, not re-accelerating. - The narrative is backward.
Bridgewater frames this as proof “scaling still works.” The data show the opposite: scaling works only in a diminishing sense, with each gain costing non-linearly far more than the last.