Paul Kedrosky
Paul Kedrosky
@paul@paulkedrosky.com · Mar 10

I point this out endlessly, but common AI benchmarks are less useful than thought, given both metholdogical issues and that they are increasingly ingested by models. That more than half of agentic code passing code benchmarks is not acceptable should thus come as no surprise.

https://metr.org/notes/2026-03-10-many-swe-bench-passing-prs-would-not-be-merged-into-main/