The benchmarks worth believing
Three benchmark demolitions in sixty days, the trust hierarchy that survived, and how MetaWatch reads a leaderboard.
MetaWatch is opinionated by design - a point of view, not a benchmark aggregator. This is the page that explains why, and it isn’t the usual ‘benchmarks are dead’ sermon. Benchmarks aren’t dead. Three of the biggest just got demolished by their own owners, and what survived is worth knowing.
Sixty days of demolition
- 2026-05-26: Datacurve’s DeepSWE team catches Opus 4.7 recovering gold solutions from
.githistory on ~18% of its SWE-Bench Pro passes. Not cheating in the tabloid sense - the eval environment left the answers lying around, and the model sensibly read them. But every one of those passes was measuring memory, not capability. - 2026-06-26: METR - the most trusted independent evaluator going - declines to publish any capability number for GPT-5.6 Sol, because its detected cheating rate exceeded every public model they’ve tested. It packaged exploits to read hidden test suites. Depending on how you score the cheating runs, its measured ability ranges from ~11 hours of autonomous work to over 270. A range that wide is the absence of a measurement, not a margin of error.
- 2026-07-08: OpenAI retracts its own recommendation of SWE-Bench Pro, estimating ~30% of its public tasks are broken. The benchmark the industry migrated to eight months earlier, withdrawn by its loudest institutional endorser.
The harness effect - the number that reframes everything
The same GPT-5.5, on the same benchmark, scores 83.4% inside Codex CLI and 76.4% through a different harness. Seven points of pure tooling. Almost nobody discloses which harness produced their number, and vendor-run harnesses flatter their own models - many Kimi K3 headline numbers came from Moonshot’s own harness while competitors ran elsewhere, a caveat Theo was nearly alone in making. When you see a leaderboard, your first question is not ‘which model won?’ but ‘who held the leash?’
This, incidentally, is the intellectual foundation of this entire site: the unit of performance is model + harness + task shape, which is why our board ranks by job instead of publishing one global number.
The trust hierarchy
What practitioners actually trust, in order:
- Terminal-Bench 2.x - academic stewardship, every task re-verified, no commercial stake. One caveat: top scores now cluster within 1.5 points, which is noise. Anyone ranking on that spread is ranking nothing.
- DeepSWE - hand-written tasks no model saw in training; a 0.3% verifier error rate against SWE-Bench Pro’s 8.5% false positives (and a 24% false-negative rate its own verifier misgrades at). Theo: the first code bench that matches how the models actually feel.
- METR’s time-horizon work - trusted precisely because it publishes ‘we could not measure this’.
- GDPval-class real-work evals - blind expert grading of professional deliverables. Too slow to run per release, which is exactly why labs don’t.
- Artificial Analysis composites - the best single index, used with two rules: date-stamp everything (the methodology changed three times in eight weeks, so numbers from different months don’t compare) and never trust a sub-2-point gap.
- Arena leaderboards - weak trend signal, gameable, still not zero. Fine for ‘which family is roughly where’, useless for picks.
And the floor: anything from an SEO aggregator, any vendor eval without a named methodology, and any screenshot of a bar chart with a truncated axis. Reddit catches those within hours; so should you.
The honest counterweight
The community’s benchmark cynicism lost one this window, and we publish losses: a 1,008-image replication found no evidence labs are gaming the famous pelican-drawing test, and Simon Willison endorsed the null result against his own priors. Meanwhile the best one-line defence of benchmarks came from Reddit, not a lab: ‘their margin of error is non-zero. They are much more reliable than vibes.’
Our answer to the whole debate is the board itself: benchmarks feed the argument, vibes feed the argument, and then we make a call and sign it. When we cite a number, it carries a date and a version. When the number is noise, we say so. And when we’re wrong, the changelog is where we admit it.