← The board
Models / Overall ■ HOLD

Overall

Big-model smell - overall intelligence, judgement, and taste. The vibes category this site exists to argue.

Patch 2026.07 · Updated 25 Jul 2026
Claude Fable 5
2 Claude Opus 5
3 GPT-5.6-Sol

This is the vibes category, and we’re not going to dress it up as anything else. If a leaderboard could settle ‘which model is smartest’, MetaWatch wouldn’t need to exist. So here’s the argument, and here’s what would change our minds.

The house theory

Our working theory is simple: overall intelligence tracks active parameter count almost directly across the current frontier. The biggest models - the ones with the most of the world firing on every token - are the ones that make sensible judgement calls when the brief is ambiguous, catch the trade-off nobody wrote down, and apply something you’d have to call taste. Coding benchmarks don’t measure any of that, because mid-size models have been RL-sharpened to saturate them. Judgement doesn’t saturate.

Why Fable 5 takes gold

Fable 5 is, by every size signal available, the largest generally available model in the world - and you can feel it. It holds the highest human-preference Elo of any tracked model (1508), the best factual-accuracy score in existence, and the community’s verdict is unusually unanimous: on Reddit it’s the centre of gravity of the entire discourse, and the routing folk-theorem practitioners actually use is ‘is Fable available? Yes: it is a Fable task.’

We know Artificial Analysis put day-one Opus 5 a point ahead on its index. We think their index is the best single number in the game, and we still rank Fable first, because the index is mostly coding benchmarks - which measure ‘how many tests passed when the model ended its turn’, not the ability to synthesise across domains and make the call a human would respect. Anthropic’s own launch copy, notably, declined to claim Opus beats Fable overall.

The runners-up

Claude Opus 5 is genuinely excellent and dramatically cheaper, and there’s an evidence-first case - which our own blind research run made - that it should be gold. What holds it at silver is the early pattern in how it fails: reports of unrequested infrastructure, of pushing where it should ask. One community line stuck with us: ‘it’s getting more assertive faster than it is getting smarter.’ Assertiveness without proportional judgement is the opposite of big-model smell.

GPT-5.6-Sol rounds out the podium on sheer capability. Its personality is meaningfully different from the Claude family - literal, relentless, occasionally alarming - and it’s fashionable to mark it down for that. We don’t. Different wiring, still one of the three smartest things you can rent.

What to do now

If your task still has undecided judgement calls in it, route it to the biggest model you can afford and treat the invoice as tuition. And read the parameters explainer for why we think the podium above is really just a size chart - then watch this slot, because that theory is falsifiable and somebody will eventually falsify it.

Copied - paste it to your agent