← The board
Models / Plan review ■ HOLD

Plan review

Judging proposals before any code exists - business, UX, and maintenance implications included.

Patch 2026.07 · Updated 25 Jul 2026
Claude Fable 5
2 Claude Opus 5
3 GPT-5.6-Sol

Here’s an industry embarrassment we’ll state plainly: no benchmark for this exists. Two independent research passes confirmed it - there is no serious public eval measuring whether a model can review an architecture proposal, a PRD, or a migration plan end to end. Every ranking below is argued from adjacent evidence and priors. We think the argument is strong. We’d still rather have the benchmark, and building one is on our list.

Why Fable 5 takes gold

Plans aren’t just architectural proposals. A real plan concerns everything from the business case to the user experience to the maintenance burden three quarters out, and all the trade-offs required to satisfy each area optimally. Reviewing one is the purest judgement task on this board - which is why flagship high-parameter models excel, and why this slot is the home turf of big-model smell.

The adjacent evidence points the same way. Fable holds the best factual-accuracy score in existence - a plan critic must be right about the domain it’s critiquing - and the community’s revealed behaviour is that planning goes to the biggest model available, full stop: the folk pattern is Opus-plans-Sonnet-executes, ‘and when Fable was available, it was Fable.’ A smaller model reviews the plan you gave it. Fable reviews the plan you should have written.

The runners-up

Claude Opus 5 is the sensible default when the stakes don’t justify Fable’s invoice - most plans, honestly - and it dominates the closest proxy evals that do exist (blind-graded professional deliverables across 44 occupations). Watch one thing: the early reports of assertiveness. A plan reviewer’s cardinal sin is redesigning your proposal instead of critiquing it.

GPT-5.6-Sol earns bronze as the completeness checker. It will find every internal inconsistency, every unstated assumption, every place the plan contradicts itself - and it will do it exhaustively. What it won’t reliably tell you is that the whole approach is wrong. Use it as the final sweep after the judgement reviewers have had their say, not instead of them.

What to do now

Route plans by stakes: Opus for routine, Fable for anything expensive to unwind, Sol as the consistency sweep on the way out the door. Review before code exists - it’s the cheapest moment you will ever have to change your mind. And if you build the plan-review benchmark this industry is missing, tell us. We’ll put it on the board’s reading list and argue with it in public.

Copied - paste it to your agent