Big-model smell - overall intelligence, judgement, and taste. The vibes category this site exists to argue.
Orchestration ■ HOLD Claude Opus 5 2 Claude Fable 5 3 GPT-5.6-SolDecomposing briefs, dispatching subagents, and owning the judgement calls at the top of the stack.
Subagent ▲ RISING Grok 4.5 2 GPT-5.6-Sol 3 GPT-5.6-TerraFast, cheap, and RL'd to write code with maximum efficiency. Just don't set them loose on their own.
Delivery engine ■ HOLD GPT-5.6-Sol 2 GPT-5.6-Terra 3 Claude Opus 5Acceptance criteria in, working code out. The slot where relentless beats tasteful, every time.
Diff review ■ HOLD GPT-5.6-Sol 2 Claude Opus 5 3 Claude Fable 5The second pair of eyes on every diff. Recall over politeness, and never the model that wrote it.
Plan review ■ HOLD Claude Fable 5 2 Claude Opus 5 3 GPT-5.6-SolJudging proposals before any code exists - business, UX, and maintenance implications included.
Frontend - No Reference ● NEW Kimi K3 2 Claude Opus 5 3 Claude Fable 5Starting from a loose brief and making the visual decisions. Taste, range, and restraint matter more than raw coding strength.
Frontend - Reference ● NEW GPT-5.6-Sol 2 Gemini-3.6-Flash 3 Kimi K3Turning an existing interface into working code. Visual reasoning, fidelity, and implementation quality decide the podium.
Writing & marketing ■ HOLD Claude Fable 5 2 Claude Opus 5 3 Gemini 3.1 ProWords a human will read - copy, positioning, long-form. A very different ranking from the code slots.
The one tool you'd roll out across the whole skill spectrum, from power users to complete beginners.
Developers ■ HOLD Claude Code 2 Codex 3 CursorFor terminal-comfortable devs who want the strongest defaults without living in their config files.
Power users ◆ NICHE Pi 2 Codex 3 Claude CodeSubstrate, not defaults - maximum extensibility for people with opinions about every layer of the stack.
Cloud agents ● NEW Cursor 2 Devin 3 Claude CodeFull coding environments away from your laptop. The environment, verification loop, and handoff matter as much as the model.
Vibe coding ● NEW Lovable 2 Replit 3 v0Idea to deployed app for people who do not write code. Ease of use wins only when the platform also protects the result from slop.
High-volume conversation where unit economics decide everything. Good enough beats brilliant here.
Cost-conscious agents ■ HOLD Grok 4.5 2 GPT-5.6-Terra 3 GLM-5.2Real agentic work inside a product that has to keep its margins. Performance per pound wins here.
High-intelligence agents ● NEW GPT-5.6-Sol 2 Kimi K3 3 Grok 4.5When your product sells intelligence, buy the most capability per pound the frontier will part with.
What parameter counts actually are, why bigger models feel smarter, and the house theory behind our Overall crown.
The tokenizer tax Guide →Why price-per-token is a lie, why cost-per-task is the only honest number, and how to read any pricing announcement.
The benchmarks worth believing Guide →Three benchmark demolitions in sixty days, the trust hierarchy that survived, and how MetaWatch reads a leaderboard.