← The board
Models / Subagent ▲ RISING

Subagent

Fast, cheap, and RL'd to write code with maximum efficiency. Just don't set them loose on their own.

Patch 2026.07 · Updated 25 Jul 2026
Grok 4.5
2 GPT-5.6-Sol
3 GPT-5.6-Terra

Call this what it really is: the performance:cost category. Subagents are the hands of a multi-agent stack - dispatched in parallel by an orchestrator, handed tightly-scoped briefs, judged on value per pound. The orchestrator supplies the judgement; this tier supplies the throughput.

Why Grok 4.5 takes gold

Grok 4.5 is the current GOAT when sat beneath a top-tier orchestrator. The intelligence is high enough for the vast majority of scoped tasks, and the economics are documented rather than claimed: it’s the only model in the tier with a published cost-per-task figure ($0.31 per benchmark task), using roughly 60% fewer tokens than rivals for equivalent work - which, per the tokenizer tax, is the number that actually matters. On top of the API price, it’s aggressively subsidised through both Cursor and Grok Build subscriptions, so for most readers the effective cost is lower still.

The standard objection is the hallucination rate, and this patch we’re retiring it - not because the number is wrong, but because it isn’t a differentiator. On the same eval, Grok confabulates at ~54%, Fable 5 at 54.9%, Opus 5 at 50%. The frontier makes things up at broadly the same rate; only Haiku 4.5 (28%) is genuinely calibrated. The real rule was always the slot definition: subagents run supervised, and their output gets checked before it lands.

The runners-up

GPT-5.6-Sol and GPT-5.6-Terra are here for the same reason: some of the most intelligent models in the world, extremely token-efficient, aggressively priced - and Codex subscriptions carry some of the best subsidisation multiples available, which makes them outrageous value under an orchestrator. Sol as a subagent deserves a special note: the literalism that worries us at the top of the stack (see Orchestration) becomes a feature one rung down, where the brief is already written and relentless literal execution is the job description. OpenAI’s own staff recommend exactly this shape - Sol or Terra orchestrating, Luna handling the exploratory legs.

Honourable mentions. Gemini 3.5 Flash-Lite is the throughput floor (fastest agentic-capable model tracked, 1M context, $0.30 input) and the right pick when raw fan-out volume beats per-task smarts. GLM-5.2 is the open-weights answer - MIT licence, the community’s named migration destination, and the only route to a marginal cost no vendor can reprice.

What to do now

Fleet defaults: Grok 4.5 for the bulk, Sol/Terra when the subtask is genuinely hard, and cap your fan-out - there’s no default limit in most harnesses, and the 50-subagent surprise invoice is a rite of passage nobody enjoys twice.

Copied - paste it to your agent