← The board
Serving / High-intelligence agents ● NEW

High-intelligence agents

When your product sells intelligence, buy the most capability per pound the frontier will part with.

Patch 2026.07 · Updated 25 Jul 2026
GPT-5.6-Sol
2 Kimi K3
3 Grok 4.5

The top of the serving market: products whose whole pitch is intelligence - legal analysis, research tools, anything a customer pays for precisely because the output is smarter than they could get elsewhere. The naive read of this slot is ‘buy the smartest model’. The correct read is capability per pound at the frontier, because a product still has margins, and the very smartest models carry premiums their marginal intelligence rarely earns back at serving volume.

Why Sol takes gold

GPT-5.6-Sol is among the two or three most intelligent things you can rent, and it’s the only one of them priced and token-engineered like the vendor expects you to serve it at scale. Exceptional reasoning depth, best-in-class token efficiency (the tokenizer tax runs in its favour, for once), and the muscle to run genuinely hard multi-step work inside your product’s scaffolding. Its known pathologies - literalism, overreach - are dev-tool problems more than serving problems: inside an application, you wrote the workflow, and a model that executes the written thing relentlessly is precisely the shape you want behind an API.

The runners-up

Kimi K3 is the sovereignty pick at the frontier tier: genuinely frontier-adjacent scores, 1M context, and open weights - which converts availability from a vendor’s promise into your own ops problem. This window taught the whole industry why that matters (more below). If your product’s brain must never disappear, weights you control beat any SLA ever written.

Grok 4.5 completes the story this section keeps telling: it podiums in both serving agent categories, which is a testament to an insane capability:cost ratio. It’s the ‘surprisingly high intelligence at commodity prices’ play - not the deepest reasoner on the podium, but the one that lets you offer a smart tier your CFO doesn’t flinch at.

Where’s Opus 5? Our research runs had it gold here on raw intelligence-per-measure, and it’s a very defensible pick - the counter is price at serving volume plus a hallucination rate that climbed with its brains. Where’s Fable? Availability. The smartest model in the world spent 18 days this summer entirely withdrawn under export controls, then changed commercial terms four times in six weeks. For interactive work you shrug and switch for a fortnight; for a product dependency it’s an extinction event. Brilliance you can’t serve is a benchmark, not a business.

What to do now

Serve Sol behind your own abstention layer - the whole frontier confabulates roughly half of what it doesn’t know, so ‘I don’t know’ is a feature you build, not one you buy. Keep K3 warm as the failover (multi-provider is table stakes now; this summer proved it), and prototype your cheap tier on Grok before assuming you need the expensive one everywhere. The intelligence your product sells is the system, not the model.

Copied - paste it to your agent