Cost-conscious agents
Real agentic work inside a product that has to keep its margins. Performance per pound wins here.
The middle of the serving market: products where the model genuinely works - multi-step tasks, tool use, generated apps - but every invocation is a line on your unit economics. Think Lovable-shaped businesses. The performance bar is real; so is the free tier eating your margin.
Why Grok 4.5 takes gold
The same capability:cost ratio that makes Grok our subagent gold makes it the serving pick here - and inside a product the fit is even better, because your pipeline is the orchestrator. The scoping a human does by judgement, your product does by design, which neutralises the supervision question entirely. The numbers are the best-documented in the tier: $2.59 per benchmark coding task at near-frontier scores, roughly 60% fewer tokens per task than rivals. Agentic loops bill by the token; a model that thinks in shorthand is a structural discount.
The runners-up
GPT-5.6-Terra is the balanced production tier from the family that owns token efficiency - most of Sol’s execution quality at prices built for exactly this segment. If your product’s tasks are hard enough that Grok’s quality ceiling shows, Terra is where you land.
GLM-5.2 is the margin-predictability pick, and that’s a different axis from price. In one 60-day window this summer, three labs changed terms under their customers’ feet - a suspension here, an overnight parameter cut there, a scheduled 50% price rise to finish. MIT weights are the only structural immunity: self-hosted, your unit cost is compute, and no vendor can reprice your product from the outside. The community’s cost-conscious builders already live here, and it’s the best open-weights performer on the long-horizon coding benchmarks this category actually exercises.
The counter-argument we take seriously: Gemini 3.6 Flash’s cached-input pricing ($0.15 - a 10× discount on the term that dominates agentic loops, since agents resend their context every step) makes it the spreadsheet winner for some workload shapes. And the revealed preference of the category’s archetype cuts against pure cost optimisation entirely: Lovable, at $400M ARR, serves a mid-frontier Claude model by default - because a regenerated broken app costs more than an expensive correct one. Model the retry rate, not just the sticker.
What to do now
Price your agentic feature at Grok rates and see if the business works. If it only works cheaper, you’ve built a chatbot with extra steps; if quality complaints eat the savings, take the Lovable lesson and move up a tier. And run the cache maths on your actual loop shape before finalising - the tokenizer tax explains why the cached number might be your real bill.