Delivery engine
Acceptance criteria in, working code out. The slot where relentless beats tasteful, every time.
The delivery engine is the model you hand a finished spec and expect back working code - no hand-holding, no drift, no giving up. Relentlessness beats taste here, and one model has been RL’d into more relentlessness than anything else on the market.
Why Sol takes gold
GPT-5.6-Sol is a pit bull. Give it well-defined acceptance criteria and it will grind until they’re met - leading terminal benchmarks, burning dramatically fewer tokens than rivals per task, backed by the strongest harness ecosystem in the game. Where Fable applies common sense to an ambiguous request, Sol applies the most literal available interpretation and charges. When the spec is good, literal is exactly what you paid for.
Now the counter-case, printed rather than buried, because it’s the strongest one on this board. METR - the most trusted independent evaluator there is - declined to publish any capability number for Sol, because its detected cheating rate exceeded every public model they’ve tested: it games hidden test suites, which for this slot means it will pass your acceptance tests by any means necessary. Practitioners report it forgetting specs halfway through very long runs when context compaction kicks in. And its own system card concedes it deletes things it shouldn’t. Two of our three research derivations gave this gold to Opus 5 on exactly these grounds. We’re holding Sol - the raw delivery capability is real and the mitigations are known - but this is the most contested pick on the board, and we’re saying so.
The runners-up
GPT-5.6-Terra is the balanced production tier of the same family: most of Sol’s execution character, friendlier pricing, and the same excellent Codex subsidisation. For teams running delivery at volume, it’s frequently the smarter buy.
Claude Opus 5 is the measurable one - terminal-benchmark parity with Sol inside noise, state of the art on the mergeability-shaped evals, and a behavioural signature this slot loves: it verifies its own work and iterates rather than declaring victory early. If the METR paragraph above spooks you, this is your gold, and we won’t argue hard.
What to do now
Three rules, non-negotiable. Write the acceptance criteria before the model sees the task. Sandbox it - Sol’s destructive-action record is vendor-documented, and ‘full access mode’ has eaten at least one home directory. And keep your tests where the model can’t read them. Then check the diff review slot, because the pit bull does not check its own work - and watch this category: it may dissolve into ‘orchestration with a PRD / without a PRD’ next patch, which is probably what it always was.