Frontend - No Reference
Starting from a loose brief and making the visual decisions. Taste, range, and restraint matter more than raw coding strength.
Designing without a reference is a taste problem disguised as a coding problem. Kimi and Claude are meaningfully better at it - they explore a visual idea, make coherent choices, and are less likely to mistake ‘premium’ for fourteen floating panels and a gradient blob.
The distinction matters because every frontier model can produce valid CSS. That is table stakes. The hard part is deciding what should dominate, what can disappear, how much copy the page actually needs, and when the design is finished. A model can be brilliant at implementation and still have dreadful instincts about the blank canvas.
The Least Bad Benchmarks
Our benchmarks of choice here are Design Arena and Arena’s WebDev leaderboard. Neither is a benchmark in the usual deterministic sense - they are Elo-style leaderboards built from real 1v1 outputs judged by humans. That is precisely why they are useful for design. A test can check whether the button exists; a person can tell you whether the page has taste.
They are not perfect. Prompts, harnesses and voter preferences all shape the result, and an arena rank cannot tell you which model fits your product. This category is necessarily vibes-based - the leaderboards give those vibes a much larger and blinder sample than one person’s screenshots folder.
Why Kimi K3 takes gold
Kimi K3 is the strongest first call when the brief gives it room to design. It has range - slick product UI, restrained editorial layouts, expressive marketing pages - without forcing every prompt through the same house style. More importantly, it tends to establish a visual thesis before decorating the page. That produces interfaces which feel composed rather than accumulated.
There is a limit. K3 does not have Sol’s appetite for hunting down every implementation fault, and a beautiful first pass can still conceal shaky responsive behaviour, loose semantics, or code that becomes expensive on the third iteration. Use it for the judgement-heavy leg, then review the build with a model that enjoys finding faults.
The Claude Pair
Claude Opus 5 takes silver because it is the most reliable balance of taste and product judgement. It is particularly good at removing things - collapsing weak sections, cutting explanatory copy, and making the working surface obvious. That restraint matters when every model has been trained on a web full of over-designed landing pages.
Claude Fable 5 takes bronze on breadth. Give it a genuinely open brief and it can produce the most considered concept of the three, especially when the interface depends on understanding the product rather than choosing a colour palette. It sits below Opus because the extra capability and cost rarely buy a proportionally better first-pass UI. For a flagship surface where the brief is still mush, that ordering can flip.
Why GPT Misses The Podium
GPT models are relentless builders. They also over-design more aggressively than any other family here - excess chrome, too many bordered regions, copy narrating what the interface already says, and a habit of treating every empty patch as unfinished work. With no reference to constrain it, Sol’s greatest strength becomes the problem: it will keep working.
That does not make it a bad frontend model. It makes it the wrong first model for this job. Sol produces excellent code on average and will often turn a tasteful but imperfect prototype into the strongest finished implementation. Asking it to invent the taste and police the implementation in the same pass wastes what it is best at.
What To Do Now
Ask Kimi, Opus or Fable for two to five genuinely different directions, each with a named visual thesis rather than a shuffled palette. Pick one. Let the same family build the first pass so the implementation preserves the idea, then hand the result to Sol with a narrow brief: find bugs, accessibility failures, responsive breakage and unnecessary code without redesigning the page.
Taste first. Teeth second.