Chatbots
High-volume conversation where unit economics decide everything. Good enough beats brilliant here.
The Serving section ranks models you build into a product, and this is its bluntest slot: high-volume conversation, invoked thousands of times a day by users who will never know or care which model answered. The only question that matters is the cheapest ‘good enough’, because every penny above that line is margin you’re donating to a lab.
Why DeepSeek V4 Flash takes gold
V4 Flash is basically unrivalled on value: crazy low pricing - $0.139 in, $0.278 out per million, roughly 9× cheaper on output than anything comparable from a Western lab - with modest but outsized capability, comfortably past the good-enough bar for support, onboarding, and everything else chat-shaped. It’s MIT-licensed too, so at real scale you can self-host and reduce your marginal cost to compute. The honest gap in the case: nobody has published latency or throughput data for it at chatbot concurrency, and for some buyers the hosting provenance is a procurement conversation. Neither has stopped the people actually shipping on it.
The runners-up
Gemini 3.5 Flash-Lite is a much more capable model at a higher price point - $0.30/$2.50, the fastest agentic-capable model tracked, a 1M context window, and the best-documented uptime of any provider. It’s arguably a worthwhile trade depending on the task, and the practitioners serving in production lean Gemini for reasons no leaderboard captures: ‘it just responds really reliably’, and the API rate limits are generous at volume. If your chatbot occasionally needs to do something rather than just say something, start here.
DeepSeek V4 proper is the bigger, considerably more capable sibling, still priced very aggressively. It’s the upgrade path when your conversations turn out to be harder than you thought - same family, same licence, no re-platforming.
Honourable mention, and a genuinely underrated one: Claude Haiku 4.5. It’s the only calibrated model in the field - a 28% hallucination rate against the frontier’s ~50% - and in support chat, where a confidently wrong answer becomes a ticket, calibration is unit economics wearing a disguise. Its constituency is small and unfashionable and completely right.
What to do now
Benchmark V4 Flash against whatever you’re currently paying for; the delta is usually your margin. Route the hard 20% of conversations up a tier rather than buying brilliance for the easy 80%. And read the tokenizer tax before comparing any of these stickers - output price is where a chatbot’s bill actually lives.