In a previous article, we explored how to improve customer support chatbot response times, with the biggest lever being inference provider selection. Our benchmark revealed a wide performance gap among providers, with Cerebras outperforming competitors.
Now we’re evaluating a new entrant in the market, Celeris. Celeris is focused on building the world’s fastest language models through diffusion (generating multiple parts of a response in parallel).
Celeris-1 is ready for customer support workflows
When Celeris first launched Celeris-1, it didn’t yet support all the features needed for a cutting-edge customer support implementation, such as native tool calling.
It has since added support for these features, and in our testing the model’s capabilities match those of our leading model, GPT OSS 120B.
Benchmark results
We tested Celeris using the same benchmark setup as our previous comparison. Each provider received the same inputs, with three requests made during every testing interval. We measured and averaged the time required to generate each complete response.
Celeris delivered on its speed claims.
Across the entire benchmark period, we observed approximately 46% lower response times from Celeris-1 compared with GPT OSS 120B on Cerebras. Celeris managed an average response time of 110 ms, as compared with 204 ms on Cerebras.

We benchmarked Celeris, Cerebras, Baseten, Cloudflare, Fireworks AI, and Together AI using the same input.
The remaining providers we tested showed similar performance to our prior test and remained well behind the leaders.
Speed comes at a price
For latency-sensitive applications, there’s a clear advantage to using Celeris. All that speed does come at a cost, however: Celeris is priced at $2 per million input tokens and $6 per million output tokens. Compare this to $0.35 and $0.75 per million, respectively, for GPT OSS 120B on Cerebras. That’s a 5-8x price difference.
The takeaway
Celeris-1 changes what is possible for customer support experiences where every millisecond matters, but it doesn’t make cost-performance tradeoffs disappear.
Dynamic model selection remains the best strategy: use a fast model when necessary, but optimize for cost and performance when speed isn’t the focus.
Valiopt can help you optimize the speed, performance, and cost of your support automation. Get in touch to see how we can help.