Nvidia and Cerebras are selling performance their customers will (probably) never see
- ID
- 18907
- Status
- summarized
- Published
- 28 Aug 2026, 6:57 AM
- Fetched
- 28 Aug 2026, 10:16 AM
- Provider
- The Register
- Category
- technology
- Original URL
- https://www.theregister.com/systems/2026/08/27/nvidia-and-cerebras-are-selling-performance-their-customers-will-probably-never-see/5293117
- Source URL
- https://www.theregister.com/headlines.atom
Summary
- Score
- 7.0
- Created
- 28 Aug 2026, 10:17 AM
- Tags
- Audience
- developersai_ml_learnerssaas_founders
What happened
Nvidia and Cerebras are trading benchmark blows at Hot Chips, with Nvidia claiming 3,400 tokens/sec on Gemma 4 31B using Groq-3-based LPX racks and Cerebras countering with its upcoming CS-4 accelerators. Both figures are measured at batch size 1 — a single concurrent request — which no production inference-as-a-service operator would run because it's economically unviable. The numbers are real but analogous to a car's top speed: technically achievable, practically irrelevant for paying workloads.
Why it matters
If you're evaluating inference hardware or picking an inference provider, don't anchor on single-request token throughput. What actually matters is throughput-per-dollar at realistic concurrency levels along the Pareto frontier. Ask vendors for batched benchmarks at the concurrency you expect to serve, not peak single-request numbers.
Discussion angle
When choosing between inference providers or hardware, what benchmark numbers should you demand instead of peak token/s — and how do you map Pareto frontier charts to your actual concurrency and cost targets?