Hot Chips 2026: Nvidia presents Groq 3 LPX architecture and unveils its first third-party inference benchmark — LP30-based rack already in production, company says
- ID
- 18271
- Status
- summarized
- Published
- 27 Aug 2026, 12:23 AM
- Fetched
- 27 Aug 2026, 1:49 AM
- Provider
- Tom's Hardware
- Category
- technology
- Original URL
- https://www.tomshardware.com/tech-industry/semiconductors/nvidia-presents-groq-3-lpx-architecture-and-unveils-its-first-third-party-inference-benchmark
- Source URL
- https://www.tomshardware.com/feeds/all
Summary
- Score
- 7.5
- Created
- 27 Aug 2026, 1:50 AM
- Tags
- Audience
- developersai_ml_learnersai_agent_userssaas_founders
What happened
At Hot Chips 2026, Nvidia's Igor Arsovski (former Groq chief architect, now Nvidia VP of hardware) presented the Groq 3 LPX inference rack architecture and released the first third-party benchmark: Artificial Analysis measured 3,431 output tokens/sec on a 100K-context Gemma 4 31B reasoning workload, roughly 4x the 870 tokens/sec of the next-fastest public endpoint. The rack is already in production, built on the LP30 chip from Nvidia's $20 billion Groq acquisition in December 2025, which displaced the Rubin CPX from Nvidia's roadmap.
Why it matters
If you build AI agent or long-context reasoning workloads, a 4x decode-speed advantage at 100K context directly changes latency budgets and per-query economics. Builders evaluating inference providers should track when Groq 3 LPX endpoints become available through cloud partners, as the gap could shift which provider is cost-competitive for high-throughput agent pipelines.
Discussion angle
Compare the 3,431 tokens/sec figure against what your current inference provider delivers for long-context agent workloads, and discuss whether the 4x gap justifies waiting for LPX endpoint availability or re-architecting around a different inference stack now.