At the Hot Chips 2026 symposium, Nvidia announced that its Groq 3 LPX inference rack has officially entered full volume production. The milestone represents a major shift in hardware design aimed directly at the growing demands of agentic AI workloads. By initiating mass manufacturing, the chipmaker is addressing the critical industry need for minimal latency and massive throughput in production environments.
The hardware architecture of the Groq 3 LPX brings together Nvidia's Vera Rubin platform and specialized inference processors. While the Vera Rubin platform manages initial model training and the prefill stage, 256 dedicated Language Processing Units handle rapid token decoding. These LPUs are manufactured using Samsung's 4-nanometer process and build directly on the technology acquired through Nvidia's Groq transaction.
Benchmark results from Artificial Analysis indicate that the new system achieves performance levels of up to 3,400 tokens per second. This high throughput directly tackles the primary bottleneck in multi-step AI agent workflows, where models must generate multiple intermediate reasoning tokens in iterative loops. Previously, slow decoding speeds severely limited the responsiveness and practical deployment of autonomous software agents.
Cloud infrastructure provider Nebius has been confirmed as the first partner to deploy the Groq 3 LPX system live in its data centers. This partnership allows enterprise customers and developers to access the high-speed inference capabilities through cloud instances without managing physical hardware. The rollout is intended to support real-time enterprise AI applications across diverse industries.
The move underscores Nvidia's strategic push to dominate inference infrastructure as operational spending on running models surpasses initial training costs. By splitting heterogeneous compute workloads between prefill and decoding stages, the Groq 3 LPX establishes a new benchmark for data center efficiency. Deliveries of the production systems are set to expand enterprise capacity for high-speed language model execution.

