NVIDIA Groq 3 LPX Enters Full Production With 3,400 Tokens Per Second on Vera Rubin
NVIDIA has confirmed that Groq 3 LPX is now in full production, adding a specialized inference accelerator to the Vera Rubin platform as the company targets increasingly demanding agentic AI workloads. According to NVIDIA, Groq 3 LPX delivered approximately 3,400 output tokens per second while running Gemma 4 31B with a 100,000 token context in testing by Artificial Analysis, which NVIDIA says was around 4x faster than the nearest alternative platform in the same benchmark.
The result is important because Groq 3 LPX is not positioned as a replacement for NVIDIA Rubin GPUs. Instead, NVIDIA designed the 2 architectures to work together. Vera Rubin NVL72 handles high throughput workloads including model prefill, long context processing and attention, while Groq 3 LPX specializes in the latency sensitive token generation stages that become increasingly important when AI agents need to reason, call tools and execute long chains of actions. NVIDIA Dynamo coordinates these different processors and routes workloads between them according to where they can run most efficiently.
At the hardware level, a complete Groq 3 LPX rack contains 256 Groq 3 LP30 accelerators delivering 315 PFLOPS of FP8 inference compute. The system includes 128 GB of total on chip SRAM, 40 PB/s of SRAM bandwidth and 640 TB/s of rack scale communication bandwidth. Unlike GPU architectures that depend heavily on HBM, the LPU architecture emphasizes extremely fast SRAM access, deterministic compiler scheduled execution and tightly controlled data movement to reduce latency variation during inference.
That architecture is particularly relevant for agentic AI because generating a response becomes more difficult as workflows become longer. A traditional chatbot may produce one relatively simple answer, while an autonomous agent can repeatedly reason, use tools, evaluate results and continue working across hundreds or thousands of inference steps. Small delays during token generation therefore accumulate quickly, creating a new infrastructure bottleneck even when enormous amounts of GPU compute are available.
NVIDIA's latest technical results show Groq 3 LPX reaching 3,431 output tokens per second in the Artificial Analysis 100K context test using Gemma 4 31B. NVIDIA also reported a median 4,767 output tokens per second in SPEED Bench coding workloads. The company says the architecture can support several configurations alongside Vera Rubin, including separating prefill from decode, dividing attention from feed forward network execution and using LPX as an external speculative decoding engine.
The production milestone follows NVIDIA's earlier positioning of Groq 3 LPX as one of the central components of its broader Vera Rubin infrastructure. A complete LPX system was previously detailed in our coverage of Groq 3 LPX and Foxconn's expanding Vera Rubin production, where supply chain reports pointed toward an aggressive 2026 manufacturing ramp. NVIDIA now officially confirming full production moves the platform beyond projections and closer to large scale commercial deployment.
Groq itself will be among the first companies deploying the new platform. The company confirmed that it is working with Dell Technologies to introduce NVIDIA Groq 3 LPX alongside Vera Rubin NVL72 inside its purpose built inference cloud. This creates an interesting relationship where technology developed around Groq's LPU architecture is becoming integrated directly into NVIDIA's broader AI factory platform while Groq continues operating its own inference infrastructure.
NVIDIA claims the combined Vera Rubin and Groq 3 LPX architecture can provide up to 35x higher inference throughput per megawatt for trillion parameter models compared with previous configurations. These remain NVIDIA performance claims and actual efficiency will depend heavily on model architecture, context length, batching, software optimization and deployment configuration. However, the new Artificial Analysis result provides an important third party data point showing that LPX can deliver extremely high token generation speeds under long context conditions.
Groq 3 LPX shows how NVIDIA's AI strategy is evolving beyond simply making each GPU generation faster. Training, prefill, attention and token generation do not necessarily benefit from exactly the same hardware architecture, so combining Rubin GPUs with specialized LPUs gives NVIDIA another route toward improving total AI factory efficiency.
The most interesting number may not even be 3,400 tokens per second. It is the combination of 40 PB/s SRAM bandwidth and deterministic execution that allows LPX to attack latency differently from conventional GPU inference. If agentic AI continues consuming dramatically more tokens and requiring longer sequences of decisions, responsiveness could become just as strategically important as total compute throughput. Vera Rubin plus Groq 3 LPX represents NVIDIA's attempt to optimize both.
Will specialized inference accelerators like Groq 3 LPX become essential for agentic AI, or can future GPUs eventually handle training and ultra low latency inference efficiently on their own?
