Nvidia’s inference rack Groq 3 LPX delivers record 3,431 tokens per second

Nothing chews through inference tasks as fast as Nvidia’s Groq. By far. (Picture: Nvidia/generated)
The never-before-seen feat was achieved running Google’s Gemma 4 31B through Artificial Analysis’ standard tests, and is way ahead of anything on the market.

The only comparable score is that of OpenAI’s GPT-Sol running on Cerebras chips, which achieved 750 tokens per second earlier in August. They called this «Ultrafast mode.»

For more normal hardware setups, Opus 5 gets 58.8 tokens per second, and GPT-5.6-Sol clocks in at 74.4 per second, while the «faster» Gemini 3.7 Flash gets 371.1 throughput tokens.

Nvidia further says that it achieved this output score while maintaining hundreds of thousands of context tokens, and that Groq is some 34X faster than today’s quickest hardware in time to generate 5,000 tokens.

A Groq chip pairs 500 MB of high speed SRAM on die, directly next to the chip, that delivers 150 TB/s throughput. They come in racks of 256 chips stacked together for a total of 40 petabytes of memory bandwidth.

While they are great for inference tasks, GPUs will still be the workhorse of AI data centers, as they can handle both training and later inference — but Jensen Huang of Nvidia recommends setting aside 25% of data center space for the new Groq chip racks, according to CNBC.

The new test scores come as Nvidia is announcing that Groq 3 chips are now in production with Samsung and are generally available.

Read more: Nvidia’s announcement, production note. Writeups on CNBC and The Register.