Nvidia has shifted its Groq 3 LPX inference accelerator dedicated to AI agents into a mass-production system. [Photo: Nvidia]

Nvidia has shifted its dedicated inference accelerator for artificial intelligence agents, Groq 3 LPX, into a mass-production system.

On Aug. 24 local time, foreign media outlets including SiliconANGLE reported that the chip extends Nvidia's Vera Rubin data center platform and targets ultra-fast token generation needed for agentic AI. Nvidia unveiled the chip at Hot Chips 2026. Nebius Group was named as the first customer to adopt Groq 3 LPX.

AI inference refers to deploying a trained model in real services. As autonomous AI agents that perform tasks on behalf of humans increase, they process large volumes of tokens while repeatedly carrying out inference and planning, writing and executing code, checking system files and calling external tools. In this process, "decode latency" has emerged as a key bottleneck that slows user response speed. Nvidia explained that reducing it requires a dedicated computing structure that separates large-context processing from token generation.

Vera Rubin NVL72, a rack-scale platform, serves that role. Multiple Vera Rubin graphics processing units handle large-context input and processing, and adding Groq 3 LPX enables decode tasks to be handled separately. Nvidia said a full rack configuration can connect up to 256 LP30 accelerators via an ultra-high-bandwidth chip interconnect. It said the GPUs and language processing units operate together as an integrated inference engine that handles all stages of AI agent inference.

Nvidia said this can eliminate the trade-off between throughput and response speed. It said Groq 3 LPX supports AI agents without latency as they scan long context windows, verify data, call external tools and repeatedly perform complex multi-step tasks in real time. In an Artificial Analysis benchmark, Groq 3 LPX set an all-time record of 3,400 tokens per second when running the open-source Gemma 4 31B agentic model with a 100,000-token context window. Nvidia said the figure shows responsiveness 4 times higher than competing platforms in latency-sensitive tasks and can cut multi-step agent work time from hours to minutes.

The chip was developed based on licensed technology from Groq, a small semiconductor startup. Nvidia secured access to the technology in December last year for $20 billion, or about 27.7 trillion won. Under the same deal, it also hired Groq founder Jonathan Ross (조너선 로스) and President Sunny Madra (서니 마드라). Groq is a different company from SpaceX's AI model Grok and has developed processors specialized in inference rather than AI training.

Nvidia CEO Jensen Huang (젠슨 황) said the company has already achieved innovation in AI inference performance and efficiency with its Grace Blackwell and NVL72 platforms. "Vera Rubin extends this with a workload-optimized AI factory configuration designed for the agentic AI era," he said. "We are expanding performance limits with LPX for ultra-fast token generation. This changes how intelligence is produced and delivers another major leap in AI throughput, efficiency and responsiveness," he said.

Nebius Group is one of the early companies to adopt the Groq 3 LPX chip. It plans to apply the chip to its production inference platform, Nebius Token Factory, to provide extremely fast token generation speeds for agentic applications where responsiveness is critical.

Danila Shtan (다닐라 시탄), Nebius' chief technology officer, said, "The generation stage is the inference process that determines the actual responsiveness of an AI system, and Groq 3 LPX is designed to accelerate exactly this point." "Through Nebius Token Factory, as the first AI cloud to commercialise it, we are making every stage of the agent loop feel instantaneous," he said.

Nvidia also said at the Hot Chips event that it has secured SpaceX as a new flagship customer. SpaceX plans to build its next-generation AI architecture around the Vera Rubin platform. It will deploy Nvidia's Vera CPUs for CPU-intensive orchestration, tool execution and simulation tasks across ground data centres and satellites in orbit.

Nvidia also unveiled several complementary technologies for AI factories. They include Spectrum-X Multiplane, an AI-optimised ethernet architecture that supports cluster scaling of up to 512,000 GPUs; Nvidia Scale-In, an infrastructure software platform that accelerates by offloading security, network and data management from main host compute nodes; and Nvidia NVLink Fusion, which connects custom CPUs and data processing devices to sixth-generation NVLink rack systems.

Keyword

#Nvidia #Groq 3 LPX #Vera Rubin #Nebius Group #SpaceX
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.