[Photo: Nvidia]

Nvidia unveiled a range of new AI inference accelerators and network infrastructure products at the Hot Chips 2026 semiconductor design conference in the United States on Aug. 25, a day before its second-quarter earnings release. The interactive AI inference accelerator Nvidia Grok 3 LPX has entered full-scale mass production. The company also introduced Spectrum-X Multiplane, an Ethernet scaling technology, and ScaleIn, an infrastructure acceleration technology. The announcements focused on the inference stage, particularly token generation speed.

Nvidia highlighted three pillars in the announcement: inference speed, network scalability and CPUs for agents. The company stressed that all three rest on a single platform called Vera Rubin, and is giving the market positive expectations for guidance in its earnings release. Nvidia will report second-quarter results on Aug. 26 local time, or early Aug. 27 in South Korea.

Mass production begins for Grok 3 LPX...up to 512,000 GPUs in a single layer

Nvidia said at Hot Chips 2026 that Grok 3 LPX has entered mass production. It is an inference accelerator that expands Vera Rubin, a next-generation AI factory platform, and Nvidia described it as the fastest performance measured so far on the model. According to Nvidia, Grok 3 LPX produced 3,400 output tokens per second in an Artificial Analysis benchmark applying a 100,000-token context to the open-source agentic model Gemma 4 31B.

AI cloud company Nebius plans to be the first to adopt Grok 3 LPX in its inference platform, Nebius Token Factory. AI inference-focused cloud company Grok also joined the list of early adopters.

Nvidia founder and CEO Jensen Huang (젠슨 황) said, "Inference is the driving force behind AI growth," and added, "Vera Rubin expands this vision with an AI factory configuration optimized for workloads for the agentic AI era, and pushes performance limits further with LPX for ultra-fast token generation."

Agentic AI is AI that completes tasks by going through multiple steps on its own without human intervention. It generates large volumes of tokens across inference processes spanning hundreds to thousands of steps, so slow generation speeds delay the entire task. Nvidia said agentic AI poses two challenges at the same time: large context processing and low-latency token generation. Grok 3 LPX is specialized in boosting token generation speed for individual users.

Nvidia also unveiled Spectrum-X Multiplane, a new Ethernet architecture technology. To scale AI factories beyond today’s largest clusters, a third network tier had been required. That approach increases latency and performance variation, and also raises costs for cables, optical components and power.

To address this, Multiplane splits a server’s network connections into multiple independent paths, or planes, and runs each plane as its own Layer 2 network. Nvidia said the approach creates a flat structure that can scale to 512,000 GPUs without a third tier.

A dedicated hardware engine inside the ConnectX SuperNIC handles path management. When a failure occurs, it immediately reroutes around the affected point, so applications and software recognize only a single connection. In an eight-plane configuration, overall bandwidth remains about 90 percent even if one plane fails, and hardware recovery is 11 times faster than software-based load balancing. Nvidia said this increases AI factory throughput by 1.6 times.

Nvidia also introduced technology that treats infrastructure services themselves as targets for acceleration. Nvidia ScaleIn, based on the BlueField-4 processor and the DOCA software platform, accelerates networking, storage, security and operations together. Nvidia said as AI factories scale, these infrastructure services must also be accelerated at the same pace as AI computing.

The newly unveiled network technology is designed to fit Vera Rubin NVL72. It includes Spectrum-X SN6000 series switches and the ConnectX-9 SuperNIC. Spectrum-XGS Ethernet, which links multiple data centers like a single AI superfactory, accelerates multi-site NCCL collective operations by 1.9 times.

SpaceXAI adopts Vera CPU...bringing the architecture into orbit

Nvidia is also sending compute devices into orbit. It said SpaceXAI will adopt the Vera CPU to accelerate next-generation agentic AI workloads. The orbital environment has different constraints from ground-based data centers in power and thermal management, bandwidth, reliability and physical integration.

Vera is the first CPU designed for AI agents. It includes 88 Nvidia-designed Olympus cores and high-bandwidth LPDDR5X memory to provide up to 1.2 TB per second of bandwidth. Nvidia said it increases task completion speed by up to 1.8 times versus x86 CPUs in agentic AI, reinforcement learning and data processing workloads.

Nvidia’s emphasis on CPUs reflects how agentic AI operates. Agentic applications are becoming more CPU-dependent as they orchestrate tools, execute code and process data between model calls. Mike Nichols (마이크 니콜스), president of SpaceXAI, said, "Vera enables GPUs to focus on what they do best while delivering the CPU performance and memory bandwidth needed for massive-scale orchestration, code execution and data processing."

SpaceXAI plans to expand AI infrastructure supporting Grok to gigawatt-scale computing capacity, based on Vera Rubin. It will apply an optimized Vera Rubin NVL72 system to its first-generation Starmind AI satellite. The idea is to move the architecture that powers ground-based AI factories into orbit.

Keyword

#Nvidia #Hot Chips 2026 #Vera Rubin #Grok 3 LPX #Spectrum-X Multiplane
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.