Search results for Hot Chips 2026
AI & Enterprise
Nvidia begins mass production of Groq 3 LPX for AI agents
Nvidia has shifted its Groq 3 LPX inference accelerator for AI agents into mass production. The chip extends Nvidia\'s Vera Rubin data center platform and targets high-speed token generation to reduce decode latency in agent workflows. In an Artificial Analysis benchmark, it produced 3,400 tokens per second running the open-source Gemma 4 31B model with a 100,000-token context window. Nebius Group is among early adopters, and SpaceX was named a new flagship customer.
Industry
Nvidia unveils four new products ahead of earnings, focusing on token generation speed
Nvidia unveiled new AI inference accelerators and network infrastructure products at Hot Chips 2026 in the United States on Aug. 25, a day before its second-quarter earnings. The company said interactive inference accelerator Nvidia Grok 3 LPX has entered mass production, alongside Spectrum-X Multiplane Ethernet scaling and ScaleIn infrastructure acceleration technology. The announcements focused on inference, particularly token generation speed, and also highlighted inference speed, network scalability and agent-focused CPUs within the Vera Rubin platform.
Industry
Post-HBM paths diverge as Samsung stacks and SK hynix connects
Samsung Electronics and SK hynix are pursuing different approaches for the next step after high-bandwidth memory, as rising AI workloads make it hard to expand bandwidth and capacity with HBM alone. Samsung is focusing on 3D vertical stacking that places HBM on top of accelerators to reduce data-movement bottlenecks. SK hynix is standardising high-bandwidth flash with SanDisk, linking different memory tiers via UCIe to raise system capacity and efficiency.