Qualcomm disclosed NPU performance for its next-generation Snapdragon mobile platform to be announced at the Snapdragon Summit. [Photo: Qualcomm]

[DigitalToday reporter Daegeon Seok (석대건)] Qualcomm has redesigned its AI processor architecture to run agentic AI on smartphones. Qualcomm on Sept. 11 disclosed design details and performance figures for the Hexagon NPU, a dedicated AI processor to be included in its next premium mobile platform.

Qualcomm said it added a new compute accelerator, judging that agentic AI is better served by using specialised models tailored to tasks, situations and users rather than a single large model. It also increased shared NPU memory by up to 50 percent.

Compute for transformers, the model structure underpinning generative and agentic AI, is handled by a new accelerator called the Element Accelerator. Large-model compute is accelerated with both vector and scalar extension functions. The vector extension supports large volumes of AI math, while the scalar extension supports an agent's logic for judgement, task allocation and coordination. Qualcomm said this allows AI agents to respond and reason faster while maintaining power efficiency.

AI models repeatedly draw on the KV cache, which stores conversation context and previously computed results while generating answers. If there is not enough space inside the NPU, data must move to and from DDR, the device's main memory, creating delays. With shared memory increased by up to 50 percent, more of this data can be kept inside the NPU. With less DDR access, latency falls and freed memory bandwidth can be used for other tasks.

With more memory headroom, the NPU can focus on inference and on-device models can handle longer context, use multiple tools and run concurrent tasks, Qualcomm said. Prefill performance, the first stage where AI reads questions and context at once, rose 50 percent for Int4 models compressed to 4-bit precision. Time to generate the first token is under 1.5 seconds on a 4 billion-parameter model.

Developers can tune the balance among performance, memory, model quality and power. Precision support includes Int2, Int4, Int8, FP8 and FP16. It also supports mixture-of-experts, or MoE, models that select only the required expert networks depending on the input. Unlike conventional models that run the entire network each time, this reduces compute and memory burden.

Qualcomm is developing a method with memory and model companies to load some experts from flash storage. The full set of experts remains available, but about 3 billion parameters are activated per task. The company said its "intelligent flash-to-memory expert management," combined with special caching techniques, delivers a larger-model-class AI experience with less power and memory.

Qualcomm said it has completed its performance previews for the next-generation Snapdragon CPU, GPU and NPU for implementing agentic AI, with the Hexagon NPU as the last. The "Oryon" CPU directs planning, reasoning and coordination at each stage and hands off tasks that require cloud APIs. It operates at 5 GHz per prime core and applies a new cache structure called "Flex Cache." The "Adreno" GPU handles AI inference in a hybrid environment where multiple models run on a device simultaneously alongside the NPU. Qualcomm said it will disclose detailed features at Snapdragon Summit 2026.

Keyword

#Qualcomm #Hexagon NPU #Snapdragon Summit 2026 #Element Accelerator #Adreno
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.