Qualcomm says its next Hexagon NPU is designed for on device AI agents, adding a transformer focused Element Accelerator, 50% more shared memory, support for contexts up to 32,000 tokens, and up to 50% faster INT4 pre... The NPU is part of a three engine strategy: Oryon CPU cores handle general purpose processing, A...
Published byEdited with GPT-5.6 TerraImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Qualcomm reveal on September 9, 2026 about the redesigned Hexagon NPU in its next-generation premium Snapdragon mobile platform—inc. Article summary: Qualcomm’s September 9 preview positioned its next premium Snapdragon platform as an on-device “agentic AI” design, not merely a faster phone chip. The redesigned Hexagon NPU adds dedicated transformer hardware and more . Topic tags: general, general web, user generated. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fa
Qualcomm’s September 9 preview makes its direction clear: the next premium Snapdragon platform is being built around longer-running, multimodal AI workloads that can execute on a phone. The centerpiece is a redesigned Hexagon NPU intended for generative and agentic AI, with new transformer-specific hardware and a larger local memory pool. 1
The new NPU adds an Element Accelerator, a dedicated block for transformer workloads. Transformers underpin many generative AI systems, including language and multimodal models. Qualcomm positions the accelerator alongside existing scalar, vector and tensor-oriented processing resources rather than as a replacement for them. 1
11
The company also says shared NPU memory is 50% larger than in the prior design. Its stated goal is to keep frequently accessed model data close to the NPU, reducing the need to fetch it repeatedly from system memory while an AI agent works across longer contexts, tools and concurrent tasks. Qualcomm has not disclosed the absolute capacity of that shared memory. 1
5
Qualcomm says the revised Hexagon architecture supports context lengths of up to 32,000 tokens and includes key-value-cache acceleration. In practical terms, a larger context lets a model retain more of a conversation, document or task history during an on-device session. 11
12
It also claims up to 50% faster prefill for INT4 models. Prefill is the stage in which a model processes the user’s initial prompt and establishes its working context before generating a response. That figure should not be read as a promise of 50% faster output-token generation in every app or model. The NPU supports precision formats spanning INT2, INT4, INT8, FP8 and FP16. 1
6
Qualcomm highlights support for Mixture-of-Experts, or MoE, models with up to 30 billion total parameters. The example activates roughly 3 billion parameters per token. 6
That distinction matters. In an MoE model, a routing mechanism selects only part of the model for a given token or task. The claim therefore does not mean a phone performs inference across all 30 billion parameters for every token; its feasibility depends on selective expert activation and the particular model, implementation and device configuration. 1
13
The NPU announcement follows Qualcomm’s previews of the next platform’s CPU and GPU. The upcoming eight-core Oryon CPU is set to include two Prime cores reaching 5GHz and six Performance cores. Qualcomm’s FlexCache design gives heterogeneous CPU cores access to a dynamically allocated shared cache pool. 21
22
For graphics, Qualcomm has announced Adreno Matrix Cores and 18MB of Adreno High Performance Memory. Its Neural Fusion technology combines neural processing, AI super resolution and frame generation within the graphics pipeline. Qualcomm says the technology can reduce power consumption by up to 40% for applicable rendering workloads; that is a vendor claim, not an independent benchmark result. 17
23
Together, these disclosures describe a heterogeneous approach to local AI:
The important takeaway is not that every phone feature will run on the NPU. Qualcomm is presenting CPU, GPU and NPU resources as complementary engines that device makers can use for different parts of an on-device experience.
These were architecture previews, not a full platform specification. Qualcomm’s Snapdragon Summit is scheduled for September 22–24 in Maui, Hawaii. 31
Until that event, buyers and developers should treat the disclosures as a technical direction rather than a complete performance picture. The company has not yet provided absolute shared-memory capacity or full, independently verifiable performance and efficiency results for the new NPU. 5 The eventual platform announcements will determine how these capabilities are configured across actual Snapdragon products and, ultimately, how they translate into shipping phones.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
Qualcomm says its next Hexagon NPU is designed for on device AI agents, adding a transformer focused Element Accelerator, 50% more shared memory, support for contexts up to 32,000 tokens, and up to 50% faster INT4 pre...
Qualcomm says its next Hexagon NPU is designed for on device AI agents, adding a transformer focused Element Accelerator, 50% more shared memory, support for contexts up to 32,000 tokens, and up to 50% faster INT4 pre... The NPU is part of a three engine strategy: Oryon CPU cores handle general purpose processing, Adreno Matrix Cores target AI assisted graphics, and Hexagon is aimed at sustained local inference.
Qualcomm’s 30 billion parameter model example is a Mixture of Experts design that activates about 3 billion parameters per token—not all 30 billion at once.