Where NPUs Fit into LLM Inference

By Jack Smith

7 min read
Share this post
Blog thumbnail titled "Where NPUs Fit into LLM Inference" with the Cascadia logo below it. On the right, a glowing chip labeled NPU sits at the center of a dark circuit-board grid.

Today, most modern laptops, computers, and mobile devices ship with a chip dedicated to AI: the NPU.

NPUs (Neural Processing Units) are specialized chips designed for AI/ML workloads. They were largely popularized due to Apple’s A11 chip, which included a neural engine to build facial maps and process biometrics on-device.

Since the rise of Large Language Models (LLMs), NPUs have moved more into focus in the desktop space. Intel popularized the term “AI PC” with their Core Ultra lineup, and AMD’s Ryzen PRO 7040 processors were some of the first to include an NPU.

NPUs are purpose-built for the math that most neural networks use, and can operate much faster (and with less power) than a CPU can for running inference.

So, where exactly do NPUs fit into the LLM scene? What are they good at, and where do they fall short? And, how can you use them for inference today?

NPUs, GPUs, and LLMs

On paper, NPUs sound like a good fit for generative AI. They are purpose-built for machine learning and are energy efficient. Yet today, GPUs dominate the AI hardware market. Why?

LLMs are rooted in the transformer architecture, and one of the main reasons for their emergence and success is scale. The computing power needed to train frontier models and perform inference is huge.

That said, the compute required to run inference is not often the bottleneck, however, memory capacity and bandwidth are.

Capable AI models can be extremely large in size, and need to be loaded into memory in order to perform inference. This means a device needs to have enough RAM (or VRAM) to hold an AI model comfortably, and the memory bandwidth to stream responses at a reasonable rate.

General Purpose GPUs (GPGPUs) act almost like a jack of all trades for this purpose, as they provide:

  • The hardware architecture for fast calculations (similar to NPUs)
  • The memory capacity to hold larger models
  • A mature ecosystem (NVIDIA’s CUDA support, for example)
  • The flexibility to perform other tasks (such as training, which NPUs aren’t designed for)

That being said, this does not mean GPUs are always better than NPUs for those looking to run LLMs. There are many scenarios, especially local and on-premises focused tasks, which make NPUs much more appealing compared to GPUs.

Where NPUs can help (prefill vs decode)

There are two different phases when it comes to LLM inference:

  • The prefill phase (processing the prompt)
  • The decode phase (generating a response)

NPUs can do especially well in the prefill phase, because it is ultimately compute-bound. When prefill takes place, all of the input tokens can be processed in parallel (which is perfect for the math that NPUs are good at).

For the decode phase, each token is generated one at a time. This means the bottleneck comes back to memory bandwidth as mentioned previously, as each token requires reading the model in memory again (something the NPU can’t particularly speed up).

While larger, more powerful models typically require more memory, smaller models can now fit in the memory of some NPU-equipped devices. This can be useful for those looking to keep AI workflows local/on-prem.

Open-weight models such as Qwen3.8-27B are drastically smaller than cloud models (~18GB in size quantized), while relatively close in performance to the older frontier models. This is much closer to what a realistic workstation could manage, although it wouldn’t be able to hold much else (and is probably the largest model a single 32GB workstation could handle).

As open models continue improving, they will become an attractive option to save on costly token expenses, or for privacy-focused workflows where an organization may want to use an LLM, but not have any data sent offsite.

Using NPUs in LLM workflows today

Like GPUs, an NPU can’t typically execute an AI model alone. You still need a host CPU for orchestration, loading models, tool calling, and more.

This is why many chipsets such as the Intel Core Ultra lineup contain a CPU, NPU, and integrated GPU on a SoC (System on a Chip), and split workloads across them. We’ve even integrated our own optimizations for NPU and CPU workloads in Cascadia.

To measure the difference, we ran Llama 3.2 1B (INT4) on an Intel Core Ultra laptop using Cascadia. We compared CPU-only inference against a split where the NPU handles prefill and the CPU handles decode, and measured the time to first token (TTFT) for three different lengths of prompt.

Config46 Tokens442 Tokens910 TokensDecode
CPU only693ms3893ms8173ms~9.5 tok/s
NPU prefill + CPU decode280ms921ms1518ms~9.3 tok/s
Speedup2.5x4.2x5.4x-

Using the NPU for prefill made the first token arrive 2.5 to 5.4 times faster, and the gain grew with prompt length (expected from the compute-heavy nature of prefill). Decode speed stayed about the same, since the CPU still handled that phase. For the 910-token prompt, the full response (64 tokens) arrived in 8.2 seconds instead of 14.9.

The split does come at a cost in memory, because it has to hold a second copy of the model weights compiled for the NPU. Peak memory usage was 10.2 GB compared with 3.2 GB for CPU only. These numbers are also from a single machine and a small model, so this is meant as an example rather than a benchmark.

Hardware is becoming much more heterogeneous this way and it’s likely that NPUs will make up part of an LLM inference system, rather than taking up the entire task themselves.

Depending on the hardware available, there are a few different inference runtimes/engines to consider.

OpenVINO

OpenVINO is the gold standard for Intel-based hardware. It has wide support for Intel CPUs, NPUs, and GPUs, and can split workloads across them with automatic multi-device execution.

OpenVINO supports the entire Intel Core Ultra processor lineup, alongside a handful of more specialized Intel NPU chips.

Intel’s newest Core Ultra chips also have up to 50 NPU TOPS (trillion operations per second), and are beginning to comfortably support plenty of operations for smaller models.

ONNX Runtime

ONNX doesn’t have native NPU support, but provides a variety of Execution Providers (EPs) which hand off the work to vendor-specific runtimes.

If you have hardware which isn’t Intel-based, ONNX Runtime GenAI supports many hardware vendors for LLMs and NPUs using their EPs.

This includes support for Qualcomm and their Snapdragon/Hexagon NPUs, AMD’s Ryzen AI NPUs, and other specialized hardware and manufacturers.

llama.cpp

llama.cpp, like ONNX, doesn’t have native NPU support but offers various “backends” which can provide NPU execution. This includes support for Intel via OpenVINO, Huawei Ascend, and a range of community-supported platforms (some of which are experimental).

In most cases, hardware vendors provide their own tooling for running AI workloads on their NPUs (such as Apple and their new Core AI framework). However, for LLM workflows in particular, support can be a little spotty and is still in development in a lot of places.

The definition of an NPU is changing

With the introduction of ASICs (Application-Specific Integrated Circuits), the line between “NPU” and “GPU” is blurring.

Google’s custom Tensor Processing Units (TPUs), for example, are highly optimized for LLM training and inference, and are essentially data-center grade NPUs (with some crossover with GPUs).

In other words, the NPUs that exist in laptops today are being adopted on a much larger scale and in a much more powerful setting. More focus coming to NPUs this way might also lead to better support and tighter integration with many LLM-based systems.

Cascadia: run AI models across multiple NPUs

NPUs are abundant and we built Cascadia to help organizations put their idle Intel NPUs (and CPUs) to work, instead of purchasing expensive AI hardware.

One of the bottlenecks mentioned in this post was memory capacity, and Cascadia solves this by pooling together the memory of the devices you own (such as desktop workstations and employee laptops). This means you can run a much more powerful model than any individual machine could.

If you’d like to learn more about what we do and how we can help your team, book a call with us.

Cascadia is also open source. You can check out our GitHub repository here and star it if you find our work useful.

Share this post

Bring a fleet. Keep everything.

We stand up your first private agent in under a week, on your hardware, on your network. You keep all of it.

Free for individuals

Download

Apply to become a design partner

Request a Pilot