NPU
An NPU is a dedicated chip block that runs AI inference at low power, rated in TOPS, separate from the CPU and the GPU.
An NPU, or neural processing unit, is a dedicated block of silicon built to run the matrix multiplications behind neural networks. It sits apart from the CPU's general-purpose logic and the GPU's graphics pipeline. It executes low-precision math, usually INT8 or INT4, across many small parallel units, trading flexibility for throughput per watt on inference.
Laptop NPUs are rated in TOPS, trillions of operations per second, at INT8. Intel's Core Ultra 200V series is rated around 48 TOPS, Qualcomm's Snapdragon X Elite around 45 TOPS, and Apple's M4 Neural Engine around 38 TOPS. Microsoft set 40 TOPS as the floor for a Copilot+ PC. A discrete GPU posts far more peak throughput, but draws tens to hundreds of watts doing it. The NPU handles a continuous background task, camera background blur or noise suppression, on a few watts and leaves the GPU free.
For workstation buyers it matters in laptops and thin clients running on-device AI features all day. It is no substitute for training or heavy inference, which still want a discrete GPU with real VRAM and a CUDA or ROCm stack behind it.
Sources
Source | Publisher |
|---|---|
Intel | |
Intel | |
AMD |
- Publisher
Intel
- Publisher
Intel
- Publisher
AMD
Last verified August 29, 2026.
- neural processing unit
- AI accelerator
- on-device AI
- TOPS INT8
- Apple Neural Engine
- Copilot+ PC
- Hexagon NPU
- Intel AI Boost
- inference accelerator