CPU vs GPU vs NPU: What’s Really Different?
Your smartphone or laptop may have a CPU, a GPU, and an NPU, but they are designed for very different kinds of work. This guide explains, with diagrams and tables, why each exists and how embedded AI engineers should divide work between them.
- The big picture
- CPU: the flexible manager
- GPU: the parallel specialist
- NPU: the AI specialist
- Side-by-side comparison
- Why neural networks suit NPUs
- How they work together on one chip
- Memory, the hidden bottleneck
- Precision and quantization
- Real workloads: who runs what
- Guidance for embedded AI engineers
- Common myths
- Summary
1. The big picture
Every processor is a trade-off between three things: how flexible it is, how fast it is on a given job, and how much energy it burns. No design wins on all three. A chip that can run any program must spend silicon on branch prediction, caches, and scheduling. A chip that only multiplies matrices can dedicate nearly all of its area to arithmetic and run far more efficiently, but it cannot run your operating system.
That is why modern system-on-chip (SoC) designs in phones, laptops, cameras, cars, and IoT devices combine several processor types. The easy mnemonic is:
2. CPU: the flexible manager
The Central Processing Unit is the processor every program ultimately depends on. It runs the operating system, schedules tasks, handles input and output, executes application logic, and performs calculations of every kind. Its defining strength is that it can run any code, including code full of unpredictable branches, pointer chasing, and irregular memory access.
How a CPU is built
- A few powerful cores. Phones typically have 4 to 10 cores, laptops 4 to 24. Many designs mix fast “performance” cores with frugal “efficiency” cores.
- Deep control logic. Out-of-order execution, speculative execution, and branch prediction keep the pipeline busy even when code is unpredictable.
- Large caches. Multi-level caches hide memory latency, so a single thread runs fast.
- Low latency focus. The CPU is optimized to finish one task quickly, not to finish a million tasks at once.
- Vector extensions. SIMD units such as Arm NEON or x86 AVX let the CPU process several values per instruction, which helps with small AI models.
A large share of a CPU core’s silicon is spent on control, not arithmetic. That is the price of flexibility, and it is exactly why a CPU is inefficient for repetitive math on huge arrays.
What CPUs do best
- Operating system, drivers, and scheduling
- Application logic with many branches and decisions
- Pre- and post-processing around an AI model (parsing, decoding, sorting results)
- Small or unusual neural-network layers that accelerators do not support
- Anything that has to start instantly with no setup cost
3. GPU: the parallel specialist
The Graphics Processing Unit began as a chip for drawing triangles and shading pixels. Rendering an image means applying the same small program to millions of pixels, so GPUs evolved into machines with a huge number of simple processing units running in lockstep. This design is called SIMT (single instruction, multiple threads).
How a GPU is built
- Hundreds to thousands of lanes. Mobile GPUs have hundreds; desktop GPUs have many thousands.
- Simple cores, little control logic. Most area goes to arithmetic, not prediction or reordering.
- Latency hiding through threads. When one group of threads waits on memory, another runs. The GPU favors throughput over single-thread speed.
- High memory bandwidth. Wide memory interfaces feed all those lanes.
- Programmable. Through shader languages and compute APIs, GPUs run general parallel code, not just graphics.
Why GPUs became AI hardware
Neural-network layers boil down to huge numbers of independent multiply-add operations, the same pattern as shading pixels. Developers found that GPUs trained and ran these networks orders of magnitude faster than CPUs. Modern GPUs also add dedicated matrix units (often called tensor cores) that blur the line with NPUs.
GPU limitations
- Branch-heavy code runs poorly because lanes in a group must follow the same path.
- Power draw is high, which matters on battery or thermally limited devices.
- Launching small jobs has overhead, so tiny workloads may not benefit.
- On a phone, the GPU is also needed to draw the screen, so AI work competes with the user interface.
4. NPU: the AI specialist
A Neural Processing Unit is designed from the start for neural-network inference. It does not try to be general. It implements the operations that dominate neural networks, mainly matrix multiplication and convolution, directly in hardware, and arranges data movement to match.
How an NPU is built
- MAC arrays. Grids of multiply-accumulate units (sometimes a systolic array) perform thousands of multiply-adds per clock cycle.
- Local on-chip memory. Weights and activations sit in scratchpad SRAM close to the compute, avoiding costly trips to external memory.
- Low-precision arithmetic. INT8, INT4, and FP16 are native formats; they need less energy and bandwidth than FP32.
- Fixed dataflow. A compiler plans data movement ahead of time instead of the hardware guessing at runtime.
- Fused operations. Activation functions, scaling, and pooling are often applied right after the matrix operation without writing intermediate results to memory.
Common NPU names
Vendors use different branding for the same idea: neural engine, AI engine, tensor processor, deep learning accelerator, or DSP with AI extensions. Performance is usually quoted in TOPS (trillions of operations per second), usually measured at INT8 precision.
NPU limitations
- Only a fixed set of operators is accelerated. Unsupported layers fall back to the CPU or GPU.
- Models usually must be quantized and compiled with the vendor toolchain.
- Toolchains and drivers differ between vendors, so portability is weaker.
- Poor fit for non-neural workloads.
5. Side-by-side comparison
| Property | CPU | GPU | NPU |
|---|---|---|---|
| Main idea | Flexible control and logic | Massive data parallelism | Fixed-function AI acceleration |
| Core count | Few (4 to 24 typical) | Hundreds to thousands of lanes | Thousands of MAC units in arrays |
| Core complexity | High | Low to medium | Very low, specialized |
| Optimized for | Low latency per task | High throughput | Throughput per watt |
| Flexibility | Highest | Medium to high | Lowest |
| Typical precision | FP32, FP64, INT | FP32, FP16, some INT8 | INT8, INT4, FP16 |
| Branchy code | Excellent | Poor | Not supported |
| Matrix and convolution | Slow, power-hungry | Fast | Fastest per watt |
| Power on AI tasks | High per result | Medium to high | Low |
| Setup effort | Easy: any compiler | Moderate: compute APIs | Higher: quantize and compile |
| Runs the OS | Yes | No | No |
| Best role | Orchestration and logic | Graphics and parallel compute | Sustained on-device AI |
Rule-of-thumb performance profile
| Workload | CPU | GPU | NPU |
|---|---|---|---|
| Running an OS and apps | ★★★★★ | ☆ | ☆ |
| 3D graphics and UI drawing | ★ | ★★★★★ | ☆ |
| Image processing filters | ★★ | ★★★★★ | ★★★ |
| CNN vision inference | ★★ | ★★★★ | ★★★★★ |
| Speech recognition | ★★ | ★★★ | ★★★★★ |
| Large language model decode | ★★ | ★★★★ | ★★★★ (memory-bound) |
| Model training | ★ | ★★★★★ | ★ (rarely supported) |
| Irregular or custom logic | ★★★★★ | ★★ | ☆ |
Ratings are qualitative and simplified. Real results depend on the exact chip, model, and software stack.
6. Why neural networks suit NPUs
Strip away the jargon and a neural-network layer is mostly this: take a vector or matrix of inputs, multiply by a matrix of weights, add a bias, and apply a simple function. A convolution is the same idea applied to sliding windows of an image.
The core operation: multiply-accumulate
Every output value is a sum of products: y = w1·x1 + w2·x2 + … + wn·xn + b. Each product-and-add is one MAC. A modest image model needs billions of MACs per inference, and a language model needs many more per generated word.
The key properties of this workload are:
- Predictable. There are no data-dependent branches. The whole sequence of operations is known before running.
- Highly parallel. Thousands of MACs are independent and can run at once.
- Reuse-heavy. The same weights or inputs are used many times, so keeping them in local memory pays off.
- Tolerant of low precision. Networks usually work well at 8 bits or fewer.
A CPU wastes its branch predictors and out-of-order engines on this predictable work. A GPU handles it well but carries graphics-oriented hardware and runs hot. An NPU removes everything that is not needed and keeps just the MAC arrays, local memory, and a compiler-planned dataflow.
7. How they work together on one chip
In a phone or laptop SoC, the three processors sit on the same die and share system memory. The CPU is the conductor. Software on the CPU decides which processor should run each piece of work, prepares the data, and collects the results.
A typical hand-off, step by step
Consider a camera app that blurs a background in real time:
- The camera sensor and image signal processor (ISP) deliver a frame into memory.
- The CPU wakes the app logic and schedules the frame for processing.
- The NPU runs a segmentation network to find the person versus the background.
- The GPU applies the blur and composites the final image using shaders.
- The CPU handles encoding settings, UI events, and saving or streaming the result.
8. Memory: the hidden bottleneck
Doing the arithmetic is rarely the hard part. Moving data is. Fetching a value from external DRAM can cost far more energy than performing a multiply on it. Much of processor design is about avoiding those fetches.
| Memory level | Speed | Energy cost | Size | Used by |
|---|---|---|---|---|
| Registers | Fastest | Lowest | Bytes | All |
| On-chip SRAM / cache | Very fast | Low | KB to MB | CPU caches, GPU shared memory, NPU scratchpad |
| System DRAM | Slower | High (often 100× a compute op) | GB | All, shared |
| Flash storage | Slowest | Highest | GB to TB | Model files |
Compute-bound vs memory-bound
A workload is compute-bound when the processor runs out of arithmetic capacity first, and memory-bound when it waits for data. Convolutional vision models with lots of weight reuse tend to be compute-bound, so NPUs shine. Large language models generating one word at a time must read nearly all their weights for each step, which makes them memory-bound. For these, memory bandwidth and model size matter more than raw TOPS.
NPUs reduce traffic by keeping weights in local SRAM, tiling large layers into pieces that fit, fusing operations so intermediate results never leave the chip, and compressing weights.
9. Precision and quantization
Training usually uses 32-bit or 16-bit floating-point numbers. Inference on edge devices can often use much smaller formats with little loss in accuracy. Smaller numbers mean less memory, less bandwidth, and less energy per operation.
| Format | Bits | Size vs FP32 | Typical use | Best on |
|---|---|---|---|---|
| FP32 | 32 | 1× | Training, reference | CPU, GPU |
| FP16 / BF16 | 16 | 0.5× | Training and inference | GPU, some NPUs |
| INT8 | 8 | 0.25× | Standard edge inference | NPU, CPU (SIMD) |
| INT4 | 4 | 0.125× | Compressed LLM weights | Newer NPUs, GPU |
Practical quantization tips
- Start with post-training quantization using a small representative dataset.
- If accuracy drops too far, use quantization-aware training.
- Keep sensitive layers (often the first and last) at higher precision.
- Always compare accuracy on real data, not only on the benchmark set.
10. Real workloads: who runs what
| Feature | Primary engine | Why |
|---|---|---|
| Face unlock | NPU | Small CNN, always-on, low power |
| Voice wake word | Tiny NPU or DSP | Milliwatt budget, runs continuously |
| Live speech-to-text | NPU | Sustained sequence model |
| Photo night mode | ISP + NPU + GPU | Multi-frame fusion plus learned denoising |
| Video background blur | NPU + GPU | Segmentation, then pixel compositing |
| On-device text generation | NPU (or GPU) | Quantized weights, memory-bound |
| Game rendering | GPU | Shaders and rasterization |
| Web browsing, app logic | CPU | Branchy, irregular code |
| Sensor fusion filters | CPU or DSP | Small, latency-critical |
| Custom or unsupported layers | CPU fallback | Accelerators lack the operator |
11. Guidance for embedded AI engineers
Understanding how these processors divide work is becoming a core skill. These practices help you make good decisions.
Choosing the target
- Profile first. Measure the model on each engine for latency, power, and memory instead of guessing from TOPS.
- Prefer the NPU for sustained or always-on AI when your model’s operators are supported.
- Use the GPU when the model has unsupported operators, when you need FP16 flexibility, or when the data is already on the GPU (for example, camera frames being rendered).
- Use the CPU for tiny models, odd operators, and glue logic, and as the guaranteed fallback.
Designing models for accelerators
- Choose operators the target NPU supports; check the vendor’s operator list early.
- Use static input shapes where possible; dynamic shapes often force CPU fallback.
- Favor efficient architectures such as depthwise-separable convolutions.
- Reduce layer width and depth until accuracy is just good enough.
- Avoid frequent switching between engines; each hand-off costs time and memory copies.
System-level concerns
| Concern | What to watch |
|---|---|
| Latency | Include data copy, preprocessing, and engine hand-off, not just the inference time |
| Power and heat | Sustained workloads trigger thermal throttling; NPUs delay this |
| Memory | Model size plus activation buffers must fit within the device’s budget |
| Fallbacks | Count how many layers run off-NPU; each split adds overhead |
| Toolchain | Compilers, quantizers, and drivers differ by vendor; budget time for debugging |
| Accuracy | Re-validate after quantization and compilation, on the real device |
| Concurrency | The GPU may be busy with the UI; the CPU may be busy with the OS |
A simple decision flow
12. Common myths
| Myth | Reality |
|---|---|
| “An NPU replaces the GPU.” | No. They overlap on AI but the GPU still handles graphics and general parallel compute. |
| “More TOPS always means faster.” | Only for ideal workloads. Memory bandwidth and operator support often decide real speed. |
| “The CPU is obsolete for AI.” | The CPU runs the surrounding logic, handles fallbacks, and is best for tiny models. |
| “Any model runs on an NPU unchanged.” | Most need quantization and compilation, and unsupported operators fall back. |
| “GPUs are always faster than CPUs.” | Not for small or branchy jobs, where launch overhead and divergence hurt. |
| “NPUs are only for phones.” | They appear in laptops, cameras, cars, industrial gateways, and microcontroller-class devices. |
13. Summary
The three processors are complementary rather than competing. Each trades flexibility for efficiency in a different proportion:
- CPU: the flexible manager. It runs the operating system, application logic, control flow, and anything irregular.
- GPU: the parallel specialist. It applies the same operation to enormous amounts of data, ideal for graphics, image processing, simulation, and many AI workloads.
- NPU: the AI specialist. It implements matrix multiplication and convolution in hardware, delivering AI inference at much lower power.
Modern devices use all three together: the CPU manages the system, the GPU handles graphics and parallel workloads, and the NPU accelerates image enhancement, speech recognition, computer vision, and generative AI features. For embedded AI engineers, the skill is matching each piece of the pipeline to the right engine, minimizing hand-offs and memory traffic, and verifying results with real measurements on the target hardware.
