CPU vs GPU vs NPU: Who Handles What?

CPU vs GPU vs NPU_ Who Handles What
CPU vs GPU vs NPU: What’s Really Different?
Embedded AI · Hardware Primer

CPU vs GPU vs NPU: What’s Really Different?

Your smartphone or laptop may have a CPU, a GPU, and an NPU, but they are designed for very different kinds of work. This guide explains, with diagrams and tables, why each exists and how embedded AI engineers should divide work between them.

CPUFlexibility. The general-purpose manager that can do many different jobs.
GPUMassive parallel processing. Thousands of simple units doing the same operation at once.
NPUEfficient AI acceleration. Built around matrix multiplication and convolution.
In this article
  1. The big picture
  2. CPU: the flexible manager
  3. GPU: the parallel specialist
  4. NPU: the AI specialist
  5. Side-by-side comparison
  6. Why neural networks suit NPUs
  7. How they work together on one chip
  8. Memory, the hidden bottleneck
  9. Precision and quantization
  10. Real workloads: who runs what
  11. Guidance for embedded AI engineers
  12. Common myths
  13. Summary

1. The big picture

Every processor is a trade-off between three things: how flexible it is, how fast it is on a given job, and how much energy it burns. No design wins on all three. A chip that can run any program must spend silicon on branch prediction, caches, and scheduling. A chip that only multiplies matrices can dedicate nearly all of its area to arithmetic and run far more efficiently, but it cannot run your operating system.

That is why modern system-on-chip (SoC) designs in phones, laptops, cameras, cars, and IoT devices combine several processor types. The easy mnemonic is:

CPU = flexibility. GPU = massive parallel processing. NPU = efficient AI acceleration.
CPU GPU NPU Runs anything Lower efficiency Same op, many data High throughput Neural nets only Best performance per watt General flexibility → Specialization More specialized = less flexible, but far more efficient on its target job Efficiency on AI workloads increases left to right
Figure 1. Specialization buys efficiency at the cost of flexibility.

2. CPU: the flexible manager

The Central Processing Unit is the processor every program ultimately depends on. It runs the operating system, schedules tasks, handles input and output, executes application logic, and performs calculations of every kind. Its defining strength is that it can run any code, including code full of unpredictable branches, pointer chasing, and irregular memory access.

How a CPU is built

  • A few powerful cores. Phones typically have 4 to 10 cores, laptops 4 to 24. Many designs mix fast “performance” cores with frugal “efficiency” cores.
  • Deep control logic. Out-of-order execution, speculative execution, and branch prediction keep the pipeline busy even when code is unpredictable.
  • Large caches. Multi-level caches hide memory latency, so a single thread runs fast.
  • Low latency focus. The CPU is optimized to finish one task quickly, not to finish a million tasks at once.
  • Vector extensions. SIMD units such as Arm NEON or x86 AVX let the CPU process several values per instruction, which helps with small AI models.

A large share of a CPU core’s silicon is spent on control, not arithmetic. That is the price of flexibility, and it is exactly why a CPU is inefficient for repetitive math on huge arrays.

CPU: few, complex cores ControlControlControlControl CacheCacheCacheCache ALU Shared last-level cache
Figure 2. Most CPU area is control and cache; arithmetic units (ALU) are a small slice.

What CPUs do best

  • Operating system, drivers, and scheduling
  • Application logic with many branches and decisions
  • Pre- and post-processing around an AI model (parsing, decoding, sorting results)
  • Small or unusual neural-network layers that accelerators do not support
  • Anything that has to start instantly with no setup cost

3. GPU: the parallel specialist

The Graphics Processing Unit began as a chip for drawing triangles and shading pixels. Rendering an image means applying the same small program to millions of pixels, so GPUs evolved into machines with a huge number of simple processing units running in lockstep. This design is called SIMT (single instruction, multiple threads).

How a GPU is built

  • Hundreds to thousands of lanes. Mobile GPUs have hundreds; desktop GPUs have many thousands.
  • Simple cores, little control logic. Most area goes to arithmetic, not prediction or reordering.
  • Latency hiding through threads. When one group of threads waits on memory, another runs. The GPU favors throughput over single-thread speed.
  • High memory bandwidth. Wide memory interfaces feed all those lanes.
  • Programmable. Through shader languages and compute APIs, GPUs run general parallel code, not just graphics.
GPU: many simple cores in lockstep Shared control + scheduler (one instruction drives many lanes) High-bandwidth memory interface
Figure 3. A GPU spends its area on arithmetic lanes driven by shared control.

Why GPUs became AI hardware

Neural-network layers boil down to huge numbers of independent multiply-add operations, the same pattern as shading pixels. Developers found that GPUs trained and ran these networks orders of magnitude faster than CPUs. Modern GPUs also add dedicated matrix units (often called tensor cores) that blur the line with NPUs.

GPU limitations

  • Branch-heavy code runs poorly because lanes in a group must follow the same path.
  • Power draw is high, which matters on battery or thermally limited devices.
  • Launching small jobs has overhead, so tiny workloads may not benefit.
  • On a phone, the GPU is also needed to draw the screen, so AI work competes with the user interface.

4. NPU: the AI specialist

A Neural Processing Unit is designed from the start for neural-network inference. It does not try to be general. It implements the operations that dominate neural networks, mainly matrix multiplication and convolution, directly in hardware, and arranges data movement to match.

How an NPU is built

  • MAC arrays. Grids of multiply-accumulate units (sometimes a systolic array) perform thousands of multiply-adds per clock cycle.
  • Local on-chip memory. Weights and activations sit in scratchpad SRAM close to the compute, avoiding costly trips to external memory.
  • Low-precision arithmetic. INT8, INT4, and FP16 are native formats; they need less energy and bandwidth than FP32.
  • Fixed dataflow. A compiler plans data movement ahead of time instead of the hardware guessing at runtime.
  • Fused operations. Activation functions, scaling, and pooling are often applied right after the matrix operation without writing intermediate results to memory.
NPU: MAC array fed by local memory WeightSRAM MAC array (multiply + accumulate) Activation+ OutputSRAM DMA engine + shared system memory (slow, power-hungry; used sparingly)
Figure 4. Data stays in local SRAM while the MAC array does the math.

Common NPU names

Vendors use different branding for the same idea: neural engine, AI engine, tensor processor, deep learning accelerator, or DSP with AI extensions. Performance is usually quoted in TOPS (trillions of operations per second), usually measured at INT8 precision.

Careful with TOPS. Peak TOPS assumes ideal conditions. Real speed depends on memory bandwidth, supported operators, and how well your model maps to the hardware.

NPU limitations

  • Only a fixed set of operators is accelerated. Unsupported layers fall back to the CPU or GPU.
  • Models usually must be quantized and compiled with the vendor toolchain.
  • Toolchains and drivers differ between vendors, so portability is weaker.
  • Poor fit for non-neural workloads.

5. Side-by-side comparison

PropertyCPUGPUNPU
Main ideaFlexible control and logicMassive data parallelismFixed-function AI acceleration
Core countFew (4 to 24 typical)Hundreds to thousands of lanesThousands of MAC units in arrays
Core complexityHighLow to mediumVery low, specialized
Optimized forLow latency per taskHigh throughputThroughput per watt
FlexibilityHighestMedium to highLowest
Typical precisionFP32, FP64, INTFP32, FP16, some INT8INT8, INT4, FP16
Branchy codeExcellentPoorNot supported
Matrix and convolutionSlow, power-hungryFastFastest per watt
Power on AI tasksHigh per resultMedium to highLow
Setup effortEasy: any compilerModerate: compute APIsHigher: quantize and compile
Runs the OSYesNoNo
Best roleOrchestration and logicGraphics and parallel computeSustained on-device AI

Rule-of-thumb performance profile

WorkloadCPUGPUNPU
Running an OS and apps★★★★★☆☆
3D graphics and UI drawing★★★★★★☆
Image processing filters★★★★★★★★★★
CNN vision inference★★★★★★★★★★★
Speech recognition★★★★★★★★★★
Large language model decode★★★★★★★★★★ (memory-bound)
Model training★★★★★★★ (rarely supported)
Irregular or custom logic★★★★★★★☆

Ratings are qualitative and simplified. Real results depend on the exact chip, model, and software stack.

6. Why neural networks suit NPUs

Strip away the jargon and a neural-network layer is mostly this: take a vector or matrix of inputs, multiply by a matrix of weights, add a bias, and apply a simple function. A convolution is the same idea applied to sliding windows of an image.

The core operation: multiply-accumulate

Every output value is a sum of products: y = w1·x1 + w2·x2 + … + wn·xn + b. Each product-and-add is one MAC. A modest image model needs billions of MACs per inference, and a language model needs many more per generated word.

The key properties of this workload are:

  • Predictable. There are no data-dependent branches. The whole sequence of operations is known before running.
  • Highly parallel. Thousands of MACs are independent and can run at once.
  • Reuse-heavy. The same weights or inputs are used many times, so keeping them in local memory pays off.
  • Tolerant of low precision. Networks usually work well at 8 bits or fewer.

A CPU wastes its branch predictors and out-of-order engines on this predictable work. A GPU handles it well but carries graphics-oriented hardware and runs hot. An NPU removes everything that is not needed and keeps just the MAC arrays, local memory, and a compiler-planned dataflow.

Illustrative energy per AI operation (lower is better) CPUhighest GPUmedium NPUlowest Relative, not measured data. Actual ratios vary by chip and model.
Figure 5. The same neural-network operation costs the least energy on an NPU.

7. How they work together on one chip

In a phone or laptop SoC, the three processors sit on the same die and share system memory. The CPU is the conductor. Software on the CPU decides which processor should run each piece of work, prepares the data, and collects the results.

System-on-Chip (SoC) CPUOS · logic · control GPUGraphics · parallel NPUAI inference System interconnect (shared bus) System memoryLPDDR / DDR ISP · Video · DSPCamera, codec Sensors · I/OMic, camera, radio
Figure 6. All three engines share an interconnect and system memory.

A typical hand-off, step by step

Consider a camera app that blurs a background in real time:

  1. The camera sensor and image signal processor (ISP) deliver a frame into memory.
  2. The CPU wakes the app logic and schedules the frame for processing.
  3. The NPU runs a segmentation network to find the person versus the background.
  4. The GPU applies the blur and composites the final image using shaders.
  5. The CPU handles encoding settings, UI events, and saving or streaming the result.
Camera+ ISP CPUschedule NPUsegment GPUblur, blend CPUencode, UI A single feature uses all three engines
Figure 7. Workload flow for a real-time background-blur feature.

8. Memory: the hidden bottleneck

Doing the arithmetic is rarely the hard part. Moving data is. Fetching a value from external DRAM can cost far more energy than performing a multiply on it. Much of processor design is about avoiding those fetches.

Memory levelSpeedEnergy costSizeUsed by
RegistersFastestLowestBytesAll
On-chip SRAM / cacheVery fastLowKB to MBCPU caches, GPU shared memory, NPU scratchpad
System DRAMSlowerHigh (often 100× a compute op)GBAll, shared
Flash storageSlowestHighestGB to TBModel files

Compute-bound vs memory-bound

A workload is compute-bound when the processor runs out of arithmetic capacity first, and memory-bound when it waits for data. Convolutional vision models with lots of weight reuse tend to be compute-bound, so NPUs shine. Large language models generating one word at a time must read nearly all their weights for each step, which makes them memory-bound. For these, memory bandwidth and model size matter more than raw TOPS.

NPUs reduce traffic by keeping weights in local SRAM, tiling large layers into pieces that fit, fusing operations so intermediate results never leave the chip, and compressing weights.

9. Precision and quantization

Training usually uses 32-bit or 16-bit floating-point numbers. Inference on edge devices can often use much smaller formats with little loss in accuracy. Smaller numbers mean less memory, less bandwidth, and less energy per operation.

FormatBitsSize vs FP32Typical useBest on
FP32321×Training, referenceCPU, GPU
FP16 / BF16160.5×Training and inferenceGPU, some NPUs
INT880.25×Standard edge inferenceNPU, CPU (SIMD)
INT440.125×Compressed LLM weightsNewer NPUs, GPU

Practical quantization tips

  • Start with post-training quantization using a small representative dataset.
  • If accuracy drops too far, use quantization-aware training.
  • Keep sensitive layers (often the first and last) at higher precision.
  • Always compare accuracy on real data, not only on the benchmark set.

10. Real workloads: who runs what

FeaturePrimary engineWhy
Face unlockNPUSmall CNN, always-on, low power
Voice wake wordTiny NPU or DSPMilliwatt budget, runs continuously
Live speech-to-textNPUSustained sequence model
Photo night modeISP + NPU + GPUMulti-frame fusion plus learned denoising
Video background blurNPU + GPUSegmentation, then pixel compositing
On-device text generationNPU (or GPU)Quantized weights, memory-bound
Game renderingGPUShaders and rasterization
Web browsing, app logicCPUBranchy, irregular code
Sensor fusion filtersCPU or DSPSmall, latency-critical
Custom or unsupported layersCPU fallbackAccelerators lack the operator

11. Guidance for embedded AI engineers

Understanding how these processors divide work is becoming a core skill. These practices help you make good decisions.

Choosing the target

  1. Profile first. Measure the model on each engine for latency, power, and memory instead of guessing from TOPS.
  2. Prefer the NPU for sustained or always-on AI when your model’s operators are supported.
  3. Use the GPU when the model has unsupported operators, when you need FP16 flexibility, or when the data is already on the GPU (for example, camera frames being rendered).
  4. Use the CPU for tiny models, odd operators, and glue logic, and as the guaranteed fallback.

Designing models for accelerators

  • Choose operators the target NPU supports; check the vendor’s operator list early.
  • Use static input shapes where possible; dynamic shapes often force CPU fallback.
  • Favor efficient architectures such as depthwise-separable convolutions.
  • Reduce layer width and depth until accuracy is just good enough.
  • Avoid frequent switching between engines; each hand-off costs time and memory copies.

System-level concerns

ConcernWhat to watch
LatencyInclude data copy, preprocessing, and engine hand-off, not just the inference time
Power and heatSustained workloads trigger thermal throttling; NPUs delay this
MemoryModel size plus activation buffers must fit within the device’s budget
FallbacksCount how many layers run off-NPU; each split adds overhead
ToolchainCompilers, quantizers, and drivers differ by vendor; budget time for debugging
AccuracyRe-validate after quantization and compilation, on the real device
ConcurrencyThe GPU may be busy with the UI; the CPU may be busy with the OS

A simple decision flow

Is it a neural network? Run on CPU Are all ops NPU-supported? Run on NPU (INT8) GPU available? Run on GPU (FP16) CPU fallback NoYesYesNoYesNo
Figure 8. A starting-point decision flow; always confirm with measurements.

12. Common myths

MythReality
“An NPU replaces the GPU.”No. They overlap on AI but the GPU still handles graphics and general parallel compute.
“More TOPS always means faster.”Only for ideal workloads. Memory bandwidth and operator support often decide real speed.
“The CPU is obsolete for AI.”The CPU runs the surrounding logic, handles fallbacks, and is best for tiny models.
“Any model runs on an NPU unchanged.”Most need quantization and compilation, and unsupported operators fall back.
“GPUs are always faster than CPUs.”Not for small or branchy jobs, where launch overhead and divergence hurt.
“NPUs are only for phones.”They appear in laptops, cameras, cars, industrial gateways, and microcontroller-class devices.

13. Summary

The three processors are complementary rather than competing. Each trades flexibility for efficiency in a different proportion:

  • CPU: the flexible manager. It runs the operating system, application logic, control flow, and anything irregular.
  • GPU: the parallel specialist. It applies the same operation to enormous amounts of data, ideal for graphics, image processing, simulation, and many AI workloads.
  • NPU: the AI specialist. It implements matrix multiplication and convolution in hardware, delivering AI inference at much lower power.

Modern devices use all three together: the CPU manages the system, the GPU handles graphics and parallel workloads, and the NPU accelerates image enhancement, speech recognition, computer vision, and generative AI features. For embedded AI engineers, the skill is matching each piece of the pipeline to the right engine, minimizing hand-offs and memory traffic, and verifying results with real measurements on the target hardware.

Remember: CPU = flexibility · GPU = massive parallel processing · NPU = efficient AI acceleration.
This article uses simplified, illustrative figures to explain concepts. Real hardware varies by vendor and generation.