What is TOPS and Why Does Everyone Talk About It?
Imagine two AI chips: one says 10 TOPS, another says 100 TOPS. Does the second chip automatically deliver 10× better AI? Not necessarily. Let’s see why.
1. TOPS in plain words
“Tera” means one trillion (1,000,000,000,000). So 10 TOPS means up to ten trillion operations every second. An “operation” here is usually a tiny arithmetic step: a multiplication or an addition on small numbers.
Modern AI models such as image recognisers, voice assistants and chatbots are made of enormous grids of numbers (matrices). Running them means multiplying and adding those numbers again and again. A chip built for AI has many small calculators working in parallel, and TOPS tells you how many operations they could do if every one of them stayed busy every moment.
Think of a restaurant kitchen. TOPS is like saying “we have 100 chefs, each able to chop one vegetable per second.” That is the peak. Whether dinner is served fast depends on ingredients arriving on time, the recipe, the kitchen layout and whether all 100 chefs actually have work. TOPS only counts the chefs.
The unit ladder
| Unit | Meaning | Operations per second |
|---|---|---|
| MOPS | Mega | 1 million (106) |
| GOPS | Giga | 1 billion (109) |
| TOPS | Tera | 1 trillion (1012) |
| POPS | Peta | 1 quadrillion (1015) |
2. Where does the number come from?
Vendors usually calculate TOPS from the hardware design, not by running a real AI model. The core building block is the MAC unit (multiply-accumulate): it multiplies two numbers and adds the result to a running total. Because that is two operations, each MAC counts as 2 ops.
Notice what is missing: no memory, no software, no real model. The figure assumes every MAC unit is busy every single cycle, which almost never happens.
3. Precision changes the number
An operation on an 8-bit integer is far cheaper than one on a 16-bit or 32-bit floating-point number. Chips can pack more low-precision calculators into the same area and power, so the same chip may quote very different TOPS depending on the number format used.
| Format | Bits | Typical use | Relative TOPS |
|---|---|---|---|
| FP32 | 32 | Training, scientific work | 1× (baseline) |
| FP16 / BF16 | 16 | Training, higher-quality inference | ~2×–4× |
| INT8 / FP8 | 8 | Most edge and phone inference | ~4×–8× |
| INT4 / FP4 | 4 | Compressed language models | ~8×–16× |
Lower precision also loses accuracy. Many models tolerate INT8 well, some need FP16, and aggressive INT4 works best with special compression techniques. So a big low-precision TOPS number only helps if your model runs well at that precision.
4. Why TOPS alone can mislead
A chip is only as fast as its slowest link. Five things routinely stop real performance from reaching the headline number.
Memory bandwidth
Calculators starve if data cannot arrive fast enough from memory. Large language models are usually limited here, not by compute.
Architecture
Some designs keep data close to the MACs and reuse it; others shuffle it around wastefully.
Software stack
Drivers, compilers and model converters decide how much of the hardware a model can actually use.
Power and heat
A laptop or phone throttles when hot. Peak TOPS may last only seconds.
The model itself
Some layers map perfectly to the chip; others fall back to a slower CPU.
Utilization
Real workloads often use 20%–60% of peak, sometimes far less.
The memory wall, visually
Engineers describe this with the roofline model: performance is capped either by compute (TOPS) or by memory (bandwidth × how many operations you do per byte fetched). Models that do lots of math per byte, like image convolutions, can approach peak TOPS. Chatbots generating one word at a time do very little math per byte, so they hit the memory ceiling first.
| Workload | Usually limited by | Does more TOPS help? |
|---|---|---|
| Camera image enhancement | Compute | Yes, strongly |
| Object detection (video) | Compute + memory | Yes, up to a point |
| Speech recognition | Balanced | Moderately |
| Chatbot text generation | Memory bandwidth | Barely; bandwidth matters more |
| Chatbot prompt reading | Compute | Yes |
5. Same TOPS, different results
Here is an illustrative example. Two chips both claim 40 TOPS at INT8.
| Property | Chip A | Chip B |
|---|---|---|
| Peak INT8 TOPS | 40 | 40 |
| Memory bandwidth | 120 GB/s | 40 GB/s |
| Mature software tools | Yes | Limited |
| Power at full load | 15 W | 35 W |
| Typical utilization on a vision model | ~65% | ~30% |
| Effective throughput | ~26 TOPS | ~12 TOPS |
Identical on the label, more than 2× apart in practice, and one draws far less power. That is why buyers should never treat TOPS as a final score.
6. Why automotive chips list hundreds of TOPS
A self-driving or driver-assistance computer must process many cameras, radar and lidar streams simultaneously, run several large neural networks, and do it with safety margins and redundancy. It has a big power budget (tens to hundreds of watts) and active cooling, so it can host huge arrays of MAC units. A phone, in contrast, has just a few watts to spare and no fan.
| Device class | Power budget | Typical TOPS claim | What it runs |
|---|---|---|---|
| Smartwatch / sensor | milliwatts | Under 1 | Wake word, activity detection |
| Smartphone | 2–8 W | Tens | Photos, translation, assistants |
| AI laptop | 5–15 W (NPU) | ~40–50+ | Background blur, transcription, local assistants |
| Edge AI box | 10–60 W | Tens to ~275 | Cameras, robots, factory vision |
| Automotive | 50–500 W | Hundreds to 1000+ | Perception, planning, driver monitoring |
| Data-center GPU | 300–1000+ W | Thousands | Training and large-scale inference |
7. TOPS per watt: efficiency matters
For battery and edge devices, TOPS/W is often more useful. It tells you how much AI work you get for each watt. A 10-TOPS chip at 2 W (5 TOPS/W) may beat a 30-TOPS chip at 20 W (1.5 TOPS/W) inside a phone, because the phone simply cannot supply or cool 20 W.
| Chip | TOPS | Watts | TOPS/W | Fits a phone? |
|---|---|---|---|---|
| X | 10 | 2 | 5.0 | Yes |
| Y | 30 | 20 | 1.5 | No, too hot |
| Z | 200 | 40 | 5.0 | No, but great for a robot |
Also check whether the quoted power is for the whole chip or just the AI block, and whether it is a sustained or brief peak figure.
8. Tricks that inflate the headline
- Sparsity: Some chips skip multiplications by zero. If a model is specially pruned, effective throughput can double, and some vendors quote that “sparse” number. Dense models see no gain.
- Combined engines: “Platform TOPS” may add CPU + GPU + NPU together, even if a real app can use only one at a time.
- Lowest precision: As shown above, INT4 figures look far bigger than INT8.
- Boost clocks: The number may assume a maximum clock that thermals do not allow for long.
- Different definitions: Whether a MAC counts as one or two operations is a convention; check it.
___ TOPS, dense, INT8, NPU only, sustained. If any blank is missing, treat the figure as a marketing headline.9. Better ways to judge AI performance
| Metric | Measures | Why it helps |
|---|---|---|
| Inferences per second | Real model throughput | Actual result on a named model |
| Latency (ms) | Time for one answer | Matters for cameras and voice |
| Tokens per second | Chatbot speed | Reflects memory limits |
| Time to first token | Delay before reply starts | Shows compute for prompt reading |
| Memory bandwidth (GB/s) | Data feed rate | Predicts language-model speed |
| Memory capacity (GB) | Model size that fits | Decides which models can run at all |
| TOPS/W, battery life | Efficiency | Vital for mobile and edge |
| Benchmark suites (e.g. MLPerf) | Standardised tests | Comparable across vendors |
A simple rule for chatbots
For a local language model, each generated word requires reading roughly the entire model from memory. So the speed ceiling is approximately:
Example: a 4 GB model on 60 GB/s memory gives about 15 tokens per second at best, whether the chip has 40 TOPS or 400 TOPS. This is why raising TOPS alone often fails to speed up chat.
10. Ask three questions
- At what precision? INT8, FP16 or INT4 changes the number by multiples.
- For which workload? Ask for results on the model you care about, not only peak.
- At what power? Sustained watts and thermal limits decide real-world use.
Buyer’s checklist
| Check | Good sign | Red flag |
|---|---|---|
| Precision stated | “INT8, dense” | Just “TOPS” |
| Scope | NPU only, clearly labelled | Combined CPU+GPU+NPU total |
| Software support | Published tools, model zoo | No documentation |
| Real benchmarks | Named models with latency | Only peak numbers |
| Memory | Bandwidth and capacity listed | Not mentioned |
| Efficiency | TOPS/W or battery data | No power figure |
11. Common myths
| Myth | Reality |
|---|---|
| 100 TOPS is 10× faster than 10 TOPS | Only if everything else is equal, which it rarely is. |
| More TOPS always means better AI PCs | Software support and memory often matter more day to day. |
| TOPS measures intelligence | It measures arithmetic capacity, not model quality. |
| TOPS is meaningless | Not true; it is a useful rough guide to compute capability. |
| GPUs and NPUs quote TOPS the same way | Formats, sparsity and scope frequently differ. |
12. Summary
So when someone says, “This chip has 100 TOPS,” smile and ask: 100 TOPS at what precision, for which workload, and at what power?
