What Does TOPS Mean In AI Chips?

What is TOPS and Why Does Everyone Talk About It
What is TOPS and Why Does Everyone Talk About It?
AI Hardware Explained

What is TOPS and Why Does Everyone Talk About It?

Imagine two AI chips: one says 10 TOPS, another says 100 TOPS. Does the second chip automatically deliver 10× better AI? Not necessarily. Let’s see why.

Quick answerTOPS stands for Tera Operations Per Second. It counts how many trillion simple math operations (mostly multiply and add) an AI processor can perform each second, at best. It is a theoretical peak, not a promise of real-world speed.

1. TOPS in plain words

“Tera” means one trillion (1,000,000,000,000). So 10 TOPS means up to ten trillion operations every second. An “operation” here is usually a tiny arithmetic step: a multiplication or an addition on small numbers.

Modern AI models such as image recognisers, voice assistants and chatbots are made of enormous grids of numbers (matrices). Running them means multiplying and adding those numbers again and again. A chip built for AI has many small calculators working in parallel, and TOPS tells you how many operations they could do if every one of them stayed busy every moment.

Think of a restaurant kitchen. TOPS is like saying “we have 100 chefs, each able to chop one vegetable per second.” That is the peak. Whether dinner is served fast depends on ingredients arriving on time, the recipe, the kitchen layout and whether all 100 chefs actually have work. TOPS only counts the chefs.

The unit ladder

UnitMeaningOperations per second
MOPSMega1 million (106)
GOPSGiga1 billion (109)
TOPSTera1 trillion (1012)
POPSPeta1 quadrillion (1015)

2. Where does the number come from?

Vendors usually calculate TOPS from the hardware design, not by running a real AI model. The core building block is the MAC unit (multiply-accumulate): it multiplies two numbers and adds the result to a running total. Because that is two operations, each MAC counts as 2 ops.

TOPS = MAC units × 2 × clock speed (Hz) ÷ 1012
MAC unitse.g. 4,096 × 2 ops / MACmultiply + add × Clocke.g. 1.25 GHz ≈ 10.2TOPS 4,096 × 2 × 1.25 × 10⁹ ≈ 10.24 × 10¹² ops/s
A worked example: the headline TOPS is just hardware arithmetic.

Notice what is missing: no memory, no software, no real model. The figure assumes every MAC unit is busy every single cycle, which almost never happens.

3. Precision changes the number

An operation on an 8-bit integer is far cheaper than one on a 16-bit or 32-bit floating-point number. Chips can pack more low-precision calculators into the same area and power, so the same chip may quote very different TOPS depending on the number format used.

FormatBitsTypical useRelative TOPS
FP3232Training, scientific work1× (baseline)
FP16 / BF1616Training, higher-quality inference~2×–4×
INT8 / FP88Most edge and phone inference~4×–8×
INT4 / FP44Compressed language models~8×–16×
Watch outA chip advertising “100 TOPS (INT4)” is not comparable with one advertising “100 TOPS (INT8)”. The first delivers roughly half the INT8 throughput. The ratios above are typical, not exact; they vary by architecture.

Lower precision also loses accuracy. Many models tolerate INT8 well, some need FP16, and aggressive INT4 works best with special compression techniques. So a big low-precision TOPS number only helps if your model runs well at that precision.

10FP32 ~30FP16 ~80INT8 ~160INT4 Same imaginary chip, same silicon. Only the number format changes.
Illustrative only: one chip, four different “TOPS” headlines.

4. Why TOPS alone can mislead

A chip is only as fast as its slowest link. Five things routinely stop real performance from reaching the headline number.

Memory bandwidth

Calculators starve if data cannot arrive fast enough from memory. Large language models are usually limited here, not by compute.

Architecture

Some designs keep data close to the MACs and reuse it; others shuffle it around wastefully.

Software stack

Drivers, compilers and model converters decide how much of the hardware a model can actually use.

Power and heat

A laptop or phone throttles when hot. Peak TOPS may last only seconds.

The model itself

Some layers map perfectly to the chip; others fall back to a slower CPU.

Utilization

Real workloads often use 20%–60% of peak, sometimes far less.

The memory wall, visually

Memoryweights + data narrow pipe limited GB/s Compute engine 100 TOPS capable mostly waiting… Resultslower thanpeak A giant engine behind a thin pipe runs at the speed of the pipe.
When data arrives slowly, extra TOPS sit idle.

Engineers describe this with the roofline model: performance is capped either by compute (TOPS) or by memory (bandwidth × how many operations you do per byte fetched). Models that do lots of math per byte, like image convolutions, can approach peak TOPS. Chatbots generating one word at a time do very little math per byte, so they hit the memory ceiling first.

WorkloadUsually limited byDoes more TOPS help?
Camera image enhancementComputeYes, strongly
Object detection (video)Compute + memoryYes, up to a point
Speech recognitionBalancedModerately
Chatbot text generationMemory bandwidthBarely; bandwidth matters more
Chatbot prompt readingComputeYes

5. Same TOPS, different results

Here is an illustrative example. Two chips both claim 40 TOPS at INT8.

PropertyChip AChip B
Peak INT8 TOPS4040
Memory bandwidth120 GB/s40 GB/s
Mature software toolsYesLimited
Power at full load15 W35 W
Typical utilization on a vision model~65%~30%
Effective throughput~26 TOPS~12 TOPS

Identical on the label, more than 2× apart in practice, and one draws far less power. That is why buyers should never treat TOPS as a final score.

6. Why automotive chips list hundreds of TOPS

A self-driving or driver-assistance computer must process many cameras, radar and lidar streams simultaneously, run several large neural networks, and do it with safety margins and redundancy. It has a big power budget (tens to hundreds of watts) and active cooling, so it can host huge arrays of MAC units. A phone, in contrast, has just a few watts to spare and no fan.

Typical peak TOPS (log-style scale, approximate, marketing INT8/sparse figures vary) Tiny microcontroller~0.5 Smartphone NPU~35–60 AI laptop NPU~40–50+ Edge AI module~20–275 Automotive computer~250–1000+ Data-center GPUthousands
Bars show relative rank, not exact scale. Ranges are indicative and change every generation.
Device classPower budgetTypical TOPS claimWhat it runs
Smartwatch / sensormilliwattsUnder 1Wake word, activity detection
Smartphone2–8 WTensPhotos, translation, assistants
AI laptop5–15 W (NPU)~40–50+Background blur, transcription, local assistants
Edge AI box10–60 WTens to ~275Cameras, robots, factory vision
Automotive50–500 WHundreds to 1000+Perception, planning, driver monitoring
Data-center GPU300–1000+ WThousandsTraining and large-scale inference

7. TOPS per watt: efficiency matters

For battery and edge devices, TOPS/W is often more useful. It tells you how much AI work you get for each watt. A 10-TOPS chip at 2 W (5 TOPS/W) may beat a 30-TOPS chip at 20 W (1.5 TOPS/W) inside a phone, because the phone simply cannot supply or cool 20 W.

Efficiency = TOPS ÷ Watts
ChipTOPSWattsTOPS/WFits a phone?
X1025.0Yes
Y30201.5No, too hot
Z200405.0No, but great for a robot

Also check whether the quoted power is for the whole chip or just the AI block, and whether it is a sustained or brief peak figure.

8. Tricks that inflate the headline

  • Sparsity: Some chips skip multiplications by zero. If a model is specially pruned, effective throughput can double, and some vendors quote that “sparse” number. Dense models see no gain.
  • Combined engines: “Platform TOPS” may add CPU + GPU + NPU together, even if a real app can use only one at a time.
  • Lowest precision: As shown above, INT4 figures look far bigger than INT8.
  • Boost clocks: The number may assume a maximum clock that thermals do not allow for long.
  • Different definitions: Whether a MAC counts as one or two operations is a convention; check it.
Sanity ruleAsk for the number in the format: ___ TOPS, dense, INT8, NPU only, sustained. If any blank is missing, treat the figure as a marketing headline.

9. Better ways to judge AI performance

MetricMeasuresWhy it helps
Inferences per secondReal model throughputActual result on a named model
Latency (ms)Time for one answerMatters for cameras and voice
Tokens per secondChatbot speedReflects memory limits
Time to first tokenDelay before reply startsShows compute for prompt reading
Memory bandwidth (GB/s)Data feed ratePredicts language-model speed
Memory capacity (GB)Model size that fitsDecides which models can run at all
TOPS/W, battery lifeEfficiencyVital for mobile and edge
Benchmark suites (e.g. MLPerf)Standardised testsComparable across vendors

A simple rule for chatbots

For a local language model, each generated word requires reading roughly the entire model from memory. So the speed ceiling is approximately:

tokens/sec ≈ memory bandwidth ÷ model size in memory

Example: a 4 GB model on 60 GB/s memory gives about 15 tokens per second at best, whether the chip has 40 TOPS or 400 TOPS. This is why raising TOPS alone often fails to speed up chat.

10. Ask three questions

  1. At what precision? INT8, FP16 or INT4 changes the number by multiples.
  2. For which workload? Ask for results on the model you care about, not only peak.
  3. At what power? Sustained watts and thermal limits decide real-world use.

Buyer’s checklist

CheckGood signRed flag
Precision stated“INT8, dense”Just “TOPS”
ScopeNPU only, clearly labelledCombined CPU+GPU+NPU total
Software supportPublished tools, model zooNo documentation
Real benchmarksNamed models with latencyOnly peak numbers
MemoryBandwidth and capacity listedNot mentioned
EfficiencyTOPS/W or battery dataNo power figure

11. Common myths

MythReality
100 TOPS is 10× faster than 10 TOPSOnly if everything else is equal, which it rarely is.
More TOPS always means better AI PCsSoftware support and memory often matter more day to day.
TOPS measures intelligenceIt measures arithmetic capacity, not model quality.
TOPS is meaninglessNot true; it is a useful rough guide to compute capability.
GPUs and NPUs quote TOPS the same wayFormats, sparsity and scope frequently differ.

12. Summary

Key takeaways TOPS is peak trillions of operations per second. It is calculated from hardware, depends heavily on precision, and ignores memory, software, heat and the actual model. Use it to compare AI compute capability roughly, then check precision, workload, power, memory bandwidth and real benchmarks before judging.

So when someone says, “This chip has 100 TOPS,” smile and ask: 100 TOPS at what precision, for which workload, and at what power?

Figures in tables and charts are illustrative and rounded; real products vary by generation and vendor.