The scene repeats at every NAiOS Labs demo: someone on the team, with more energy than a movie scientist about to shout "this is it!", points at a rack full of glowing modules and drops a figure: "forty TOPSTOPS (Tera Operations Per Second)Unit that measures the raw performance of an AI accelerator: trillions of operations per second at three watts". Half the room nods while the other half thinks: what on earth is a TOPS? Today we answer that half — and they are right to ask.
The one-sentence definition
TOPS stands for Tera Operations Per Second: trillions of operations per second. It is the unit used to measure the raw power of AI accelerators, especially NPUs. A "40 TOPS" chip executes forty trillion elementary multiplications and additions every second — exactly the kind of arithmetic neural networks are made of. It is to AI hardware what horsepower is to a car: the figure that imperfectly summarises how much power sits under the hood.
The fine print: INT8, peaks and memory
Three caveats separate the informed buyer from the one who takes home the wrong chip. First, precision: TOPS are almost always quoted at INT8 (8-bit integers); the same silicon delivers roughly double at INT4 and half at FP16, so comparing figures without checking precision is comparing apples to oranges. Second, TOPS are a theoretical peak: the number assumes compute units never wait for data, which in practice happens constantly. Third — and this one hurts most with LLMs — memory: a large model has to fit in memory and stream through it at full speed. That is why a chip with fewer TOPS but 80 GB of well-fed memory can run a model that chokes a nominally more powerful one. With large models, bandwidth rules as much as compute.
Real 2026 reference points
To calibrate your eye: the NPUNPU (Neural Processing Unit)A chip specialized in accelerating neural network inference with maximum energy efficiency in an "AI PC" laptop sits around 40–50 TOPS (40 is the CopilotAI CopilotAI Assistant integrated into work tools+ threshold), a Hailo-10H delivers 40 TOPS at 3 watts, an Axelera Metis reaches 214, an NVIDIA Jetson AGX Orin 275, and the Tiiny AI Pocket Lab combines ~190 TOPS with 80 GB of memory. At datacenter scale the per-chip figure loses meaning and people talk about per-rack performance, as with the Qualcomm AI200/AI250. All five units are analysed in our top 5 NPUs of 2026–2027.
The number that truly separates generations
If you can only look at one metric, look at TOPS per watt. Efficiency is what brought AI down from the datacenter to the pocket, and it is the metric we are designing around at NAiOS Labs for our rack-mounted NPU unit for dedicated datacenters: maximising inference per watt means more sovereign AISovereign AIStrategic priority for countries to have their own AI infrastructures for every euro on the power bill. The practical rule for choosing hardware is three questions together: can it run the model (TOPS)? does the model fit (memory)? and at what cost (TOPS/watt)?
Continue in the wiki
We have published the TOPS entry in the NAiOS Wiki as the reference explanation, alongside its companion NPU with the full unit comparison and twenty use cases. And if you want to know what fits your case — from a private office assistant to industrial vision — let's talk: at NAiOS Labs we have TOPS and enthusiasm to spare.






