An NPU (Neural Processing Unit) is a chip —or a block within a processor— specifically designed to execute the calculations inherent to artificial intelligence, above all the matrix multiplications that underpin neural networks. Unlike a CPU (general-purpose) or a GPU (born for graphics and later adapted to AI), the NPU is optimized from the ground up for model inference: it performs those operations in a massively parallel way with far lower energy consumption. Its raw performance is usually expressed in TOPS (trillions of operations per second).
Its practical relevance is direct: it makes it feasible to run AI locally —on a laptop, a phone, an edge device or a server on the company's premises— without depending on the cloud. For an organization that translates into lower latency, a lower cost per inference and, above all, data sovereignty: the models and the information never physically leave home.
The 5 best NPU units of 2026 and 2027
In 2026–2027 the market has settled into a clear ladder, from the pocket to the datacenter. These five references cover it in full.
1. Tiiny AI Pocket Lab — personal pocket AI

Unveiled at CES 2026, the Pocket Lab by Tiiny AI packs into 300 grams a 12-core ARM v9.2 CPU, a discrete NPU of ~160 TOPS (plus ~30 TOPS in the SoC), 80 GB of LPDDR5X and a 1 TB SSD. It runs LLMs of up to 120B parameters fully offline, with one-click deployment of open models (GPT-OSS, Llama, Qwen, Mistral), for around €1,400.
2. NVIDIA Jetson AGX Orin / Thor — the embedded standard

In robotics, computer vision and industry, the Jetson family by NVIDIA remains the reference: up to 275 TOPS in the AGX Orin and a generational leap with Thor for autonomous machines. Its advantage is not just the silicon, but the most mature software ecosystem on the market (CUDA, TensorRT, Isaac).
3. Hailo-10H — generative AI at the edge with 3 watts

The Hailo-10H delivers 40 TOPS INT8 with 8 GB of dedicated RAM while consuming ~3 W. Available in M.2 and USB modules and as the brain of the Raspberry Pi AI HAT+ 2, it brings small LLMs and vision models to devices where previously only a microcontroller could fit. Today it's the TOPS/watt ratio to beat at the edge.
4. Axelera Metis — industrial vision at 214 TOPS

The European Axelera AI tackles industrial vision with Metis, an AIPU in M.2/PCIe format that delivers up to 214 TOPS INT8 with 3.5–9 W and latencies below 100 ms without stealing RAM from the host system: exactly what production lines, logistics and quality control demand.
5. Qualcomm AI200 and AI250 — the NPU reaches the rack

Qualcomm's AI200 (2026) and AI250 (2027) accelerators, based on rack-scale Hexagon NPUs, bring generative inference to the datacenter with 768 GB of LPDDR per card, direct liquid cooling and up to 160 kW per rack; the AI250 adds near-memory computing with more than 10× the effective bandwidth.
The sixth: NAiOS's rack NPU unit
NAiOS Labs is preparing its own custom NPU-based unit in rack format, designed to be installed in its own datacenters and serve the platform's Dedicated on-site modality: AI hardware within the customer's (or NAiOS's) infrastructure, with data always under control.
Quick comparison
| Unit | Format | TOPS (INT8) | Memory | Power | Availability |
|---|---|---|---|---|---|
| Tiiny AI Pocket Lab | ~190 | 80 GB LPDDR5X | ~65 W | 2026 | |
| NVIDIA Jetson AGX Orin | embedded module | 275 | 64 GB | 15–60 W | available |
| Hailo-10H | M.2 / USB | 40 | 8 GB dedicated | ~3 W | available |
| Axelera Metis | M.2 / PCIe | 214 | dedicated on card | 3.5–9 W | available |
| Qualcomm AI200 / AI250 | datacenter rack | rack scale | 768 GB/card | up to 160 kW/rack | 2026 / 2027 |
The AI250 promises more than 10× the effective bandwidth with near-memory computing.
10 general-purpose use cases
1. Private office productivity assistant
The most immediate case: an assistant that drafts emails, summarizes reports, prepares proposals and answers internal questions running entirely within the company. A mid-range NPU is enough to serve a 7–14B parameter model to dozens of employees with latencies under one second. The difference compared to a cloud service is not just the monthly bill, which disappears, but that contract drafts, business figures and sensitive communications never leave the corporate perimeter. For departments that handle confidential information daily —management, finance, human resources— it is the shortest path to adopting generative AI without opening a new compliance front or depending on a third party's data retention policy.
2. Semantic search and RAG over internal documentation
Every organization accumulates manuals, minutes, contracts, wikis and emails that no one can find when needed. A RAG system (retrieval-augmented generation) indexes that knowledge with embeddings and answers questions in natural language citing the sources. It is an ideal case for NPUs because it combines two light and constant workloads: vectorizing documents as they arrive and generating short answers on demand. A team can ask "what warranty did we sign with this supplier in 2024?" and get the exact clause in seconds, without the question or the contract passing through external servers. The hardware investment pays off quickly: internal knowledge is no longer buried and the cost per query is practically zero once deployed.
3. Customer service with private chatbots
A support chatbot that knows the catalog, the return policies and the incident history can handle 60–80% of repetitive queries without human intervention. Running it on your own NPUs changes the economics of the case: the marginal cost per conversation is zero, so it can be offered on the website, on WhatsApp and on the phone without watching the token counter. In addition, customer data —orders, addresses, complaints— is processed on your own infrastructure, which simplifies GDPR compliance and avoids international data transfers. With small models fine-tuned on the company's historical conversations, the perceived quality rivals much more expensive cloud solutions, and scaling during campaign peaks is a matter of adding another card.
4. Meeting transcription and minutes
Modern speech-to-text models work wonderfully on modest NPUs: a Hailo-10H or the NPU of a recent laptop transcribe in real time with professional accuracy. The complete workflow —transcribing the meeting, identifying speakers, extracting decisions and tasks, drafting the minutes and sending them— can be fully automated locally. This matters because meetings are among the most sensitive things a company produces: they discuss layoffs, prices, strategy and clients by name. Uploading that audio to an external service is a risk that many boards of directors don't want to take; processing it in the room, literally on the device that records it, eliminates that risk entirely and turns every meeting into searchable knowledge.
5. Multilingual corporate translation
For companies operating in several markets, translation is a constant trickle: product sheets, contracts, support, technical documentation. Today's translation models fit comfortably in a desktop NPU and produce professional quality in the usual language pairs. Running them locally makes it possible to translate confidential documents —contracts under negotiation, patents, bids— without exposing them to external services, and to do so in massive batches at no cost per word. At NAiOS we experience this firsthand: our own engine automatically translates the blog and website into four languages. Combined with corporate terminology glossaries, the result maintains the brand voice and technical consistency that generic cloud translators don't guarantee.
6. Conversational business intelligence
Asking "how are sales in the eastern region doing against the quarterly target?" and receiving the answer with its chart, instead of waiting for someone to build the dashboard, is the leap that generative AI brings to BI. A local model with access to the analytical database translates natural language into SQL, runs the query and explains the result. That this happens on your own NPUs is almost mandatory: aggregated financial and commercial data is the crown jewel, and no data manager wants to send it to an external API query by query. With a Pocket Lab-type unit or an edge server, every executive has an analyst available 24 hours a day that also never leaks information.
7. Local coding copilot
Code assistants are already the standard productivity tool in development, but sending the repository to a third party is unacceptable for banking, defense, healthcare, or any company with valuable intellectual property in its software. Open-source code models with 7–30B parameters offer autocompletion, test generation, and code explanation with highly competitive quality, and an NPU with enough memory —the Pocket Lab with its 80 GB, a well-sized Jetson— serves them to an entire development team. The copilot knows the proprietary code without it ever leaving the building, can be tuned to the team's internal conventions, and works the same in the office as in an isolated environment with no internet access.
8. Document automation: OCR and extraction
Invoices, delivery notes, ID cards, payslips, forms: the intake bureaucracy of any company. The combination of modern OCR and vision language models makes it possible to read the document, classify it, extract the relevant fields, and dump them into the ERP without rigid templates, tolerating never-before-seen formats. It's a perfect workload for NPUs because the volume is high but each document is small, and overnight batch processing takes advantage of hardware that serves other cases during the day. By processing locally, the personal data of customers and employees traveling in those documents is not exposed, something directly required when special categories of data are involved. The return is among the most measurable: hours of data entry that vanish.
9. Intelligent video surveillance and access control
Cameras generate the largest volume of data in a company and almost no one looks at it. A vision NPU —a Metis, a Hailo— turns that stream into useful events: after-hours intrusion, missing PPE in a construction zone, exceeded capacity, unknown vehicle at the loading dock. Processing video at the edge, next to the camera, avoids saturating the network by uploading streams to the cloud and keeps the footage —potential biometric data, especially sensitive— within the premises. Alerts reach the security team in real time with the relevant clip, and aggregated analytics (flows, occupancy, patterns) feed operations decisions without compromising the privacy of employees and visitors.
10. Internal process agents
The step after the chatbot is the agent: an AI that not only responds but executes —it processes a supplier registration, reconciles a payment, prepares an employee's onboarding— chaining calls to internal systems. Precisely because it touches ERP, CRM and databases with write permissions, it is the case where local execution makes the most sense: credentials and operational data never leave the corporate network, and every action is audited on proprietary infrastructure. A dedicated NPU also guarantees stable latencies for flows with dozens of steps. This is the natural terrain of platforms like NAiOS: agents with governed access to the company's tools, running on hardware the company controls.
10 niche use cases
1. Visual inspection on the production line
In manufacturing, detecting a millimetric defect at line speed requires inference in milliseconds: there is no time to go to the cloud and back. An industrial NPU like the Axelera Metis analyzes hundreds of parts per minute next to the camera itself, detecting cracks, burrs, cold welds or poorly printed labels with precision superior to human inspection and without fatigue. The model is trained with the company's own parts and improves with each newly labeled defect. The impact is twofold: less defective product reaches the customer and quality traceability is documented automatically. For plants with old lines, retrofitting is affordable: an industrial camera, an M.2 module and an embedded PC are enough to give vision to a line from the nineties.
2. Predictive maintenance in the plant
Vibration, temperature, electrical consumption, sound: machinery warns before it breaks, but you have to listen to it continuously. Lightweight time-series models running on low-power NPUs next to the sensor detect the drift of a bearing or the cavitation of a pump weeks before failure, when the repair is still cheap and schedulable. Edge processing is key because the volume of raw sensor data is unmanageable to upload entirely to the cloud, and because many plants have limited connectivity or are directly isolated for OT security. The result is measured in avoided unplanned downtime: in continuous-process industries, a single avoided stoppage pays for the entire hardware deployment several times over.
3. Diagnostic support in healthcare at the edge
In healthcare, patient data is the most protected category that exists, and latency can be clinical: an ultrasound that highlights structures in real time, a retinal analysis in the consultation room itself, the prioritization of urgent X-rays in the hospital's PACS. NPUs make it possible to run these certified models inside the medical device or the hospital data center, without any image leaving the facility, which radically simplifies compliance with GDPR and the medical device regulation. For rural healthcare and medical NGOs, a portable device with an NPU brings reference diagnostic capability to clinics without reliable connectivity. It does not replace the doctor: it brings the expert second opinion closer to the point of care.
4. Precision agriculture
A drone flying over a plot or a tractor with multispectral cameras generate gigabytes per hour that cannot fit through the field's mobile coverage. With an NPU on board —a Jetson is the de facto standard— the analysis happens in flight: water stress detection, plant counting, identification of pests and weeds plant by plant. This enables the selective application of phytosanitary products, which reduces chemical use to double-digit percentage figures, with the savings and environmental benefit it entails. For cooperatives and medium-sized farms, the same hardware serves season after season, and the accumulated vigor and yield maps become a proprietary agronomic asset that no external provider holds.
5. Physical store analytics
Online commerce knows everything about its visitors; the physical store, almost nothing. Cameras with vision NPUs even the balance anonymously: entry flows, routes, hot zones, queues that exceed the threshold, shelves with stock gaps. Everything is processed at the edge and only aggregated metrics come out, without identifying people or storing images, which keeps the system within the safe territory of GDPR. Retail operates on thin margins and dozens or hundreds of locations: having each store need only a small NPU-equipped device, without recurring per-camera cloud costs, makes economically viable what remote processing did not make worthwhile. Decisions about layout, staffing, and restocking go from intuition to data.
6. ADAS and fleet management
Advanced driver-assistance systems are the extreme case of latency: braking for a pedestrian doesn't allow a round trip to the server. Vehicular NPUs —with Jetson Thor as the high-end benchmark— process cameras, radar and lidar within the vehicle itself, making decisions in milliseconds. In commercial fleets, the same hardware adds operational value: driver fatigue and distraction detection, smart incident logging that saves only the relevant seconds, and efficient-driving coaching. For a transport company, that translates into fewer accidents, insurance premiums renegotiable with data, and fuel savings. Cabin video, which is especially sensitive, never leaves the vehicle except for a justified event.
7. Collaborative robotics
A cobot that shares space with people needs to perceive and react in real time: to see the hand entering its zone, adjust the force, replan the trajectory. That perception runs on NPUs embedded in the robot itself, because functional safety cannot depend on the factory's wifi. The new generation of manipulation models —visuomotor policies trained by demonstration— also makes it possible to reprogram the cobot by teaching it the task instead of programming it, and executing those policies requires precisely the kind of dense, efficient inference that an NPU offers. For industrial SMEs, the cobot-plus-NPU combination makes it cheaper to automate picking, assembly and palletizing tasks that previously only paid off in large production runs.
8. Operations in isolated environments
Ships on the high seas, oil & gas platforms and plants, underground mining, defense facilities: environments where connectivity is expensive, intermittent or prohibited by security policy. There, AI either runs locally or doesn't exist. Ruggedized NPUs make it possible to deploy safety vision, predictive maintenance, technical assistants with the plant's documentation and sensor-data analysis on-site, syncing only summaries with the mainland when there is a communication window. An assistant that answers questions about the manuals of all the equipment on board, without internet, changes how a chief engineer operates. This is also the case where hardware with temperature, vibration and waterproofing certifications justifies its premium price.
9. Sectors with professional confidentiality
Law firms, auditors, private banking, notaries: professions where confidentiality is not a preference but a deontological and legal obligation. For them, sending client documentation to an external API is, in many cases, directly unfeasible, no matter how encrypted it travels. An NPU unit in the office —the size of a book— enables contract analysis, assisted due diligence, semantic case-law search over one's own archive and draft writing, all without a single document crossing the door. The commercial argument is powerful: the firm can guarantee in writing to its clients that their information has never left its premises. AI stops being a reputational risk and becomes a differentiator.
10. Smart cities and traffic
Modern urban management relies on thousands of cameras and sensors whose continuous transmission to a central data center is extremely expensive and poses obvious privacy risks. The winning pattern is edge processing: NPUs in the traffic cabinets themselves analyze flows, detect incidents, adjust traffic lights adaptively, prioritize emergencies and count transport modes —car, bike, pedestrian— sending only events and anonymous statistics to the control center. Images are discarded on site, which facilitates compliance with the GDPR and public video surveillance regulations. For medium-sized municipalities, the distributed model also reduces the communications bill and makes the system resilient: a fiber cut does not leave the city blind.
In summary
From the 300 grams of the Pocket Lab to the 160 kW of an AI200 rack, NPUs already cover the entire spectrum between the pocket and the datacenter, and the use cases —from office work to the operating room— share a common denominator: AI performs better, costs less and compromises less data the closer to the data it runs. You have the full analysis of the five units in our blog post, and the detail of the key metric in the TOPS entry.