What is an NPU — how it differs from a GPU, and why it wins at inference
An NPU, or neural processing unit, is a processor designed to do one thing: the multiply-and-add arithmetic that neural networks run billions of times, usually on low-precision numbers. A GPU can do the same math, but it is built for graphics and general parallel computing as well, which costs extra circuitry and power. By giving up that flexibility, an NPU runs already-trained models, known as inference, with more work per watt and per dollar. That is why NPUs sit in smartphones, AI laptops and data-center inference servers, while GPUs still dominate training. Korean startups such as Rebellions, FuriosaAI and DeepX focus on NPUs, and in October 2026 AMD agreed with Korea's government to test Korean NPUs alongside its own GPUs and CPUs
The three lines
- Definition — a chip for neural-network math only; it trades generality for efficiency
- Strength — inference, not training: more answers per watt in phones, laptops and servers
- Weak spot — software; CUDA is the default, so NPUs need tools and ecosystems to catch up
Key questions
- NPU vs GPU difference
- **A GPU is a general parallel engine; an NPU is a dedicated neural-network engine.** | | GPU | NPU | |---|---|---| | Origin | Graphics → general parallel compute | Neural networks only | | Best at | Training and inference; flexible | Efficient inference | | Drawback | Power and price | Slower to support new model designs | | Software | Nvidia CUDA, mature | Vendor-specific, smaller |
- Do I need an NPU in my laptop
- **Only for AI features that run on the device itself.** | Use | Example | |---|---| | Video calls | Background blur, noise removal | | Photos | Subject cutouts, upscaling | | Documents | On-device search and summaries | | Benchmark | Microsoft Copilot+ PCs require 40+ TOPS NPUs (2024) | Cloud chatbots run on servers, so an NPU does not speed them up.
- Which companies make NPUs
- **Big platforms and specialist startups.** | Type | Examples | |---|---| | In-house | Google TPU, Apple Neural Engine, phone chipmakers | | Korean data-center inference | Rebellions, FuriosaAI, HyperAccel | | Korean edge devices | DeepX, Mobilint |
Say "AI chip" and most people picture an Nvidia GPU. But phones, laptops and parts of data centers increasingly carry a different processor: the NPU. The term came back into the news on October 7, 2026, when AMD's Lisa Su met 14 Korean AI chip firms and agreed to test infrastructure that mixes AMD CPUs and GPUs with Korean NPUs. Here is what an NPU is, how it differs from a GPU, and why so many startups have bet on it.
1. What an NPU does
An NPU (neural processing unit) is a chip made only for neural-network math. When an AI model produces an answer, it multiplies and adds matrices of numbers billions of times. An NPU packs its silicon with multiply-accumulate (MAC) units and arranges memory so data moves as little as possible between them.
| Chip | Analogy | Best at |
|---|---|---|
| CPU | A few highly skilled generalists | Complex logic, sequential tasks, coordination |
| GPU | Thousands of workers doing the same move at once | Graphics, big parallel jobs, AI training and inference |
| NPU | A dedicated production line | Running neural networks on little power |
Specialization wins because no silicon or power is spent on unused features. NPUs mostly use low-precision numbers (8-bit integers, 8- or 4-bit floating point). Trained models tolerate less precision with little change in output, and smaller numbers mean more math per square millimeter and per watt.
2. Training vs inference: where each chip fits
| Training | Inference | |
|---|---|---|
| Task | Adjust the model's weights from data | Compute answers with fixed weights |
| Frequency | Once per model (weeks to months) | Every user request, nonstop |
| Needs | Higher precision, flexibility, thousands of linked chips | Low latency, energy and cost efficiency |
| Main chips | GPUs, Google TPUs | GPUs plus NPUs and inference chips |
| GPU | NPU | |
|---|---|---|
| Design | General parallel compute | Neural networks only |
| New model architectures | Handled quickly in software | May lag, depending on chip design |
| Inference per watt | Relatively lower | Relatively higher |
| Software | Nvidia CUDA is the de facto standard; AMD ROCm catching up | Vendor-specific, small ecosystems |
As AI services scale, spending shifts from training to inference: training a chatbot once costs less than answering hundreds of millions of questions every day. Cutting that bill is the NPU's reason to exist.
3. Where NPUs already live
| Location | Examples | Job |
|---|---|---|
| Data centers | Google TPU (announced 2016, used internally since 2015), Korean inference chips | Search, translation, chatbots |
| Smartphones | Apple Neural Engine (since the A11 in 2017), Android processors | Face unlock, photo processing, speech |
| Laptops | Microsoft Copilot+ PCs require 40+ TOPS NPUs (2024) | Video-call effects, on-device AI |
| Cars, robots, cameras | Edge NPUs | Object detection, real-time decisions |
TOPS (trillions of operations per second) is the usual headline figure, but it depends on the precision measured, so the same chip can show figures two or three times apart.
4. Why Korean firms bet on NPUs
Nvidia controls both GPU hardware and CUDA software, so a startup cannot easily attack it head-on. Inference NPUs left room: buyers began asking not for the fastest chip but for the one that produces the most answers per watt and per dollar.
| Segment | Company | Notes |
|---|---|---|
| Data-center inference | Rebellions | ATOM, next-generation REBEL; works with Korean clouds and telcos |
| Data-center inference | FuriosaAI | First-gen Warboy, second-gen RNGD (Renegade) |
| LLM inference | HyperAccel | Designed for language-model serving |
| Edge | DeepX, Mobilint | Low-power chips for cameras, robots, industrial gear |
The hard part is software. Code written for CUDA must be converted and optimized to run on an NPU, and every new model architecture must be supported again. That is why the AMD tie-up matters: if Korean NPUs plug into AMD's open-source ROCm ecosystem, data centers can mix GPUs and NPUs instead of choosing between Nvidia and a local chip.
5. Frequently asked
| Question | Answer |
|---|---|
| Will NPUs replace GPUs? | No. GPUs remain the workhorse for training and experiments; NPUs complement them by cutting inference costs |
| Does a laptop NPU make ChatGPT faster? | No. Cloud services run on servers; the NPU only helps AI that runs locally |
| Is higher TOPS always better? | Roughly, at equal precision, but memory bandwidth and software decide real speed |
| Is a TPU an NPU? | Broadly, yes — both are neural-network accelerators; Google uses TPUs for training and inference |
6. What remains unconfirmed
- NPU-versus-GPU efficiency differs too much by model, precision and generation to give one multiple.
- Deployment volumes for Korean NPUs in data centers are not consistently disclosed.
- When and at what scale the AMD testbed begins will be clearer after a November steering committee.