Skip to content
TEN Brief Ten verified stories a day 2026.10.08 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 3 min read · Explainer

What is an NPU — how it differs from a GPU, and why it wins at inference

An NPU, or neural processing unit, is a processor designed to do one thing: the multiply-and-add arithmetic that neural networks run billions of times, usually on low-precision numbers. A GPU can do the same math, but it is built for graphics and general parallel computing as well, which costs extra circuitry and power. By giving up that flexibility, an NPU runs already-trained models, known as inference, with more work per watt and per dollar. That is why NPUs sit in smartphones, AI laptops and data-center inference servers, while GPUs still dominate training. Korean startups such as Rebellions, FuriosaAI and DeepX focus on NPUs, and in October 2026 AMD agreed with Korea's government to test Korean NPUs alongside its own GPUs and CPUs

Gloved hands checking a chip on a circuit board in a bright test lab

The three lines

  • Definition — a chip for neural-network math only; it trades generality for efficiency
  • Strength — inference, not training: more answers per watt in phones, laptops and servers
  • Weak spot — software; CUDA is the default, so NPUs need tools and ecosystems to catch up

Key questions

NPU vs GPU difference
**A GPU is a general parallel engine; an NPU is a dedicated neural-network engine.** | | GPU | NPU | |---|---|---| | Origin | Graphics → general parallel compute | Neural networks only | | Best at | Training and inference; flexible | Efficient inference | | Drawback | Power and price | Slower to support new model designs | | Software | Nvidia CUDA, mature | Vendor-specific, smaller |
Do I need an NPU in my laptop
**Only for AI features that run on the device itself.** | Use | Example | |---|---| | Video calls | Background blur, noise removal | | Photos | Subject cutouts, upscaling | | Documents | On-device search and summaries | | Benchmark | Microsoft Copilot+ PCs require 40+ TOPS NPUs (2024) | Cloud chatbots run on servers, so an NPU does not speed them up.
Which companies make NPUs
**Big platforms and specialist startups.** | Type | Examples | |---|---| | In-house | Google TPU, Apple Neural Engine, phone chipmakers | | Korean data-center inference | Rebellions, FuriosaAI, HyperAccel | | Korean edge devices | DeepX, Mobilint |

Say "AI chip" and most people picture an Nvidia GPU. But phones, laptops and parts of data centers increasingly carry a different processor: the NPU. The term came back into the news on October 7, 2026, when AMD's Lisa Su met 14 Korean AI chip firms and agreed to test infrastructure that mixes AMD CPUs and GPUs with Korean NPUs. Here is what an NPU is, how it differs from a GPU, and why so many startups have bet on it.

1. What an NPU does

An NPU (neural processing unit) is a chip made only for neural-network math. When an AI model produces an answer, it multiplies and adds matrices of numbers billions of times. An NPU packs its silicon with multiply-accumulate (MAC) units and arranges memory so data moves as little as possible between them.

ChipAnalogyBest at
CPUA few highly skilled generalistsComplex logic, sequential tasks, coordination
GPUThousands of workers doing the same move at onceGraphics, big parallel jobs, AI training and inference
NPUA dedicated production lineRunning neural networks on little power

Specialization wins because no silicon or power is spent on unused features. NPUs mostly use low-precision numbers (8-bit integers, 8- or 4-bit floating point). Trained models tolerate less precision with little change in output, and smaller numbers mean more math per square millimeter and per watt.

2. Training vs inference: where each chip fits

TrainingInference
TaskAdjust the model's weights from dataCompute answers with fixed weights
FrequencyOnce per model (weeks to months)Every user request, nonstop
NeedsHigher precision, flexibility, thousands of linked chipsLow latency, energy and cost efficiency
Main chipsGPUs, Google TPUsGPUs plus NPUs and inference chips
GPUNPU
DesignGeneral parallel computeNeural networks only
New model architecturesHandled quickly in softwareMay lag, depending on chip design
Inference per wattRelatively lowerRelatively higher
SoftwareNvidia CUDA is the de facto standard; AMD ROCm catching upVendor-specific, small ecosystems

As AI services scale, spending shifts from training to inference: training a chatbot once costs less than answering hundreds of millions of questions every day. Cutting that bill is the NPU's reason to exist.

3. Where NPUs already live

LocationExamplesJob
Data centersGoogle TPU (announced 2016, used internally since 2015), Korean inference chipsSearch, translation, chatbots
SmartphonesApple Neural Engine (since the A11 in 2017), Android processorsFace unlock, photo processing, speech
LaptopsMicrosoft Copilot+ PCs require 40+ TOPS NPUs (2024)Video-call effects, on-device AI
Cars, robots, camerasEdge NPUsObject detection, real-time decisions

TOPS (trillions of operations per second) is the usual headline figure, but it depends on the precision measured, so the same chip can show figures two or three times apart.

4. Why Korean firms bet on NPUs

Nvidia controls both GPU hardware and CUDA software, so a startup cannot easily attack it head-on. Inference NPUs left room: buyers began asking not for the fastest chip but for the one that produces the most answers per watt and per dollar.

SegmentCompanyNotes
Data-center inferenceRebellionsATOM, next-generation REBEL; works with Korean clouds and telcos
Data-center inferenceFuriosaAIFirst-gen Warboy, second-gen RNGD (Renegade)
LLM inferenceHyperAccelDesigned for language-model serving
EdgeDeepX, MobilintLow-power chips for cameras, robots, industrial gear

The hard part is software. Code written for CUDA must be converted and optimized to run on an NPU, and every new model architecture must be supported again. That is why the AMD tie-up matters: if Korean NPUs plug into AMD's open-source ROCm ecosystem, data centers can mix GPUs and NPUs instead of choosing between Nvidia and a local chip.

5. Frequently asked

QuestionAnswer
Will NPUs replace GPUs?No. GPUs remain the workhorse for training and experiments; NPUs complement them by cutting inference costs
Does a laptop NPU make ChatGPT faster?No. Cloud services run on servers; the NPU only helps AI that runs locally
Is higher TOPS always better?Roughly, at equal precision, but memory bandwidth and software decide real speed
Is a TPU an NPU?Broadly, yes — both are neural-network accelerators; Google uses TPUs for training and inference

6. What remains unconfirmed

  • NPU-versus-GPU efficiency differs too much by model, precision and generation to give one multiple.
  • Deployment volumes for Korean NPUs in data centers are not consistently disclosed.
  • When and at what scale the AMD testbed begins will be clearer after a November steering committee.

Sources

  1. IBM — What is a neural processing unit (NPU)?
  2. Google Cloud — An in-depth look at Google's first Tensor Processing Unit (TPU)
  3. Microsoft — Introducing Copilot+ PCs
  4. SBS News — Lisa Su joins hands with 14 Korean AI chip companies

Verification

Published
Last modified
Cross-check
Checked against 4 independent sources.
Unverified
  • The efficiency gap between NPUs and GPUs varies by model, precision and generation, so no single multiple is given.
  • Actual data-center deployment volumes of Korean NPUs are not consistently disclosed.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related