Skip to content
TEN Brief Ten verified stories a day 2026.10.06 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 2 min read · Reference

What mixture of experts (MoE) is — how a 125B model runs on a gaming GPU

Mixture of experts (MoE) is a way of building an AI model from many smaller sub-networks, called experts, plus a router that sends each token to only a few of them. Because only part of the model runs for each token, an MoE model can hold far more total parameters than a dense model with the same compute cost. Mistral's Mixtral 8x7B used 2 of 8 experts per token, and DeepSeek-V3 activated about 37 billion of its 671 billion parameters. The trade-off is memory: idle experts still have to be stored. In October 2026 the open-source Strata engine ran a 125-billion-parameter MoE model on a single 12GB gaming GPU by keeping frequently used experts on the GPU and the rest in system RAM

A bright workshop with many workbenches, two lit by sunbeams where craftspeople work

The three lines

  • Structure — many 'experts' plus a router that picks a few per token: big in total, small in compute
  • Examples — Mixtral used 2 of 8; DeepSeek-V3 used ~37B of 671B; most top open models are MoE
  • Trade-off — compute falls, memory doesn't; Strata spread experts across GPU, RAM and SSD to run 125B on 12GB

Key questions

What is mixture of experts in AI
**A model built from many 'experts' with a router that uses only a few per token.** | Part | Role | |---|---| | Expert | A small neural network; a model may have a few or tens of thousands | | Router (gate) | Chooses which experts handle each token | | Active parameters | The share actually computed per token | | Total parameters | All experts combined | Experts are not assigned human topics like 'math'; the router learns the split during training.
MoE vs dense model
**At the same size, MoE needs less compute but the same memory.** | | Dense | MoE | |---|---|---| | Compute per token | All parameters | Chosen experts only | | Speed and cost | Scale with size | Scale with active share | | Memory | All | **All** (idle experts still stored) | | Training | More stable | Must manage expert imbalance |
Examples of mixture of experts models
**Most leading open-weight models now use MoE.** | Model | Total parameters | Active per token | |---|---|---| | Mixtral 8x7B (Dec 2023) | ~46.7B | ~12.9B (2 of 8 experts) | | DeepSeek-V3 (Dec 2024) | 671B | ~37B | | Qwen3.8-Flash-Next (2026) | 125B | 10 of 24,576 experts |

Mixture of experts is the design that separates an AI model's size from its running cost. In early October 2026, developers passed around news that a 125-billion-parameter model was running on a single gaming graphics card: the open-source Strata engine had put the MoE model Qwen3.8-Flash-Next on a PC with a 12GB GPU and 32GB of RAM. Why that works is the MoE idea itself.

1. A model where not everyone works

A conventional, or dense, model uses every parameter for every token. MoE splits much of the model into "experts" and puts a router in front. For each token, the router picks a few experts; the rest sit that token out.

AnalogyDense modelMoE model
CompanyAll 100 staff attend every meeting2–8 relevant people per meeting
Company size (knowledge)100100
Meeting cost (compute)100 peopleA few people
Office space (memory)100 desksStill 100 desks

The idea is old. Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton proposed adaptive mixtures of local experts in 1991; in 2017 Google researchers led by Noam Shazeer added sparsely gated MoE layers to large neural networks. It went mainstream with Mistral's Mixtral 8x7B in December 2023.

2. Total versus active

ModelReleasedTotal parametersActive per tokenActive share
Mixtral 8x7BDec 2023~46.7B~12.9B (2 of 8)~28%
DeepSeek-V3Dec 2024671B~37B~5.5%
Qwen3.8-Flash-Next2026125B10 of 24,576 expertsVery low

The lower the active share, the cheaper the same amount of knowledge is to serve. That is a large part of how Chinese labs compete on price, and why the latest efficiency-focused DeepSeek model sits within about 3% of the top US model on LiveBench. Alongside distillation, MoE is one of the two main forces pushing AI prices down.

StrengthsWeaknesses
Less compute and power per token at a given sizeIdle experts must still be held in memory
Easy to grow total knowledgeRouter imbalance can hurt quality
Lower serving cost → lower API pricesCross-GPU communication overhead

3. What Strata did: counter and pantry

MoE's weak spot is memory: 125 billion parameters normally need several data-center GPUs. Strata turned MoE's own property around — only some experts run per token, so only the frequently used ones need to sit on the GPU.

LocationHoldsSpeed
GPU memory (12GB)"Hot" experts called most oftenFastest
System RAM (32GB)All 24,576 expertsMedium
SSDLookup tableSlow
Small draft modelGuesses next tokens; big model verifies in batch1.6–1.8x throughput (developer claim)

The project describes it as keeping everyday tools on the counter and the rest in the pantry. It is MIT-licensed, supports NVIDIA and AMD cards with 12GB or more, and exposes OpenAI- and Anthropic-compatible APIs on the local machine. Reported speeds range from about 60 to 140 tokens per second depending on setup.

The takeaway: the bar for AI that runs entirely on your own PC dropped again. It does not mean local output matches data-center quality.

4. What remains unclear

  • Strata's speeds lack published test conditions.
  • Qwen3.8-Flash-Next's total active parameters per token were not confirmed.
  • Whether closed frontier models use MoE is undisclosed.

Sources

  1. Shazeer et al. — Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (2017)
  2. Mistral AI — Mixtral of Experts
  3. DeepSeek-AI — DeepSeek-V3 Technical Report
  4. Startup Fortune — Strata lets a 125 billion parameter model run on a gaming GPU
  5. AI Weekly — Strata Runs 125B Qwen3.8-Flash-Next on a 12GB Gaming GPU

Verification

Published
Last modified
Cross-check
Checked against 5 independent sources.
Unverified
  • Reported Strata speeds vary (60–95 tokens per second; 100–140 on an RTX 3090); test conditions were not published, so we give a range.
  • Total active parameters per token for Qwen3.8-Flash-Next were not confirmed; only the 10-of-24,576-experts figure was.
  • Whether closed frontier models from OpenAI, Anthropic or Google use MoE has not been officially disclosed.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related