What mixture of experts (MoE) is — how a 125B model runs on a gaming GPU
Mixture of experts (MoE) is a way of building an AI model from many smaller sub-networks, called experts, plus a router that sends each token to only a few of them. Because only part of the model runs for each token, an MoE model can hold far more total parameters than a dense model with the same compute cost. Mistral's Mixtral 8x7B used 2 of 8 experts per token, and DeepSeek-V3 activated about 37 billion of its 671 billion parameters. The trade-off is memory: idle experts still have to be stored. In October 2026 the open-source Strata engine ran a 125-billion-parameter MoE model on a single 12GB gaming GPU by keeping frequently used experts on the GPU and the rest in system RAM
The three lines
- Structure — many 'experts' plus a router that picks a few per token: big in total, small in compute
- Examples — Mixtral used 2 of 8; DeepSeek-V3 used ~37B of 671B; most top open models are MoE
- Trade-off — compute falls, memory doesn't; Strata spread experts across GPU, RAM and SSD to run 125B on 12GB
Key questions
- What is mixture of experts in AI
- **A model built from many 'experts' with a router that uses only a few per token.** | Part | Role | |---|---| | Expert | A small neural network; a model may have a few or tens of thousands | | Router (gate) | Chooses which experts handle each token | | Active parameters | The share actually computed per token | | Total parameters | All experts combined | Experts are not assigned human topics like 'math'; the router learns the split during training.
- MoE vs dense model
- **At the same size, MoE needs less compute but the same memory.** | | Dense | MoE | |---|---|---| | Compute per token | All parameters | Chosen experts only | | Speed and cost | Scale with size | Scale with active share | | Memory | All | **All** (idle experts still stored) | | Training | More stable | Must manage expert imbalance |
- Examples of mixture of experts models
- **Most leading open-weight models now use MoE.** | Model | Total parameters | Active per token | |---|---|---| | Mixtral 8x7B (Dec 2023) | ~46.7B | ~12.9B (2 of 8 experts) | | DeepSeek-V3 (Dec 2024) | 671B | ~37B | | Qwen3.8-Flash-Next (2026) | 125B | 10 of 24,576 experts |
Mixture of experts is the design that separates an AI model's size from its running cost. In early October 2026, developers passed around news that a 125-billion-parameter model was running on a single gaming graphics card: the open-source Strata engine had put the MoE model Qwen3.8-Flash-Next on a PC with a 12GB GPU and 32GB of RAM. Why that works is the MoE idea itself.
1. A model where not everyone works
A conventional, or dense, model uses every parameter for every token. MoE splits much of the model into "experts" and puts a router in front. For each token, the router picks a few experts; the rest sit that token out.
| Analogy | Dense model | MoE model |
|---|---|---|
| Company | All 100 staff attend every meeting | 2–8 relevant people per meeting |
| Company size (knowledge) | 100 | 100 |
| Meeting cost (compute) | 100 people | A few people |
| Office space (memory) | 100 desks | Still 100 desks |
The idea is old. Robert Jacobs, Michael Jordan, Steven Nowlan and Geoffrey Hinton proposed adaptive mixtures of local experts in 1991; in 2017 Google researchers led by Noam Shazeer added sparsely gated MoE layers to large neural networks. It went mainstream with Mistral's Mixtral 8x7B in December 2023.
2. Total versus active
| Model | Released | Total parameters | Active per token | Active share |
|---|---|---|---|---|
| Mixtral 8x7B | Dec 2023 | ~46.7B | ~12.9B (2 of 8) | ~28% |
| DeepSeek-V3 | Dec 2024 | 671B | ~37B | ~5.5% |
| Qwen3.8-Flash-Next | 2026 | 125B | 10 of 24,576 experts | Very low |
The lower the active share, the cheaper the same amount of knowledge is to serve. That is a large part of how Chinese labs compete on price, and why the latest efficiency-focused DeepSeek model sits within about 3% of the top US model on LiveBench. Alongside distillation, MoE is one of the two main forces pushing AI prices down.
| Strengths | Weaknesses |
|---|---|
| Less compute and power per token at a given size | Idle experts must still be held in memory |
| Easy to grow total knowledge | Router imbalance can hurt quality |
| Lower serving cost → lower API prices | Cross-GPU communication overhead |
3. What Strata did: counter and pantry
MoE's weak spot is memory: 125 billion parameters normally need several data-center GPUs. Strata turned MoE's own property around — only some experts run per token, so only the frequently used ones need to sit on the GPU.
| Location | Holds | Speed |
|---|---|---|
| GPU memory (12GB) | "Hot" experts called most often | Fastest |
| System RAM (32GB) | All 24,576 experts | Medium |
| SSD | Lookup table | Slow |
| Small draft model | Guesses next tokens; big model verifies in batch | 1.6–1.8x throughput (developer claim) |
The project describes it as keeping everyday tools on the counter and the rest in the pantry. It is MIT-licensed, supports NVIDIA and AMD cards with 12GB or more, and exposes OpenAI- and Anthropic-compatible APIs on the local machine. Reported speeds range from about 60 to 140 tokens per second depending on setup.
The takeaway: the bar for AI that runs entirely on your own PC dropped again. It does not mean local output matches data-center quality.
4. What remains unclear
- Strata's speeds lack published test conditions.
- Qwen3.8-Flash-Next's total active parameters per token were not confirmed.
- Whether closed frontier models use MoE is undisclosed.
Sources
- Shazeer et al. — Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (2017)
- Mistral AI — Mixtral of Experts
- DeepSeek-AI — DeepSeek-V3 Technical Report
- Startup Fortune — Strata lets a 125 billion parameter model run on a gaming GPU
- AI Weekly — Strata Runs 125B Qwen3.8-Flash-Next on a 12GB Gaming GPU