DeepSeek V4.1 Flash scores 81.1 on LiveBench — US-China AI gap hits 3%
DeepSeek V4.1 Flash, released in September 2026, scored 81.1 on the LiveBench leaderboard as of October 4, against 83.4 for Anthropic's Claude Fable 5.1, the top US model. That 2.3-point, roughly 3% gap is the smallest Bloomberg Intelligence has tracked between the leading US and Chinese models; it was about 15% early in 2026 and about 9% in May. On LiveBench's agentic coding category DeepSeek scored 77.3 against Claude's 66.1. Analyst Robert Lea credits efficiency work and tuning for Chinese hardware. Only three of the top 15 LiveBench entries are Chinese, so the catch-up is concentrated at the very top
The three lines
- Scores — DeepSeek V4.1 Flash 81.1 vs Claude Fable 5.1 83.4 on LiveBench; a 2.3-point (~3%) gap
- Trend — the gap shrank from ~15% early in 2026 to ~9% in May to ~3%; DeepSeek leads agentic coding 77.3 to 66.1
- Limits — only 3 Chinese models in the top 15; Bloomberg sees China's AI sector unprofitable until 2030
Key questions
- How good is DeepSeek V4.1 Flash
- **81.1 overall on LiveBench, 2.3 points behind the top US model.** | Model | Overall | Agentic coding | |---|---|---| | Claude Fable 5.1 (max effort) | **83.4** | 66.1 | | DeepSeek V4.1 Flash (max effort) | 81.1 | **77.3** | | Gap | 2.3 (~3%) | DeepSeek +11.2 |
- How big is the US China AI gap
- **About 3% at the top, per Bloomberg Intelligence — a record low.** | Time | Top-model gap | |---|---| | Early 2026 | ~15% | | May 2026 | ~9% | | Oct 2026 | **~3%** |
- What is LiveBench
- **An AI benchmark that keeps replacing its questions to limit memorized answers.** | Feature | Detail | |---|---| | Questions | Built from recent contests, papers and news | | Updates | Questions refreshed regularly | | Scoring | Automatic, against fixed answers | | Aim | Reduce training-data contamination |
The distance between the best American and Chinese AI models is now a single stride. According to Bloomberg Intelligence, DeepSeek's V4.1 Flash, released in September, scored 81.1 on the LiveBench leaderboard — 2.3 points behind Anthropic's Claude Fable 5.1 at 83.4. That is the narrowest gap Bloomberg has recorded.
1. Nine months, 15% to 3%
| Time | US-China top-model gap | Note |
|---|---|---|
| Early 2026 | ~15% | — |
| May 2026 | ~9% | — |
| Sept 2026 | — | DeepSeek V4.1 Flash released |
| Oct 4, 2026 | ~3% | 81.1 vs 83.4 |
The category detail is starker. On agentic coding — models fixing and running code on their own to finish a task — DeepSeek scored 77.3 to Claude's 66.1. That is the capability US labs are betting most heavily on.
The "Flash" name matters too. DeepSeek's paper focuses on compressing the KV cache, the memory a model uses to track a conversation, and using only part of its parameters per token. The gap closed with a fast, lean model, not a giant one.
2. Why it narrowed
Bloomberg Intelligence senior analyst Robert Lea points to better engineering at Chinese labs and optimization for domestic hardware rather than US GPUs. He says that raises questions about how much US chip export controls are slowing China. Usage data already pointed the same way: in late September Chinese models took 57–67% of token traffic on the OpenRouter marketplace, mostly on price. Now the benchmark numbers are close as well.
3. What not to overclaim
| Signal | Reading |
|---|---|
| 3 Chinese models in LiveBench top 15 | The catch-up is concentrated in a few frontier models |
| Benchmark scores | They do not prove national AI leadership |
| Enterprise use | Marketplace data under-counts big-company workloads, where US models still dominate |
| Profitability | Lea expects China's AI industry to stay unprofitable until 2030, with 1,100+ large models competing on price |
"China has caught up" overstates it. "The top one or two models now score almost the same on one test" is accurate.
4. What remains unclear
- Whether other leaderboards show the same gap.
- Where OpenAI's new GPT-6.1 Sol lands on LiveBench.
- For Korean and other non-Chinese firms, data-governance questions about using DeepSeek models are separate from these scores.