Skip to content
TEN Brief Ten verified stories a day 2026.10.06 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 2 min read · Breaking

DeepSeek V4.1 Flash scores 81.1 on LiveBench — US-China AI gap hits 3%

DeepSeek V4.1 Flash, released in September 2026, scored 81.1 on the LiveBench leaderboard as of October 4, against 83.4 for Anthropic's Claude Fable 5.1, the top US model. That 2.3-point, roughly 3% gap is the smallest Bloomberg Intelligence has tracked between the leading US and Chinese models; it was about 15% early in 2026 and about 9% in May. On LiveBench's agentic coding category DeepSeek scored 77.3 against Claude's 66.1. Analyst Robert Lea credits efficiency work and tuning for Chinese hardware. Only three of the top 15 LiveBench entries are Chinese, so the catch-up is concentrated at the very top

Two runners a stride apart on a sunlit athletics track, seen from behind

The three lines

  • Scores — DeepSeek V4.1 Flash 81.1 vs Claude Fable 5.1 83.4 on LiveBench; a 2.3-point (~3%) gap
  • Trend — the gap shrank from ~15% early in 2026 to ~9% in May to ~3%; DeepSeek leads agentic coding 77.3 to 66.1
  • Limits — only 3 Chinese models in the top 15; Bloomberg sees China's AI sector unprofitable until 2030

Key questions

How good is DeepSeek V4.1 Flash
**81.1 overall on LiveBench, 2.3 points behind the top US model.** | Model | Overall | Agentic coding | |---|---|---| | Claude Fable 5.1 (max effort) | **83.4** | 66.1 | | DeepSeek V4.1 Flash (max effort) | 81.1 | **77.3** | | Gap | 2.3 (~3%) | DeepSeek +11.2 |
How big is the US China AI gap
**About 3% at the top, per Bloomberg Intelligence — a record low.** | Time | Top-model gap | |---|---| | Early 2026 | ~15% | | May 2026 | ~9% | | Oct 2026 | **~3%** |
What is LiveBench
**An AI benchmark that keeps replacing its questions to limit memorized answers.** | Feature | Detail | |---|---| | Questions | Built from recent contests, papers and news | | Updates | Questions refreshed regularly | | Scoring | Automatic, against fixed answers | | Aim | Reduce training-data contamination |

The distance between the best American and Chinese AI models is now a single stride. According to Bloomberg Intelligence, DeepSeek's V4.1 Flash, released in September, scored 81.1 on the LiveBench leaderboard — 2.3 points behind Anthropic's Claude Fable 5.1 at 83.4. That is the narrowest gap Bloomberg has recorded.

1. Nine months, 15% to 3%

TimeUS-China top-model gapNote
Early 2026~15%—
May 2026~9%—
Sept 2026—DeepSeek V4.1 Flash released
Oct 4, 2026~3%81.1 vs 83.4

The category detail is starker. On agentic coding — models fixing and running code on their own to finish a task — DeepSeek scored 77.3 to Claude's 66.1. That is the capability US labs are betting most heavily on.

The "Flash" name matters too. DeepSeek's paper focuses on compressing the KV cache, the memory a model uses to track a conversation, and using only part of its parameters per token. The gap closed with a fast, lean model, not a giant one.

2. Why it narrowed

Bloomberg Intelligence senior analyst Robert Lea points to better engineering at Chinese labs and optimization for domestic hardware rather than US GPUs. He says that raises questions about how much US chip export controls are slowing China. Usage data already pointed the same way: in late September Chinese models took 57–67% of token traffic on the OpenRouter marketplace, mostly on price. Now the benchmark numbers are close as well.

3. What not to overclaim

SignalReading
3 Chinese models in LiveBench top 15The catch-up is concentrated in a few frontier models
Benchmark scoresThey do not prove national AI leadership
Enterprise useMarketplace data under-counts big-company workloads, where US models still dominate
ProfitabilityLea expects China's AI industry to stay unprofitable until 2030, with 1,100+ large models competing on price

"China has caught up" overstates it. "The top one or two models now score almost the same on one test" is accurate.

4. What remains unclear

  • Whether other leaderboards show the same gap.
  • Where OpenAI's new GPT-6.1 Sol lands on LiveBench.
  • For Korean and other non-Chinese firms, data-governance questions about using DeepSeek models are separate from these scores.

Sources

  1. Implicator — DeepSeek Cuts US AI Benchmark Lead to About 3%
  2. AI Weekly — DeepSeek V4.1 Flash narrows US-China AI gap to 3% on LiveBench
  3. Startup Fortune — DeepSeek Narrows AI Gap With US to Just 3 Percent, Bloomberg Says
  4. arXiv — DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Verification

Published
Last modified
Cross-check
Checked against 4 independent sources.
Unverified
  • The ~3% figure is Bloomberg Intelligence's calculation from LiveBench; we did not confirm the gap on other leaderboards such as LMArena.
  • LiveBench scores for OpenAI's GPT-6.1 Sol and GPT-6 Astra on the same date were not confirmed.
  • DeepSeek V4.1 Flash API pricing was not verified for this article.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related