Leading Chinese AI lab ~6 months behind US frontier

Last updated
They are about six months behind the best OpenBrain models.

A dated public-model reconstruction supports an approximately six-month mid-2026 gap. Later Chinese progress and task-specific competition remain important qualifications.

At a glance

  • Assessment: On Track
  • Confidence in assessment: 75%
  • Outcome: unresolved
  • Timing: Original window passed; outcome unresolved
  • Evidence: proxy
  • Predicted timing: Mid 2026
  • Primary source: ai-2027.com, Mid 2026: China Wakes Up

What AI 2027 Predicted

China’s leading models are approximately six months behind the US frontier around mid-2026. This is a dated, approximate comparison, not a permanent gap on every task. AI 2027

How We Track This

Use a common capability index to compare the Chinese frontier at the target date with earlier US frontier scores. Treat the resulting time bracket as our reconstruction, not a published confidence interval or a direct measure of internal systems.

Current Evidence

Epoch’s September 1 ECI table places the highest listed Chinese score available by June 30 at 153.55, Qwen 3.7-Max. This lies between US frontier scores of 153.06 on November 18 and 155.07 on December 11, 2025. The implied point-score bracket is 201-224 days, approximately 6.6-7.4 months.

By July 16, Kimi K3’s 157.33 falls between US scores dated February 5 and March 5: 133-161 days, approximately 4.4-5.3 months. The improvement does not erase the earlier dated comparison. Epoch ECI table

Independent task-specific research also finds a mixed picture. Artificial Analysis placed Kimi K3 close to leading US systems on its July intelligence and knowledge-work tests. A joint UK AISI/CAISI cyber study found a substantial gap on a longer operational task. Artificial Analysis · Joint government study

Counterevidence & Limitations

ECI scores have uncertainty and combine selected benchmarks. The brackets use point estimates; they do not include that full uncertainty. Public releases can lag internal capability. The September 1 table predates Astra’s September 3 release, so these calculations are not a September 7 frontier estimate.

The dated comparison supports the approximate scenario claim more directly than a single narrow benchmark. It leaves considerable uncertainty about a general, exact six-month lag.

What Would Change Our Assessment

  • Strengthen: Multiple comparable indices reproduce the midyear gap after accounting for measurement uncertainty.
  • Weaken: Broader contemporary evidence establishes substantially closer general capability in the specified window.
  • Update current context: New common evaluations incorporate recent releases while retaining the original midyear comparison.

Update History

DateUpdate
2026-09-07Reconstructed a midyear ECI comparison: approximately 6.6-7.4 months on point scores, narrowing by July. The approximate original gap receives affirmative support; internal capability and score uncertainty remain unresolved.
2026-09-06Assessment revised from behind (0.70) to emerging (0.65). The cited seven-month Epoch mean is historical, not a September 2026 gap estimate. A dated mid-2026 comparison remains unresolved.
2026-08-24Aikido’s same-harness comparison found DeepSeek V4 Pro leading tested public closed models on pooled recall across 32 fresh vulnerability-rediscovery tasks. This narrows the gap on one cyber workflow, but does not support a general months-based estimate. Status and confidence remain unchanged.
2026-07-20Moonshot AI released Kimi K3 and reported mixed head-to-head results against current US closed models. Moonshot says the model still trails the strongest proprietary systems overall, while leading on some agentic coding evaluations. Confidence and status remain unchanged because the results are vendor-reported and do not quantify the gap in months.
2026-06-22Z.ai released GLM-5.2 and reported large gains on coding and long-horizon task benchmarks, including 81.0 on Terminal-Bench versus 63.5 for GLM-5.1. This adds counterevidence to a stable six-month US model lead, while still requiring independent benchmark comparisons against US frontier systems. Confidence adjusted 0.65 -> 0.70.
2026-04-13DeepSeek V4 delayed twice, now expected late April 2026. First frontier model on Chinese domestic chips (Huawei Ascend 950PR), confirmed by Reuters (Apr 4) and AFP/HKFP reporting (HKFP, findskill.ai). Specs: ~1T params (MoE, ~37B active), 1M context, multimodal. Training cost reportedly ~$5.2M. However, the 16-month gap between V3 (Dec 2024) and V4 (late April 2026) vs US labs’ 2-4 month frontier cadence suggests the gap may be widening to 8-12 months rather than the predicted 6. Transition to Huawei chips required “substantial re-engineering” per Counterpoint Research, contributing to delays. Prediction remains behind — gap may be larger than the 6 months forecast.
2026-04-06DeepSeek V4 confirmed to run on Huawei chips — Reuters and The Information report (Apr 3) that the new model will operate on Huawei-designed chips, with Alibaba, ByteDance, and Tencent placing bulk orders for hundreds of thousands of Huawei AI chips ahead of launch (Reuters, The Information). This is the most concrete evidence yet of China building a self-sufficient AI hardware stack independent of Nvidia. Separately, MiniMax released M2.5 (80.2% SWE-bench) and M2.7 (56.22% on SWE-Pro, matching GPT-5.3-Codex) — Chinese models competitive on coding benchmarks. However, US models continue advancing: Gemini 3.1 Pro leads 13 of 16 benchmarks, Anthropic Mythos in testing. Net assessment: gap narrowing on standard benchmarks but may be widening on novel evaluations (ARC-AGI 2). No status change.
2026-03-30DeepSeek V4 Lite appeared on the DeepSeek web interface on March 9, 2026, with improved coding performance and a knowledge cutoff of May 2025 — but the full V4 remains unreleased despite multiple predicted launch windows passing (PromptZone). The full model (1T parameters, MoE, Huawei Ascend training) was expected in February but is still delayed. Leaked benchmarks (unverified) claim 80%+ SWE-bench and 90% HumanEval. Meanwhile, a mystery model “Hunter Alpha” was found in testing and initially attributed to DeepSeek but revealed to be Xiaomi’s model (Reuters). With US models (GPT-5.4, Anthropic Mythos in testing) advancing while China’s next major release slips, the gap assessment remains uncertain but may be widening slightly.
2026-03-16DeepSeek V4 withheld from US chipmakers (Nvidia, AMD), granting Huawei exclusive early access (Reuters, Feb 26). Signals growing US-China AI decoupling in model ecosystem. DeepSeek V4 benchmarks not yet public but preferential treatment of domestic hardware suggests strategic positioning. Meanwhile US frontier models (GPT-5.4, Opus 4.6, Gemini 3.1 Pro) pushed benchmarks further — gap assessment depends on V4 performance data. No status change pending V4 benchmarks.
2026-03Gap appears variable: 3 months on standard benchmarks, 8+ months on novel evaluations like ARC-AGI 2. The predicted ‘6 months’ may overstate Chinese capabilities on harder evaluations.
2025-12DeepSeek-V3 and Qwen models competitive on standard benchmarks, narrowing measured gap to ~3 months on some tasks.
2025-09DeepSeek reveals R1 training cost of $294,000 — a dramatic demonstration that China can train competitive models without U.S.-level compute. Alibaba’s Qwen3-Max (1T+ parameters, September 24) claims to outperform Claude and DeepSeek-V3.1 on certain agentic benchmarks. China’s frontier model development appears resilient under compute constraints.
2025-08DeepSeek releases upgraded V3.1 optimized for domestic chips using UE8M0 FP8 precision format, signaling active adaptation to export control constraints. Combined with June 2025 METR finding of ~6-month capability gap, China’s strategy appears to be efficiency-focused rather than compute-scaling-focused.
2025-06METR publishes evaluation of mid-2025 DeepSeek and Qwen models (June 27): “autonomous capabilities of mid-2025 DeepSeek models similar to capabilities of frontier models from late 2024.” METR data confirms approximately 6-month gap on agentic tasks — real but not as large as some export-control-based forecasts assumed.
2025-04Alibaba releases Qwen3 (April 28): 235B-A22B MoE model with benchmark scores competitive with some closed frontier models. Apache 2.0. Suggests China’s open-source AI ecosystem is not standing still under compute constraints.