SWE-bench-Verified score reaches 85%

Last updated
We forecast that mid-2025 agents will score 85% on SWEBench-Verified. (page 3, footnote 3. Note: the 72% starting point is tracker-added context, not from the source.)

The mid-2025 deadline was missed. Anthropic documented a later 87.6% result in April 2026; Epoch’s separate 484-task setup returned 83.47%.

At a glance

  • Assessment: Behind
  • Confidence in assessment: 90%
  • Outcome: observed
  • Timing: late
  • Evidence: direct
  • Predicted timing: Mid 2025
  • Primary source: ai-2027.com, Late 2025 section; AI Futures grading

What AI 2027 Predicted

The benchmark prediction was 85% on SWE-bench Verified by mid-2025. Later achievement does not restore the original timing. AI Futures self-grading

How We Track This

Record the model, publication date, task set and evaluation configuration. A documented vendor evaluation can establish a reported crossing without being an independent replication. Different scaffolds and task subsets must remain distinguishable.

Current Evidence

The authors’ self-grading records 74.5% for the mid-2025 comparison and 81% by the end of 2025. The 74.5% figure was not a February 2026 frontier score.

A later crossing is documented. Anthropic’s April 16, 2026 Opus 4.7 release and system card report 87.6% on the 500-problem SWE-bench Verified subset, averaged over five trials. The standard configuration uses adaptive thinking at maximum effort. The relevant section does not provide a complete Verified-specific scaffold and token budget. This is a primary vendor result. Dated release · System card, pages 191-192

Epoch’s independently run benchmark data records 83.47% for Opus 4.7 at maximum effort, from a run beginning April 20. Its denominator is 484 validated tasks, with 404 solved. Epoch dataset · Evaluation methodology

Counterevidence & Limitations

The results establish substantial progress and sensitivity to evaluation setup. They are not directly interchangeable. Neither a vendor result nor passing benchmark tests establishes acceptance of every change in real software development.

Behind describes the original deadline miss. The outcome is recorded as observed late, qualified by the vendor-reported configuration; independent replication at the same score remains separate.

What Would Change Our Assessment

  • Strengthen later achievement: Independent evaluation reproduces the crossing under a fully documented comparable setup.
  • Reassess timing: Credible contemporary records establish a qualifying result by mid-2025.
  • Weaken the later result: Evaluation errors or task leakage materially invalidate the reported score.

Update History

DateUpdate
2026-09-07Verified the primary Opus 4.7 system card: 87.6% averaged over five trials, published April 16, 2026. Outcome observed late; the missed mid-2025 deadline and Epoch’s different 83.47% setup remain explicit.
2026-09-06Assessment revised from on-track (0.85) to behind (0.85). The original mid-2025 timing was missed. Later vendor reports require a comparable standardized evaluation before resolving the numerical threshold.
2026-07-14OpenAI estimated that about 30% of SWE-Bench Pro’s public tasks are broken and retracted its recommendation to use the benchmark. Confidence adjusted from 0.90 to 0.85 because measurement uncertainty increased, while status remains on-track.
2026-05-05The May 4 LLM Stats evidence was judged sufficient to move the prediction from behind to on-track, while keeping the methodology caveat intact: the 85% numerical threshold appears crossed, but roughly 10–11 months late and not yet under standardized Epoch verification. Confidence unchanged at 0.90.
2026-05-04LLM Stats listed Claude Opus 4.7 at 87.6% and Claude Mythos Preview at 93.9% on SWE-bench Verified, apparently crossing the 85% numerical threshold roughly 10–11 months late. Because LLM Stats labels these as self-reported rather than independently verified results, methodology comparability remains important before treating the result as fully confirmed. Confidence adjusted 0.85 → 0.90.
2026-04-27Anthropic released Claude Opus 4.7 with claimed improvements over Opus 4.6 across SWE-bench Verified, SWE-bench Pro, and related coding evaluations. At this update, the assessment had not established a comparable public SWE-bench Verified score above the 85% threshold. Anthropic also noted some internal implementations are not directly comparable to public leaderboard scores (Anthropic). Status remains behind pending a comparable public benchmark.
2026-04-06Leaderboard data as of April 4: Claude Opus 4.5 still leads at 80.9% SWE-bench Verified, followed by Opus 4.6 at 80.8%, Gemini 3.1 Pro at 80.6%, MiniMax M2.5 at 80.2%, and GPT-5.2 at 80.0% (llm-stats, marc0.dev). Plateau around 80% continues — no model has breached 81% on the standard benchmark. MiniMax M2.7 just released (Apr 6) claims 56.22% on SWE-Pro matching GPT-5.3-Codex, and leads the new MM-ClawBench at 62.7% (MiniMax, benchlm.ai). The competitive Chinese model landscape is compressing the frontier from below. METR study continues to cast doubt on benchmark significance: ~50% of SWE-bench-passing PRs would not be merged by maintainers. The 85% target remains ~4 points away — achievable but persistent. No status change.
2026-03-30GPT-5.4 released March 5 (Thinking and Pro variants) — METR SWE-bench evaluation pending. Unverified DeepSeek V4 leaked benchmarks claim 81% SWE-bench but have not been independently confirmed (nxcode.io). Anthropic’s leaked Mythos/Capybara announcement states it “gets dramatically higher scores on tests of software coding” vs. Opus 4.6 — suggesting the 85% target may be within reach of an imminent release. The bar of 85% now appears plausibly achievable within Q2 2026, ~12 months past the original mid-2025 deadline. No status change; confidence unchanged as the target remains unmet by a verified, publicly released model.
2026-03-23Multiple sources confirm top SWE-bench Verified scores: Claude Opus 4.5 at ~80.9%, Opus 4.6 at ~80.8%, MiniMax M2.5 at 80.2% (Morphllm, llm-stats). However, METR study (Mar 10) found ~50% of SWE-bench-passing PRs would not actually be merged by maintainers, suggesting the benchmark overstates real-world coding capability (METR). Scale AI’s harder SWE-Bench Pro shows best models at only ~57% (GPT-5.4). The 85% SWE-bench Verified target appears reachable but its significance as a proxy for coding capability is increasingly questioned. No status or confidence change.
2026-03-16New benchmarks show Opus 4.6 at 80.8% SWE-Bench Verified, Gemini 3.1 Pro at ~80.6%. Gap to 85% target closing but still not met, now ~11 months past the predicted mid-2025 deadline. Confidence adjusted 0.80 → 0.85 as 85% appears reachable within months.
2026-03Best score reaches ~74.5%, significantly behind the 85% target. Progress slower than predicted despite substantial investment in coding agent capabilities.
2026-02AI Futures grading confirmed: best SWE-bench Verified score 74.5% (Opus 4.1 without extended thinking), “surprisingly slow” vs. 85% mid-2025 target. Epoch AI methodology upgrade (v2.0.0) improved scaffolding but top models cluster 70–75% on standardized eval.
2026-01AI Futures Project clarification post confirmed SWE-bench progress was one of the primary factors in revised (longer) timelines. Daniel Kokotajlo’s median for full coding automation: December 2030.
2025-12AI Futures Project Dec 2025 model update assessed SWE-bench trajectory as behind pace. The best non-extended-thinking score (74.5% from Claude Opus 4.1) used as the cleanest comparable. Status: behind (confidence: 0.45).
2025-11Rapid convergence: GPT-5.1-Codex-Max hits 77.9%, Gemini 3 hits 76.2%, and Claude Opus 4.5 hits 80.9% — the first model above 80% — all within 6 days. The 85% target is now clearly achievable in early 2026, roughly 6-9 months late relative to the AI 2027 prediction.
2025-09Claude Sonnet 4.5 reaches 77.2% (September 29) — progress is real but slow. Scale AI’s SWE-Bench Pro (September 19) shows models dropping to ~23% on harder long-horizon tasks, raising questions about whether the original benchmark’s 85% threshold remains a meaningful milestone.
2025-08Claude Opus 4.1 reaches 74.5% (August 5). The mid-2025 target of 85% appears approximately 6 months behind pace. No model from any lab has publicly disclosed surpassing 75% at this point.
2025-06AI 2027 predicted 85% SWE-bench-Verified by mid-2025. Best score at that time was ~72%, establishing a gap.
2025-05Claude Opus 4 scores 72.5% on SWE-bench Verified (May 22). Compared to Claude 3.7 Sonnet’s 70.3% in February, pace is slower than needed to reach 85% by mid-2025. The prediction is tracking behind schedule.