OSWorld benchmark reaches 65% by mid-2025

Last updated
Specifically, we forecast that they score 65% on the OSWorld benchmark of basic computer tasks (compared to 38% for Operator and 70% for a typical skilled non-expert human).

The numerical milestone has supporting later evidence; achievement after the original mid-2025 window must not be read as an on-time forecast.

At a glance

  • Assessment: Confirmed
  • Confidence in assessment: 90%
  • Outcome: observed
  • Timing: late
  • Evidence: direct
  • Predicted timing: Mid 2025 (65%), Early 2026 (80%)
  • Primary source: AI 2027, page 3, footnote 2 (Mid 2025: Stumbling Agents)

What AI 2027 Predicted

The scenario forecasted a rapid progression in computer-using AI agents, measured against the OSWorld benchmark — a suite of 369 real computer tasks involving web browsing, desktop applications, file management, and multi-app workflows in actual OS environments (Ubuntu, Windows, macOS). The authors predicted agents would reach 65% by mid-2025 (approaching but not matching a typical skilled non-expert human at 70%) and 80% by early 2026 (matching or exceeding that human baseline).

A second, higher target appears in the Early 2026 section: “Specifically, we predict a score of 80% on OSWorld (equivalent to a skilled but non-expert human).” At the time of writing (April 2025), OpenAI’s Operator scored 38%: setting the baseline for the prediction.

How We Track This

  • OSWorld official leaderboard at os-world.github.io
  • OSWorld-Verified results, recording the task version and evaluation changes
  • Model release announcements from Anthropic, OpenAI, Google DeepMind with benchmark scores
  • XLANG Lab (CMU) publications and leaderboard updates

Current Evidence

Mid-2025 target (65%): Reported crossed after the target window

The 65% threshold was crossed in late 2025. By early 2026, multiple frontier models surpass it:

  • Claude Opus 4.6 (Feb 2026): 72.7% on OSWorld: a dated result, up from 66.3% for Opus 4.5
  • CoACT-1: 60.76% (agentic framework approach)
  • Historical framework and model results depend on scaffold, task version and evaluation setup; they are not a current leaderboard census.

An average across a changing collection of systems is not used to resolve the frontier threshold.

Early 2026 target (80%): Reached after the target window

Google’s July 2026 launch materials report Gemini 3.6 Flash at 83.0% on OSWorld-Verified, compared with 78.4% for Gemini 3.5 Flash. This is the first cited result on this page above the scenario’s 80% early-2026 target. The score is vendor-reported and arrived several months after that target window.

Scaffolding: Specialized agent systems can change benchmark performance. Comparisons should state the scaffold and task version rather than interpreting scores as properties of the underlying model alone.

Counterevidence & Limitations

  • OSWorld scores are sensitive to agent scaffolding and prompting strategies — raw model capability vs. engineered system scores can differ significantly
  • The benchmark was introduced in 2024; task difficulty calibration may shift as the community identifies easy vs. hard subsets
  • Human baselines depend on participant selection, task version and instructions and need to be matched to the cited comparison.
  • OSWorld and OSWorld-Verified results must identify task and evaluation revisions; a version name alone does not establish a uniformly stricter or easier benchmark.

What Would Change Our Assessment

  • Maintain the reported milestone: Comparable evidence continues to support the numerical threshold, with vendor and independent results distinguished.
  • Resolve timing: Record each threshold’s first qualifying date against its original window; a later result cannot establish an on-time success.
  • Reassess comparability: Substantial task or scoring changes require an explicit bridge to the original benchmark, not silent recalibration.

Update History

DateUpdate
2026-09-06Current evidence and assessment explanation reviewed. The numerical milestone has supporting later evidence; achievement after the original mid-2025 window must not be read as an on-time forecast.
2026-07-27Google reported Gemini 3.6 Flash at 83.0% on OSWorld-Verified, crossing the scenario’s 80% early-2026 target several months late. Status and confidence remain unchanged because the 65% milestone was already confirmed and the new score is vendor-reported.
2026-03GPT-5.4 scores 75.0% on OSWorld-Verified — first model to surpass human baseline. 80% target likely within months. Status: Confirmed for 65%; 80% target running ~3-6 months late. Confidence 0.90.
2026-02Claude Opus 4.6 scores 72.7% — matching the human baseline (~72.4%). 80% not yet reached but progress accelerating.
2026-01UiPath Screen Agent (Claude Opus 4.5) achieves #1 on OSWorld-Verified at 67.1%. Enterprise agents now match research agents. 80% target unmet as “early 2026” window opens.
2025-12GPT-5.2 scores 47.3% on OSWorld-Verified — surprisingly below Claude’s 66.3%. 80% target still looks ambitious.
2025-11Claude Opus 4.5 scores 66.3% — 65% target confirmed, approximately 5 months behind AI-2027’s mid-2025 timeline. AskUI VisionAgent independently confirms 65%+ is achievable (66.2% on OSWorld-Verified). Confidence rises for 80% target.
2025-09Claude Sonnet 4.5 scores 61.4% on OSWorld. UiPath Screen Agent debuts at #2 on OSWorld-Verified. 65% target nearly met, ~3 months behind schedule.
2025-08GPT-5 and Claude Opus 4.1 release (~44%). CoACT-1 breaks 60% (60.76%) via hybrid GUI+coding approach. 65% threshold now within striking distance.
2025-07OSWorld-Verified launches with 300+ task fixes and stricter evaluation. GTA1 (Salesforce) reaches #1 at 45.2%. Mid-2025 65% target officially missed.
2025-05Prediction begins tracking. Best production agent: OpenAI CUA at 38.1%. Research agents (PC Agent-E at 14.9%, ARPO at 29.9%) show progress but 65% mid-2025 target looks unlikely.