OSWorld benchmark reaches 65% by mid-2025
Specifically, we forecast that they score 65% on the OSWorld benchmark of basic computer tasks (compared to 38% for Operator and 70% for a typical skilled non-expert human).
The numerical milestone has supporting later evidence; achievement after the original mid-2025 window must not be read as an on-time forecast.
At a glance
- Assessment: Confirmed
- Confidence in assessment: 90%
- Outcome: observed
- Timing: late
- Evidence: direct
- Predicted timing: Mid 2025 (65%), Early 2026 (80%)
- Primary source: AI 2027, page 3, footnote 2 (Mid 2025: Stumbling Agents)
What AI 2027 Predicted
The scenario forecasted a rapid progression in computer-using AI agents, measured against the OSWorld benchmark — a suite of 369 real computer tasks involving web browsing, desktop applications, file management, and multi-app workflows in actual OS environments (Ubuntu, Windows, macOS). The authors predicted agents would reach 65% by mid-2025 (approaching but not matching a typical skilled non-expert human at 70%) and 80% by early 2026 (matching or exceeding that human baseline).
A second, higher target appears in the Early 2026 section: “Specifically, we predict a score of 80% on OSWorld (equivalent to a skilled but non-expert human).” At the time of writing (April 2025), OpenAI’s Operator scored 38%: setting the baseline for the prediction.
How We Track This
- OSWorld official leaderboard at os-world.github.io
- OSWorld-Verified results, recording the task version and evaluation changes
- Model release announcements from Anthropic, OpenAI, Google DeepMind with benchmark scores
- XLANG Lab (CMU) publications and leaderboard updates
Current Evidence
Mid-2025 target (65%): Reported crossed after the target window
The 65% threshold was crossed in late 2025. By early 2026, multiple frontier models surpass it:
- Claude Opus 4.6 (Feb 2026): 72.7% on OSWorld: a dated result, up from 66.3% for Opus 4.5
- CoACT-1: 60.76% (agentic framework approach)
- Historical framework and model results depend on scaffold, task version and evaluation setup; they are not a current leaderboard census.
An average across a changing collection of systems is not used to resolve the frontier threshold.
Early 2026 target (80%): Reached after the target window
Google’s July 2026 launch materials report Gemini 3.6 Flash at 83.0% on OSWorld-Verified, compared with 78.4% for Gemini 3.5 Flash. This is the first cited result on this page above the scenario’s 80% early-2026 target. The score is vendor-reported and arrived several months after that target window.
Scaffolding: Specialized agent systems can change benchmark performance. Comparisons should state the scaffold and task version rather than interpreting scores as properties of the underlying model alone.
Counterevidence & Limitations
- OSWorld scores are sensitive to agent scaffolding and prompting strategies — raw model capability vs. engineered system scores can differ significantly
- The benchmark was introduced in 2024; task difficulty calibration may shift as the community identifies easy vs. hard subsets
- Human baselines depend on participant selection, task version and instructions and need to be matched to the cited comparison.
- OSWorld and OSWorld-Verified results must identify task and evaluation revisions; a version name alone does not establish a uniformly stricter or easier benchmark.
What Would Change Our Assessment
- Maintain the reported milestone: Comparable evidence continues to support the numerical threshold, with vendor and independent results distinguished.
- Resolve timing: Record each threshold’s first qualifying date against its original window; a later result cannot establish an on-time success.
- Reassess comparability: Substantial task or scoring changes require an explicit bridge to the original benchmark, not silent recalibration.
Update History
| Date | Update |
|---|---|
| 2026-09-06 | Current evidence and assessment explanation reviewed. The numerical milestone has supporting later evidence; achievement after the original mid-2025 window must not be read as an on-time forecast. |
| 2026-07-27 | Google reported Gemini 3.6 Flash at 83.0% on OSWorld-Verified, crossing the scenario’s 80% early-2026 target several months late. Status and confidence remain unchanged because the 65% milestone was already confirmed and the new score is vendor-reported. |
| 2026-03 | GPT-5.4 scores 75.0% on OSWorld-Verified — first model to surpass human baseline. 80% target likely within months. Status: Confirmed for 65%; 80% target running ~3-6 months late. Confidence 0.90. |
| 2026-02 | Claude Opus 4.6 scores 72.7% — matching the human baseline (~72.4%). 80% not yet reached but progress accelerating. |
| 2026-01 | UiPath Screen Agent (Claude Opus 4.5) achieves #1 on OSWorld-Verified at 67.1%. Enterprise agents now match research agents. 80% target unmet as “early 2026” window opens. |
| 2025-12 | GPT-5.2 scores 47.3% on OSWorld-Verified — surprisingly below Claude’s 66.3%. 80% target still looks ambitious. |
| 2025-11 | Claude Opus 4.5 scores 66.3% — 65% target confirmed, approximately 5 months behind AI-2027’s mid-2025 timeline. AskUI VisionAgent independently confirms 65%+ is achievable (66.2% on OSWorld-Verified). Confidence rises for 80% target. |
| 2025-09 | Claude Sonnet 4.5 scores 61.4% on OSWorld. UiPath Screen Agent debuts at #2 on OSWorld-Verified. 65% target nearly met, ~3 months behind schedule. |
| 2025-08 | GPT-5 and Claude Opus 4.1 release (~44%). CoACT-1 breaks 60% (60.76%) via hybrid GUI+coding approach. 65% threshold now within striking distance. |
| 2025-07 | OSWorld-Verified launches with 300+ task fixes and stricter evaluation. GTA1 (Salesforce) reaches #1 at 45.2%. Mid-2025 65% target officially missed. |
| 2025-05 | Prediction begins tracking. Best production agent: OpenAI CUA at 38.1%. Research agents (PC Agent-E at 14.9%, ARPO at 29.9%) show progress but 65% mid-2025 target looks unlikely. |