AI research capability becomes vastly better than the best humans
A superintelligent AI researcher is vastly better than the best human researcher at AI research.
No qualifying replicated multi-subfield tenfold research advantage has been identified.
At a glance
- Assessment: Not Yet Testable
- Confidence in assessment: 90%
- Outcome: unresolved
- Timing: pending
- Evidence: insufficient
- Predicted timing: November 2027, racing branch
- Primary source: Appendix J, SIAR milestone; AI 2027 archived scenario
What AI 2027 Predicted
A superintelligent AI researcher is vastly better than the best human researcher at AI research. The source is Appendix J, SIAR milestone in the AI 2027 scenario.
Archived scenario version and section-page references. The linked live scenario may change; these definitions use the archived version identified here.
How We Track This
Extend the complete-project research protocol with at least six tasks across three subfields. Candidate operational threshold: independently judged outputs match or exceed the best human output on every task and achieve at least a 10× reduction in time-to-equivalent-quality on four of six tasks under disclosed, comparable experimental compute. Require independent replication and assessment of novel problem selection, not just execution.
These operational thresholds specify this tracker’s test. Where the scenario uses qualitative language, they are interpretations rather than numbers quoted from the original.
Current Evidence
Existing research-automation studies cover narrower settings than the complete-project comparison below. They inform the trajectory but do not resolve the racing-branch milestone.
- Automated Researchers Can Mitigate Well-Characterized Alignment Failures
- Long-horizon AI research evaluation
Counterevidence & Limitations
The 10× proxy is the tracker’s interpretation of “vastly better,” not an original number. Finite tasks cannot establish unrestricted superiority; show the operational test and broad-claim uncertainty separately. The November date belongs to the racing branch.
Missing qualifying public evidence does not establish that the event has not happened privately. The assessment concerns the documented test and its observable evidence.
What Would Change Our Assessment
A substantial research domain remains human-led, replication fails, or the advantage depends on unmatched resources or benchmark familiarity.
Evidence satisfying the full test would change the assessment; early qualifying evidence can be considered before the scenario window. Partial proxies are recorded with their scope.
Update History
| Date | Update |
|---|---|
| 2026-09-06 | No qualifying replicated multi-subfield tenfold research advantage has been identified. Relevant primary evidence and limitations are described above. |