AI research capability exceeds the strongest human researchers

Last updated
An individual Agent-4 copy operating at human speed is qualitatively better at AI research than any human.

No qualifying broad, independently replicated complete-project comparison has been identified.

At a glance

  • Assessment: Not Yet Testable
  • Confidence in assessment: 85%
  • Outcome: unresolved
  • Timing: pending
  • Evidence: insufficient
  • Predicted timing: September 2027
  • Primary source: September 2027: Agent-4, the Superhuman AI Researcher; Appendix J; AI 2027 archived scenario

What AI 2027 Predicted

An individual Agent-4 copy operating at human speed is qualitatively better at AI research than any human. The source is September 2027: Agent-4, the Superhuman AI Researcher; Appendix J in the AI 2027 scenario.

Archived scenario version and section-page references. The linked live scenario may change; these definitions use the archived version identified here.

How We Track This

For a strong test, preregister at least six complete research tasks across at least three ML subfields and recruit at least three independently credentialed leading human researchers. Match experimental compute, access and wall-clock budgets; report inference resources separately. Use blinded external judges for validity, novelty and importance, include failed tasks, and require the system to outperform the best participating human on at least four of six tasks without a systematic deficit in one subfield.

These operational thresholds specify this tracker’s test. Where the scenario uses qualitative language, they are interpretations rather than numbers quoted from the original.

Current Evidence

Narrow researcher-baseline wins coexist with gaps in judging research proposals. TASTE measures agreement with a constructed human preference reference, not objective research success. These results cannot establish superiority to the strongest humans across complete research programs.

Counterevidence & Limitations

A finite contest cannot prove “better than any human.” Treat success as strong operational evidence, keeping that limitation in the headline assessment. To mark the broad claim confirmed requires convergent independently replicated evidence across complete research programs, not this threshold alone. Narrative deadline is September; alternate supplement milestones are tracked separately.

Missing qualifying public evidence does not establish that the event has not happened privately. The assessment concerns the documented test and its observable evidence.

What Would Change Our Assessment

Top humans remain superior on research selection, novel insights, error detection or reproducible complete projects; wins depend on unmatched compute.

Evidence satisfying the full test would change the assessment; early qualifying evidence can be considered before the scenario window. Partial proxies are recorded with their scope.

Update History

DateUpdate
2026-09-06No qualifying broad, independently replicated complete-project comparison has been identified. Relevant primary evidence and limitations are described above.