AI substantially outperforms the best humans across cognitive domains

Last updated
Artificial superintelligence is much better than the best human at every cognitive task.

The broad cognitive-domain test remains unresolved; a collection of benchmark wins is insufficient.

At a glance

  • Assessment: Not Yet Testable
  • Confidence in assessment: 90%
  • Outcome: unresolved
  • Timing: pending
  • Evidence: insufficient
  • Predicted timing: December 2027, racing branch
  • Primary source: Appendix J, ASI milestone; December 2027 racing branch; AI 2027 archived scenario

What AI 2027 Predicted

Artificial superintelligence is much better than the best human at every cognitive task. The source is Appendix J, ASI milestone; December 2027 racing branch in the AI 2027 scenario.

Archived scenario version and section-page references. The linked live scenario may change; these definitions use the archived version identified here.

How We Track This

Use a frozen domain matrix: AI research, software engineering, mathematics, scientific reasoning, strategic planning, communication and practical knowledge work. For each, require at least three unfamiliar complete tasks with leading-human comparators, blinded quality evaluation and disclosed resources. Candidate proxy: better quality on all domain aggregates and at least 2× faster time to equivalent quality in each domain, independently replicated. Record any material counterexample.

These operational thresholds specify this tracker’s test. Where the scenario uses qualitative language, they are interpretations rather than numbers quoted from the original.

Current Evidence

The available research evidence measures particular tasks and workflows. It does not provide the frozen, independently replicated domain matrix needed for this assessment. This is a limitation of the available evidence, not proof that a hidden system lacks the capability.

Counterevidence & Limitations

No finite matrix proves “every cognitive task.” Passing supports broad superhuman capability under the tested scope; it does not conclusively establish the universal claim. Physical embodiment, affordability and deployment are separate qualifiers. The date is branch-conditional.

Missing qualifying public evidence does not establish that the event has not happened privately. The assessment concerns the documented test and its observable evidence.

What Would Change Our Assessment

At least one substantial cognitive domain has reproducible leading-human advantages, or gains fail on unfamiliar tasks.

Evidence satisfying the full test would change the assessment; early qualifying evidence can be considered before the scenario window. Partial proxies are recorded with their scope.

Update History

DateUpdate
2026-09-06The broad cognitive-domain test remains unresolved; a collection of benchmark wins is insufficient. Relevant primary evidence and limitations are described above.