AI teams conduct complete research projects with little human direction
AI collectives produce most research progress while most humans cease to contribute usefully; leading humans retain some research judgment.
Automated research is substantive in constrained settings; three qualifying complete projects have not been verified.
At a glance
- Assessment: Emerging
- Confidence in assessment: 80%
- Outcome: unresolved
- Timing: pending
- Evidence: proxy
- Predicted timing: June 2027
- Primary source: June 2027: Self-improving AI; AI 2027 archived scenario
What AI 2027 Predicted
AI collectives produce most research progress while most humans cease to contribute usefully; leading humans retain some research judgment. The source is June 2027: Self-improving AI in the AI 2027 scenario.
Archived scenario version and section-page references. The linked live scenario may change; these definitions use the archived version identified here.
How We Track This
Require three independently scoped machine-learning research projects lasting at least 14 calendar days each, covering question selection, experiment design, execution, interpretation and accepted output. Human interventions must account for no more than 10% of logged research decision time and must not supply the central hypothesis or conclusion. At least three of five blinded qualified reviewers judge each output technically valid and meaningfully novel relative to a frozen literature cutoff.
These operational thresholds specify this tracker’s test. Where the scenario uses qualitative language, they are interpretations rather than numbers quoted from the original.
Current Evidence
Automated alignment researchers improved well-characterized objectives over short research runs. Long-horizon research evaluations also reveal variable results and limited novelty. The required project duration, intervention records and independent judgments remain unresolved.
- Automated Researchers Can Mitigate Well-Characterized Alignment Failures
- Long-horizon AI research evaluation
OpenAI’s September 6 research account reports agents completing defined research tasks spanning several human workdays, with humans directing and judging the work. Anthropic’s September 4 formalization report describes an 11-day, largely autonomous Lean formalization of an existing proof of Fermat’s Last Theorem. It reports 13 million lines and 29,500 intermediate theorems used, with occasional high-level human instructions.
Google’s August 27 detailed Teamwork report, added here as earlier evidence, attributes seven mathematical results to Gemini 3.1 Pro, with three reproduced using Gemini 3.7 Flash. Google describes expert review and a Lean-verified result. These participant reports extend the evidence for coordinated research work. They do not establish three qualifying machine-learning projects with the required duration, intervention logs and blinded review. The mathematical artifacts have not been independently validated by this tracker.
OpenAI’s September 8 research release presents a 165-page proof and Lean formalization claiming finite-time blowup for forced three-dimensional Navier-Stokes. OpenAI says an internal system significantly more capable than Astra produced the result through a large parallel-agent effort. The Clay Mathematics Institute said on September 11 that the problem had “apparently been settled” and that its evaluation process would be unhurried. The Isaac Newton Institute described the result as an OpenAI claim. This is a substantial coordinated research artifact under external review, not an accepted result. It also falls outside this dossier’s test because it is not one of three machine-learning projects and the public intervention record does not establish the 10% human-decision ceiling.
Anthropic’s September 17 biomolecular-modeling report describes a general-purpose internal research model, supervised by two scientists, optimizing more than 30 open-source models in just under four weeks. Anthropic released the optimized code and reports output checks plus substantial inference-speed improvements. This is a multi-week research-engineering project with concrete artifacts. It does not satisfy this dossier’s complete-project test: humans supplied the objective and supervision, logged human decision time is not reported, and the required three independently scoped machine-learning projects and blinded novelty reviews are absent.
Counterevidence & Limitations
The 14-day and 10% rules are tracker thresholds. Logs must include failed experiments, compute budgets and human assistance. Parallel coding, narrow benchmark optimization or nominal elapsed time do not suffice. Passing this test establishes autonomous projects, not the scenario’s entire country-scale research economy.
Missing qualifying public evidence does not establish that the event has not happened privately. The assessment concerns the documented test and its observable evidence.
What Would Change Our Assessment
Humans provide decisive ideas or repairs, independent reviewers reject the work, or outputs cannot be reproduced.
Evidence satisfying the full test would change the assessment; early qualifying evidence can be considered before the scenario window. Partial proxies are recorded with their scope.
Update History
| Date | Update |
|---|---|
| 2026-09-21 | Anthropic reported a supervised, multi-week project in which Claude optimized more than 30 biomolecular models and released the code. The artifact is material but does not meet the three-project, human-intervention and blinded-review thresholds; status and confidence unchanged. |
| 2026-09-14 | OpenAI released a proposed Navier-Stokes result and Lean formalization produced through a large parallel-agent effort. Clay said the problem had apparently been settled but its evaluation remains pending. The artifact is material and does not satisfy the three-project machine-learning research test. Status and confidence unchanged. |
| 2026-09-07 | Added attributed lab reports of research agents, mathematical formalization and coordinated research. These extend project-level evidence without resolving the complete autonomous-research test. Status and confidence unchanged. |
| 2026-09-06 | Automated research is substantive in constrained settings; three qualifying complete projects have not been verified. Relevant primary evidence and limitations are described above. |