Leading humans still add substantial value to AI-led research

Last updated
Most humans contribute little, but the strongest retain research taste and planning value; the appendix estimates removing humans would slow research by roughly half.

The marginal contribution of humans to whole-project research speed remains unmeasured under this test.

At a glance

  • Assessment: Not Yet Testable
  • Confidence in assessment: 80%
  • Outcome: unresolved
  • Timing: pending
  • Evidence: insufficient
  • Predicted timing: June 2027
  • Primary source: June 2027: Self-improving AI; Appendix I; AI 2027 archived scenario

What AI 2027 Predicted

Most humans contribute little, but the strongest retain research taste and planning value; the appendix estimates removing humans would slow research by roughly half. The source is June 2027: Self-improving AI; Appendix I in the AI 2027 scenario.

Archived scenario version and section-page references. The linked live scenario may change; these definitions use the archived version identified here.

How We Track This

Compare matched AI-led complete projects with versus without expert human input over at least three projects, holding model, tools and experiment compute comparable. Primary measure: accepted progress per calendar time. A 40-60% reduction on removal operationalizes the “roughly half” clause; report the estimate and interval, intervention logs and the distribution across experts. Separately assess whether most participating humans add little value.

These operational thresholds specify this tracker’s test. Where the scenario uses qualitative language, they are interpretations rather than numbers quoted from the original.

Current Evidence

Proposal-judgment evidence supports some human advantages; automated methods outperform human suggestions in other constrained settings. Neither quantifies the change in total research progress after removing expert input.

Counterevidence & Limitations

The 40-60% band is a proposed interpretation. Evidence that one human helps is not evidence that most do not. Keep marginal contribution separate from the total R&D multiplier and from fully autonomous project completion.

Missing qualifying public evidence does not establish that the event has not happened privately. The assessment concerns the documented test and its observable evidence.

What Would Change Our Assessment

Removing experts has little effect, or humans remain indispensable across the entire research workflow rather than a narrow research-taste role.

Evidence satisfying the full test would change the assessment; early qualifying evidence can be considered before the scenario window. Partial proxies are recorded with their scope.

Update History

DateUpdate
2026-09-06The marginal contribution of humans to whole-project research speed remains unmeasured under this test. Relevant primary evidence and limitations are described above.