Coding fully automated but research taste still requires humans

Last updated
Now that coding has been fully automated... research taste has proven difficult to train due to longer feedback loops and less data availability.

Autonomous implementation is advancing while difficult research judgments remain a bottleneck in tested settings. The full March 2027 combination remains unestablished.

At a glance

  • Assessment: Emerging
  • Confidence in assessment: 80%
  • Outcome: unresolved
  • Timing: pending
  • Evidence: proxy
  • Predicted timing: Early-Mid 2027
  • Primary source: ai-2027.com, March 2027: Algorithmic Breakthroughs

What AI 2027 Predicted

By March 2027, coding is fully automated while research taste and strategic judgment remain important human contributions. The scenario does not require AI to have no useful research judgment. AI 2027

How We Track This

Assess implementation autonomy and research judgment separately. A coding benchmark cannot by itself establish sound problem selection, and success on a well-scored research task need not transfer to open-ended inquiry.

Current Evidence

Anthropic’s August 28 TASTE study compared judgments on 92 pairs of AI-safety proposals. The best model agreed with researcher preferences about 60% of the time, versus an estimated 77% for researchers. Model uncertainty was approximately ±10 percentage points. This identifies a judgment gap in one curated setting. TASTE study

An independent research study ran two six-day shadow evaluations. Agents completed substantial engineering but failed to answer the central research questions; the original researchers rejected the outputs. That supports a distinction between implementation and research success, with a very small sample. Kirgis et al.

Counterevidence & Limitations

Automated researchers also produce useful methods. Anthropic tested ten well-characterized alignment failures and found automated method search could outperform one-shot proposals from 28 researchers, including on held-out evaluations. Unequal iteration budgets and measurable objectives limit broader comparisons. Automated alignment researchers

Preference agreement is not an objective measure of scientific value. Two difficult research projects cannot characterize every field. Current evidence supports an emerging division of labor, but neither comprehensive coding automation nor a durable human advantage in taste is established.

What Would Change Our Assessment

  • Strengthen: Independent longitudinal evaluations show reliable autonomous implementation alongside persistent strategic research failures.
  • Confirm: Both components are demonstrated at the specified March 2027 scope.
  • Weaken: Agents reliably choose and execute valuable open-ended research, or coding itself remains a substantial human bottleneck.

Update History

DateUpdate
2026-09-07Added proposal-judgment and open-ended research evaluations, balanced against successful automated alignment-method search. The evidence supports a developing implementation/judgment split, not the full 2027 conjunction.
2026-09-06Assessment revised from not-yet-testable (0.55) to emerging (0.80). Coding automation and research-judgment limitations have supporting evidence in particular settings; full coding automation and a universal judgment barrier remain unestablished.
2026-09-06Removed the inference from developers’ reluctance to work without AI to full coding automation. METR’s study update does not establish that milestone. Status and confidence were unchanged.
2026-08-31Two long-horizon AI-R&D evaluations found strong implementation and experiment execution but persistent limits in strategy revision, novelty, stability, and experience reuse. The larger trajectory analysis covered 1,338 post-training runs. This strengthens early evidence for a research-judgment bottleneck, while full coding automation remains unproven. Status and confidence remain unchanged.
2026-08-03Two six-day shadow evaluations found frontier agents completed all engineering without human help but made no substantial progress on the central research questions. This supports the predicted research-taste bottleneck, while the two-case sample and absence of full coding automation keep the prediction not yet testable.
2026-03Prediction timeframe not yet reached. Coding automation progressing rapidly but far from complete. Early signals support the predicted asymmetry — routine coding tasks automate faster than research taste and architectural judgment.
2026-02METR notes developers refuse to work without AI (coding automated). METR time horizons showed Opus 4.6 making a major jump on general software tasks, but RE-Bench-specific research engineering scores still lag. Gap narrows but persists.
2025-11Claude Code at $1B ARR confirms massive coding automation adoption. SWE-bench crosses 80%. But RE-Bench (research engineering) scores remain 0.5-0.8 — coding ahead, research taste behind.
2025-08Anthropic internal study: engineers can only “fully delegate” 0-20% of work. Task complexity rising (3.2→3.8). Coding increasingly automated; judgment and design remain human.