IDA achieves superhuman performance at coding

Last updated
Emerging · Model Capability · 40% confidence in assessment
Predicted: Early 2027 · Updated: 2026-09-06 · Source: ai-2027.com, Appendix F (pages 49-50): IDA achieving superhuman performance at coding specifically
Now, the models have become sufficiently good at verifying more subjective things, allowing the use of IDA to improve the model at many tasks.

At a glance

  • Assessment: Emerging
  • Confidence in assessment: 40%
  • Predicted timing: Early 2027
  • Primary source: ai-2027.com, Appendix F (pages 49-50): IDA achieving superhuman performance at coding specifically

What AI 2027 Predicted

The scenario describes Iterated Distillation and Amplification (IDA) being deployed at scale around early 2027. IDA, originally proposed by Paul Christiano, involves a cycle: amplify a model’s capability through careful decomposition and verification, then distill the amplified capability back into the model. AI 2027 predicts that by 2027, models become good enough at verifying “subjective” outputs — not just math proofs or game outcomes — enabling IDA to improve performance across many real-world tasks. This is framed as one of several algorithmic breakthroughs that accelerate the path to superhuman AI.

How We Track This

We monitor:

  • Self-play and self-improvement methods in frontier model training
  • RL-based training with model-generated verification (constitutional AI, RLAIF)
  • AlphaProof-style methods applied beyond formal mathematics
  • Lab announcements about recursive self-improvement in training pipelines
  • Academic work on IDA, debate, and amplification approaches

Current Evidence

Related AI-training techniques supply precursor evidence, but the reviewed sources do not establish the specified broad IDA deployment.

AlphaProof: Reinforcement learning with formal proof verification demonstrates improvement in a well-verifiable domain. It should not automatically be classified as the decomposition, amplification and distillation process specified here.

Constitutional AI and RLAIF: Training from AI feedback is relevant, but the use of model judgments alone does not establish the full IDA cycle or its broad deployment.

RL post-training: Reinforcement learning can improve model capability. RL, distillation and iterated amplification are distinct mechanisms and require separate technical evidence.

The gap: The claim concerns scaling the specified mechanism to harder-to-verify outputs. Research on related techniques does not by itself resolve that claim.

Sources:

Counterevidence & Limitations

  • Results in formally verifiable domains cannot establish broad performance on subjective tasks.
  • The reviewed sources do not establish the specified subjective-verification breakthrough; they also do not prove there is no viable path.
  • Self-play / RL-from-model-feedback can lead to reward hacking and mode collapse, limiting its generality
  • It’s unclear whether constitutional AI / RLAIF truly counts as IDA or is a more limited form of the concept
  • The mechanism and domain coverage must be explicit; a generic self-improvement label is insufficient.

What Would Change Our Assessment

  • Strengthen the assessment: Technical evidence demonstrates the specified amplification-and-distillation cycle on open-ended tasks.
  • Confirm: Production use establishes that mechanism and broad task coverage, rather than generic iterative improvement.
  • Reassess timing after early 2027: Keep the original window while distinguishing missing evidence from a demonstrated failure.
  • Watch for: Better verification of subjective quality, with generalization tested separately.

Update History

DateUpdate
2026-09-06Current evidence and assessment explanation reviewed. Clarified measurement scope, source interpretation and limitations.
2026-03-13IDA-adjacent techniques (AlphaProof for math, RLAIF, RL post-training) are widespread but limited to domains with clear verifiers. The predicted ‘subjective verification’ breakthrough — enabling IDA for open-ended tasks — remains undemonstrated.