TL;DR: A new paper argues recursive self-improvement isn't imminent because today's agents cannot carry out open-ended ML research.
Summary: A new arXiv paper tests whether frontier coding agents can independently reproduce the work of accepted-but-unpublished NeurIPS papers, with the original authors grading the results. The agents failed to replicate the research, and the authors use that failure to argue that current systems cannot perform open-ended ML research and therefore cannot recursively self-improve. The claim is contested by definition — it rests on a single benchmark design and a specific snapshot of agent capability.
Why it matters: It gives builders a concrete, graded benchmark for agent research autonomy rather than vibes-based claims about self-improvement timelines. Worth reading the task design before citing it: the strength of the conclusion depends entirely on how much scaffolding the agents were given.
Source: reddit