1,200 volunteers reproduced 2,226 ICML 2026 papers in 19 days; nearly a quarter had claims falsified
A Hugging Face hackathon used coding agents to reproduce a third of the papers accepted at ICML 2026, and found that 23 percent of the papers examined had at least one claim falsified or contested.
A hackathon organised by Hugging Face during July and early August had more than 1,200 community members use their own coding agents to reproduce papers accepted at ICML 2026, one of the field's largest machine-learning conferences. Over 19 days, participants published 6,816 reproduction logbooks covering 2,226 papers — about a third of the conference's accepted body.
ICML 2026 accepted 6,352 out of 23,918 submissions, roughly double the previous year. The organisers said this exponential growth is at least partly driven by AI agents making it faster to run experiments and write them up, while reviewing capacity has not doubled along with it — one accepted spotlight paper carried a reviewer note that read, 'My low confidence score is because I did not check all the proofs carefully.'
Participants used coding agents including Claude Code, Codex, Cursor and OpenResearch's orx. Every attempted reproduction produced a Trackio logbook that combined the write-up, code and artifacts. An automated judge then re-read every logbook and issued a per-claim verdict of verified, falsified, toy (evidence at reduced scale) or inconclusive. In all, 35,908 individual claims were judged.
Of the 2,226 examined papers, 51 percent (1,103) had at least one claim independently verified. Of those, 266 papers were fully reproduced, and 632 more were partially reproduced with nothing falsified — a total of 3,978 individual claims were confirmed with real experiments. Meanwhile, 23 percent (496) of examined papers had at least one claim falsified or contested; 49 of those had every claim falsified. Perhaps most notably, 242 papers had independent reproduction teams reach opposite verdicts on the same claims.
Where a full reproduction was impossible — for example, when a paper's dataset was proprietary or its checkpoints were unreleased — participants ran toy reproductions on synthetic data mimicking the original's properties. 502 papers ended up with toy-scale evidence only, and 280 papers had no evidence one way or the other, most often because artifacts were missing.
