ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
… In this paper, we present the first systematic study of benchmark contamination in LRMs, structured around two points where contamination can happen. In particular, Stage I (pre-LRM) …
Benchmark contamination—where test data leaks into training data—has long plagued machine learning evaluation, but its impact on large reasoning models (LRMs) is particularly insidious. These models, designed to perform multi-step logical and mathematical reasoning, are often trained on vast internet-scale corpora that may inadvertently include benchmark questions. As LRMs become more capable, the risk of inflated performance due to contamination grows, undermining the validity of leaderboards and scientific conclusions. This paper is the first to systematically study contamination in LRMs, providing a structured analysis that is urgently needed as these models are deployed in high-stakes domains.
The paper's two-stage framework (pre-LRM and post-LRM) offers a clear lens for understanding where contamination can enter. Pre-LRM contamination occurs when benchmark data is included in the pretraining corpus, while post-LRM contamination can happen during fine-tuning or alignment. This distinction is crucial because it informs where detection and mitigation efforts should focus. By highlighting the fragility of existing detection methods, the authors challenge the community to rethink evaluation protocols for reasoning models.
The abstract does not provide specific quantitative results, but the key finding is that existing contamination detection methods are fragile when applied to LRMs. This suggests that many reported performance numbers for reasoning models may be inflated due to undetected contamination. The paper likely includes experiments showing how detection methods fail, but these details are not in the abstract.
This paper has profound implications for AI evaluation. As reasoning models become more prevalent, the integrity of benchmarks is paramount. The fragility of contamination detection means that the AI community cannot trust current evaluations of LRMs, potentially leading to overestimation of capabilities and misguided research directions. This work serves as a wake-up call, urging researchers to develop more rigorous evaluation standards and contamination detection techniques. It also highlights the need for transparency in model training data and benchmark curation. Ultimately, this paper contributes to the broader goal of ensuring that AI progress is measured accurately and honestly.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba