Role of RAG Noise in LLMs
Jinyang Wu, Feihu Che, Mingkuan Feng, et al.
Defines seven noise types in RAG, builds NoiserBench benchmark, and shows some noise can benefit LLMs while other noise harms them.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Jinyang Wu, Feihu Che, Mingkuan Feng, et al.
Defines seven noise types in RAG, builds NoiserBench benchmark, and shows some noise can benefit LLMs while other noise harms them.
Yu Wang, Shiwan Zhao, Zhihu Wang, et al.
SCoT improves LLM reasoning by first eliciting a problem-solving strategy before generating Chain-of-Thought steps, achieving significant gains on reasoning benchmarks.
Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li, et al.
The LiTS benchmark evaluates liver and tumor segmentation algorithms on a diverse CT dataset, finding no single algorithm excels at both tasks and highlighting the need for improved tumor detection.
Ajay N. Jain, Anthony Nicholls
This paper proposes standards for evaluating computational methods in drug design, including statistical reporting, data sharing, and benchmark best practices.
Paul Bergmann, Kilian Batzner, Michael Fauser, et al.
Introduces MVTec AD, a comprehensive dataset with 5354 high-resolution images across 15 categories for unsupervised anomaly detection, and benchmarks state-of-the-art methods.
Kai Li, Christian Bienia
This paper analyzes how pre-intervention exercise habits and baseline depression levels predict adherence, contamination, and dropout rates in walking and control groups.
Simon Baker, Daniel Scharstein, John Lewis, et al.
Proposes a new benchmark and evaluation methodology for optical flow algorithms, including diverse datasets and improved error metrics, to address challenges in complex natural scenes.
David Lewis, Yiming Yang, Tony Rose, et al.
This paper introduces RCV1, a large-scale benchmark of over 800,000 categorized newswire stories for text categorization, detailing its coding policy, taxonomy semantics, and corrections, while providing baseline results for supervised learning metho
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, et al.
MoleculeNet provides a standardized benchmark for molecular machine learning, curating datasets and evaluating featurization and learning algorithms.
Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, et al.
DeepLab advances semantic segmentation via atrous convolution, ASPP, and fully connected CRFs, achieving state-of-the-art on multiple benchmarks.
Farieda Gaber, Maqsood Shaik, Fabio Allega, et al.
Benchmarks LLMs and a RAG workflow on 2000 MIMIC-derived medical cases for triage, referral, and diagnosis support.
Nourhan Ibrahim, Samar AboulEla, Ahmed Ibrahim, et al.
This survey classifies LLM-KG integration into three paradigms—KG-augmented LLMs, LLM-augmented KGs, and synergized frameworks—and evaluates their methodologies, metrics, benchmarks, and challenges.