Preprint2024
Can LLMs Do Retrieval and Reasoning in 1 Million Context Window?
Mo Li, Songyang Zhang, Taolin Zhang, et al.
NeedleBench is a synthetic framework for evaluating LLMs' retrieval and reasoning in long contexts, revealing that reasoning models like Deepseek-R1 and o3 struggle with continuous retrieval in information-dense scenarios due to 'under-thinking'.
18Jul 16, 2024Large Language ModelsRetrieval Augmented Generation
arXiv