Benchmarking llms via uncertainty quantification
Unknown
Introduces a new benchmarking approach for LLMs that integrates uncertainty quantification, evaluated across nine LLMs spanning five series.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
Introduces a new benchmarking approach for LLMs that integrates uncertainty quantification, evaluated across nine LLMs spanning five series.
Unknown
This survey provides a comprehensive overview of neural network-based uncertainty quantification methodologies and their applications in decision making.
Unknown
This survey categorizes uncertainty quantification methods for deep learning and demonstrates their application across various domains.
Unknown
This paper enhances uncertainty-based hallucination detection in LLMs by focusing on sentence-level detection, improving performance on accessible datasets.
Unknown
This paper investigates whether large language models can accurately assess their own knowledge and uncertainty, proposing methods to improve self-awareness and calibration.
Unknown
This paper evaluates uncertainty quantification methods for reasoning models without fine-tuning, using self-verbalized approaches to assess when models know they don't know.
Sebastian Farquhar, Jannik Kossen, Lorenz Kuhn, et al.
Proposes semantic entropy to detect LLM confabulations by measuring uncertainty at the meaning level, enabling task-agnostic hallucination detection.
Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, et al.
A comprehensive survey of uncertainty estimation in deep neural networks, covering sources, modeling approaches, calibration, and practical applications.
Craig Boutilier, Taraneh Dean, Steve Hanks
This paper surveys how structural properties of MDPs can be exploited via AI-style representations to ease the computational burden of planning under uncertainty.
Dong-Sheng Guo, Jikun Wu, S. Yiu
ReaLM-Retrieve introduces a reasoning-aware retrieval framework that injects external evidence at step-level uncertainty during multi-step reasoning, improving F1 by 10.1% over standard RAG while reducing retrieval calls by 47%.
Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, et al.
Comprehensive review of uncertainty quantification methods in deep learning, covering Bayesian and ensemble techniques across vision, NLP, and reinforcement learning.