Threats in LLM-Powered AI Agents Workflows
M. Ferrag, N. Tihanyi, Djallel Hamouda, et al.
A unified threat model for LLM-agent ecosystems covering host-to-tool and agent-to-agent attacks, with over 30 techniques and mitigation strategies.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
M. Ferrag, N. Tihanyi, Djallel Hamouda, et al.
A unified threat model for LLM-agent ecosystems covering host-to-tool and agent-to-agent attacks, with over 30 techniques and mitigation strategies.
Manuel Cossio
This paper provides a comprehensive taxonomy of LLM hallucinations, arguing their theoretical inevitability and emphasizing the need for robust detection, mitigation, and human oversight.
Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, et al.
Proposes Lookback Lens, a simple hallucination detector using attention weight ratios, effective across tasks and models, and reduces hallucinations via classifier-guided decoding.
Adi Simhi, Jonathan Herzig, Idan Szpektor, et al.
This paper distinguishes between two types of LLM hallucinations—HK- (model lacks knowledge) and HK+ (model has knowledge but answers incorrectly)—and shows that distinguishing them improves mitigation.
Roberto Navigli, Simone Conia, Björn Roß
This paper surveys bias in LLMs from data selection to social biases like gender, ethnicity, and religion, and outlines measurement and mitigation directions.
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, et al.
A comprehensive survey of bias evaluation and mitigation techniques for LLMs, proposing taxonomies for metrics, datasets, and mitigation methods.
Rabeeh Karimi Mahabadi, Yonatan Belinkov, James Henderson
Proposes two learning strategies to train neural models that are more robust to dataset biases and transfer better to out-of-domain datasets.
Yue Zhang, Yafu Li, Leyang Cui, et al.
A comprehensive survey of hallucination in large language models, covering detection, explanation, and mitigation with taxonomies and benchmarks.
Lei Huang, Weijiang Yu, Weitao Ma, et al.
This survey provides a comprehensive taxonomy, analysis of causes, detection methods, and mitigation strategies for hallucination in large language models, highlighting challenges and future directions.