Journal Article
Large Language Models

Biases in Large Language Models: Origins, Inventory, and Discussion

Roberto Navigli(Sapienza University of Rome), Simone Conia(Sapienza University of Rome), Björn Roß(University of Edinburgh)
May 16, 2023Journal of Data and Information Quality531 citations

531

Citations

19

Influential Citations

Journal of Data and Information Quality

Venue

2023

Year

Abstract

In this article, we introduce and discuss the pervasive issue of bias in the large language models that are currently at the core of mainstream approaches to Natural Language Processing (NLP). We first introduce data selection bias, that is, the bias caused by the choice of texts that make up a training corpus. Then, we survey the different types of social bias evidenced in the text generated by language models trained on such corpora, ranging from gender to age, from sexual orientation to ethnicity, and from religion to culture. We conclude with directions focused on measuring, reducing, and tackling the aforementioned types of bias.

Analysis

Why This Paper Matters

This paper is a critical resource for the AI community as large language models become increasingly deployed in real-world applications. Bias in LLMs can perpetuate and amplify societal inequalities, making it essential to understand its origins and manifestations. By systematically cataloging bias types from data selection to social categories, the authors provide a structured framework that helps practitioners identify and address bias in their own models. The paper's emphasis on measurement and reduction directions is particularly timely given the rapid adoption of LLMs in sensitive domains like hiring, healthcare, and content moderation.

Technical Contributions

The paper's main technical contributions are:

  • Taxonomy of bias origins: Distinguishes data selection bias (caused by training corpus composition) from social biases that emerge in generated text.
  • Comprehensive bias inventory: Covers gender, age, sexual orientation, ethnicity, religion, and cultural biases, providing a clear categorization for researchers.
  • Mitigation directions: Outlines strategies for measuring bias (e.g., via benchmark datasets) and reducing it (e.g., through data augmentation, debiasing techniques, and fairness-aware training).

Results

As a survey paper, no new experimental results are presented. The paper synthesizes existing findings from the literature, noting that bias is pervasive across LLMs and that current mitigation methods have limited effectiveness. The authors call for more rigorous evaluation frameworks and interdisciplinary collaboration to address the problem.

Significance

This paper has become a widely cited reference (531 citations) in the fairness and ethics subfield of NLP. It provides a common vocabulary and conceptual map for researchers working on bias in LLMs, facilitating clearer communication and more targeted research. The work also has practical implications for AI practitioners who need to audit and improve the fairness of their models, and for policymakers developing guidelines for responsible AI deployment.