Preprint
Machine Learning

Natural language processing: an introduction

Prakash M. Nadkarni(Yale University), Lucila Ohno‐Machado(University of California San Diego), Wendy W. Chapman(University of California San Diego)
August 16, 2011Journal of the American Medical Informatics Association1,988 citations

2.0k

Citations

27

Influential Citations

Journal of the American Medical Informatics Association

Venue

2011

Year

Abstract

OBJECTIVES: To provide an overview and tutorial of natural language processing (NLP) and modern NLP-system design. TARGET AUDIENCE: This tutorial targets the medical informatics generalist who has limited acquaintance with the principles behind NLP and/or limited knowledge of the current state of the art. SCOPE: We describe the historical evolution of NLP, and summarize common NLP sub-problems in this extensive field. We then provide a synopsis of selected highlights of medical NLP efforts. After providing a brief description of common machine-learning approaches that are being used for diverse NLP sub-problems, we discuss how modern NLP architectures are designed, with a summary of the Apache Foundation's Unstructured Information Management Architecture. We finally consider possible future directions for NLP, and reflect on the possible impact of IBM Watson on the medical field.

Analysis

Why This Paper Matters

This paper is a seminal tutorial that has introduced natural language processing (NLP) to a broad audience of medical informatics generalists. Published in the Journal of the American Medical Informatics Association (JAMIA) in 2011, it has accumulated nearly 2000 citations, indicating its widespread use as an educational resource. The paper matters because it addresses a critical need: making NLP accessible to clinicians and informaticians who may not have a background in computational linguistics. By explaining the historical evolution, common sub-problems, and modern architectures, it empowers practitioners to understand and apply NLP in healthcare settings, such as extracting information from clinical notes or supporting decision-making.

Technical Contributions

The paper's technical contributions are primarily pedagogical and architectural:

  • Historical evolution: Traces NLP from early rule-based systems to statistical and machine-learning approaches.
  • Common sub-problems: Covers tokenization, part-of-speech tagging, parsing, named entity recognition, and relation extraction.
  • Medical NLP highlights: Summarizes key efforts like the Unified Medical Language System (UMLS) and clinical information extraction systems.
  • Machine-learning approaches: Describes supervised, unsupervised, and semi-supervised methods used for NLP tasks.
  • Modern architectures: Introduces the Apache Unstructured Information Management Architecture (UIMA), a framework for building modular NLP pipelines.
  • Future directions: Discusses the potential of IBM Watson and other advanced systems in medicine.

Results

As a tutorial, the paper does not present experimental results or quantitative comparisons. Its impact is measured by its citation count (1988) and its role in educating a generation of medical informatics researchers. The paper's value lies in its clear exposition and synthesis of a complex field, rather than in novel empirical findings.

Significance

The broader impact of this paper is substantial. It has helped demystify NLP for the medical community, fostering interdisciplinary collaboration and accelerating the adoption of NLP in clinical research and practice. By providing a structured overview and pointing to resources like UIMA, it has enabled practitioners to build and deploy NLP systems more effectively. The paper's forward-looking discussion of IBM Watson also anticipated the growing role of AI in healthcare, making it a prescient and influential work.