Preprint
Large Language Models

ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports

Katharina Jeblick(LMU Klinikum), Balthasar Schachtner(LMU Klinikum), Jakob Dexl(LMU Klinikum), Andreas Mittermeier(LMU Klinikum), Anna Theresa Stüber(LMU Klinikum), Johanna Topalis(LMU Klinikum), Tobias Weber(LMU Klinikum), Philipp Wesp(LMU Klinikum), Bastian O. Sabel(LMU Klinikum), Jens Ricke(LMU Klinikum), Michael Ingrisch(LMU Klinikum)
October 5, 2023European Radiology554 citations

554

Citations

12

Influential Citations

European Radiology

Venue

2023

Year

Abstract

OBJECTIVES: To assess the quality of simplified radiology reports generated with the large language model (LLM) ChatGPT and to discuss challenges and chances of ChatGPT-like LLMs for medical text simplification. METHODS: In this exploratory case study, a radiologist created three fictitious radiology reports which we simplified by prompting ChatGPT with "Explain this medical report to a child using simple language." In a questionnaire, we tasked 15 radiologists to rate the quality of the simplified radiology reports with respect to their factual correctness, completeness, and potential harm for patients. We used Likert scale analysis and inductive free-text categorization to assess the quality of the simplified reports. RESULTS: Most radiologists agreed that the simplified reports were factually correct, complete, and not potentially harmful to the patient. Nevertheless, instances of incorrect statements, missed relevant medical information, and potentially harmful passages were reported. CONCLUSION: While we see a need for further adaption to the medical field, the initial insights of this study indicate a tremendous potential in using LLMs like ChatGPT to improve patient-centered care in radiology and other medical domains. CLINICAL RELEVANCE STATEMENT: Patients have started to use ChatGPT to simplify and explain their medical reports, which is expected to affect patient-doctor interaction. This phenomenon raises several opportunities and challenges for clinical routine. KEY POINTS: • Patients have started to use ChatGPT to simplify their medical reports, but their quality was unknown. • In a questionnaire, most participating radiologists overall asserted good quality to radiology reports simplified with ChatGPT. However, they also highlighted a notable presence of errors, potentially leading patients to draw harmful conclusions. • Large language models such as ChatGPT have vast potential to enhance patient-centered care in radiology and other medical domains. To realize this potential while minimizing harm, they need supervision by medical experts and adaption to the medical field.

Analysis

Why This Paper Matters

This paper addresses a rapidly emerging real-world phenomenon: patients using ChatGPT to simplify and understand their medical reports. As LLMs become more accessible, their use in healthcare—without formal validation—poses both opportunities and risks. The study provides one of the first systematic assessments of the quality of such simplified outputs, specifically in radiology. Its findings are directly relevant to clinicians, patients, and AI developers, as they underscore the need for careful integration of LLMs into clinical workflows.

The significance lies in its pragmatic, case-study approach. Rather than proposing a new model, it evaluates an existing, widely-used tool (ChatGPT) in a realistic scenario. This makes the results immediately actionable for healthcare professionals who may encounter patients using such tools. The paper also sets a precedent for how to evaluate LLM-generated medical content, combining quantitative Likert ratings with qualitative free-text analysis to capture both overall quality and specific failure modes.

Technical Contributions

  • Prompt Engineering for Simplification: The study uses a single, simple prompt ("Explain this medical report to a child using simple language") to generate simplified reports, demonstrating that even basic prompting can yield useful outputs.
  • Evaluation Framework: A structured questionnaire with Likert scales (factual correctness, completeness, potential harm) and inductive free-text categorization provides a reproducible method for assessing LLM-generated medical text.
  • Error Taxonomy: The qualitative analysis identifies three key error types: incorrect statements, missed relevant information, and potentially harmful passages, offering a foundation for future safety evaluations.
  • Domain-Specific Insight: The study highlights that while LLMs can produce generally accurate simplifications, they may omit critical details (e.g., specific findings or follow-up recommendations) that are essential for patient safety.

Results

The study reports that the majority of 15 participating radiologists rated the simplified reports as factually correct, complete, and not potentially harmful. However, the free-text responses revealed notable exceptions: some radiologists pointed out incorrect statements (e.g., misinterpretation of medical terms), missing relevant information (e.g., absence of specific measurements or comparisons), and passages that could lead patients to harmful conclusions (e.g., downplaying the need for follow-up). No quantitative metrics (e.g., inter-rater agreement, average Likert scores) are provided in the abstract; the results are qualitative and exploratory.

Significance

This paper is a timely contribution to the growing literature on LLMs in healthcare. It provides early evidence that ChatGPT can generate simplified radiology reports that are generally acceptable to radiologists, but it also sounds a cautionary note about the risks of unsupervised use. The findings support the need for domain-specific fine-tuning, expert oversight, and clear guidelines for patients and clinicians. For the AI field, it underscores the importance of evaluating LLMs not just on benchmark tasks but on real-world, high-stakes applications where errors can have serious consequences. The study also opens avenues for future work on automated quality assurance, personalized simplification, and integration with electronic health records.