ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
554
Citations
12
Influential Citations
European Radiology
Venue
2023
Year
OBJECTIVES: To assess the quality of simplified radiology reports generated with the large language model (LLM) ChatGPT and to discuss challenges and chances of ChatGPT-like LLMs for medical text simplification. METHODS: In this exploratory case study, a radiologist created three fictitious radiology reports which we simplified by prompting ChatGPT with "Explain this medical report to a child using simple language." In a questionnaire, we tasked 15 radiologists to rate the quality of the simplified radiology reports with respect to their factual correctness, completeness, and potential harm for patients. We used Likert scale analysis and inductive free-text categorization to assess the quality of the simplified reports. RESULTS: Most radiologists agreed that the simplified reports were factually correct, complete, and not potentially harmful to the patient. Nevertheless, instances of incorrect statements, missed relevant medical information, and potentially harmful passages were reported. CONCLUSION: While we see a need for further adaption to the medical field, the initial insights of this study indicate a tremendous potential in using LLMs like ChatGPT to improve patient-centered care in radiology and other medical domains. CLINICAL RELEVANCE STATEMENT: Patients have started to use ChatGPT to simplify and explain their medical reports, which is expected to affect patient-doctor interaction. This phenomenon raises several opportunities and challenges for clinical routine. KEY POINTS: • Patients have started to use ChatGPT to simplify their medical reports, but their quality was unknown. • In a questionnaire, most participating radiologists overall asserted good quality to radiology reports simplified with ChatGPT. However, they also highlighted a notable presence of errors, potentially leading patients to draw harmful conclusions. • Large language models such as ChatGPT have vast potential to enhance patient-centered care in radiology and other medical domains. To realize this potential while minimizing harm, they need supervision by medical experts and adaption to the medical field.
This paper addresses a rapidly emerging real-world phenomenon: patients using ChatGPT to simplify and understand their medical reports. As LLMs become more accessible, their use in healthcare—without formal validation—poses both opportunities and risks. The study provides one of the first systematic assessments of the quality of such simplified outputs, specifically in radiology. Its findings are directly relevant to clinicians, patients, and AI developers, as they underscore the need for careful integration of LLMs into clinical workflows.
The significance lies in its pragmatic, case-study approach. Rather than proposing a new model, it evaluates an existing, widely-used tool (ChatGPT) in a realistic scenario. This makes the results immediately actionable for healthcare professionals who may encounter patients using such tools. The paper also sets a precedent for how to evaluate LLM-generated medical content, combining quantitative Likert ratings with qualitative free-text analysis to capture both overall quality and specific failure modes.
The study reports that the majority of 15 participating radiologists rated the simplified reports as factually correct, complete, and not potentially harmful. However, the free-text responses revealed notable exceptions: some radiologists pointed out incorrect statements (e.g., misinterpretation of medical terms), missing relevant information (e.g., absence of specific measurements or comparisons), and passages that could lead patients to harmful conclusions (e.g., downplaying the need for follow-up). No quantitative metrics (e.g., inter-rater agreement, average Likert scores) are provided in the abstract; the results are qualitative and exploratory.
This paper is a timely contribution to the growing literature on LLMs in healthcare. It provides early evidence that ChatGPT can generate simplified radiology reports that are generally acceptable to radiologists, but it also sounds a cautionary note about the risks of unsupervised use. The findings support the need for domain-specific fine-tuning, expert oversight, and clear guidelines for patients and clinicians. For the AI field, it underscores the importance of evaluating LLMs not just on benchmark tasks but on real-world, high-stakes applications where errors can have serious consequences. The study also opens avenues for future work on automated quality assurance, personalized simplification, and integration with electronic health records.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba