ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
1.0k
Citations
56
Influential Citations
Nature Medicine
Venue
2025
Year
Large language models (LLMs) have shown promise in medical question answering, with Med-PaLM being the first to exceed a 'passing' score in United States Medical Licensing Examination style questions. However, challenges remain in long-form medical question answering and handling real-world workflows. Here, we present Med-PaLM 2, which bridges these gaps with a combination of base LLM improvements, medical domain fine-tuning and new strategies for improving reasoning and grounding through ensemble refinement and chain of retrieval. Med-PaLM 2 scores up to 86.5% on the MedQA dataset, improving upon Med-PaLM by over 19%, and demonstrates dramatic performance increases across MedMCQA, PubMedQA and MMLU clinical topics datasets. Our detailed human evaluations framework shows that physicians prefer Med-PaLM 2 answers to those from other physicians on eight of nine clinical axes. Med-PaLM 2 also demonstrates significant improvements over its predecessor across all evaluation metrics, particularly on new adversarial datasets designed to probe LLM limitations (P < 0.001). In a pilot study using real-world medical questions, specialists preferred Med-PaLM 2 answers to generalist physician answers 65% of the time. While specialist answers were still preferred overall, both specialists and generalists rated Med-PaLM 2 to be as safe as physician answers, demonstrating its growing potential in real-world medical applications.
This paper marks a significant milestone in the application of large language models to medicine. While previous work like Med-PaLM achieved a passing score on USMLE-style questions, Med-PaLM 2 pushes the frontier to expert-level performance, with physicians preferring its answers over those from other physicians on most clinical axes. This is a critical step toward real-world deployment, as it addresses both accuracy and trustworthiness—two key barriers for clinical adoption.
The paper also introduces a rigorous human evaluation framework that goes beyond automated metrics, capturing nuanced aspects like safety, reasoning, and clinical utility. The finding that both specialists and generalists rated Med-PaLM 2 as safe as physician answers is particularly noteworthy, as safety is the paramount concern in healthcare AI.
This work demonstrates that LLMs can achieve expert-level performance in medical question answering, challenging the notion that human expertise is irreplaceable in complex clinical reasoning. The implications for healthcare are profound: such models could assist clinicians by providing rapid, accurate answers, reducing diagnostic errors, and democratizing access to medical knowledge.
However, the paper also underscores the importance of rigorous evaluation and safety testing. The fact that specialists still preferred human expert answers overall indicates that LLMs are not yet a replacement for human judgment. The adversarial dataset approach provides a template for stress-testing models before deployment.
For the AI field, this paper advances the state of the art in domain-specific LLM adaptation, particularly through the combination of fine-tuning, retrieval augmentation, and ensemble methods. The human evaluation framework could become a standard for assessing clinical AI systems.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba