VeriSim: Evaluating Medical AI Under Realistic Patient Noise (April 2026)
FreeTruth-preserving patient simulation framework injecting controllable, clinically evidence-grounded noise — evaluates medical AI robustness under realistic imperfect patient data conditions
About VeriSim: Evaluating Medical AI Under Realistic Patient Noise (April 2026)
VeriSim is an open-source, configurable framework for evaluating medical AI under realistic patient noise. It injects controllable, clinically evidence-grounded noise into simulated patient responses while preserving medical ground truth through a hybrid UMLS-LLM verification mechanism. The framework operationalizes six noise dimensions derived from peer-reviewed medical communication literature, including patient recall limitations, health literacy barriers, anxiety, and stigma-driven non-disclosure. Experiments across seven open-weight LLMs revealed diagnostic accuracy drops of 15–25% and conversation length increases of 34–55% under realistic noise, with smaller models (7B) showing 40% greater degradation than larger models (70B+). Medical fine-tuning on standard corpora provided limited robustness benefits. The simulation quality was validated by board-certified clinicians with strong inter-annotator agreement (kappa 0.80). VeriSim establishes a rigorous testbed for identifying the Sim-to-Real gap in clinical AI robustness.
Key Features
Pros & Cons
- Open-source and free to use
- Noise dimensions grounded in peer-reviewed medical communication literature
- Hybrid UMLS-LLM verification ensures medical accuracy while injecting realistic noise
- Clinician-validated simulations with strong inter-annotator reliability
- Configurable framework allows researchers to adjust noise levels and dimensions
- Includes LLM-as-a-Judge for scalable, automated evaluation
- Currently focuses only on the six noise dimensions derived from literature; may not cover all real-world patient communication barriers
- Requires access to LLMs and UMLS for full setup
- Evaluation demonstrated significant performance degradation, which may discourage adoption if not addressed