VeriSim: Evaluating Medical AI Under Realistic Patient Noise (April 2026) logo

VeriSim: Evaluating Medical AI Under Realistic Patient Noise (April 2026)

Free

Truth-preserving patient simulation framework injecting controllable, clinically evidence-grounded noise — evaluates medical AI robustness under realistic imperfect patient data conditions

FreeFree tier
Type
Open Source

About VeriSim: Evaluating Medical AI Under Realistic Patient Noise (April 2026)

VeriSim is an open-source, configurable framework for evaluating medical AI under realistic patient noise. It injects controllable, clinically evidence-grounded noise into simulated patient responses while preserving medical ground truth through a hybrid UMLS-LLM verification mechanism. The framework operationalizes six noise dimensions derived from peer-reviewed medical communication literature, including patient recall limitations, health literacy barriers, anxiety, and stigma-driven non-disclosure. Experiments across seven open-weight LLMs revealed diagnostic accuracy drops of 15–25% and conversation length increases of 34–55% under realistic noise, with smaller models (7B) showing 40% greater degradation than larger models (70B+). Medical fine-tuning on standard corpora provided limited robustness benefits. The simulation quality was validated by board-certified clinicians with strong inter-annotator agreement (kappa 0.80). VeriSim establishes a rigorous testbed for identifying the Sim-to-Real gap in clinical AI robustness.

Key Features

Controllable injection of six noise dimensions from medical communication literature (recall limitations, health literacy, anxiety, stigma-driven non-disclosure, etc.)
Hybrid UMLS-LLM verification mechanism to maintain strict adherence to medical ground truth
Configurable patient simulation framework tailored for evaluating clinical AI robustness
Open-source release for reproducible benchmarking
Validated by board-certified clinicians with strong inter-annotator agreement (kappa 0.80)
LLM-as-a-Judge auxiliary evaluator for scalable assessment

Pros & Cons

Pros
  • Open-source and free to use
  • Noise dimensions grounded in peer-reviewed medical communication literature
  • Hybrid UMLS-LLM verification ensures medical accuracy while injecting realistic noise
  • Clinician-validated simulations with strong inter-annotator reliability
  • Configurable framework allows researchers to adjust noise levels and dimensions
  • Includes LLM-as-a-Judge for scalable, automated evaluation
Cons
  • Currently focuses only on the six noise dimensions derived from literature; may not cover all real-world patient communication barriers
  • Requires access to LLMs and UMLS for full setup
  • Evaluation demonstrated significant performance degradation, which may discourage adoption if not addressed

Best For

Stress-testing medical LLMs under realistic patient communication noiseIdentifying robustness gaps between standardized benchmarks and real clinical encountersTraining and fine-tuning clinical AI models to handle imperfect patient dataBenchmarking open-weight LLMs (7B to 70B+) for diagnostic accuracy and conversation efficiency

FAQ

What is VeriSim?
VeriSim is a truth-preserving patient simulation framework that injects controllable, clinically evidence-grounded noise into patient responses to evaluate the robustness of medical AI models.
How does VeriSim maintain medical accuracy?
It uses a hybrid UMLS-LLM verification mechanism that ensures injected noise adheres to medical ground truth while simulating realistic patient communication barriers.
Is VeriSim open-source?
Yes, the framework is released as an open-source noise-injection tool for research and benchmarking.
What noise dimensions does VeriSim simulate?
It operationalizes six noise dimensions from peer-reviewed medical communication literature, including patient recall limitations, health literacy barriers, anxiety, and stigma-driven non-disclosure.
What were the key findings from experiments with VeriSim?
All evaluated LLMs degraded significantly under realistic noise, with diagnostic accuracy dropping 15-25% and conversation length increasing 34-55%. Smaller models (7B) showed 40% greater degradation than larger models (70B+).