Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models
Published in MedReasoner Workshop, CVPR 2026, 2026
Consistency metrics can hide a failure mode: a medical Vision-Language Model (VLM) that gives the same answer with or without the image. We classify safety per sample and find that a large share of “consistent” predictions are image-invariant, so aggregate consistency overstates how much a model can be trusted in the clinic.
Recommended citation: Sadanandan, B., & Behzadan, V. (2026). Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models. MedReasoner Workshop, CVPR 2026.
Download Paper
