Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models

Published in UNSURE Workshop, MICCAI 2026 (poster), 2026

PaperCode

We test whether a single predictive-entropy signal can flag two failure modes in medical Vision-Language Models (VLMs). On Targeted LoRA over the PadChest flip bank, entropy ranks paraphrase flips at AUROC 0.823 and errors at 0.862, with the flip result replicated across architectures. Softmax entropy, temperature-scaled entropy, and absolute margin are rank-equivalent; confidence ranking still does not certify that a prediction is image-grounded.

Recommended citation: Sadanandan, B., & Behzadan, V. (2026). Predictive entropy as a joint screen for error and paraphrase instability in medical vision-language models. UNSURE Workshop, MICCAI 2026 (poster).
Download Paper