Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs
Published in iMIMIC Workshop, MICCAI 2026, 2026
Saliency and attention maps are often offered as evidence that a medical Vision-Language Model (VLM) looked at the right region. We evaluate these explanations causally and find that a plausible-looking map does not imply the highlighted region drove the answer, which matters when explanations are used to justify clinical trust.
Recommended citation: Sadanandan, B., & Behzadan, V. (2026). Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs. iMIMIC Workshop, MICCAI 2026.
