Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs

Published in iMIMIC Workshop, MICCAI 2026, 2026

arXivCode and data instructions

Saliency and attention maps are often offered as evidence that a medical Vision-Language Model (VLM) looked at the right region. Bounding-box coverage, patch-rank agreement with occlusion, and region-based causal scoring show that attention only marginally beats a shifted box and does not track causal patch importance. Attention is a coarse localizer, short of a faithful explanation, and should not be treated as safety evidence.

Recommended citation: Sadanandan, B., & Behzadan, V. (2026). Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs. iMIMIC Workshop, MICCAI 2026.
Download Paper