Trustworthiness Evaluation of Medical Vision-Language Models: A Scoping Review of Robustness, Grounding, Hallucination, and Uncertainty

Published in JMIR AI preprint; manuscript under review, 2026

JMIR preprint

Following PRISMA-ScR, this review maps how medical Vision-Language Model (VLM) research evaluates trustworthiness across robustness, visual grounding, hallucination, and uncertainty. It identifies which methods and datasets are commonly used, where evaluation coverage remains thin, and the paraphrase-invariance gap that motivates the dissertation. The manuscript is under review and has not yet been peer reviewed or edited; it should not guide clinical practice.

Recommended citation: Sadanandan, B., Karimi, A., Upadhayay, B., & Behzadan, V. (2026). Trustworthiness evaluation of medical vision-language models: A scoping review of robustness, grounding, hallucination, and uncertainty. Manuscript under review. doi:10.2196/preprints.102330.
Download Paper