Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
Published in Conference on Health, Inference, and Learning (CHIL) 2026, 2026
We use sparse-autoencoder interpretability to identify candidate features and the layers where a medical Vision-Language Model (VLM) commits to an answer, then apply a Low-Rank Adaptation (LoRA) on layers 15 to 19, touching 0.1% of parameters. On a patient-disjoint test the pairwise flip rate drops by about 59% (8.5% to 3.5% over five seeds) with no observed accuracy reduction. A safety re-audit finds that the adapter gains consistency by relying more heavily on question text, so lower flip rates do not imply better visual grounding.
Recommended citation: Sadanandan, B., & Behzadan, V. (2026). Mechanistically guided LoRA improves paraphrase consistency in medical vision-language models. Conference on Health, Inference, and Learning (CHIL) 2026. arXiv:2603.00148.
Download Paper
