Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
Published in CHIL 2026, 2026
We use sparse-autoencoder interpretability to locate the layers where a medical Vision-Language Model (VLM) commits to an answer, then apply a Low-Rank Adaptation (LoRA) on layers 15 to 19, touching 0.1% of parameters. On a patient-disjoint test the pairwise flip rate drops by about 59% (8.5% to 3.5% over five seeds) with no observed accuracy reduction.
Recommended citation: Sadanandan, B., & Behzadan, V. (2026). Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models. CHIL 2026.
Download Paper
