Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs
Sadanandan, B., & Behzadan, V. (2026). Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs. iMIMIC Workshop, MICCAI 2026.
My final dissertation and supporting materials are available through the dissertation companion.
I’m a Technical Fellow and senior data science leader with 10+ years building production machine learning, generative AI, and enterprise data platforms in regulated healthcare, on a 19-year engineering foundation. I advise senior leadership on machine learning strategy at Medtronic. I successfully defended my Ph.D. dissertation in medical AI safety, mechanistic interpretability, and Large Language Model (LLM) and Vision-Language Model (VLM) evaluation on August 26, 2026, and have 12 publications and preprints and two patents.
Medtronic, Surgical Innovation (North Haven, CT, 2015 - Present)
| Role | Years |
|---|---|
| Senior Principal R&D Engineer & Technical Fellow | 2023 - Present |
| Senior Principal R&D Engineer | 2021 - 2023 |
| Principal R&D Applications Engineer | 2016 - 2021 |
| Enterprise Solutions Consultant | 2015 - 2016 |
As Technical Fellow, I advise senior leadership on machine learning and data science strategy and mentor engineering teams across the organization. My primary focus is device data from the Signia powered stapler: predictive models and generative AI for product development and device safety on AWS and Snowflake, retrieval-augmented generation (RAG) pipelines with vector databases for R&D document search, and analysis of MedDRA-coded adverse event data for post-market device safety surveillance. I also built a foundation model for predicting non-small cell lung cancer (NSCLC) recurrence by fine-tuning MedGemma on SEER-Medicare data, and my team’s generative models for early detection of device failures were published at ICMHI 2024.
In earlier roles here, I led scientific and clinical evidence strategies for minimally invasive surgical staplers, built data pipelines with Dataiku, Python, Redshift, and Snowflake for clinical evidence generation, delivered Power BI dashboards for stakeholders, and architected Windchill product lifecycle management (PLM) solutions for R&D.
Ph.D. Researcher, SAIL Lab, University of New Haven (2021 - 2026)
My dissertation, Paraphrase Sensitivity in Medical Vision-Language Models: Measurement, Mechanisms, Mitigation, and Deployment Safety, studies clinically equivalent questions that produce contradictory diagnoses. I built the 92,856-pair PSF-Med benchmark, used sparse-autoencoder and residual-stream analyses to identify candidate internal mechanisms, designed a targeted Low-Rank Adaptation (LoRA) intervention that cut pairwise flips about 59%, and showed why consistency, visual grounding, correctness, and calibration must be audited jointly. The final dissertation, benchmarks, models, and code are available through the dissertation companion and on Hugging Face and GitHub.
Earlier career
| Role | Company | Years |
|---|---|---|
| Technical Architect (PLM) & Application Consultant | Barry-Wehmiller Design Group | 2012 - 2015 |
| Product Specialist & Enterprise Support Engineer | PTC | 2010 - 2012 |
| Infrastructure Engineer | Hewlett Packard Enterprise | 2009 - 2010 |
| Technical Associate | Minacs | 2007 - 2009 |
| Degree | Institution | Years |
|---|---|---|
| Ph.D., Engineering and Applied Science (Data Science) | University of New Haven | 2021 - 2026; dissertation successfully defended August 26, 2026 |
| M.S., Data Science (GPA 3.92) | University of Connecticut School of Business | 2017 - 2019 |
| B.E., Electronics and Communication | Cochin University of Science and Technology | 2004 - 2008 |
Certifications: Tableau Desktop Specialist; Deep Learning Nanodegree, Udacity (2017); Certificate in Project Management.
| Area | Tools and methods |
|---|---|
| Machine Learning & GenAI | Production ML (batch and real-time inference), predictive modeling, deep learning, LLM/VLM evaluation, retrieval-augmented generation (RAG), vector databases, prompt engineering |
| Responsible AI | Robustness and bias evaluation, failure-mode analysis, model monitoring, clinical AI safety |
| Data Platforms | AWS, Snowflake, Redshift, Databricks, Azure ML, Dataiku, lakehouse and pipeline architecture, data governance, PLM/Windchill, APIs and data services |
| Analytics & Visualization | Python, SQL, R, statistics, exploratory data analysis, anomaly detection, Power BI, Tableau, Plotly/Dash |
| Health Data | DICOM, MedDRA, clinical evidence generation |
I serve as a peer reviewer for ICLR 2026, MICCAI 2026, Machine Learning for Healthcare (MLHC) 2026, and IEEE ICMLA 2026, and I was Lead Reviewer for the International Medical Devices Safety Conference in 2024 and 2025.
Sadanandan, B., & Behzadan, V. (2026). Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs. iMIMIC Workshop, MICCAI 2026.
Sadanandan, B., Karimi, A., Upadhayay, B., & Behzadan, V. (2026). Trustworthiness evaluation of medical vision-language models: A scoping review of robustness, grounding, hallucination, and uncertainty. Manuscript under review. doi:10.2196/preprints.102330.
Sadanandan, B., & Behzadan, V. (2026). Mechanistically guided LoRA improves paraphrase consistency in medical vision-language models. Conference on Health, Inference, and Learning (CHIL) 2026. arXiv:2603.00148.
Sadanandan, B., & Behzadan, V. (2026). Predictive entropy as a joint screen for error and paraphrase instability in medical vision-language models. UNSURE Workshop, MICCAI 2026 (poster).
Sadanandan, B., & Behzadan, V. (2026). When chain-of-thought backfires: Evaluating prompt sensitivity in medical language models. 2AI Conference 2026. arXiv:2603.25960.
Sadanandan, B., & Behzadan, V. (2026). Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models. MedReasoner Workshop, CVPR 2026.
Sadanandan, B., Behzadan, V., Jayan, L., & Kurup, A. G. (2026). PSF-Med: A clinician-audited benchmark for paraphrase sensitivity in medical vision-language models. MMFM-BIOMED Workshop, CVPR 2026. arXiv:2602.21428.
Sadanandan, B., & Behzadan, V. (2025). VSF-Med: A vulnerability scoring framework for medical vision-language models. IEEE ISBI 2026 (pilot abstract). arXiv:2507.00052.
Multimodal Deep Learning for Early Prediction of Patient Deterioration in the ICU. Journal of Data Science, 2025.
Sadanandan, B., & Behzadan, V. (2025). Promise of Data-Driven Modeling and Decision Support for Precision Oncology and Theranostics. arXiv preprint arXiv:2505.09899.
Kumar, B., Bidarahalli, P.K., & Miesse, A.M. (2025). System and method for machine learning method to verify consistent staple line delivery. WO Patent WO2025027493A1.
Sadanandan, B., Arghavani Nobar, B., & Behzadan, V. (2024). Comparative Study of Generative Models for Early Detection of Failures in Medical Devices. 2024 8th International Conference on Medical and Health Informatics.
Sadanandan, B. (2024). Machine Learning for Anastomotic Leak Prediction: A Systematic Review and Experimental Validation. International Journal of Trend in Scientific Research and Development, 8(3), 12.
Miesse, A.M., Sadanandan, B.K., Evans, C.K., Knapp, R.H., & Jalaja, N.L.V. (2023). Systems and methods for machine learning control of a surgical device. WO Patent WO2023223258A2.