Binesh Sadanandan PhD Dissertation Companion
Read thesis

Evidence explorer

Every result, traceable to its source

45 results and 20 charts from the dissertation, each one traceable to the sample it covers and the file it was computed from. Start with the five that carry the argument, or browse everything and filter by thrust, model, dataset, metric, or population.

Flip rates depend on the evaluation set The thrusts run on different evaluation sets, so a flip rate is only comparable within its set: MedGemma-4B flips 8.3% of pairs on the benchmark and 42.1% on the curated diagnostic set. Every panel names its population, and the population reference lists them.

The five results the thesis rests on

Read in order, they are the argument: the failure is real and general, consistency says nothing about grounding, most consistent answers ignore the image, engineering the consistency made that worse, and the gate built on it admits questions the text alone can answer.

  1. 6.4% to 54.7%

    Paraphrase flip rates span 6.4% to 54.7% across six medical vision-language models

    Rephrasing a clinical yes/no question in a way that preserves its meaning changes the answer on anywhere from 1 in 16 to more than half of paraphrase pairs, depending on which model and which patient population you test. Every model tested flips, and the spread between the best and worst model-dataset cell is 8.5-fold.

    Multiple models · Multiple datasets · A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%). · Sample and source

  2. 96.3%

    LLaVA-Rad gives the same answer 96.3% of the time with the image removed

    Delete the chest X-ray entirely and ask LLaVA-Rad the same question, and it returns its original answer on 96 of every 100 pairs. It is also the most paraphrase-consistent of the three backends tested here (6.5% flip rate). Its consistency is almost entirely a property of the question text: the image is nearly decorative.

    LLaVA-Rad · Multiple datasets · n = 107 · Sample and source

  3. 81% (range 45-100%)

    81% of consistent predictions are image-invariant

    Averaged across ten model-dataset settings, 81% of each model's consistent predictions are unchanged when the image is removed (range 45-100%), so a low paraphrase-flip rate is not evidence of grounded visual reasoning.

    Multiple models · Multiple datasets · n = 10 · Sample and source

  4. 76.8%

    The adapter buys consistency by leaning harder on the question text

    On the clean patient-disjoint audit, text-only agreement rises from 53.5% for the base model to 76.8% (+/- 3.1) after targeted adaptation: the adapted model returns its original answer with the image deleted 23 points more often. Both adapters get much of their new consistency by relying more on the question text, with no steadier use of the image. Targeted and Full LoRA are statistically tied on image reliance (76.8% +/- 3.1 vs 76.6% +/- 1.7).

    Targeted LoRA · MIMIC-CXR · n = 241 · Sample and source

  5. 2.9%

    2.9% accuracy on the slice where the image should matter

    Targeted LoRA scores only 2.9% accuracy on the PadChest text-disagrees slice - the cases where the image-conditioned answer must depart from the text-only answer - so gates raise admitted-case accuracy by selecting text-answerable cases over image-grounded ones (RQ5.3 = no).

    Targeted LoRA · PadChest · n = 241 · Sample and source

See how they connect

Cite

Permanent link

Type to search. Press Escape to close.