Evidence explorer
Every result, traceable to its source
47 results and 20 charts from the dissertation, each one traceable to the sample it covers and the file it was computed from. Start with the five that carry the argument, or browse everything and filter by thrust, model, dataset, metric, or population.
The five results the thesis rests on
Read in order, they are the argument: the failure is real and general, consistency says nothing about grounding, most consistent answers ignore the image, engineering the consistency made that worse, and the gate built on it admits questions the text alone can answer.
-
6.4% to 54.7%
Paraphrase flip rates span 6.4% to 54.7% across six medical vision-language models
Rephrasing a clinical yes/no question in a way that preserves its meaning changes the answer on anywhere from 1 in 16 to more than half of paraphrase pairs, depending on which model and which patient population you test. Every model tested flips, and the spread between the best and worst model-dataset cell is 8.5-fold.
Multiple models · Multiple datasets · A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%). · Sample and source
-
96.3%
LLaVA-Rad gives the same answer 96.3% of the time with the image removed
Delete the chest X-ray entirely and ask LLaVA-Rad the same question, and it returns its original answer on 96 of every 100 pairs. It is also the most paraphrase-consistent of the three backends tested here (6.5% flip rate). Its consistency is almost entirely a property of the question text: the image is nearly decorative.
LLaVA-Rad · Multiple datasets · n = 107 · Sample and source
-
81% (range 45-100%)
81% of consistent predictions are image-invariant
Averaged across ten model-dataset settings, 81% of each model's consistent predictions are unchanged when the image is removed (range 45-100%), so a low paraphrase-flip rate is not evidence of grounded visual reasoning.
Multiple models · Multiple datasets · n = 10 · Sample and source
-
76.8%
The adapter buys consistency by leaning harder on the question text
On the clean patient-disjoint audit, text-only agreement rises from 53.5% for the base model to 76.8% (+/- 3.1) after targeted adaptation: the adapted model returns its original answer with the image deleted 23 points more often. Both adapters get much of their new consistency by relying more on the question text, with no steadier use of the image. Targeted and Full LoRA are statistically tied on image reliance (76.8% +/- 3.1 vs 76.6% +/- 1.7).
Targeted LoRA · MIMIC-CXR · n = 241 · Sample and source
-
2.9%
2.9% accuracy on the slice where the image should matter
Targeted LoRA scores only 2.9% accuracy on the PadChest text-disagrees slice - the cases where the image-conditioned answer must depart from the text-only answer - so gates raise admitted-case accuracy by selecting text-answerable cases over image-grounded ones (RQ5.3 = no).
Targeted LoRA · PadChest · n = 241 · Sample and source
Showing all 20 charts and 47 results
No result matches every filter.
Some combinations do not exist in the dissertation: PadChest has no MedGemma-27B query-level cell, and MIMIC carries no demographics. Try removing one filter.
Charts
Pairwise paraphrase flip rate by model and dataset
Six base vision-language models on the equivalence-filtered PSF-Med binary yes/no subset: MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, VinDr-CXR 2,807 / 8,612.
- MedGemma-1.5-4B
- MedGemma-4B
- LLaVA-Rad
- MedGemma-27B
- CheXone
- RadFM
| Series | Category | percent% | n | Note |
|---|---|---|---|---|
| MedGemma-1.5-4B | MIMIC-CXR | 7.4 | 5,076 | |
| MedGemma-1.5-4B | PadChest | 17.8 | 36,244 | dissertation Table 3.3 value; the v2 recompute file gives 17.47% |
| MedGemma-1.5-4B | VinDr-CXR | 0.8 | 8,612 | degenerate yes-bias cell (over 98% positive on originals) |
| MedGemma-4B | MIMIC-CXR | 8.3 | 5,076 | |
| MedGemma-4B | PadChest | 13.4 | 36,244 | |
| MedGemma-4B | VinDr-CXR | 15.3 | 8,612 | |
| LLaVA-Rad | MIMIC-CXR | 15.6 | 5,076 | |
| LLaVA-Rad | PadChest | 0.8 | 36,244 | degenerate yes-bias cell (over 98% positive on originals) |
| LLaVA-Rad | VinDr-CXR | 11.6 | 8,612 | |
| MedGemma-27B | MIMIC-CXR | 6.4 | 5,076 | |
| MedGemma-27B | PadChest | 13.9 | 36,244 | no cell in the v2 recompute file; dissertation table only |
| MedGemma-27B | VinDr-CXR | 8.1 | 8,612 | |
| CheXone | MIMIC-CXR | 8.2 | 5,076 | |
| CheXone | PadChest | 12.0 | 36,244 | dissertation Table 3.3 value; the v2 recompute file gives 13.04% |
| CheXone | VinDr-CXR | 11.1 | 8,612 | |
| RadFM | MIMIC-CXR | 13.7 | 5,076 | |
| RadFM | PadChest | 32.8 | 34,733 | run covers 8,092 of the table's questions (34,733 pairs) |
| RadFM | VinDr-CXR | 54.7 | 8,612 |
What this shows. Every model changes its yes/no answer on a meaningful share of clinically equivalent rephrasings, and the spread across models is 8.5-fold (6.4% for MedGemma-27B on MIMIC-CXR to 54.7% for RadFM on VinDr-CXR). No model is stable on all three populations, and rank order changes with the dataset.
Population. Six base models (MedGemma-1.5-4B, MedGemma-4B, MedGemma-27B, LLaVA-Rad, CheXone, RadFM) on the equivalence-filtered PSF-Med binary yes/no subset. Denominator is the paraphrase pair: MIMIC-CXR 5,076 pairs over 1,539 questions, PadChest 36,244 over 8,445, VinDr-CXR 8,612 over 2,807. RadFM's PadChest cell covers 8,092 questions / 34,733 pairs.
- Thesis
- Chapter 3, Table 3.3 (tab:t1_flip_rates)
- Paper
- PSF-Med: A Clinician-Audited Benchmark for Paraphrase Sensitivity in Medical Vision-Language Models
- Source artifact
results/uai/revision/psf_binary_recompute_v2.jsonin the medical-vlm-robustness repository- Related results
- flip-rate-range, flip-rate-medgemma4b-mimic, flip-rate-medgemma4b-padchest
Query-level flip rate: at least one paraphrase flips
Same PSF-Med binary subset as the pairwise chart, re-scored per question (MIMIC-CXR n=1,539, PadChest n=8,445, VinDr-CXR n=2,807 questions). MedGemma-27B on PadChest is absent from the recompute file and is not plotted.
- MedGemma-1.5-4B
- MedGemma-4B
- LLaVA-Rad
- MedGemma-27B
- CheXone
- RadFM
| Series | Category | percent% | n | Note |
|---|---|---|---|---|
| MedGemma-1.5-4B | MIMIC-CXR | 11.0 | 1,539 | |
| MedGemma-1.5-4B | PadChest | 43.6 | 8,445 | |
| MedGemma-1.5-4B | VinDr-CXR | 1.9 | 2,807 | degenerate yes-bias cell (over 98% positive on originals) |
| MedGemma-4B | MIMIC-CXR | 18.1 | 1,539 | |
| MedGemma-4B | PadChest | 32.2 | 8,445 | dissertation Table 3.4 value; the recompute file gives 32.1% |
| MedGemma-4B | VinDr-CXR | 28.5 | 2,807 | |
| LLaVA-Rad | MIMIC-CXR | 30.1 | 1,539 | |
| LLaVA-Rad | PadChest | 2.5 | 8,445 | degenerate yes-bias cell (over 98% positive on originals) |
| LLaVA-Rad | VinDr-CXR | 17.7 | 2,807 | |
| MedGemma-27B | MIMIC-CXR | 13.6 | 1,539 | |
| MedGemma-27B | VinDr-CXR | 17.0 | 2,807 | |
| CheXone | MIMIC-CXR | 14.6 | 1,539 | |
| CheXone | PadChest | 30.2 | 8,445 | |
| CheXone | VinDr-CXR | 24.9 | 2,807 | |
| RadFM | MIMIC-CXR | 25.1 | 1,539 | |
| RadFM | PadChest | 68.4 | 8,092 | |
| RadFM | VinDr-CXR | 60.7 | 2,807 |
What this shows. Counting a question as unstable if any of its roughly 3.3 paraphrases flips roughly doubles every rate: MedGemma-4B moves from 8.3% pairwise to 18.1% on MIMIC-CXR and from 13.4% to 32.2% on PadChest. Per-question exposure is what a clinician meets, and it is far higher than the pairwise number.
Population. Six base models on the equivalence-filtered PSF-Med binary yes/no subset. Denominator is the question: MIMIC-CXR 1,539, PadChest 8,445, VinDr-CXR 2,807. RadFM's PadChest cell covers 8,092 questions. MedGemma-27B / PadChest has no cell in the source file.
- Thesis
- Chapter 3, Table 3.4 (tab:t1_denominators)
- Paper
- PSF-Med: A Clinician-Audited Benchmark for Paraphrase Sensitivity in Medical Vision-Language Models
- Source artifact
results/uai/revision/psf_binary_recompute_v2.jsonin the medical-vlm-robustness repository- Related results
- pairwise-vs-query-medgemma4b-padchest, flip-rate-medgemma4b-padchest
Accuracy against pairwise flip rate
17 model-dataset cells (six base models x three datasets; MedGemma-27B / PadChest absent) on the PSF-Med binary subset, 5,076 / 36,244 / 8,612 pairs.
- MIMIC-CXR
- PadChest
- VinDr-CXR
| Series | Pairwise flip rate | percent% | n | Note |
|---|---|---|---|---|
| MIMIC-CXR | 7.41 | 57.1 | 5,076 | medgemma-15-4b |
| PadChest | 17.47 | 73.1 | 36,244 | medgemma-15-4b |
| VinDr-CXR | 0.78 | 100.0 | 8,612 | medgemma-15-4b — degenerate yes-bias cell (over 98% positive on originals) |
| MIMIC-CXR | 8.29 | 82.7 | 5,076 | medgemma-4b |
| PadChest | 13.35 | 41.4 | 36,244 | medgemma-4b |
| VinDr-CXR | 15.27 | 50.0 | 8,612 | medgemma-4b |
| MIMIC-CXR | 15.62 | 64.9 | 5,076 | llava-rad |
| PadChest | 0.74 | 83.3 | 36,244 | llava-rad — degenerate yes-bias cell (over 98% positive on originals) |
| VinDr-CXR | 11.57 | 85.8 | 8,612 | llava-rad |
| MIMIC-CXR | 6.4 | 76.8 | 5,076 | medgemma-27b |
| VinDr-CXR | 8.07 | 66.5 | 8,612 | medgemma-27b |
| MIMIC-CXR | 8.18 | 68.6 | 5,076 | chexone |
| PadChest | 13.04 | 41.3 | 36,244 | chexone |
| VinDr-CXR | 11.03 | 28.8 | 8,612 | chexone |
| MIMIC-CXR | 13.73 | 49.0 | 5,076 | radfm |
| PadChest | 32.76 | 40.6 | 34,733 | radfm |
| VinDr-CXR | 54.71 | 60.2 | 8,612 | radfm |
What this shows. Consistency and correctness do not travel together. The most consistent cells include the degenerate ones, and on VinDr-CXR (all-positive) a model that always answers yes scores 100% accuracy at a 0.8% flip rate. Reading either axis alone ranks models wrongly.
Population. Six base models x three datasets on the equivalence-filtered PSF-Med binary yes/no subset (MedGemma-27B / PadChest has no cell). x = pairwise flip rate over paraphrase pairs; y = weak-proxy accuracy over original questions in the same run.
- Thesis
- Chapter 3, Table 3.3 / Table 3.4
- Paper
- PSF-Med: A Clinician-Audited Benchmark for Paraphrase Sensitivity in Medical Vision-Language Models
- Source artifact
results/uai/revision/psf_binary_recompute_v2.jsonin the medical-vlm-robustness repository- Related results
- flip-rate-range, scale-no-consistency
How often the answer survives deleting the image
Two separate populations: the 107-pair three-backend diagnostic set (MedGemma-4B, MedGemma-27B, LLaVA-Rad) and the clean patient-disjoint safety re-audit, presence-of-finding endpoint, n=241 questions (base, Targeted LoRA over 5 seeds, Full LoRA over 3 seeds).
- MedGemma-4B
- MedGemma-27B
- LLaVA-Rad
- Targeted LoRA
- Full LoRA
| Series | Category | percent% | 95% interval | n | Note |
|---|---|---|---|---|---|
| MedGemma-4B | Three-backend diagnostic set (n=107 pairs) | 66.4 | — | 107 | |
| MedGemma-27B | Three-backend diagnostic set (n=107 pairs) | 85.0 | — | 107 | |
| LLaVA-Rad | Three-backend diagnostic set (n=107 pairs) | 96.3 | — | 107 | |
| MedGemma-4B | Patient-disjoint safety re-audit (n=241 questions) | 53.5 | — | 241 | base model, presence-of-finding endpoint |
| Targeted LoRA | Patient-disjoint safety re-audit (n=241 questions) | 76.8 | 73.7 to 80.0 | 241 | mean +/- 1 sd over 5 seeds |
| Full LoRA | Patient-disjoint safety re-audit (n=241 questions) | 76.6 | 74.9 to 78.4 | 241 | mean +/- 1 sd over 3 seeds |
What this shows. A high bar means the model would give the same answer with no image at all. On the 107-pair set LLaVA-Rad reaches 96.3%, so it is effectively a text-only model that accepts an image. On the patient-disjoint audit both adapters buy their consistency by leaning harder on the question text: text-only agreement rises about 23 points above the base model, and Targeted and Full LoRA are statistically tied.
Population. Group 1: the 107-pair three-backend diagnostic set (MedGemma-4B, MedGemma-27B, LLaVA-Rad), denominator = pair. Group 2: the patient- and image-disjoint MIMIC-CXR safety re-audit, presence-of-finding endpoint, 241 questions, 239 held-out subjects, zero subject and zero image overlap with training; Targeted LoRA is the mean over seeds 42/123/456/789/2024 and Full LoRA over seeds 42/123/456.
- Thesis
- Chapter 4, Table 4.x (tab:backend_summary) and the Chapter 5 patient-disjoint safety audit
- Paper
- Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models
- Source artifact
results/lora_fleet_patient_disjoint/audit_targeted_s42.jsonin the medical-vlm-robustness repository- Related results
- text-only-llavarad-107, text-only-medgemma4b-107, lora-textonly-increase-pd
Four-quadrant safety screen: consistency against image reliance
Ten model-dataset settings on the curated behavioural flip banks: MedGemma variants on MIMIC-CXR n=98 and PadChest n=861; LLaVA-Rad variants on MIMIC-CXR n=88 and PadChest n=732 (its text-only-baseline subset).
Each bar is one model-dataset setting, split by how its predictions land on the two axes. The red segment, consistent but image-invariant, is the failure this dissertation is about: taller red is worse, and the adapters have the tallest.
- Ideal (consistent, image-reliant)
- Fragile (inconsistent, image-reliant)
- Dangerous (consistent, text-reliant)
- Worst (inconsistent, text-reliant)
| Series | Category | percent% | n |
|---|---|---|---|
| Ideal (consistent, image-reliant) | Base MedGemma-4B / MIMIC-CXR | 31.6 | 98 |
| Fragile (inconsistent, image-reliant) | Base MedGemma-4B / MIMIC-CXR | 13.3 | 98 |
| Dangerous (consistent, text-reliant) | Base MedGemma-4B / MIMIC-CXR | 25.5 | 98 |
| Worst (inconsistent, text-reliant) | Base MedGemma-4B / MIMIC-CXR | 29.6 | 98 |
| Ideal (consistent, image-reliant) | Base MedGemma-4B / PadChest | 3.5 | 861 |
| Fragile (inconsistent, image-reliant) | Base MedGemma-4B / PadChest | 19.5 | 861 |
| Dangerous (consistent, text-reliant) | Base MedGemma-4B / PadChest | 22.5 | 861 |
| Worst (inconsistent, text-reliant) | Base MedGemma-4B / PadChest | 54.5 | 861 |
| Ideal (consistent, image-reliant) | Targeted LoRA / MIMIC-CXR | 21.4 | 98 |
| Fragile (inconsistent, image-reliant) | Targeted LoRA / MIMIC-CXR | 13.3 | 98 |
| Dangerous (consistent, text-reliant) | Targeted LoRA / MIMIC-CXR | 59.2 | 98 |
| Worst (inconsistent, text-reliant) | Targeted LoRA / MIMIC-CXR | 6.1 | 98 |
| Ideal (consistent, image-reliant) | Targeted LoRA / PadChest | 2.7 | 861 |
| Fragile (inconsistent, image-reliant) | Targeted LoRA / PadChest | 7.7 | 861 |
| Dangerous (consistent, text-reliant) | Targeted LoRA / PadChest | 84.1 | 861 |
| Worst (inconsistent, text-reliant) | Targeted LoRA / PadChest | 5.6 | 861 |
| Ideal (consistent, image-reliant) | Full LoRA / MIMIC-CXR | 17.3 | 98 |
| Fragile (inconsistent, image-reliant) | Full LoRA / MIMIC-CXR | 2.0 | 98 |
| Dangerous (consistent, text-reliant) | Full LoRA / MIMIC-CXR | 78.6 | 98 |
| Worst (inconsistent, text-reliant) | Full LoRA / MIMIC-CXR | 2.0 | 98 |
| Ideal (consistent, image-reliant) | Full LoRA / PadChest | 16.3 | 861 |
| Fragile (inconsistent, image-reliant) | Full LoRA / PadChest | 7.7 | 861 |
| Dangerous (consistent, text-reliant) | Full LoRA / PadChest | 60.6 | 861 |
| Worst (inconsistent, text-reliant) | Full LoRA / PadChest | 15.4 | 861 |
| Ideal (consistent, image-reliant) | LLaVA-Rad base / MIMIC-CXR | 2.3 | 88 |
| Fragile (inconsistent, image-reliant) | LLaVA-Rad base / MIMIC-CXR | 18.2 | 88 |
| Dangerous (consistent, text-reliant) | LLaVA-Rad base / MIMIC-CXR | 62.5 | 88 |
| Worst (inconsistent, text-reliant) | LLaVA-Rad base / MIMIC-CXR | 17.1 | 88 |
| Ideal (consistent, image-reliant) | LLaVA-Rad base / PadChest | 0.0 | 732 |
| Fragile (inconsistent, image-reliant) | LLaVA-Rad base / PadChest | 10.4 | 732 |
| Dangerous (consistent, text-reliant) | LLaVA-Rad base / PadChest | 79.5 | 732 |
| Worst (inconsistent, text-reliant) | LLaVA-Rad base / PadChest | 10.1 | 732 |
| Ideal (consistent, image-reliant) | LLaVA-Rad LoRA / MIMIC-CXR | 21.6 | 88 |
| Fragile (inconsistent, image-reliant) | LLaVA-Rad LoRA / MIMIC-CXR | 10.2 | 88 |
| Dangerous (consistent, text-reliant) | LLaVA-Rad LoRA / MIMIC-CXR | 58.0 | 88 |
| Worst (inconsistent, text-reliant) | LLaVA-Rad LoRA / MIMIC-CXR | 10.2 | 88 |
| Ideal (consistent, image-reliant) | LLaVA-Rad LoRA / PadChest | 16.0 | 732 |
| Fragile (inconsistent, image-reliant) | LLaVA-Rad LoRA / PadChest | 7.1 | 732 |
| Dangerous (consistent, text-reliant) | LLaVA-Rad LoRA / PadChest | 56.7 | 732 |
| Worst (inconsistent, text-reliant) | LLaVA-Rad LoRA / PadChest | 20.2 | 732 |
What this shows. The Dangerous cell (consistent but text-reliant) is the largest in eight of the ten settings, and it grows when consistency is optimised: the base model on MIMIC-CXR is 25.5% Dangerous while Targeted LoRA is 59.2% and Full LoRA 78.6%. Averaged across the ten settings, 81% of each model's consistent predictions are image-invariant (range 45-100%).
Population. Curated behavioural flip banks, question-level. MedGemma-4B base / Targeted LoRA / Full LoRA on MIMIC-CXR (n=98, the subset with a text-only baseline) and PadChest (n=861); LLaVA-Rad base and LLaVA-Rad LoRA on MIMIC-CXR (n=88) and PadChest (n=732). These flip banks are distinct from the PSF-Med benchmark and must not be pooled with it.
- Thesis
- Chapter 6, Table 6.x (tab:quadrant_counts)
- Paper
- Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models
- Source artifact
dissertation/tables/thrust4/table_quadrant_counts.texin the medical-vlm-robustness repository
Image-swap sensitivity: how often a contradicting image changes the answer
One population only: the 861-question PadChest flip bank, four models (base MedGemma-4B, Targeted LoRA, Full LoRA, LLaVA-Rad). MIMIC-CXR cells are excluded.
- MedGemma-4B
- Targeted LoRA
- Full LoRA
- LLaVA-Rad
| Series | Category | percent% | n |
|---|---|---|---|
| MedGemma-4B | PadChest flip bank | 39.5 | 861 |
| Targeted LoRA | PadChest flip bank | 31.5 | 861 |
| Full LoRA | PadChest flip bank | 28.2 | 861 |
| LLaVA-Rad | PadChest flip bank | 14.5 | 861 |
What this shows. Replacing the radiograph with one showing the opposite finding changes the answer for at most 39.5% of questions, and for LLaVA-Rad only 14.5%. Even the most image-sensitive model here leaves roughly six answers in ten unchanged when the visual evidence is contradicted.
Population. The 861-question PadChest flip bank, question-level denominator. Swap sensitivity = fraction of questions whose answer changes when the image is replaced with one showing the opposite finding, computed by the safety-evaluation pipeline (results/miccai) alongside flip rate and text-only agreement.
- Thesis
- Chapter 6, Table 6.x (tab:backend_safety)
- Paper
- Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models
- Source artifact
dissertation/tables/thrust4/table_backend_safety.texin the medical-vlm-robustness repository- Related results
- swap-sensitivity-medgemma4b-107
LoRA flip reduction, seed by seed
Clean patient- and image-disjoint MIMIC-CXR test, presence-of-finding endpoint: 238 questions / 991 pairs, 236 held-out subjects, zero subject and zero image overlap with training. Targeted LoRA seeds 42/123/456/789/2024; Full LoRA seeds 42/123/456.
- MedGemma-4B
- Targeted LoRA
- Full LoRA
| Series | Seed | percent% | n | Note |
|---|---|---|---|---|
| MedGemma-4B | 42 | 8.5 | 991 | base model; identical in every seed (the adapter is the only thing that varies) |
| MedGemma-4B | 123 | 8.5 | 991 | base model; identical in every seed (the adapter is the only thing that varies) |
| MedGemma-4B | 456 | 8.5 | 991 | base model; identical in every seed (the adapter is the only thing that varies) |
| MedGemma-4B | 789 | 8.5 | 991 | base model; identical in every seed (the adapter is the only thing that varies) |
| MedGemma-4B | 2024 | 8.5 | 991 | base model; identical in every seed (the adapter is the only thing that varies) |
| Targeted LoRA | 42 | 2.7 | 991 | McNemar exact two-sided p = 2.5e-07 |
| Targeted LoRA | 123 | 3.6 | 991 | McNemar exact two-sided p = 1.2e-03 |
| Targeted LoRA | 456 | 3.3 | 991 | McNemar exact two-sided p = 1.8e-04 |
| Targeted LoRA | 789 | 3.4 | 991 | McNemar exact two-sided p = 5.5e-06 |
| Targeted LoRA | 2024 | 4.4 | 991 | McNemar exact two-sided p = 1.8e-03 |
| Full LoRA | 42 | 2.6 | 991 | McNemar exact two-sided p = 5.7e-05 |
| Full LoRA | 123 | 3.1 | 991 | McNemar exact two-sided p = 1.3e-03 |
| Full LoRA | 456 | 3.0 | 991 | McNemar exact two-sided p = 2.4e-05 |
What this shows. Every adapter seed lands well below the base model's 8.5% pairwise flip rate: Targeted LoRA averages 3.5% (about a 59% reduction) and Full LoRA 2.9%, with McNemar exact p between 2.5e-7 and 1.8e-3 on every seed. The seed spread is small relative to the base-to-adapter gap.
Population. Base MedGemma-4B versus Targeted LoRA (layers 15-19, r=16, alpha=32, 4.38M params) and Full LoRA (all 34 layers, 30.5M params) on the clean patient-disjoint MIMIC-CXR test. Primary endpoint = presence-of-finding questions: 238 questions, 991 paraphrase pairs; denominator is the paraphrase pair. Base is a single run and is identical across seeds by construction.
- Thesis
- Chapter 5, Table 5.x (tab:pd_safety_audit)
- Paper
- Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
- Source artifact
results/lora_fleet_patient_disjoint/eval_targeted_s42.jsonin the medical-vlm-robustness repository- Related results
- lora-flip-reduction-pd, lora-param-fraction, lora-textonly-increase-pd
Accuracy on the clean patient-disjoint endpoint, with 95% Wilson intervals
Presence-of-finding endpoint, n=238 questions per arm on the patient- and image-disjoint MIMIC-CXR test (236 held-out subjects, zero subject and zero image overlap with training). One point per adapter seed.
- MedGemma-4B
- Targeted LoRA
- Full LoRA
| Series | Category | percent% | 95% interval | n | Note |
|---|---|---|---|---|---|
| MedGemma-4B | Base (no adapter) | 84.5 | 79.3 to 88.5 | 238 | 201/238 correct |
| Targeted LoRA | Targeted LoRA seed 42 | 85.3 | 80.2 to 89.2 | 238 | 203/238 correct |
| Targeted LoRA | Targeted LoRA seed 123 | 85.7 | 80.7 to 89.6 | 238 | 204/238 correct |
| Targeted LoRA | Targeted LoRA seed 456 | 85.3 | 80.2 to 89.2 | 238 | 203/238 correct |
| Targeted LoRA | Targeted LoRA seed 789 | 83.6 | 78.4 to 87.8 | 238 | 199/238 correct |
| Targeted LoRA | Targeted LoRA seed 2024 | 83.6 | 78.4 to 87.8 | 238 | 199/238 correct |
| Full LoRA | Full LoRA seed 42 | 83.2 | 77.9 to 87.4 | 238 | 198/238 correct |
| Full LoRA | Full LoRA seed 123 | 83.2 | 77.9 to 87.4 | 238 | 198/238 correct |
| Full LoRA | Full LoRA seed 456 | 83.6 | 78.4 to 87.8 | 238 | 199/238 correct |
What this shows. Every arm's interval overlaps every other arm's: the base model scores 84.5% and Targeted LoRA 84.7% on average, so the large flip reduction is not paid for by a large accuracy loss.
Population. Base MedGemma-4B, Targeted LoRA (5 seeds) and Full LoRA (3 seeds) on the clean patient- and image-disjoint MIMIC-CXR test, presence-of-finding primary endpoint. Denominator = 238 original questions per arm (question-level accuracy on the original phrasing). Interval method: 95% Wilson score interval on the binomial proportion of correct originals, computed per arm from the per-question block, z = 1.96, no continuity correction, unpaired.
- Thesis
- Chapter 5, Table 5.x (tab:pd_safety_audit)
- Paper
- Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
- Source artifact
results/lora_fleet_patient_disjoint/eval_targeted_s42.jsonin the medical-vlm-robustness repository- Related results
- lora-accuracy-delta-pd, lora-flip-reduction-pd
Where the answer is committed: residual transplant flip rate by layer
1,396 detecting/missing residual-transplant pairs spanning 91 findings, MedGemma-4B. Only the four layer anchors stated in Chapter 5 are plotted (8% at 14, 35% at 15, 73% at 16, about 100% by 20); the curve is not interpolated.
- Answer-position transplant
- Image-token control
| Series | Layer | percent% | n | Note |
|---|---|---|---|---|
| Answer-position transplant | 14 | 8.0 | 1,396 | anchors stated in the text; no dense per-layer trajectory file exists |
| Answer-position transplant | 15 | 35.0 | 1,396 | anchors stated in the text; no dense per-layer trajectory file exists |
| Answer-position transplant | 16 | 73.0 | 1,396 | median commit layer; 95% bootstrap CI [16, 16] |
| Answer-position transplant | 20 | 100.0 | 1,396 | text states the curve reaches about 100% by layer 20 |
| Image-token control | 14 | 0.1 | 1,396 | upper bound: the control never exceeds 0.0007 (0.07%) at any layer |
| Image-token control | 15 | 0.1 | 1,396 | upper bound: the control never exceeds 0.0007 (0.07%) at any layer |
| Image-token control | 16 | 0.1 | 1,396 | upper bound: the control never exceeds 0.0007 (0.07%) at any layer |
| Image-token control | 20 | 0.1 | 1,396 | upper bound: the control never exceeds 0.0007 (0.07%) at any layer |
What this shows. Swapping a single layer's residual state at the answer position flips the model's committed answer on 73% of pairs at layer 16 and saturates by layer 20, while the image-token control never exceeds 0.07%. The commit is narrow-band: the median per-pair commit layer is 16 with a 95% bootstrap CI of [16, 16].
Population. MedGemma-4B, 1,396 detecting/missing minimal pairs spanning 91 findings drawn from the PadChest paraphrase population. Denominator = transplant pair. Control series = a pre-question image-token position, which a causal mask makes unable to receive the later donor state.
- Thesis
- Chapter 5, Figure 5.x (fig:t3_commit_curve)
- Paper
- Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
- Source artifact
dissertation/chapters/04_thrust3.texin the medical-vlm-robustness repository- Related results
- layer16-commit-rate
The diagnosis layer is not the best intervention layer
Five LoRA layer-range configurations (rank 16, alpha=32, combined loss, lambda=1.0) scored on the 355-question MIMIC-CXR validation split against the 1.87 no-adapter baseline.
- No adapter
- LoRA layer window
| Series | Category | logit margin difference | n | Note |
|---|---|---|---|---|
| No adapter | Baseline (no LoRA) | 1.87 | 355 | reference margin difference |
| LoRA layer window | Early (0-10) | 0.260 | 355 | 86% reduction from the 1.87 baseline |
| LoRA layer window | Random block (5-9) | 0.300 | 355 | 84% reduction |
| LoRA layer window | All (0-33) | 0.340 | 355 | 82% reduction |
| LoRA layer window | Middle (15-19) | 0.380 | 355 | 80% reduction; the mechanistically identified window |
| LoRA layer window | Late (25-33) | 0.700 | 355 | 63% reduction |
What this shows. Adapting early layers 0-10 cuts the paraphrase margin difference furthest (0.26, an 86% reduction), beating the mechanistically identified middle window 15-19 (0.38, 80%) by 6 points, and even a randomly chosen block at layers 5-9 (0.30) beats it. Where paraphrase sensitivity is diagnosed is not where it is best fixed.
Population. MedGemma-4B with five LoRA layer ranges (early 0-10, random block 5-9, all 0-33, middle 15-19, late 25-33), rank 16, alpha=32, combined loss lambda=1.0, evaluated on the 355-question MIMIC-CXR validation split. Denominator = validation question. Baseline = the un-adapted model, margin difference 1.87.
- Thesis
- Chapter 5, Table 5.x (tab:thrust3_layer_ablation)
- Paper
- Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
- Source artifact
dissertation/tables/thrust3/table_layer_ablation.texin the medical-vlm-robustness repository- Related results
- layer-ablation-early-vs-middle
Exploratory: layer-17 sparse-autoencoder features most associated with flipping
Top 8 GemmaScope 2 features at layer 17 by absolute rank-biserial association with flipping, Targeted LoRA on the 861-question PadChest flip bank (5,112 records, 305 flipped questions).
| Series | Category | rank-biserial correlation | n | Note |
|---|---|---|---|---|
| Layer 17 feature | #3818 | -0.462 | 861 | Feature 3818 — the candidate operator/register gate; prevalence 0.85 |
| Layer 17 feature | #21 | -0.393 | 861 | prevalence 0.61 |
| Layer 17 feature | #534 | -0.335 | 861 | prevalence 1.00 |
| Layer 17 feature | #2882 | -0.321 | 861 | prevalence 0.74 |
| Layer 17 feature | #6871 | -0.313 | 861 | prevalence 0.86 |
| Layer 17 feature | #4103 | -0.301 | 861 | prevalence 0.34 |
| Layer 17 feature | #12762 | 0.290 | 861 | prevalence 0.98 |
| Layer 17 feature | #2913 | 0.263 | 861 | prevalence 0.42 |
What this shows. Feature 3818 has the strongest association with flipping of any layer-17 feature (rank-biserial -0.462): when it is active the model flips less often. It encodes presence-versus-exclusion phrasing, which is the candidate mechanism the dissertation follows up.
Population. Targeted LoRA (layers 15-19) on the 861-question PadChest flip bank, 5,112 original/paraphrase records, 305 flipped questions (35.4% question-level flip rate in this screen). GemmaScope 2 sparse autoencoder at layer 17; denominator = feature. Association is the rank-biserial correlation between a feature's activation and whether the pair flipped.
- Thesis
- Chapter 5, Table 5.x (tab:f3818_controls)
- Paper
- Mechanistically Guided LoRA Improves Paraphrase Consistency in Medical Vision-Language Models
- Source artifact
results/cross_modal/chil_lora/padchest_feature_ranking.jsonin the medical-vlm-robustness repository- Related results
- feature3818-operator-preserving-top1, feature3818-patch-recovery
Predictive entropy of stable versus flipped predictions
Targeted LoRA on the 861-question PadChest flip bank: 747 stable questions and 114 that flipped under paraphrase.
- Stable
- Flipped
| Series | Predictive entropy (nats) | question count | n |
|---|---|---|---|
| Stable | 0.017 | 0.000 | 747 |
| Flipped | 0.017 | 0.000 | 114 |
| Stable | 0.052 | 0.000 | 747 |
| Flipped | 0.052 | 0.000 | 114 |
| Stable | 0.087 | 0.000 | 747 |
| Flipped | 0.087 | 0.000 | 114 |
| Stable | 0.121 | 0.000 | 747 |
| Flipped | 0.121 | 0.000 | 114 |
| Stable | 0.156 | 0.000 | 747 |
| Flipped | 0.156 | 0.000 | 114 |
| Stable | 0.191 | 10.0 | 747 |
| Flipped | 0.191 | 0.000 | 114 |
| Stable | 0.225 | 32.0 | 747 |
| Flipped | 0.225 | 0.000 | 114 |
| Stable | 0.26 | 55.0 | 747 |
| Flipped | 0.26 | 2.00 | 114 |
| Stable | 0.295 | 84.0 | 747 |
| Flipped | 0.295 | 1.00 | 114 |
| Stable | 0.329 | 80.0 | 747 |
| Flipped | 0.329 | 4.00 | 114 |
| Stable | 0.364 | 81.0 | 747 |
| Flipped | 0.364 | 4.00 | 114 |
| Stable | 0.399 | 53.0 | 747 |
| Flipped | 0.399 | 0.000 | 114 |
| Stable | 0.433 | 83.0 | 747 |
| Flipped | 0.433 | 7.00 | 114 |
| Stable | 0.468 | 43.0 | 747 |
| Flipped | 0.468 | 10.0 | 114 |
| Stable | 0.503 | 57.0 | 747 |
| Flipped | 0.503 | 4.00 | 114 |
| Stable | 0.537 | 25.0 | 747 |
| Flipped | 0.537 | 5.00 | 114 |
| Stable | 0.572 | 32.0 | 747 |
| Flipped | 0.572 | 6.00 | 114 |
| Stable | 0.607 | 35.0 | 747 |
| Flipped | 0.607 | 12.0 | 114 |
| Stable | 0.641 | 28.0 | 747 |
| Flipped | 0.641 | 10.0 | 114 |
| Stable | 0.676 | 49.0 | 747 |
| Flipped | 0.676 | 49.0 | 114 |
What this shows. Questions whose answer flips under paraphrase sit at visibly higher single-pass predictive entropy than stable ones, which is what lets one forward pass rank flip risk (AUROC 0.823).
Population. Targeted LoRA (MedGemma-4B, layers 15-19) on the 861-question PadChest flip bank, one question per row, single-pass softmax entropy over the yes/no answer distribution.
- Thesis
- Chapter 7, Table 7.x (tab:bridge_auroc)
- Paper
- Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models
- Source artifact
results/uai/chil_lora/padchest/softmax_entropy.jsonlin the medical-vlm-robustness repository- Related results
- entropy-flip-auroc
Predicting a paraphrase flip from uncertainty on the original question
Targeted LoRA on the 861-question PadChest flip bank (temperature scaling on the 732-question subset that survives its 15% calibration holdout). AUROC for predicting whether the question will flip.
Each bar is a model: the height is how well entropy on the original question ranks which answers will flip when reworded. 0.5 is a coin flip, 1.0 is perfect, so higher is better.
| Series | Category | AUROC | n | Note |
|---|---|---|---|---|
| Targeted LoRA, PadChest | Softmax entropy | 0.823 | 861 | p = 4.3e-29 |
| Targeted LoRA, PadChest | Monte Carlo dropout | 0.823 | 861 | p = 4.6e-29 |
| Targeted LoRA, PadChest | Deep ensemble | 0.552 | 861 | p = 3.7e-2; the only genuinely distinct signal, and it fails |
| Targeted LoRA, PadChest | Temperature scaling | 0.813 | 732 | n=732: temperature is fitted on a 15% calibration holdout, so 129 rows are withheld |
| Targeted LoRA, PadChest | Absolute margin | 0.823 | 861 | rank-equivalent to softmax entropy |
What this shows. Cheap predictive entropy on the original question predicts whether that question will flip under paraphrasing at AUROC 0.823, so uncertainty and paraphrase sensitivity are measuring related failure. The expensive deep ensemble is the worst predictor at 0.552.
Population. Targeted LoRA (layers 15-19) on the 861-question PadChest flip bank; denominator = question. Positive class = the question flips under at least one paraphrase. Temperature scaling is scored on the 732 questions outside its 15% calibration holdout.
- Thesis
- Chapter 7, Table 7.x (tab:bridge_auroc)
- Paper
- Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models
- Source artifact
dissertation/tables/thrust4/table_bridge_auroc.texin the medical-vlm-robustness repository- Related results
- entropy-flip-auroc, entropy-error-auroc, ensemble-ood-failure
Risk-coverage: what accepting fewer predictions buys you
PadChest flip bank, n=861 questions per model. Selective risk (error rate among accepted) against coverage, ranking by softmax confidence, for base MedGemma-4B, Targeted LoRA and Full LoRA. Curves recomputed from the post-image-fix run and downsampled to 50 points each.
Moving right along a line answers more questions automatically; the height is the error rate among those answered. Lower lines are better, and the gap between models widens as coverage grows.
- MedGemma-4B
- Targeted LoRA
- Full LoRA
| Series | Coverage | selective risk (error rate on accepted) | n |
|---|---|---|---|
| MedGemma-4B | 0.0012 | 0.000 | 861 |
| MedGemma-4B | 0.0221 | 0.000 | 861 |
| MedGemma-4B | 0.0418 | 0.000 | 861 |
| MedGemma-4B | 0.0627 | 0.000 | 861 |
| MedGemma-4B | 0.0825 | 0.000 | 861 |
| MedGemma-4B | 0.1034 | 0.000 | 861 |
| MedGemma-4B | 0.1231 | 0.000 | 861 |
| MedGemma-4B | 0.144 | 0.000 | 861 |
| MedGemma-4B | 0.1638 | 0.000 | 861 |
| MedGemma-4B | 0.1847 | 0.000 | 861 |
| MedGemma-4B | 0.2056 | 0.000 | 861 |
| MedGemma-4B | 0.2253 | 0.000 | 861 |
| MedGemma-4B | 0.2462 | 0.000 | 861 |
| MedGemma-4B | 0.266 | 0.000 | 861 |
| MedGemma-4B | 0.2869 | 0.000 | 861 |
| MedGemma-4B | 0.3066 | 0.000 | 861 |
| MedGemma-4B | 0.3275 | 0.000 | 861 |
| MedGemma-4B | 0.3473 | 0.003 | 861 |
| MedGemma-4B | 0.3682 | 0.003 | 861 |
| MedGemma-4B | 0.3879 | 0.003 | 861 |
| MedGemma-4B | 0.4088 | 0.006 | 861 |
| MedGemma-4B | 0.4297 | 0.016 | 861 |
| MedGemma-4B | 0.4495 | 0.018 | 861 |
| MedGemma-4B | 0.4704 | 0.022 | 861 |
| MedGemma-4B | 0.4901 | 0.026 | 861 |
| MedGemma-4B | 0.511 | 0.027 | 861 |
| MedGemma-4B | 0.5308 | 0.028 | 861 |
| MedGemma-4B | 0.5517 | 0.029 | 861 |
| MedGemma-4B | 0.5714 | 0.033 | 861 |
| MedGemma-4B | 0.5923 | 0.033 | 861 |
| MedGemma-4B | 0.6132 | 0.034 | 861 |
| MedGemma-4B | 0.633 | 0.037 | 861 |
| MedGemma-4B | 0.6539 | 0.046 | 861 |
| MedGemma-4B | 0.6736 | 0.057 | 861 |
| MedGemma-4B | 0.6945 | 0.070 | 861 |
| MedGemma-4B | 0.7143 | 0.085 | 861 |
| MedGemma-4B | 0.7352 | 0.100 | 861 |
| MedGemma-4B | 0.7549 | 0.115 | 861 |
| MedGemma-4B | 0.7758 | 0.127 | 861 |
| MedGemma-4B | 0.7956 | 0.139 | 861 |
| MedGemma-4B | 0.8165 | 0.154 | 861 |
| MedGemma-4B | 0.8374 | 0.161 | 861 |
| MedGemma-4B | 0.8571 | 0.168 | 861 |
| MedGemma-4B | 0.878 | 0.172 | 861 |
| MedGemma-4B | 0.8978 | 0.181 | 861 |
| MedGemma-4B | 0.9187 | 0.188 | 861 |
| MedGemma-4B | 0.9384 | 0.197 | 861 |
| MedGemma-4B | 0.9593 | 0.202 | 861 |
| MedGemma-4B | 0.9791 | 0.203 | 861 |
| MedGemma-4B | 1 | 0.212 | 861 |
| Targeted LoRA | 0.0012 | 0.000 | 861 |
| Targeted LoRA | 0.0221 | 0.000 | 861 |
| Targeted LoRA | 0.0418 | 0.000 | 861 |
| Targeted LoRA | 0.0627 | 0.000 | 861 |
| Targeted LoRA | 0.0825 | 0.000 | 861 |
| Targeted LoRA | 0.1034 | 0.000 | 861 |
| Targeted LoRA | 0.1231 | 0.009 | 861 |
| Targeted LoRA | 0.144 | 0.008 | 861 |
| Targeted LoRA | 0.1638 | 0.007 | 861 |
| Targeted LoRA | 0.1847 | 0.006 | 861 |
| Targeted LoRA | 0.2056 | 0.006 | 861 |
| Targeted LoRA | 0.2253 | 0.005 | 861 |
| Targeted LoRA | 0.2462 | 0.005 | 861 |
| Targeted LoRA | 0.266 | 0.004 | 861 |
| Targeted LoRA | 0.2869 | 0.004 | 861 |
| Targeted LoRA | 0.3066 | 0.004 | 861 |
| Targeted LoRA | 0.3275 | 0.007 | 861 |
| Targeted LoRA | 0.3473 | 0.007 | 861 |
| Targeted LoRA | 0.3682 | 0.006 | 861 |
| Targeted LoRA | 0.3879 | 0.006 | 861 |
| Targeted LoRA | 0.4088 | 0.009 | 861 |
| Targeted LoRA | 0.4297 | 0.008 | 861 |
| Targeted LoRA | 0.4495 | 0.008 | 861 |
| Targeted LoRA | 0.4704 | 0.007 | 861 |
| Targeted LoRA | 0.4901 | 0.009 | 861 |
| Targeted LoRA | 0.511 | 0.009 | 861 |
| Targeted LoRA | 0.5308 | 0.011 | 861 |
| Targeted LoRA | 0.5517 | 0.011 | 861 |
| Targeted LoRA | 0.5714 | 0.010 | 861 |
| Targeted LoRA | 0.5923 | 0.012 | 861 |
| Targeted LoRA | 0.6132 | 0.015 | 861 |
| Targeted LoRA | 0.633 | 0.015 | 861 |
| Targeted LoRA | 0.6539 | 0.018 | 861 |
| Targeted LoRA | 0.6736 | 0.017 | 861 |
| Targeted LoRA | 0.6945 | 0.020 | 861 |
| Targeted LoRA | 0.7143 | 0.026 | 861 |
| Targeted LoRA | 0.7352 | 0.028 | 861 |
| Targeted LoRA | 0.7549 | 0.028 | 861 |
| Targeted LoRA | 0.7758 | 0.030 | 861 |
| Targeted LoRA | 0.7956 | 0.034 | 861 |
| Targeted LoRA | 0.8165 | 0.037 | 861 |
| Targeted LoRA | 0.8374 | 0.043 | 861 |
| Targeted LoRA | 0.8571 | 0.045 | 861 |
| Targeted LoRA | 0.878 | 0.049 | 861 |
| Targeted LoRA | 0.8978 | 0.052 | 861 |
| Targeted LoRA | 0.9187 | 0.053 | 861 |
| Targeted LoRA | 0.9384 | 0.059 | 861 |
| Targeted LoRA | 0.9593 | 0.065 | 861 |
| Targeted LoRA | 0.9791 | 0.075 | 861 |
| Targeted LoRA | 1 | 0.086 | 861 |
| Full LoRA | 0.0012 | 1.00 | 861 |
| Full LoRA | 0.0221 | 1.00 | 861 |
| Full LoRA | 0.0418 | 0.944 | 861 |
| Full LoRA | 0.0627 | 0.833 | 861 |
| Full LoRA | 0.0825 | 0.662 | 861 |
| Full LoRA | 0.1034 | 0.551 | 861 |
| Full LoRA | 0.1231 | 0.481 | 861 |
| Full LoRA | 0.144 | 0.427 | 861 |
| Full LoRA | 0.1638 | 0.390 | 861 |
| Full LoRA | 0.1847 | 0.358 | 861 |
| Full LoRA | 0.2056 | 0.328 | 861 |
| Full LoRA | 0.2253 | 0.320 | 861 |
| Full LoRA | 0.2462 | 0.302 | 861 |
| Full LoRA | 0.266 | 0.293 | 861 |
| Full LoRA | 0.2869 | 0.271 | 861 |
| Full LoRA | 0.3066 | 0.273 | 861 |
| Full LoRA | 0.3275 | 0.266 | 861 |
| Full LoRA | 0.3473 | 0.264 | 861 |
| Full LoRA | 0.3682 | 0.256 | 861 |
| Full LoRA | 0.3879 | 0.264 | 861 |
| Full LoRA | 0.4088 | 0.270 | 861 |
| Full LoRA | 0.4297 | 0.276 | 861 |
| Full LoRA | 0.4495 | 0.282 | 861 |
| Full LoRA | 0.4704 | 0.281 | 861 |
| Full LoRA | 0.4901 | 0.282 | 861 |
| Full LoRA | 0.511 | 0.282 | 861 |
| Full LoRA | 0.5308 | 0.289 | 861 |
| Full LoRA | 0.5517 | 0.288 | 861 |
| Full LoRA | 0.5714 | 0.291 | 861 |
| Full LoRA | 0.5923 | 0.288 | 861 |
| Full LoRA | 0.6132 | 0.290 | 861 |
| Full LoRA | 0.633 | 0.294 | 861 |
| Full LoRA | 0.6539 | 0.295 | 861 |
| Full LoRA | 0.6736 | 0.293 | 861 |
| Full LoRA | 0.6945 | 0.291 | 861 |
| Full LoRA | 0.7143 | 0.289 | 861 |
| Full LoRA | 0.7352 | 0.291 | 861 |
| Full LoRA | 0.7549 | 0.291 | 861 |
| Full LoRA | 0.7758 | 0.298 | 861 |
| Full LoRA | 0.7956 | 0.301 | 861 |
| Full LoRA | 0.8165 | 0.300 | 861 |
| Full LoRA | 0.8374 | 0.305 | 861 |
| Full LoRA | 0.8571 | 0.310 | 861 |
| Full LoRA | 0.878 | 0.316 | 861 |
| Full LoRA | 0.8978 | 0.317 | 861 |
| Full LoRA | 0.9187 | 0.320 | 861 |
| Full LoRA | 0.9384 | 0.327 | 861 |
| Full LoRA | 0.9593 | 0.331 | 861 |
| Full LoRA | 0.9791 | 0.332 | 861 |
| Full LoRA | 1 | 0.338 | 861 |
What this shows. Targeted LoRA dominates: its risk stays near zero out to high coverage (AUGRC 0.014 against 0.047 for the base model and 0.153 for Full LoRA), so confidence-based deferral works for it and barely works for Full LoRA, whose curve rises almost immediately.
Population. Base MedGemma-4B, Targeted LoRA and Full LoRA on the 861-question PadChest flip bank, softmax-entropy method, one prediction per question. Selection score = max(p_yes, 1-p_yes); at coverage k/n the selective risk is errors among the k most confident, divided by k. Denominator = question.
- Thesis
- Chapter 7, Figure 7.x (risk-coverage)
- Paper
- Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models
- Source artifact
results/uai/base/padchest/softmax_entropy.jsonlin the medical-vlm-robustness repository- Related results
- augrc-targeted-padchest, augrc-fulllora-padchest
Offline readiness audit: admission rate against accuracy
Six cells: base MedGemma-4B, Targeted LoRA and Full LoRA on the MIMIC-CXR flip bank (n=98) and the PadChest flip bank (n=861). The audit admits a prediction only if the first two paraphrases agree with the original and the answer shows image reliance.
- Admitted by the audit
- Accuracy on admitted
- Accuracy on all questions
| Series | Category | percent% | n | Note |
|---|---|---|---|---|
| Admitted by the audit | Base MedGemma-4B / MIMIC-CXR | 22.4 | 98 | 22 of 98 questions admitted |
| Accuracy on admitted | Base MedGemma-4B / MIMIC-CXR | 95.5 | 98 | |
| Accuracy on all questions | Base MedGemma-4B / MIMIC-CXR | 83.7 | 98 | residual flip on uninspected paraphrases: 40.9% |
| Admitted by the audit | Base MedGemma-4B / PadChest | 41.6 | 861 | 358 of 861 questions admitted |
| Accuracy on admitted | Base MedGemma-4B / PadChest | 96.9 | 861 | |
| Accuracy on all questions | Base MedGemma-4B / PadChest | 78.6 | 861 | residual flip on uninspected paraphrases: 72.9% |
| Admitted by the audit | Targeted LoRA / MIMIC-CXR | 30.6 | 98 | 30 of 98 questions admitted |
| Accuracy on admitted | Targeted LoRA / MIMIC-CXR | 76.7 | 98 | |
| Accuracy on all questions | Targeted LoRA / MIMIC-CXR | 77.6 | 98 | residual flip on uninspected paraphrases: 23.3% |
| Admitted by the audit | Targeted LoRA / PadChest | 33.0 | 861 | 284 of 861 questions admitted |
| Accuracy on admitted | Targeted LoRA / PadChest | 96.8 | 861 | |
| Accuracy on all questions | Targeted LoRA / PadChest | 91.5 | 861 | residual flip on uninspected paraphrases: 4.2% |
| Admitted by the audit | Full LoRA / MIMIC-CXR | 14.3 | 98 | 14 of 98 questions admitted |
| Accuracy on admitted | Full LoRA / MIMIC-CXR | 92.9 | 98 | |
| Accuracy on all questions | Full LoRA / MIMIC-CXR | 81.6 | 98 | residual flip on uninspected paraphrases: 0.0% |
| Admitted by the audit | Full LoRA / PadChest | 26.7 | 861 | 230 of 861 questions admitted |
| Accuracy on admitted | Full LoRA / PadChest | 83.9 | 861 | |
| Accuracy on all questions | Full LoRA / PadChest | 66.0 | 861 | residual flip on uninspected paraphrases: 10.4% |
What this shows. Filtering on agreement plus image reliance lifts accuracy in five of six cells (base on PadChest: 78.6% to 96.9%), but admits only 14-42% of questions. The price of a trustworthy answer is refusing most of them.
Population. Base MedGemma-4B, Targeted LoRA and Full LoRA on the curated behavioural flip banks: MIMIC-CXR n=98 and PadChest n=861 questions. Denominator = question. Admission rule: agree(p1, p2) AND (swap_changed OR roi_matters), where roi_matters (region-only versus background-only crops disagree) is PadChest-only.
- Thesis
- Chapter 7, Table 7.x (tab:gate_results)
- Paper
- None
- Source artifact
results/miccai/deployment_gate.jsonin the medical-vlm-robustness repository- Related results
- gate-admission-targeted-padchest, gate-grounded-slice-failure
How many generated paraphrases survive the equivalence audit
All 122,778 candidate pairs adjudicated by the rubric audit: MIMIC-CXR 12,259, PadChest 79,378, VinDr-CXR 31,141. Overall 61,761 retained (50.3%), 59,788 rejected (48.7%), 1,229 uncertain (1.0%).
- Retained as equivalent
- Rejected as not equivalent
- Uncertain
| Series | Category | paraphrase pairs | n | Note |
|---|---|---|---|---|
| Retained as equivalent | MIMIC-CXR | 8933.0 | 12,259 | 72.87% retention of 12,259 adjudicated pairs |
| Rejected as not equivalent | MIMIC-CXR | 3113.0 | 12,259 | |
| Uncertain | MIMIC-CXR | 213.0 | 12,259 | |
| Retained as equivalent | PadChest | 28364.0 | 79,378 | 35.73% retention of 79,378 adjudicated pairs |
| Rejected as not equivalent | PadChest | 50200.0 | 79,378 | |
| Uncertain | PadChest | 814.0 | 79,378 | |
| Retained as equivalent | VinDr-CXR | 24464.0 | 31,141 | 78.56% retention of 31,141 adjudicated pairs |
| Rejected as not equivalent | VinDr-CXR | 6475.0 | 31,141 | |
| Uncertain | VinDr-CXR | 202.0 | 31,141 |
What this shows. Roughly half of everything an LLM generates as a 'paraphrase' of a clinical question is not clinically equivalent, and the rate depends strongly on the source: PadChest keeps only 35.7% against 78.6% for VinDr-CXR. Without this audit a benchmark would be scoring models on questions that genuinely changed meaning.
Population. All 122,778 construction-stage candidate paraphrase pairs adjudicated against the equivalence rubric, by source dataset. Denominator = candidate pair. Verdicts recomputed on 2026-07-15 from data/judge_full_results.json and data/vindr_judge/gpt_results.json; the total reproduces the 122,778 reported in Chapter 3 exactly.
- Thesis
- Chapter 3, Table 3.x (rubric audit retention)
- Paper
- PSF-Med: A Clinician-Audited Benchmark for Paraphrase Sensitivity in Medical Vision-Language Models
- Source artifact
scripts/analysis/recompute_audit_retention.pyin the medical-vlm-robustness repository- Related results
- audit-retention-overall, negation-rejection-rate, benchmark-eval-pairs
Which kinds of rephrasing break the model
Mean pairwise flip rate over the six base models on the PSF-Med binary yes/no subset, by transformation type. Per-cell n is binary-subset pairs per model (11 to 14,662). PadChest has no negation cell.
- MIMIC-CXR
- PadChest
- VinDr-CXR
| Series | Category | percent% | n | Note |
|---|---|---|---|---|
| MIMIC-CXR | Lexical substitution | 11.1 | 1,511 | |
| PadChest | Lexical substitution | 17.2 | 14,662 | |
| VinDr-CXR | Lexical substitution | 12.5 | 2,816 | |
| MIMIC-CXR | Syntactic restructuring | 8.7 | 1,400 | |
| PadChest | Syntactic restructuring | 16.5 | 8,235 | |
| VinDr-CXR | Syntactic restructuring | 18.6 | 1,145 | |
| MIMIC-CXR | Scope quantification | 9.8 | 1,183 | |
| PadChest | Scope quantification | 17.2 | 8,095 | |
| VinDr-CXR | Scope quantification | 14.1 | 1,842 | |
| MIMIC-CXR | Specificity modulation | 10.0 | 971 | |
| PadChest | Specificity modulation | 5.3 | 5,252 | |
| VinDr-CXR | Specificity modulation | 14.3 | 1,715 | |
| MIMIC-CXR | Negation pattern | 18.2 | 11 | only 11 pairs survive the audit on MIMIC-CXR; reported for completeness only |
| VinDr-CXR | Negation pattern | 34.7 | 1,126 | highest cell in the table; the VinDr regeneration is the only source with a usable negation pool |
What this shows. No transformation type is safe: even plain lexical substitution flips 11-17% of pairs. Negation-pattern rephrasings are the worst where a usable pool survives the audit (34.7% on VinDr-CXR), roughly triple the lexical rate on the same dataset.
Population. Mean over six base models (MedGemma-4B, MedGemma-1.5-4B, MedGemma-27B, LLaVA-Rad, CheXone, RadFM) on the equivalence-filtered PSF-Med binary yes/no subset, split by paraphrase transformation type. Denominator = paraphrase pair; per-cell n is pairs per model.
- Thesis
- Chapter 3, Table 3.x (tab:t1_para_types)
- Paper
- PSF-Med: A Clinician-Audited Benchmark for Paraphrase Sensitivity in Medical Vision-Language Models
- Source artifact
dissertation/tables/thrust1/table_paraphrase_types.texin the medical-vlm-robustness repository- Related results
- negation-rejection-rate, flip-rate-range
Accuracy by patient age band
PadChest flip bank only, n=861 questions per model split into four age bands (under 40 n=48, 40-60 n=233, 60-80 n=416, over 80 n=164), for base MedGemma-4B, Targeted LoRA and Full LoRA.
- MedGemma-4B
- Targeted LoRA
- Full LoRA
| Series | Category | percent% | n | Note |
|---|---|---|---|---|
| MedGemma-4B | Under 40 | 79.2 | 48 | smallest bin in the analysis |
| MedGemma-4B | 40 to 60 | 76.8 | 233 | |
| MedGemma-4B | 60 to 80 | 78.8 | 416 | |
| MedGemma-4B | Over 80 | 81.1 | 164 | |
| Targeted LoRA | Under 40 | 83.3 | 48 | smallest bin in the analysis |
| Targeted LoRA | 40 to 60 | 87.6 | 233 | |
| Targeted LoRA | 60 to 80 | 92.5 | 416 | |
| Targeted LoRA | Over 80 | 96.3 | 164 | |
| Full LoRA | Under 40 | 66.7 | 48 | smallest bin in the analysis |
| Full LoRA | 40 to 60 | 63.1 | 233 | |
| Full LoRA | 60 to 80 | 65.9 | 416 | |
| Full LoRA | Over 80 | 71.3 | 164 |
What this shows. Targeted LoRA is the most accurate arm in every band but has the steepest age gradient: 83.3% for patients under 40 against 96.3% for those over 80, a 13-point spread. The arm with the smallest between-sex calibration gap is therefore also the one with the widest age disparity, so no model is uniformly the most equitable.
Population. Base MedGemma-4B, Targeted LoRA and Full LoRA on the 861-question PadChest flip bank, softmax-entropy method, stratified by PatientAge from the PadChest master table. Denominator = question. Bands: under 40 (n=48), 40-60 (n=233), 60-80 (n=416), over 80 (n=164).
- Thesis
- Chapter 6, Table 6.x (tab:fairness)
- Paper
- Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models
- Source artifact
results/uai/fairness_analysis.jsonin the medical-vlm-robustness repository- Related results
- fairness-ece-sex-gap-targeted, fairness-age-gradient-targeted
Does failure detection survive image corruption?
AUGRC against corruption severity (0 = clean, then 1/3/5 averaged over five corruption types) on the 861-question PadChest flip bank, softmax entropy, for base MedGemma-4B, Targeted LoRA and Full LoRA. Bars are bootstrap 95% intervals. Lower is better.
- MedGemma-4B
- Targeted LoRA
- Full LoRA
| Series | Corruption severity | AUGRC | 95% interval | n | Note |
|---|---|---|---|---|---|
| MedGemma-4B | 0 | 0.047 | 0.038 to 0.056 | 861 | clean, uncorrupted images |
| MedGemma-4B | 1 | 0.049 | 0.041 to 0.059 | 861 | mean over 5 corruption types |
| MedGemma-4B | 3 | 0.059 | 0.049 to 0.070 | 861 | mean over 5 corruption types |
| MedGemma-4B | 5 | 0.088 | 0.076 to 0.102 | 861 | mean over 5 corruption types |
| Targeted LoRA | 0 | 0.014 | 0.010 to 0.019 | 861 | clean, uncorrupted images |
| Targeted LoRA | 1 | 0.015 | 0.010 to 0.021 | 861 | mean over 5 corruption types |
| Targeted LoRA | 3 | 0.018 | 0.013 to 0.024 | 861 | mean over 5 corruption types |
| Targeted LoRA | 5 | 0.027 | 0.021 to 0.035 | 861 | mean over 5 corruption types |
| Full LoRA | 0 | 0.153 | 0.136 to 0.172 | 861 | clean, uncorrupted images |
| Full LoRA | 1 | 0.154 | 0.136 to 0.172 | 861 | mean over 5 corruption types |
| Full LoRA | 3 | 0.162 | 0.144 to 0.180 | 861 | mean over 5 corruption types |
| Full LoRA | 5 | 0.185 | 0.167 to 0.204 | 861 | mean over 5 corruption types |
What this shows. Targeted LoRA's undetected-failure risk stays low and nearly flat under corruption (0.014 clean to 0.027 at severity 5) while the base model nearly doubles (0.047 to 0.088). Full LoRA is the worst at every severity (0.153 to 0.186): it is flat only because it starts badly.
Population. Base MedGemma-4B, Targeted LoRA and Full LoRA on the 861-question PadChest flip bank under synthetic corruption, softmax-entropy method. Denominator = question. AUGRC (Traub et al., NeurIPS 2024) is bounded in [0, 0.5] and reads as average undetected-failure risk; severity 0 is one clean run, severities 1/3/5 are the mean over five corruption types.
- Thesis
- Chapter 7, Table 7.x (augrc_degradation)
- Paper
- Predictive Entropy as a Joint Screen for Error and Paraphrase Instability in Medical Vision-Language Models
- Source artifact
results/uai/week2_analysis/degradation_summary.jsonin the medical-vlm-robustness repository- Related results
- augrc-targeted-padchest, augrc-fulllora-padchest, conformal-coverage-sev5-base
Paraphrase flip rates for frontier general-purpose models
Four configurations of three frontier models with no medical finetuning, scored on the same clinically equivalent paraphrase pairs: Claude Opus 5 on 6,209 pairs with reasoning on and 6,194 with it off, GPT-5.6 Sol on 6,195, all across MIMIC-CXR, PadChest and VinDr-CXR; Kimi K3 on 2,274 PadChest pairs only. Bars are Wilson 95% confidence intervals.
Each dot is one model configuration's pairwise flip rate, with its Wilson 95% confidence interval; lower is better and zero would mean the model never contradicts itself on a rephrasing. The four configurations sit at their reasoning or effort setting on the horizontal axis.
- Claude Opus 5
- GPT-5.6 Sol
- Kimi K3
| Series | Reasoning or effort setting | percent% | 95% interval | n | Note |
|---|---|---|---|---|---|
| Claude Opus 5 | Reasoning on | 9.5 | 8.8 to 10.3 | 6,209 | 592 flips of 6,209 pairs; all three datasets |
| Claude Opus 5 | Reasoning off | 8.4 | 7.7 to 9.1 | 6,194 | 518 flips of 6,194 pairs; all three datasets |
| GPT-5.6 Sol | Effort low | 5.3 | 4.8 to 5.9 | 6,195 | 329 flips of 6,195 pairs; all three datasets |
| Kimi K3 | Effort low | 3.1 | 2.4 to 3.9 | 2,274 | 70 flips of 2,274 pairs; PadChest only, so not comparable with the all-three-dataset runs |
What this shows. Frontier models that were never medically finetuned still change their yes/no answer on 3.1% to 9.5% of clinically equivalent rephrasings. That is lower than the 6.4% to 54.7% span the six medical models post on the binary subset, but it is nowhere near zero, so paraphrase sensitivity is not a defect of small domain-tuned models that frontier scale has removed. Two details carry most of the weight. Reasoning does not help: Claude Opus 5 flips more with reasoning on (9.53%) than with it off (8.36%), a 1.17 percentage-point increase in the wrong direction (two-proportion z = 2.29, p = 0.022, 95% CI on the difference 0.17 to 2.18 points). And on PadChest, the one dataset where a frontier run and the published medical cells cover the same population, Kimi K3's 3.08% sits against 12.0% to 32.8% for the medical models, so the frontier advantage on a like-for-like population is real and large.
Population. Three frontier general-purpose multimodal models with no medical finetuning, evaluated in a post-draft addendum on clinically equivalent paraphrase pairs from the PSF-Med binary yes/no subset. Denominator is the paraphrase pair. Claude Opus 5 covers 6,209 pairs with reasoning on and 6,194 with reasoning off, GPT-5.6 Sol 6,195, each spanning MIMIC-CXR, PadChest and VinDr-CXR; Kimi K3 covers 2,274 PadChest pairs only.
- Thesis
- Chapter 3
- Paper
- None
- Source artifact
results/frontier/frontier_flip_eval.jsonin the medical-vlm-robustness repository
Results
| Result | Value | Model | Dataset | n | Denominator | Strength |
|---|---|---|---|---|---|---|
| Paraphrase flip rates span 6.4% to 54.7% across six medical vision-language models Rephrasing a clinical yes/no question in a way that preserves its meaning changes the answer on anywhere from 1 in 16 to more than half of paraphrase pairs, depending on which model and which patient population you test. Every model tested flips, and the spread between the best and worst model-dataset cell is 8.5-fold. Source, denominator, and limits
|
6.4% to 54.7% | Multiple models | Multiple datasets | A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%). | pairwise | Confirmed |
| MedGemma-4B flips on 8.3% of MIMIC-CXR paraphrase pairs On in-distribution US chest X-ray questions, the mechanistic subject model gives a different yes/no answer to about 1 in 12 clinically equivalent rephrasings. Source, denominator, and limits
|
8.3% | MedGemma-4B | MIMIC-CXR | 5,076 | pairwise | Confirmed |
| MedGemma-4B flips on 13.4% of PadChest paraphrase pairs On Spanish chest X-ray questions, out of distribution for a model trained on US data, MedGemma-4B changes its yes/no answer on about 1 in 7 equivalent rephrasings, noticeably more often than on MIMIC. Source, denominator, and limits
|
13.4% | MedGemma-4B | PadChest | 36,244 | pairwise | Confirmed |
| Query-level flip rate is 32.2% where the pairwise rate is 13.4% Counting questions instead of pairs more than doubles the apparent failure rate on identical data. Asking 'how often does one rephrasing change the answer?' gives 13.4%; asking 'how many questions break under at least one of their roughly 3.3 rephrasings?' gives 32.2%. Neither is wrong, but a benchmark that does not say which one it reports is uninterpretable. Source, denominator, and limits
|
32.2% | MedGemma-4B | PadChest | 8,445 | query-level | |
| PSF-Med evaluates 92,856 question-paraphrase pairs The benchmark that every headline flip rate is computed on contains 92,856 (question, paraphrase) pairs across three continents: 8,938 from MIMIC-CXR, 59,573 from second-iteration PadChest and 24,345 from VinDr-CXR, averaging 3.5 paraphrases per question. Source, denominator, and limits
|
92,856 | Not model-specific | Multiple datasets | 92,856 | pairwise | Confirmed |
| PSF-Med covers 26,850 original clinical questions The benchmark is built from 26,850 original chest X-ray questions drawn from three healthcare systems: 3,266 from MIMIC-CXR (United States), 14,935 from PadChest (Spain) and 8,649 from VinDr-CXR (Vietnam). Source, denominator, and limits
|
26,850 | Not model-specific | Multiple datasets | 26,850 | question-level | |
| The equivalence audit rejects nearly half of generated paraphrases A rubric-based GPT-5-mini audit of 122,778 candidate pairs kept only 50.3% (61,761) as clinically equivalent, rejected 48.7% (59,788) as adversarial or meaning-changing, and flagged 1.0% (1,229) as uncertain. Roughly half of what an LLM generates as a 'paraphrase' of a clinical question does not preserve the question. Source, denominator, and limits
|
50.3% | Not model-specific | Multiple datasets | 122,778 | pairwise | |
| Two LLM judge families agree with each other on 91.6% to 94.4% of pairs Re-scoring paraphrase pairs with Claude Haiku instead of GPT-5-mini reproduces the equivalence verdict 91.6% to 94.4% of the time, so the audit's filtering is not an artifact of one model family's idiosyncrasies. Disagreements concentrate in scope and quantification cases. Source, denominator, and limits
|
91.6% to 94.4% | Not model-specific | Multiple datasets | Reported as a range across judge runs without a single pooled denominator. The dissertation states 91.6% on a stratified 500-pair re-scoring subset covering MIMIC + PadChest; the 94.4% upper end is quoted as the top of the cross-family range and its denominator is not separately stated in the source. | pairwise | |
| The audit rejects 93.3% of generated negation paraphrases Negation rewrites are almost never valid paraphrases: 93.3% were rejected by the audit as polarity- or operator-changing, meaning the rewritten question has a different correct answer. Once these are removed, negation's standing as the most fragile transformation type largely dissolves: the model had been marked wrong for correctly noticing that the question had changed. Source, denominator, and limits
|
93.3% | Not model-specific | Multiple datasets | The source reports the rejection rate for the negation-pattern category without stating the category's pre-audit denominator. The surviving counts are given instead: 11 negation pairs per model on MIMIC, almost none on PadChest, and roughly 1,126 on VinDr. | pairwise | |
| Scaling from 4B to 27B does not reliably buy paraphrase consistency MedGemma-27B is the most consistent model on MIMIC at 6.4%, beating the 4B variant's 8.3%, but the advantage does not hold out of distribution: on PadChest the 27B model is marginally worse (13.9% vs 13.4%). Six times the parameters produces a small in-distribution gain that reverses on a different population, so scale is not a fix for paraphrase sensitivity. Source, denominator, and limits
|
6.4% | MedGemma-27B | MIMIC-CXR | 5,076 | pairwise | Supported |
| Frontier models with no medical finetuning still flip on 5.3% to 9.5% of equivalent rephrasings Three frontier general-purpose models, none of them medically finetuned, were scored on the same clinically equivalent rephrasings that the medical models fail. They flip less, but they still flip: between 1 in 19 and 1 in 10 paraphrase pairs get a contradictory yes/no answer. Paraphrase sensitivity is therefore not an artefact of small domain-tuned models that frontier scale has already solved, which is what makes a benchmark that measures it still worth having. Source, denominator, and limits
|
5.3% to 9.5% | Multiple models | Multiple datasets | 18,598 | pairwise | Exploratory |
| Turning reasoning on made Claude Opus 5 flip more, not less The same frontier model, on the same pairs, contradicted itself on 9.53% of rephrasings with reasoning enabled and 8.36% with it disabled: 1.17 percentage points in the wrong direction. Whatever extra deliberation buys, it does not buy consistency under rewording, which is evidence that the failure is not one a model can reason its way out of at inference time. Source, denominator, and limits
|
+1.17 pp with reasoning on | Claude Opus 5 | Multiple datasets | 12,403 | pairwise | Exploratory |
| On PadChest, the one like-for-like comparison, a frontier model flips 3.1% against 12.0-32.8% for the medical models Kimi K3 was scored on PadChest alone, which makes it the only frontier run that can be set beside published per-dataset medical cells without a dataset-mix confound. It contradicts itself on 3.1% of rephrasings where the medical models range from 12.0% to 32.8%. On a population both were measured on, the frontier advantage is real and roughly fourfold at the best medical model. Source, denominator, and limits
|
3.08% | Kimi K3 | PadChest | 2,274 | pairwise | Exploratory |
| LLaVA-Rad gives the same answer 96.3% of the time with the image removed Delete the chest X-ray entirely and ask LLaVA-Rad the same question, and it returns its original answer on 96 of every 100 pairs. It is also the most paraphrase-consistent of the three backends tested here (6.5% flip rate). Its consistency is almost entirely a property of the question text: the image is nearly decorative. Source, denominator, and limits
|
96.3% | LLaVA-Rad | Multiple datasets | 107 | pairwise | Supported |
| MedGemma-4B is the most image-dependent backend at 66.4% text-only agreement MedGemma-4B changes its answer on a third of pairs when the image is taken away, far more than LLaVA-Rad (96.3% agreement) or MedGemma-27B (85.0%). It is the most image-dependent of the three backends, and it is also the most paraphrase-fragile. Being harder to fool with a missing image goes together with being easier to shake with a rephrasing. Source, denominator, and limits
|
66.4% | MedGemma-4B | Multiple datasets | 107 | pairwise | Supported |
| MedGemma-4B changes its answer on 30.8% of pairs when given the wrong image Swap in a different patient's chest X-ray and MedGemma-4B changes its answer about 31% of the time, the highest of the three backends. LLaVA-Rad barely notices (10.3%). This is the complement of the text-only test: a grounded model should react when the evidence changes. Source, denominator, and limits
|
30.8% | MedGemma-4B | Multiple datasets | 107 | pairwise | Supported |
| Across three backends, flip rate and text-only agreement are almost perfectly opposed Ordering the three backends by paraphrase fragility exactly reverses their ordering by image-independence (Pearson r = -0.997): the model that flips least ignores the image most. This trade-off motivates the thesis, though it rests on only three model points. Source, denominator, and limits
|
-0.997 | Multiple models | Multiple datasets | 3 | model-dataset | Supported |
| Flipping cases put 41% less attention on the pathology region When the model flips its answer under paraphrasing, it was already directing less attention into the radiologist-annotated pathology box: 8.4% of attention mass versus 14.2% for cases that stay consistent. Weak spatial grounding and paraphrase fragility travel together. But note the absolute level: even the consistent cases put only about one-seventh of their attention on the region that matters. Source, denominator, and limits
|
8.4% | MedGemma-4B | PadChest | 200 | case-level | |
| Targeted LoRA cuts the flip rate from 8.5% to 3.5%, about 59% Training a small adapter on five transformer layers cuts the paraphrase flip rate by roughly 59%, from 8.5% to 3.5%, on a held-out test that shares no patient and no image with training. The effect is significant in every one of five seeds. Source, denominator, and limits
|
3.5% | Targeted LoRA | MIMIC-CXR | 238 | pairwise | Mixed |
| No observed accuracy reduction from the targeted adapter Accuracy on the held-out presence endpoint moves from 84.5% to 84.7% (+/- 1.0), a paired per-question difference of +0.25 percentage points with a 95% interval of [-1.00, +1.51]. There is no observed accuracy reduction: the flip-rate gain does not come out of correctness. Source, denominator, and limits
|
0.25 pp | Targeted LoRA | MIMIC-CXR | 238 | question-level | Mixed |
| The targeted adapter trains 0.10% of the model's parameters The whole intervention is 4.38M trainable parameters out of roughly 4.4 billion, adapters on layers 15-19 only (rank 16, alpha 32), with the vision encoder frozen. It trains in under 20 minutes on a single A100. Full LoRA touches all 34 layers for 30.5M parameters (0.72%) and buys no better outcome. Source, denominator, and limits
|
0.10% | Targeted LoRA | Not dataset-specific | Not an evaluation statistic and so has no sample: it is a parameter count over a single model configuration (4.38M trainable adapter parameters against roughly 4.4B in MedGemma-4B). Adapters sit in attention (q, k, v, o) and MLP (gate, up, down) projections on layers 15-19. | model-dataset | Mixed |
| The adapter buys consistency by leaning harder on the question text On the clean patient-disjoint audit, text-only agreement rises from 53.5% for the base model to 76.8% (+/- 3.1) after targeted adaptation: the adapted model returns its original answer with the image deleted 23 points more often. Both adapters get much of their new consistency by relying more on the question text, with no steadier use of the image. Targeted and Full LoRA are statistically tied on image reliance (76.8% +/- 3.1 vs 76.6% +/- 1.7). Source, denominator, and limits
|
76.8% | Targeted LoRA | MIMIC-CXR | 241 | question-level | Mixed |
| Transplanting the residual at layer 16 flips the answer on 73% of pairs Swapping a single layer's residual stream from a 'detecting' pair member into a 'missing' one flips the model's committed answer on 73% of pairs at layer 16, up from 35% at layer 15 and 8% at layer 14, reaching 100% by layer 20. The answer commits in a narrow band at layer 16 (median commit layer 16, 95% bootstrap CI [16, 16]). An image-token control never exceeds 0.0007, so this is not a generic patching artifact. Source, denominator, and limits
|
73% | MedGemma-4B | PadChest | 1,396 | pairwise | Exploratory |
| Feature 3818 is the largest layer-17 change in 37 of 76 same-polarity flips Restricting to the 141 FlipBank pairs that preserve the question's operator (dropping 17 polarity switches where changing the answer is correct), the model flips on 76. On those, Feature 3818 is the single largest layer-17 activation change in 37 of them (48.7%) and in the top five in 59. It is a prominent contributor to same-polarity paraphrase flips. Source, denominator, and limits
|
37 of 76 (48.7%) | MedGemma-4B | Multiple datasets | 76 | pairwise | Exploratory |
| Patching Feature 3818 recovers 28% of the answer margin, 3.5x controls On a pleural effusion example, rephrasing 'Is there pleural effusion?' as 'Is pleural fluid present?' drives the yes-minus-no margin from +8.75 to -0.625, flipping the answer to No. Writing Feature 3818's original activation back in recovers the margin to +2.0 and restores the Yes, about 28% of the lost margin, against 8% for ten control features. Source, denominator, and limits
|
28% | MedGemma-4B | Multiple datasets | 1 | case-level | Exploratory |
| The best layers to fix are not the layers where the mechanism lives Adapting early layers (0-10) reduces the margin difference to 0.26, an 86% cut from the 1.87 baseline, while the mechanistically indicated layers 15-19 reach only 0.38 (80%). Even a randomly chosen block (layers 5-9) does better at 0.30. The SAE analysis says where paraphrase sensitivity manifests, but the best place to intervene is upstream, before the register-sensitive representation forms. RQ3.3 is answered no. Source, denominator, and limits
|
0.26 | Targeted LoRA | MIMIC-CXR | 355 | question-level | Negative result |
| The MIMIC-trained adapter gains 6.4 accuracy points on Spanish data An adapter trained only on US MIMIC-CXR data raises accuracy on balanced PadChest questions from 85.2% to 91.6%, and sharpens the margin difference from 1.02 to 0.25 (a 75% reduction). The benefit transfers across institution, equipment and patient population, but it shows up as accuracy and confidence, with flips barely reduced. Source, denominator, and limits
|
91.6% | Targeted LoRA | PadChest | 250 | question-level | |
| The training-phrasing distribution sets the flip rate: 4.8% to 88.4% In a controlled MedGemma-faithful probe where every paraphrase of a question sees identical frozen image features, the binary flip rate is set by the training-phrasing distribution alone: 4.8% when trained on every paraphrase, 67.1% on one fixed phrasing, and 88.4% when the register is tied to the answer, with architecture, parameter budget, seed, and evaluation held fixed (eight seeds per regime, Cliff's delta 1.00). Because text-only accuracy stays at 50.0-50.3% throughout, this identifies the training-phrasing distribution as a cause, not a correlate, and the flip is rank-1 patchable at early layers: a difference-direction transplant at the answer position gives net recovery 0.98-1.00 across layers 0-4 for the adversarial regime's injected flips and 0.77 rising to 1.00 over the same layers for the augmented regime's naturally occurring flips, against 0.00-0.02 for a norm-matched random direction. Source, denominator, and limits
|
4.8% / 67.1% / 88.4% | FlipLens probe | Multiple datasets | Binary flip rate over eight seeds per regime on the balanced held-out evaluation; per-finding balancing pins text-only accuracy at 50.0-50.3% in every regime. | question-level | |
| The controlled probe transfers to two unseen hospitals at AUROC 0.743 and 0.756 The FlipLens probe genuinely uses the radiograph: with the image removed its accuracy is exactly 50.0% (the balanced-design floor) against 74.8% with it, and on two hospitals that contributed nothing to training it separates findings at AUROC 0.743 (MIMIC-CXR) and 0.756 (VinDr-CXR). This is what licenses reading a paraphrase flip in this model as a language-side failure rather than as a model that never consulted the image. Source, denominator, and limits
|
0.827 in-distribution / 0.743 MIMIC / 0.756 VinDr | FlipLens probe | VinDr-CXR | 5,478 | question-level | |
| 81% of consistent predictions are image-invariant Averaged across ten model-dataset settings, 81% of each model's consistent predictions are unchanged when the image is removed (range 45-100%), so a low paraphrase-flip rate is not evidence of grounded visual reasoning. Source, denominator, and limits
|
81% (range 45-100%) | Multiple models | Multiple datasets | 10 | model-dataset | Confirmed |
| Targeted LoRA on PadChest: 84.1% of questions in the Dangerous quadrant After targeted LoRA adaptation, 84.1% of PadChest questions land in the Dangerous quadrant (consistent under paraphrase but not image-reliant): the adapter buys consistency largely by leaning on the text prior. Source, denominator, and limits
|
84.1% | Targeted LoRA | PadChest | 861 | question-level | Confirmed |
| Base MedGemma-4B on PadChest: 22.5% Dangerous Base MedGemma-4B places only 22.5% of PadChest questions in the Dangerous quadrant, but this is low only because the model is too inconsistent to land in any consistent cell at all: 74% of its predictions fall in Fragile or Worst. Source, denominator, and limits
|
22.5% | MedGemma-4B | PadChest | 861 | question-level | |
| Full LoRA on MIMIC: 78.6% Dangerous Full LoRA, the variant with the lowest flip rate in-distribution, places 78.6% of MIMIC questions in the Dangerous quadrant: consistent answers that do not change when the image is removed. Source, denominator, and limits
|
78.6% | Full LoRA | MIMIC-CXR | 98 | question-level | Confirmed |
| The flip-rate vs Dangerous-fraction correlation is largely definitional The headline r = -0.86 correlation between flip rate and the Consistent-Text-driven fraction is largely a definitional consequence of that fraction being a subset of consistent predictions; conditioning on consistency reduces it to r = -0.15. Source, denominator, and limits
|
r = -0.15 (conditioned) | Multiple models | Multiple datasets | 10 | model-dataset | Confirmed |
| Attention rank does not predict causal patch importance (rho = -0.09) Patch-rank Spearman correlation between attention weight and occlusion importance is -0.09 for Base MedGemma-4B: attention magnitude does not indicate which image patches the model causally uses. Source, denominator, and limits
|
rho = -0.09 (Base) | MedGemma-4B | PadChest | 637 | case-level | Negative result |
| Attention coverage barely beats a shifted box (0.296 vs 0.261) Base MedGemma-4B places 29.6% of attention mass inside the radiologist-annotated pathology box, only marginally more than the 26.1% it places in a same-sized shifted box: attention is a coarse localizer, short of a faithful one. Source, denominator, and limits
|
0.296 true box vs 0.261 shifted | MedGemma-4B | PadChest | 637 | case-level | Negative result |
| Targeted LoRA has the smallest between-sex ECE gap (0.012) Targeted LoRA shows the smallest between-sex calibration gap of the three MedGemma variants (ECE 0.106 male vs 0.118 female, gap 0.012) on PadChest. Source, denominator, and limits
|
0.012 | Targeted LoRA | PadChest | 861 | question-level | Mixed |
| Targeted LoRA accuracy climbs 13 pp from youngest to oldest patients Targeted LoRA accuracy on PadChest rises from 83.3% for patients under 40 to 96.3% for patients over 80, a 13 percentage-point age gradient with younger patients least accurate - the steepest of the three MedGemma variants. Source, denominator, and limits
|
13 pp (83.3% under-40 to 96.3% over-80) | Targeted LoRA | PadChest | 861 | question-level | Mixed |
| Single-pass entropy predicts paraphrase flips (AUROC 0.823) Softmax predictive entropy from one forward pass predicts whether a prediction will flip under paraphrasing with AUROC 0.823 for Targeted LoRA on PadChest: uncertainty is a deployable single-pass proxy for paraphrase fragility. Source, denominator, and limits
|
0.823 | Targeted LoRA | PadChest | 861 | question-level | Confirmed |
| Single-pass entropy detects errors (AUROC 0.862) The same single-pass softmax entropy signal detects wrong answers with AUROC 0.862 for Targeted LoRA on PadChest, matching Monte Carlo dropout (0.856) and beating the deep ensemble (0.751). Source, denominator, and limits
|
0.862 | Targeted LoRA | PadChest | 861 | question-level | Confirmed |
| Targeted LoRA has the best selective-prediction quality OOD (AUGRC 0.016) With temperature scaling, Targeted LoRA reaches AUGRC 0.016 on PadChest, the best generalized risk-coverage score of the MedGemma variants and roughly a third of Base's 0.047. AUGRC is the area under the generalized risk-coverage curve: it runs 0 to 0.5, and lower is better. Source, denominator, and limits
|
0.016 | Targeted LoRA | PadChest | 732 | question-level | |
| Full LoRA collapses OOD (AUGRC 0.156) Full LoRA, despite the lowest in-distribution flip rate, has the worst out-of-distribution selective prediction: AUGRC 0.156 on PadChest, more than three times Base's 0.047, alongside ECE 0.229. AUGRC is the area under the generalized risk-coverage curve: it runs 0 to 0.5, and lower is better. Source, denominator, and limits
|
0.156 | Full LoRA | PadChest | 732 | question-level | |
| The readiness audit admits 33.0% of Targeted LoRA PadChest predictions at 96.8% accuracy The deployment-readiness rule (paraphrase agreement AND image reliance) admits 33.0% of Targeted LoRA's PadChest predictions, and the admitted slice is 96.8% accurate versus 91.5% unfiltered, with a 4.2% residual flip rate on uninspected paraphrases. Source, denominator, and limits
|
33.0% admitted at 96.8% accuracy | Targeted LoRA | PadChest | 861 | question-level | Negative result |
| 2.9% accuracy on the slice where the image should matter Targeted LoRA scores only 2.9% accuracy on the PadChest text-disagrees slice - the cases where the image-conditioned answer must depart from the text-only answer - so gates raise admitted-case accuracy by selecting text-answerable cases over image-grounded ones (RQ5.3 = no). Source, denominator, and limits
|
2.9% | Targeted LoRA | PadChest | 241 | question-level | Negative result |
| Adapter-only Monte Carlo dropout detects almost no epistemic uncertainty (MI = 0.0001 nats) Monte Carlo dropout on the LoRA adapters yields mean mutual information of 0.0001 nats for Targeted LoRA on PadChest: the perturbation barely moves the model, so the probe detects almost none of the predictive entropy as epistemic. Source, denominator, and limits
|
0.0001 | Targeted LoRA | PadChest | 861 | question-level | Negative result |
| Deep ensemble fails out-of-distribution: worst member at 52.0% accuracy The five-seed LoRA deep ensemble fails on PadChest: per-member accuracy spans 52.0% (seed 456) to 92.1% (seed 42), and the ensemble's selective prediction collapses to AURC 0.130 versus 0.019 for single-pass entropy, admitting only 23.8% of cases at 5% risk versus 88.5%. Source, denominator, and limits
|
52.0% worst member vs 92.1% best | Targeted LoRA | PadChest | 861 | question-level | Negative result |
| Conformal coverage slips to 88.7% at severity-5 corruption (90% target) Base MedGemma-4B's split-conformal prediction sets hold 91.8% empirical coverage clean against a 90% target but slip to 88.7% under severity-5 image corruption. Source, denominator, and limits
|
88.7% | MedGemma-4B | PadChest | 732 | question-level |
Population reference
These evaluation sets are not interchangeable. Each row states what a population covers and why it does not pool with the others.
| Population | Size | What it is | Why it does not pool |
|---|---|---|---|
| PSF-Med main benchmark | 92,856 pairs | The full evaluated benchmark: 92,856 final evaluation pairs built from 26,850 original clinical questions (MIMIC-CXR 8,938 pairs, second-iteration PadChest 59,573, VinDr-CXR 24,345; mean 3.5 paraphrases per question). Source of the headline flip rates. | The 92,856 evaluation pairs are numerically close to the ~92,000 construction-stage candidate paraphrases by coincidence and are not a subset of the 61,761 audit-retained core (the audited first-iteration PadChest pairs were superseded by the regenerated second iteration). Pipeline stages are distinct counts, and these rates are not comparable to curated flip-bank rates. |
| Binary yes/no subset | 49,932 pairs | The binary presence-question subset of the main benchmark used for headline pairwise flip rates: MIMIC-CXR 1,539 questions / 5,076 pairs; PadChest 8,445 / 36,244; VinDr-CXR 2,807 / 8,612. | Pairwise flip rates on this subset are roughly half the query-level rates on the same questions (MedGemma-4B PadChest: 13.4% pairwise vs 32.2% query-level). The two denominators are not interchangeable, and these rates are not comparable to curated flip-bank rates. |
| Mechanistic FlipBank (158 pairs) | 158 pairs | A 158-pair curated set of confirmed flips (94 yes-to-no, 64 no-to-yes, 142 images, 23 findings) used for the sparse-autoencoder mechanistic analyses on MedGemma-4B. | Every pair in this set flips by construction, so any flip statistic computed on it is 100% by design; it exists to study mechanisms, never to estimate flip prevalence or to compare against benchmark rates. |
| Three-backend diagnostic set (107 pairs) | 107 pairs | A curated 107-pair set of semantically close question-paraphrase pairs (BioClinicalBERT cosine similarity > 0.95) on which MedGemma-4B, MedGemma-27B, and LLaVA-Rad were all evaluated for the text-only, image-swap, and attention diagnostics. Constructed before the equivalence audit. | MedGemma-4B flips 42.1% on this curated near-paraphrase set versus 8.3% on the benchmark; the two rates measure different things and are not comparable. The set also predates the equivalence audit, so its flip labels contain operator changes. |
| PadChest flip bank (861 questions) | 861 questions | An 861-question PadChest bank curated to be flip-prone, used for the four-quadrant audit, uncertainty quantification, the entropy bridge, and the deployment gate on MedGemma variants. | This bank was curated to be flip-prone: Base MedGemma-4B flips 73.9% of it versus 13.4% pairwise on the PadChest benchmark. Quadrant shares, AUROCs, and gate numbers computed here are not benchmark statistics. |
| MIMIC uncertainty set (98 questions) | 98 questions | The full 98-question MIMIC-CXR presence-question test set used for the four-quadrant audit and uncertainty analyses of the MedGemma variants; not truncated. | Only 98 questions with no demographic attributes on disk: quadrant and calibration estimates carry wide intervals and cannot be pooled with the 861-question PadChest bank or the main benchmark. |
| LLaVA-Rad PadChest set (732 questions) | 732 questions | The 732-question subset of the PadChest flip bank on which LLaVA-Rad variants are evaluated (the temperature-scaling calibration holdout also leaves 732 of 861 for evaluation). | A 732-question subset of the 861-question PadChest bank: LLaVA-Rad numbers computed here are not directly comparable to MedGemma numbers computed on the full 861 questions. |
| Patient-disjoint mitigation endpoint (238 questions) | 238 questions | The clean mitigation evaluation set: 238 presence-of-finding questions / 991 pairs over 236 held-out subjects and 238 held-out images, with zero subject and zero image overlap with LoRA training data. | This is the only clean mitigation endpoint: its ~59% pairwise flip reduction (8.5% to 3.5%) supersedes the retracted 79.5% (a margin-difference reduction) and 69.6% (an older flip reduction) figures, and must not be compared with numbers from the original non-disjoint pipeline. |
| Attention analysis subsample (200 pairs) | 200 pairs | A 200-pair attention-analysis subsample balanced at 100 flip / 100 no-flip pairs, drawn from a 49,522-pair annotated pool. | Balanced by construction at 50% flips: any aggregate flip statistic on this set is a design choice. It is not an observation, and attention contrasts here cannot be projected onto benchmark populations. |
| PadChest grounding subset (637 boxes) | 637 boxes | The 637 all-positive PadChest cases with radiologist bounding boxes used for attention-grounding coverage (true vs shifted vs random-far box) and ROI causal scoring. | All 637 cases are positive findings with radiologist boxes: coverage and ROI scores say nothing about negative findings, and boxes exist for only 637 of the 861 PadChest bank questions and none of MIMIC, which is also why the gate's region term is not deployable online. |
| Canonical SAE screen (41,822 pairs) | 41,822 pairs | The canonical sparse-autoencoder screening pool: 4,542 MIMIC pairs (436 flips, 9.6%) and 37,280 PadChest pairs (2,833 flips, 7.6%) used for the two-stage L17/#3818 to L29/#12139 validation, mediation, and composition analyses. | 171 of the 436 verified MIMIC flips (39%) are negation-pattern or operator changes, so restoration and mediation rates conflate correct operator handling with genuine same-polarity paraphrase sensitivity; do not read them as benchmark-level causal effect sizes. |
| Answer-commitment transplant pool (1,396 pairs) | 1,396 pairs | The 1,396-pair / 91-finding pool used for the residual-transplant study that locates the answer-commitment locus at layer 16 (flip rate 8% at L14, 73% at L16, 100% by L20; median commit layer 16, 95% bootstrap CI [16, 16]). | The pool is 94% positive (1,314 yes vs 82 no), 7.7% of pairs carry one-sided qualifiers, and the pair-pool construction script was not retained; commit-layer statistics are specific to this pool and are evidence of a general commitment locus. They are not evidence of a paraphrase-specific mechanism. |
| Layer-ablation validation split (355 questions) | 355 questions | The 355-question MIMIC-CXR validation split used to compare LoRA layer ranges (early 0-10, random 5-9, all 0-33, middle 15-19, late 25-33) on margin difference. | A validation split used for model selection. It is not the held-out test set or the patient-disjoint endpoint, so margin-difference results here (early layers 0.26 beating the mechanistic middle 0.38) must not be quoted alongside patient-disjoint endpoint numbers. |
| PadChest transfer set (250 questions) | 250 questions | A 250-question class-balanced PadChest set (125 yes / 125 no) used as the out-of-distribution transfer test for the targeted LoRA (accuracy 85.2% to 91.6%, margin difference 1.02 to 0.25). | Balanced 50/50 by construction, so the base model's 7.6% flip rate here sits far below its 13.4% pairwise PadChest benchmark rate; the transfer gain shows up in accuracy and margin sharpening, with no flip reduction, and neither number generalizes to the unbalanced benchmark. |
| Four-quadrant audit settings (10 model-dataset settings) | Ten model-dataset settings; per-setting question counts differ (MIMIC 98 or 88, PadChest 861 or 732), so there is no single pooled n. | The ten model-dataset settings of the four-quadrant safety screen, pooling the MIMIC flip banks (n=98 for MedGemma variants, n=88 for LLaVA-Rad variants) and the PadChest flip banks (n=861 and n=732) across five models. | The 81% image-invariant share is a level averaged over ten settings whose underlying flip banks were curated to be flip-prone; it is not a benchmark statistic, and the r=-0.86 consistency-grounding correlation over the same ten points is largely definitional (conditioned on consistency it falls to r=-0.15). |
| FlipLens training-regime experiment (3 regimes x 8 seeds) | 2,000 evaluation clusters | The controlled probe's regime experiment: each of three training-phrasing regimes is trained over 8 seeds on 86,288 balanced questions and scored on 2,000 held-out evaluation clusters. Balancing per finding pins the text-only floor at 0.500. | This is the probe's own controlled evaluation, not a deployed-model benchmark: its early-layer locus does not map onto MedGemma-4B's layer 16, and its per-finding balancing removes the answer prior by design, so its flip rates must not be mixed with the deployed-model PSF-Med numbers. |
| Frontier addendum, all three datasets (about 6,200 pairs) | Per-run coverage varies from 6,194 to 6,209 pairs; 18,598 pairs across the three all-three-dataset runs. | The paraphrase-pair subsample used for the post-draft frontier evaluation across MIMIC-CXR, PadChest and VinDr-CXR. Pair counts differ slightly by run because a small number of responses could not be parsed to a binary answer: Claude Opus 5 with reasoning on covers 6,209 pairs, with reasoning off 6,194, and GPT-5.6 Sol 6,195, for 18,598 pairs in total. | This is a roughly 6,200-pair subsample, not the 49,932-pair binary subset the six medical models were scored on, and its per-dataset mixture is not the binary subset's 10/73/17 split across MIMIC-CXR, PadChest and VinDr-CXR. Because flip rates differ sharply by dataset, these rates must not be pooled with the binary-subset rates, and a single frontier-versus-medical ratio computed across the two would be comparing two different dataset mixtures rather than two model classes. |
| Frontier addendum, PadChest only (2,274 pairs) | 2,274 pairs | The PadChest-only paraphrase pairs used for the Kimi K3 run in the post-draft frontier evaluation. Because PadChest is one of the three datasets the published per-model cells report separately, this run is the only frontier result that can be set beside a published medical cell without a dataset-mix confound. | PadChest only. These pairs must not be pooled with the all-three-dataset frontier runs, and this rate must not be compared with any all-three-dataset pooled value; it is comparable only with the published PadChest cells of the binary subset, which are themselves a different and much larger sample of PadChest. |
Models and datasets
| Model | Family | Role |
|---|---|---|
| MedGemma-4B | MedGemma (Gemma 3) | Google's 4B-parameter medical Gemma 3 instruction-tuned VLM (34 transformer layers, 256 image tokens); the dissertation's primary mechanistic subject and the base for both LoRA adapters. |
| MedGemma-1.5-4B | MedGemma (Gemma 3) | The 4B-parameter MedGemma 1.5 release, evaluated as one of the six base models in the PSF-Med benchmark (its VinDr cell is a degenerate yes-bias cell and is set aside). |
| MedGemma-27B | MedGemma (Gemma 3) | The 27B-parameter MedGemma model; the largest model evaluated, showing that scale does not buy paraphrase consistency (6.4-13.9% pairwise flips) and serving as one of the three diagnostic backends. |
| CheXone | Qwen2.5-VL | A 3B-parameter chest X-ray VLM built on Qwen2.5-VL, evaluated as one of the six base models in the PSF-Med benchmark. |
| LLaVA-Rad | LLaVA | A 7B-parameter LLaVA-based radiology VLM whose low flip rates coincide with near-total text reliance, making it the central example of consistency without image use. |
| RadFM | RadFM (custom) | A 14B-parameter custom-architecture radiology foundation model; the least consistent model in the benchmark (up to 54.7% pairwise flips on VinDr-CXR). |
| Targeted LoRA | MedGemma (Gemma 3) | A mechanistically-placed LoRA adapter on MedGemma-4B layers 15-19 (rank 16, alpha 32, 4.38M parameters = 0.10%) that cuts pairwise flips about 59% without observed accuracy reduction, but raises text reliance, and grounding is not restored. |
| Full LoRA | MedGemma (Gemma 3) | A LoRA adapter across all 34 MedGemma-4B layers (30.5M parameters = 0.72%); slightly fewer flips than Targeted LoRA but statistically tied on image reliance, with worse accuracy, calibration, and demographic disparity. |
| LLaVA-Rad LoRA | LLaVA | A LoRA-adapted LLaVA-Rad used as the cross-architecture check that the consistency-versus-text-reliance pattern and the entropy bridge (flip AUROC up to 0.905) are not MedGemma-specific. |
| Qwen2-VL | Qwen2-VL | A general-domain Qwen2-VL model used in the grounded-slice analysis of the deployment gate, where it answers 0 of 279 text-disagrees cases correctly under every gating rule. |
| FlipLens probe | MedGemma-faithful (Gemma 3) | A controlled MedGemma-faithful probe: a frozen MedSigLIP-448 encoder feeding a 14.1M-parameter Gemma-3 decoder that reads yes/no from the tied language-model head. Built to isolate the origin of paraphrase sensitivity, with every paraphrase seeing identical image features; released as saillab/FlipLens. |
| Claude Opus 5 | Claude (Anthropic) | A frontier general-purpose multimodal model with no medical finetuning, evaluated in the post-draft frontier addendum under both its reasoning-on and reasoning-off settings on the same paraphrase pairs as the medical models. |
| GPT-5.6 Sol | GPT (OpenAI) | A frontier general-purpose multimodal model with no medical finetuning, evaluated in the post-draft frontier addendum at its default reasoning setting. |
| Kimi K3 | Kimi (Moonshot AI) | A frontier general-purpose multimodal model with no medical finetuning, evaluated in the post-draft frontier addendum at its default reasoning setting on PadChest only, which makes it the one frontier run directly comparable with a published per-dataset cell. |
| Muse-Glimmer-30B | Muse-Glimmer (Meta) | A 30B-parameter open-weight general-purpose multimodal model run locally at a pinned Hugging Face snapshot with greedy decoding, evaluated in the post-draft frontier addendum; it sits outside the pooled frontier headline and has no image-reliance controls yet. |
| Dataset | Access | Licence |
|---|---|---|
| MIMIC-CXR | Credentialed PhysioNet access (CITI training plus a signed data use agreement); redistribution of images or reports is prohibited, so the benchmark releases only question/paraphrase text and identifiers. | PhysioNet Credentialed Health Data License 1.5.0 |
| PadChest | Registration with the Medical Imaging Databank of the Valencia Region (BIMCV) is required before download. The PadChest Dataset Research Use Agreement grants research use only and prohibits reproducing or publishing any portion of the dataset without prior written permission, so the benchmark releases only question and paraphrase text. | PadChest Dataset Research Use Agreement (BIMCV; research use only, redistribution prohibited without written permission) |
| VinDr-CXR | Credentialed PhysioNet access (CITI training plus a signed data use agreement); redistribution prohibited. | PhysioNet Credentialed Health Data License 1.5.0 |
| CheXlocalize | Free registration with Stanford AIMI is required before download; released for research use. | Stanford AIMI dataset terms (research use) |