From the synthesis table in Chapter 8. The chip states how strongly the evidence supports the answer.
ConfirmedChapter 3
Paraphrase sensitivity is common in medical vision-language models: pairwise flip rates on the binary yes/no subset span 6.4% to 54.7% across six models and three chest X-ray datasets, an 8.5x spread.
flip-rate-range · flip-rate-medgemma4b-mimic · flip-rate-medgemma4b-padchest · benchmark-eval-pairs
SupportedChapter 3
Scale does not buy paraphrase consistency: MedGemma-27B is the most consistent of the three MedGemma variants on MIMIC (6.4%) and VinDr (8.1%) but not on PadChest, where MedGemma-4B edges it (13.4% vs 13.9%), and no stable consistency ordering by parameter count exists across datasets.
scale-no-consistency · flip-rate-range
SupportedChapter 4
Low paraphrase sensitivity can coexist with output-level image-removal invariance: LLaVA-Rad reaches 96% text-only agreement at low flip rates while MedGemma-4B is more image-dependent but flips more, so a low flip rate is insufficient evidence of grounded visual reasoning.
text-only-llavarad-107 · text-only-medgemma4b-107 · swap-sensitivity-medgemma4b-107 · flip-textonly-correlation-107
ExploratoryChapter 5
A candidate two-stage mechanistic account localizes paraphrase sensitivity: Feature 3818 (layer 17) behaves as an operator/register feature and Feature 12139 (layer 29) as a downstream answer-selection feature, with layer 16 as a general answer-commitment locus.
feature3818-operator-preserving-top1 · feature3818-patch-recovery · layer16-commit-rate
MixedChapter 5
Partly. Targeted LoRA on layers 15 to 19 touches 0.1% of parameters and cuts pairwise flips from 8.5% to 3.5% over five seeds, about 59%. McNemar p < 0.002 in every seed, and no observed accuracy reduction on a patient- and image-disjoint split. But it does not preserve visual grounding: both adapters raise text-only agreement from 53.5% to about 77%.
lora-flip-reduction-pd · lora-accuracy-delta-pd · lora-param-fraction · lora-textonly-increase-pd
Negative resultChapter 5
No: the best intervention layer does not match the diagnosis layer. Ablating early layers 0-10 reduces margin difference to 0.26 (86% reduction), beating the mechanistically identified layers 15-19 (0.38, 80%) by six points.
layer-ablation-early-vs-middle
ConfirmedChapter 6
Reducing flips can create a false sense of safety: averaged across ten model-dataset settings, 81% of each model's consistent predictions are image-invariant (range 45-100%), so consistency and grounding routinely come apart.
image-invariant-share · quadrant-dangerous-targeted-padchest · quadrant-dangerous-fulllora-mimic · definitional-correlation
Negative resultChapter 6
No: attention does not reliably indicate grounding. True-box coverage only marginally exceeds a displaced box (0.296 vs 0.261 for Base), and patch-rank correlation between attention and causal occlusion importance is near zero across all three MedGemma variants.
attention-occlusion-rho · attention-true-vs-shifted
MixedChapter 6
Partially: the worst-calibrated model also shows the largest demographic disparities. Full LoRA has the worst calibration (ECE of about 0.25 or higher for every group) and the largest sex gaps, while Targeted LoRA has the smallest sex ECE gap (0.012) but the steepest age accuracy gradient (13 pp).
fairness-ece-sex-gap-targeted · fairness-age-gradient-targeted
ConfirmedChapter 7
Yes: single-pass predictive entropy predicts paraphrase flips (AUROC 0.823) as well as errors (AUROC 0.862) on PadChest, and the flip bridge replicates across architectures (LLaVA-Rad LoRA 0.830 PadChest, 0.905 MIMIC).
entropy-flip-auroc · entropy-error-auroc
Negative resultChapter 7
No. Gates admit the cases the text already answers over the ones grounded in the image. Admitted-case accuracy does rise, to 96.8% against 91.5% for Targeted LoRA on PadChest. Yet that same model scores 2.9% on the slice where text and image disagree, and Qwen2-VL gets 0 of 279 such questions right under every rule tried.
gate-grounded-slice-failure · gate-admission-targeted-padchest
Negative resultChapter 7
No: no single internal monitor transfers across model families. Gemma, LLaVA-Rad, and Qwen2-VL each require different monitors, and the multi-pass alternatives fail in family-specific ways: adapter-only Monte Carlo dropout is an uninformative epistemic probe and the MedGemma deep ensemble collapses out of distribution.
mcdropout-mi · ensemble-ood-failure · gate-grounded-slice-failure
SupportedChapter 8
A low paraphrase-flip rate is not sufficient evidence of reliable visual reasoning: safe deployment requires joint evaluation of semantic invariance, correctness, image-dependence, and calibration, because optimizing for consistency alone creates a false sense of safety.
image-invariant-share · flip-rate-range · lora-flip-reduction-pd · lora-textonly-increase-pd · quadrant-dangerous-targeted-padchest · entropy-flip-auroc
ExploratoryChapter 3
Paraphrase sensitivity persists in frontier general-purpose models that were never medically finetuned: four configurations of three such models flip on 3.1% to 9.5% of clinically equivalent rephrasings, and enabling reasoning on Claude Opus 5 raised its flip rate rather than lowering it (9.53% against 8.36%, p = 0.022). On PadChest, the one population where a frontier run and the published medical cells are directly comparable, the frontier rate is 3.08% against 12.0% to 32.8%.
frontier-flip-persists · frontier-reasoning-no-benefit · frontier-padchest-like-for-like