Start with what the failure looks like. Then go inside the model: what the model is thinking at every layer of one question, and where the answer commits across 1,396 pairs. Then try the two signals meant to catch the failure before a reader ever sees an answer. Everything after the first section runs on the saved evaluation rows the dissertation reports. Move a threshold, switch a rule off, or drag through the layers, and the numbers recompute in your browser.
These demos replay saved offline evaluation on a research dataset. They are not a clinical tool, they do not run a model, and nothing here is medical advice.
1. What a paraphrase flip looks like
Five curated cases show the failure the later signals are meant to catch: questions that mean the same thing get contradictory answers. Within a case the image never changes; only the wording does.
GPT-5-mini and Claude Haiku 4.5 appear as general-purpose reference models for contrast; the dissertation's benchmark covers six medical vision-language models. Each verdict chip compares a model's answers with the reference read: whether the original answer was right, and whether any rewording changed it. In practice, a flipped presence answer is the difference between a finding flagged for review and one missed.
Case 01 · Curated demonstration set (MIMIC-CXR)
Pneumothorax
3 flipped answers3 of 3 models flip
Illustrative pictogram, standing in for the study image
The radiograph is not shown
This case uses a MIMIC-CXR study. The PhysioNet licence for MIMIC-CXR prohibits redistributing images, so the radiograph is not published here. Anyone with credentialed access can obtain it from PhysioNet.
The question and the answers are below. The image never changes within a case; only the wording does.
Original question
Is there evidence of pneumothorax in this image?
Reference answer
No
Why this case mattersThe image shows no pneumothorax, yet a single more specific rewording about pleural air pushes all three models into the same false positive.
Why these questions count as equivalentEach paraphrase keeps the same clinical operator (presence of a finding) and the same finding (pneumothorax), varying only syntax, scope wording, specificity, or synonyms, which is the class of rewrite the equivalence audit rubric treats as clinically equivalent.
MedGemma-4B
Correct until rephrased
Answer to the original question: No
Does this image demonstrate a pneumothorax?
Syntactic restructuringNo Same as the originalMatches the reference answer
Is there evidence of pneumothorax, unilateral or bilateral, on this image?
Scope quantificationNo Same as the originalMatches the reference answer
Is there radiographic evidence of pleural air consistent with pneumothorax in this chest radiograph?
Specificity modulationYes Flipped from the originalConflicts with the reference answer
Is there air in the pleural space suggesting pneumothorax on this image?
Lexical substitutionNo Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
GPT-5-mini
Correct until rephrased
Answer to the original question: No
Does this image demonstrate a pneumothorax?
Syntactic restructuringNo Same as the originalMatches the reference answer
Is there evidence of pneumothorax, unilateral or bilateral, on this image?
Scope quantificationNo Same as the originalMatches the reference answer
Is there radiographic evidence of pleural air consistent with pneumothorax in this chest radiograph?
Specificity modulationYes Flipped from the originalConflicts with the reference answer
Is there air in the pleural space suggesting pneumothorax on this image?
Lexical substitutionNo Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
Claude Haiku 4.5
Correct until rephrased
Answer to the original question: No
Does this image demonstrate a pneumothorax?
Syntactic restructuringNo Same as the originalMatches the reference answer
Is there evidence of pneumothorax, unilateral or bilateral, on this image?
Scope quantificationNo Same as the originalMatches the reference answer
Is there radiographic evidence of pleural air consistent with pneumothorax in this chest radiograph?
Specificity modulationYes Flipped from the originalConflicts with the reference answer
Is there air in the pleural space suggesting pneumothorax on this image?
Lexical substitutionNo Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
Illustrative pictogram, standing in for the study image
The radiograph is not shown
This case uses a MIMIC-CXR study. The PhysioNet licence for MIMIC-CXR prohibits redistributing images, so the radiograph is not published here. Anyone with credentialed access can obtain it from PhysioNet.
The question and the answers are below. The image never changes within a case; only the wording does.
Original question
Is there scarring in the lung apices?
Reference answer
Yes
Why this case mattersOn one image the three models disagree with each other and with themselves: one recovers the correct answer only after a rewording, and another changes on every rewording.
Why these questions count as equivalentEach paraphrase keeps the presence-of-finding operator and the same finding in the same anatomical region (apical scarring, with fibrosis as an accepted synonym), varying only syntax, scope wording, or lexical choice. The equivalence audit rubric treats these rewrite classes as clinically equivalent.
MedGemma-4B
Wrong and unstable
Answer to the original question: No
Is there fibrosis in the apical regions of the lungs?
Lexical substitutionNo Same as the originalConflicts with the reference answer
Can scarring be seen in the apices of the lungs?
Syntactic restructuringYes Flipped from the originalMatches the reference answer
Is there any scarring in the apices of either lung?
Scope quantificationNo Same as the originalConflicts with the reference answer
Do the lung apices show evidence of scarring?
Syntactic restructuringNo Same as the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
GPT-5-mini
Correct until rephrased
Answer to the original question: Yes
Is there fibrosis in the apical regions of the lungs?
Lexical substitutionNo Flipped from the originalConflicts with the reference answer
Can scarring be seen in the apices of the lungs?
Syntactic restructuringYes Same as the originalMatches the reference answer
Is there any scarring in the apices of either lung?
Scope quantificationYes Same as the originalMatches the reference answer
Do the lung apices show evidence of scarring?
Syntactic restructuringYes Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
Claude Haiku 4.5
Wrong and unstable
Answer to the original question: No
Is there fibrosis in the apical regions of the lungs?
Lexical substitutionYes Flipped from the originalMatches the reference answer
Can scarring be seen in the apices of the lungs?
Syntactic restructuringYes Flipped from the originalMatches the reference answer
Is there any scarring in the apices of either lung?
Scope quantificationYes Flipped from the originalMatches the reference answer
Do the lung apices show evidence of scarring?
Syntactic restructuringYes Flipped from the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
Illustrative pictogram, standing in for the study image
The radiograph is not shown
This case uses a MIMIC-CXR study. The PhysioNet licence for MIMIC-CXR prohibits redistributing images, so the radiograph is not published here. Anyone with credentialed access can obtain it from PhysioNet.
The question and the answers are below. The image never changes within a case; only the wording does.
Original question
Is there evidence of atelectasis in this image?
Reference answer
Yes
Why this case mattersSubstituting the synonym pulmonary collapse for atelectasis changes the answer for two of the three models in this illustration, even though the question asks the same clinical thing.
Why these questions count as equivalentEach paraphrase keeps the presence-of-finding operator and the same finding (atelectasis, with pulmonary collapse as an accepted synonym), varying only syntax, scope wording, specificity, or lexical choice. The equivalence audit rubric treats these rewrite classes as clinically equivalent.
MedGemma-4B
Correct until rephrased
Answer to the original question: Yes
Is there evidence of pulmonary collapse on this image?
Lexical substitutionNo Flipped from the originalConflicts with the reference answer
Does this image show evidence of atelectasis?
Syntactic restructuringYes Same as the originalMatches the reference answer
Is there any atelectasis visible on this image?
Scope quantificationYes Same as the originalMatches the reference answer
Is there evidence of atelectasis in the lungs on this image?
Specificity modulationYes Same as the originalMatches the reference answer
Are there radiographic signs of pulmonary collapse on this image?
Lexical substitutionNo Flipped from the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
GPT-5-mini
Correct until rephrased
Answer to the original question: Yes
Is there evidence of pulmonary collapse on this image?
Lexical substitutionNo Flipped from the originalConflicts with the reference answer
Does this image show evidence of atelectasis?
Syntactic restructuringYes Same as the originalMatches the reference answer
Is there any atelectasis visible on this image?
Scope quantificationYes Same as the originalMatches the reference answer
Is there evidence of atelectasis in the lungs on this image?
Specificity modulationYes Same as the originalMatches the reference answer
Are there radiographic signs of pulmonary collapse on this image?
Lexical substitutionNo Flipped from the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
Claude Haiku 4.5
Consistent and correct
Answer to the original question: Yes
Is there evidence of pulmonary collapse on this image?
Lexical substitutionYes Same as the originalMatches the reference answer
Does this image show evidence of atelectasis?
Syntactic restructuringYes Same as the originalMatches the reference answer
Is there any atelectasis visible on this image?
Scope quantificationYes Same as the originalMatches the reference answer
Is there evidence of atelectasis in the lungs on this image?
Specificity modulationYes Same as the originalMatches the reference answer
Are there radiographic signs of pulmonary collapse on this image?
Lexical substitutionYes Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
Illustrative pictogram, standing in for the study image
The radiograph is not shown
This case uses a MIMIC-CXR study. The PhysioNet licence for MIMIC-CXR prohibits redistributing images, so the radiograph is not published here. Anyone with credentialed access can obtain it from PhysioNet.
The question and the answers are below. The image never changes within a case; only the wording does.
Original question
Is there hyperinflation in the lungs?
Reference answer
Yes
Why this case mattersOne model never changes its answer yet is wrong every time: a low flip rate on its own is not evidence of image-grounded reasoning.
Why these questions count as equivalentEach paraphrase keeps the presence-of-finding operator and the same finding (hyperinflation, with hyperexpansion and overinflation as accepted synonyms), varying only syntax, scope wording, specificity, or lexical choice. The equivalence audit rubric treats these rewrite classes as clinically equivalent.
MedGemma-4B
Correct until rephrased
Answer to the original question: Yes
Is there pulmonary hyperexpansion in the lung fields?
Lexical substitutionYes Same as the originalMatches the reference answer
Does the imaging show hyperinflation of the lungs?
Syntactic restructuringNo Flipped from the originalConflicts with the reference answer
Is there any hyperinflation in either lung?
Scope quantificationNo Flipped from the originalConflicts with the reference answer
Is there hyperinflation in the pulmonary parenchyma?
Specificity modulationNo Flipped from the originalConflicts with the reference answer
Is there overinflation of the lungs?
Lexical substitutionYes Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
GPT-5-mini
Consistent but wrong
Answer to the original question: No
Is there pulmonary hyperexpansion in the lung fields?
Lexical substitutionNo Same as the originalConflicts with the reference answer
Does the imaging show hyperinflation of the lungs?
Syntactic restructuringNo Same as the originalConflicts with the reference answer
Is there any hyperinflation in either lung?
Scope quantificationNo Same as the originalConflicts with the reference answer
Is there hyperinflation in the pulmonary parenchyma?
Specificity modulationNo Same as the originalConflicts with the reference answer
Is there overinflation of the lungs?
Lexical substitutionNo Same as the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
Claude Haiku 4.5
Correct until rephrased
Answer to the original question: Yes
Is there pulmonary hyperexpansion in the lung fields?
Lexical substitutionYes Same as the originalMatches the reference answer
Does the imaging show hyperinflation of the lungs?
Syntactic restructuringNo Flipped from the originalConflicts with the reference answer
Is there any hyperinflation in either lung?
Scope quantificationYes Same as the originalMatches the reference answer
Is there hyperinflation in the pulmonary parenchyma?
Specificity modulationYes Same as the originalMatches the reference answer
Is there overinflation of the lungs?
Lexical substitutionYes Same as the originalMatches the reference answer
No paraphrase flipped this model's answer in this case.
Illustrative pictogram, standing in for the study image
The radiograph is not shown
This case uses a MIMIC-CXR study. The PhysioNet licence for MIMIC-CXR prohibits redistributing images, so the radiograph is not published here. Anyone with credentialed access can obtain it from PhysioNet.
The question and the answers are below. The image never changes within a case; only the wording does.
Original question
Is there evidence of tortuosity of the thoracic aorta in this image?
Reference answer
No
Why this case mattersThe three models fail in three different ways on one image: repeated false positives after rewording, stable but wrong throughout, and a single late flip.
Why these questions count as equivalentEach paraphrase keeps the presence-of-finding operator and the same finding in the same structure (tortuosity of the thoracic aorta), varying only syntax, scope wording, or lexical choice. The equivalence audit rubric treats these rewrite classes as clinically equivalent.
MedGemma-4B
Correct until rephrased
Answer to the original question: No
Is there imaging evidence of thoracic aortic tortuosity on this study?
Lexical substitutionNo Same as the originalMatches the reference answer
Does the image show tortuosity of the thoracic aorta?
Syntactic restructuringNo Same as the originalMatches the reference answer
Is there focal or diffuse tortuosity of the thoracic aorta visible on this image?
Scope quantificationNo Same as the originalMatches the reference answer
On this image, is the thoracic aorta tortuous?
Syntactic restructuringYes Flipped from the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
GPT-5-mini
Correct until rephrased
Answer to the original question: No
Is there imaging evidence of thoracic aortic tortuosity on this study?
Lexical substitutionNo Same as the originalMatches the reference answer
Does the image show tortuosity of the thoracic aorta?
Syntactic restructuringYes Flipped from the originalConflicts with the reference answer
Is there focal or diffuse tortuosity of the thoracic aorta visible on this image?
Scope quantificationYes Flipped from the originalConflicts with the reference answer
On this image, is the thoracic aorta tortuous?
Syntactic restructuringYes Flipped from the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
Claude Haiku 4.5
Consistent but wrong
Answer to the original question: Yes
Is there imaging evidence of thoracic aortic tortuosity on this study?
Lexical substitutionYes Same as the originalConflicts with the reference answer
Does the image show tortuosity of the thoracic aorta?
Syntactic restructuringYes Same as the originalConflicts with the reference answer
Is there focal or diffuse tortuosity of the thoracic aorta visible on this image?
Scope quantificationYes Same as the originalConflicts with the reference answer
On this image, is the thoracic aorta tortuous?
Syntactic restructuringYes Same as the originalConflicts with the reference answer
No paraphrase flipped this model's answer in this case.
You have just seen answers flip. This opens one pair up and reads the model's mind on the way down. The Jacobian lens asks, at every layer and every word of the prompt: if the model had to answer from this point, what would it say? Two phrasings of one question, side by side, on one chest X-ray. The default case is a true paraphrase, where the two columns agree all the way down; switch cases to watch the answer column split. Hover any cell to see the words.
The section above reads one pair. This one asks whether that split is a property of the model in general or of that pair alone. Take two versions of the same question, one the model answers correctly and one it does not, and copy the model's internal state from the first into the second at a single layer. If the answer changes, that layer is carrying the decision. Drag through the 34 layers of MedGemma-4B and watch where it happens, across 1,396 pairs at once.
Of 1,396 pairs flip when the state is swapped here
—
—
Flip when an image token is swapped instead, at the same layer
The control: flips here would mean patching any position changes the answer.
—
What this layer is doing
—
The answer commits in a narrow band
Flip rate by transplant layer. The line is the layer you picked.
Median commit layer 16, 95% bootstrap interval [16, 16]. 71% of pairs commit within layers 15 to 17.
The same curve as a table
Layer
Answer-position transplant
Image-token control
The answer is not assembled gradually. Through layer 13 the transplant does almost nothing, under 0.2%. Then it goes 8.2% at layer 14, 35.2% at 15, 73.2% at 16, and 100% by 20. The image-token control never exceeds 0.07%, one pair of 1,396: the flips are specific to the answer position. That is what makes layer 16 the point where this model has decided.
Knowing where the answer commits does not help at inference, because you cannot open the model up on a live case. So: is there a signal you could actually read? The model produces a probability for yes and no on every question, and predictive entropy measures how close that distribution is to a coin flip. The claim in Chapter 7 is that this single number, from one forward pass, ranks both errors and paraphrase flips. Drag the threshold: everything at or below it is answered automatically, everything above is escalated to a person.
Stable and flipped questions by predictive entropy. The line is your threshold.
Answer stayed the same under paraphrase
Answer flipped under paraphrase
The same numbers as a table
Measure
At this threshold
If you answer everything
One forward pass ranks flip risk. Entropy separates flipped from stable questions well enough to score an area under the curve of 0.823 for flips and 0.862 for errors on this model, on a scale where 0.5 is a coin flip and 1.0 is perfect. Tightening the threshold buys accuracy on what remains.
5. The margin gate, and what it cannot catch
The entropy filter above is the deployed model. This is the controlled probe of Chapter 5, a small Gemma-3 decoder over a frozen MedSigLIP-448 encoder that reads a single yes or no from its own language-model head. Its only inference-time signal is the yes-minus-no margin, and a flip is a sign change of that margin across paraphrases, so it is likeliest when the margin already sits near zero. Drag the threshold: every answer with a margin at least this wide is answered automatically, everything narrower is escalated. Switch between the training hospitals and two the probe never saw to see both what one forward pass can catch and what it cannot.
Stable and flipped answers by absolute margin. The line is your threshold; everything to its right is answered automatically.
Answer stayed the same under paraphrase
Answer flipped under paraphrase
The same numbers as a table
Measure
At this threshold
If you answer everything
One forward pass ranks flip risk. The absolute margin separates flipped from stable answers at an area under the curve of — on this set, where 0.5 is a coin flip and 1.0 is perfect, and it holds on the two held-out hospitals (0.974 each). It beats re-asking under several paraphrases and a trained hidden-state probe, at no extra inference.
The margin catches flips, not blind answers
The same margin is at chance (0.47 to 0.52) for the other failure: an answer that ignores the image. A confident answer that never read the radiograph looks exactly like a confident grounded one, so only a second pass that swaps in another patient's image catches it. A finer question, whether a stable answer quietly changes how much it uses the image from one phrasing to another, is real at the population level but not identifiable per case: zero grounded-to-unreliant transitions over 345 patient-seed evaluations, a patient-level upper bound near 3%. Per-case image reliance is a population property here, not a per-prediction label.
6. The deployment-readiness audit
Entropy cannot tell you whether the image was used. So Chapter 7 builds a rule out of several signals at once, and asks it to admit only predictions that are both stable under rephrasing and dependent on the image. Switch the conditions off and on to see what each one is doing.
This is an audit; it cannot be deployed as a gate
The region condition needs a radiologist's bounding box around the finding, which by definition does not exist for a patient who has not been read yet (it is available for 637 of these 861 questions and for none of MIMIC). The swap condition needs a second image known to show the opposite label. Both are offline evaluation checks; neither could run in a clinical workflow. The rule below is the one the dissertation reports. The "text-only" condition is not part of that rule; turn it on to see why.
The rule
Admit when the paraphrases agreeand (swapping the image changes the answerorthe region and the background disagree)
Six real questions, and what the rule did with them
Each is a row from the evaluation, showing the paraphrases the rule compared, what happened when the image was swapped or masked to the annotated region, and whether the question was answerable from the text alone.
The rule works on its own terms: on the targeted adapter it admits 33.0% of questions at 96.8% accuracy, against 91.5% if you answer everything, and every condition you remove admits more and is right less often. What it does not buy is grounding. Watch the fourth number: most of what the rule admits is exactly what the question text alone already answers. On the slice where the model's text-only answer disagrees with its full answer, the same targeted adapter scores 2.9% accuracy (n=241), and Qwen2-VL gets 0 of 279 right under every rule tried. High admitted accuracy here comes from high finding base rates making the text prior usually correct. The model is not reading the radiograph. The dissertation's answer to "can a gate admit image-grounded cases?" is no.