Binesh Sadanandan PhD Dissertation Companion
Read thesis

Thrust 1 · Chapter 3

Measurement

How often do equivalent questions change model answers?

PSF-Med evaluates six medical vision-language models on 92,856 evaluated paraphrase pairs built from 26,850 chest X-ray questions spanning three countries. Pairwise flip rates on the binary yes/no subset range from 6.4% to 54.7% (an 8.5x spread), and scale does not buy consistency. Two model-dataset cells with degenerate yes-bias (>98% positive on originals) are set aside, and no consistency is claimed for them.

How this thrust was run

Generate roughly 92,000 candidate paraphrases, adjudicate 122,778 pairs for semantic equivalence with an LLM judge validated against clinician review (50.3% retained), regenerate the PadChest paraphrases to remove laterality and scope leakage and judge-filter that set separately, then measure answer flips at both pairwise and query-level denominators on the resulting 92,856 pairs: audit-retained MIMIC-CXR (8,938) and VinDr-CXR (24,345) plus PadChest v2 (59,573), which is not a subset of the 61,761-pair audited core.

Datasets
MIMIC-CXR, PadChest, VinDr-CXR
Models
MedGemma-27B, MedGemma-1.5-4B, MedGemma-4B, CheXone, LLaVA-Rad, RadFM
Thesis tables
Table 3.3 (tab:t1_flip_rates), Table 3.4 (tab:t1_denominators)
Papers
PSF-Med

Confirmed Paraphrase sensitivity is common in medical vision-language models: pairwise flip rates on the binary yes/no subset span 6.4% to 54.7% across six models and three chest X-ray datasets, an 8.5x spread.

Supported Scale does not buy paraphrase consistency: MedGemma-27B is the most consistent of the three MedGemma variants on MIMIC (6.4%) and VinDr (8.1%) but not on PadChest, where MedGemma-4B edges it (13.4% vs 13.9%), and no stable consistency ordering by parameter count exists across datasets.

Exploratory Paraphrase sensitivity persists in frontier general-purpose models that were never medically finetuned: four configurations of three such models flip on 3.1% to 9.5% of clinically equivalent rephrasings, and enabling reasoning on Claude Opus 5 raised its flip rate rather than lowering it (9.53% against 8.36%, p = 0.022). On PadChest, the one population where a frontier run and the published medical cells are directly comparable, the frontier rate is 3.08% against 12.0% to 32.8%.

Primary results

6.4% to 54.7%

pairwise

Paraphrase flip rates span 6.4% to 54.7% across six medical vision-language models

Rephrasing a clinical yes/no question in a way that preserves its meaning changes the answer on anywhere from 1 in 16 to more than half of paraphrase pairs, depending on which model and which patient population you test. Every model tested flips, and the spread between the best and worst model-dataset cell is 8.5-fold.

Multiple models · Multiple datasets · A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%).

Source, denominator, and limits
Metric
pairwise paraphrase flip rate (range across six models x three datasets)
Denominator
pairwise (per question-paraphrase pair)
Sample
A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%).
Model
Multiple models
Dataset
Multiple datasets
Split
eval
Comparison
An 8.5x spread between the most consistent cell (MedGemma-27B on MIMIC) and the least consistent (RadFM on VinDr).
Uncertainty
not reported for this value
Thesis
Chapter 3, Table 3.3
Source artifact
results/uai/revision/psf_binary_recompute_v2.json
Last verified
2026-07-15

8.3%

pairwise

MedGemma-4B flips on 8.3% of MIMIC-CXR paraphrase pairs

On in-distribution US chest X-ray questions, the mechanistic subject model gives a different yes/no answer to about 1 in 12 clinically equivalent rephrasings.

MedGemma-4B · MIMIC-CXR · n = 5,076

Source, denominator, and limits
Metric
pairwise paraphrase flip rate
Denominator
pairwise (per question-paraphrase pair)
Sample
n = 5,076
Model
MedGemma-4B
Dataset
MIMIC-CXR
Split
eval
Comparison
13.4% for the same model on PadChest and 15.3% on VinDr-CXR; 18.1% query-level on this same MIMIC set.
Uncertainty
not reported for this value
Thesis
Chapter 3, Table 3.3
Source artifact
results/uai/revision/psf_binary_recompute_v2.json
Last verified
2026-07-15

92,856

pairwise

PSF-Med evaluates 92,856 question-paraphrase pairs

The benchmark that every headline flip rate is computed on contains 92,856 (question, paraphrase) pairs across three continents: 8,938 from MIMIC-CXR, 59,573 from PadChest v2 and 24,345 from VinDr-CXR, averaging 3.5 paraphrases per question.

Not model-specific · Multiple datasets · n = 92,856

Source, denominator, and limits
Metric
post-audit evaluation pairs in the PSF-Med benchmark
Denominator
pairwise (per question-paraphrase pair)
Sample
n = 92,856
Model
Not model-specific
Dataset
Multiple datasets
Split
eval
Comparison
Distinct from the roughly 92,000 construction-stage candidate paraphrases and from the 122,778 pairs sent through the rubric audit.
Uncertainty
not reported for this value
Thesis
Chapter 3, Table 3.1
Source artifact
dissertation/tables/thrust1/table_dataset_stats.tex
Last verified
2026-07-15

50.3%

pairwise

The equivalence audit rejects nearly half of generated paraphrases

A rubric-based GPT-5-mini audit of 122,778 candidate pairs kept only 50.3% (61,761) as clinically equivalent, rejected 48.7% (59,788) as adversarial or meaning-changing, and flagged 1.0% (1,229) as uncertain. Roughly half of what an LLM generates as a 'paraphrase' of a clinical question does not preserve the question.

Not model-specific · Multiple datasets · n = 122,778

Source, denominator, and limits
Metric
share of audited paraphrase pairs retained as core-equivalent
Denominator
pairwise (per question-paraphrase pair)
Sample
n = 122,778
Model
Not model-specific
Dataset
Multiple datasets
Split
eval
Comparison
Retention varies sharply by dataset: 72.9% MIMIC, 35.7% PadChest, 78.6% VinDr.
Uncertainty
not reported for this value
Thesis
Chapter 3, Chapter 3, Rubric-Based Equivalence Audit
Source artifact
scripts/analysis/recompute_audit_retention.py
Last verified
2026-07-15

72.3%

pairwise

The automated judge agrees with clinicians on only 72.3% of pairs

Three reviewers (a radiologist, a clinician, and a device-research reviewer) adjudicated a 1,200-pair sample and agreed with each other almost perfectly (98.5% raw, kappa 0.97 on the 400 triple-reviewed pairs), but the GPT-5-mini equivalence judge matched their consensus on just 72.3% of pairs (kappa 0.52). The judge and the clinicians agree on easy cases and diverge exactly on the adversarial and uncertain ones the sample was designed to over-represent.

Not model-specific · Multiple datasets · n = 1,200 · Cohen's kappa 0.52 (moderate); consensus labels 689 equivalent / 409 not equivalent / 102 uncertain

Source, denominator, and limits
Metric
raw agreement between clinician consensus and the GPT-5-mini equivalence judge
Denominator
pairwise (per question-paraphrase pair)
Sample
n = 1,200
Model
Not model-specific
Dataset
Multiple datasets
Split
eval
Comparison
Inter-reviewer agreement on the same task is 98.5% raw (394/400, Cohen's kappa 0.97); judge-vs-judge cross-family agreement is 91.6-94.4%.
Uncertainty
Cohen's kappa 0.52 (moderate); consensus labels 689 equivalent / 409 not equivalent / 102 uncertain
Thesis
Chapter 3, Chapter 3, Clinician Adjudication
Source artifact
results/clinician_review/run_v2/kappa_report.json
Last verified
2026-07-15

Charts from this thrust

Open any of these in the evidence explorer to filter by model, dataset, or metric.

Pairwise paraphrase flip rate by model and dataset

Six base vision-language models on the equivalence-filtered PSF-Med binary yes/no subset: MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, VinDr-CXR 2,807 / 8,612.

Query-level flip rate: at least one paraphrase flips

Same PSF-Med binary subset as the pairwise chart, re-scored per question (MIMIC-CXR n=1,539, PadChest n=8,445, VinDr-CXR n=2,807 questions). MedGemma-27B on PadChest is absent from the recompute file and is not plotted.

Accuracy against pairwise flip rate

17 model-dataset cells (six base models x three datasets; MedGemma-27B / PadChest absent) on the PSF-Med binary subset, 5,076 / 36,244 / 8,612 pairs.

How many generated paraphrases survive the equivalence audit

All 122,778 candidate pairs adjudicated by the rubric audit: MIMIC-CXR 12,259, PadChest 79,378, VinDr-CXR 31,141. Overall 61,761 retained (50.3%), 59,788 rejected (48.7%), 1,229 uncertain (1.0%).

Which kinds of rephrasing break the model

Mean pairwise flip rate over the six base models on the PSF-Med binary yes/no subset, by transformation type. Per-cell n is binary-subset pairs per model (11 to 14,662). PadChest has no negation cell.

Clinicians agree with each other; the automated judge is the weak link

1,200-pair stratified adjudication sample, three reviewers, 400 pairs triple-reviewed. Reviewer-versus-reviewer agreement on the 400; clinician-consensus-versus-GPT-5-mini agreement on all 1,200 and on the 400.

Paraphrase flip rates for frontier general-purpose models

Four configurations of three frontier models with no medical finetuning, scored on the same clinically equivalent paraphrase pairs: Claude Opus 5 on 6,209 pairs with reasoning on and 6,194 with it off, GPT-5.6 Sol on 6,195, all across MIMIC-CXR, PadChest and VinDr-CXR; Kimi K3 on 2,274 PadChest pairs only. Bars are Wilson 95% confidence intervals.

Cite

Permanent link

Type to search. Press Escape to close.