How often do equivalent questions change model answers?
PSF-Med evaluates six medical vision-language models on 92,856 evaluated paraphrase pairs built from 26,850 chest X-ray questions spanning three countries. Pairwise flip rates on the binary yes/no subset range from 6.4% to 54.7% (an 8.5x spread), and scale does not buy consistency. Two model-dataset cells with degenerate yes-bias (>98% positive on originals) are set aside, and no consistency is claimed for them.
How this thrust was run
Generate roughly 92,000 candidate paraphrases, adjudicate 122,778 pairs for semantic equivalence with an LLM judge validated against clinician review (50.3% retained), regenerate the PadChest paraphrases to remove laterality and scope leakage and judge-filter that set separately, then measure answer flips at both pairwise and query-level denominators on the resulting 92,856 pairs: audit-retained MIMIC-CXR (8,938) and VinDr-CXR (24,345) plus PadChest v2 (59,573), which is not a subset of the 61,761-pair audited core.
Confirmed Paraphrase sensitivity is common in medical vision-language models: pairwise flip rates on the binary yes/no subset span 6.4% to 54.7% across six models and three chest X-ray datasets, an 8.5x spread.
Supported Scale does not buy paraphrase consistency: MedGemma-27B is the most consistent of the three MedGemma variants on MIMIC (6.4%) and VinDr (8.1%) but not on PadChest, where MedGemma-4B edges it (13.4% vs 13.9%), and no stable consistency ordering by parameter count exists across datasets.
Exploratory Paraphrase sensitivity persists in frontier general-purpose models that were never medically finetuned: four configurations of three such models flip on 3.1% to 9.5% of clinically equivalent rephrasings, and enabling reasoning on Claude Opus 5 raised its flip rate rather than lowering it (9.53% against 8.36%, p = 0.022). On PadChest, the one population where a frontier run and the published medical cells are directly comparable, the frontier rate is 3.08% against 12.0% to 32.8%.
Primary results
6.4% to 54.7%
pairwise
Paraphrase flip rates span 6.4% to 54.7% across six medical vision-language models
Rephrasing a clinical yes/no question in a way that preserves its meaning changes the answer on anywhere from 1 in 16 to more than half of paraphrase pairs, depending on which model and which patient population you test. Every model tested flips, and the spread between the best and worst model-dataset cell is 8.5-fold.
Multiple models · Multiple datasets · A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%).
Source, denominator, and limits
Metric
pairwise paraphrase flip rate (range across six models x three datasets)
Denominator
pairwise (per question-paraphrase pair)
Sample
A range over 18 model-dataset cells, so it has no single denominator. Per-dataset binary-subset denominators are MIMIC-CXR 1,539 questions / 5,076 pairs, PadChest 8,445 / 36,244, and VinDr-CXR 2,807 / 8,612, each evaluated on every model. The endpoints are MedGemma-27B on MIMIC (6.4%) and RadFM on VinDr (54.7%).
MedGemma-4B flips on 8.3% of MIMIC-CXR paraphrase pairs
On in-distribution US chest X-ray questions, the mechanistic subject model gives a different yes/no answer to about 1 in 12 clinically equivalent rephrasings.
The benchmark that every headline flip rate is computed on contains 92,856 (question, paraphrase) pairs across three continents: 8,938 from MIMIC-CXR, 59,573 from PadChest v2 and 24,345 from VinDr-CXR, averaging 3.5 paraphrases per question.
Not model-specific · Multiple datasets · n = 92,856
Source, denominator, and limits
Metric
post-audit evaluation pairs in the PSF-Med benchmark
The equivalence audit rejects nearly half of generated paraphrases
A rubric-based GPT-5-mini audit of 122,778 candidate pairs kept only 50.3% (61,761) as clinically equivalent, rejected 48.7% (59,788) as adversarial or meaning-changing, and flagged 1.0% (1,229) as uncertain. Roughly half of what an LLM generates as a 'paraphrase' of a clinical question does not preserve the question.
Not model-specific · Multiple datasets · n = 122,778
Source, denominator, and limits
Metric
share of audited paraphrase pairs retained as core-equivalent
The automated judge agrees with clinicians on only 72.3% of pairs
Three reviewers (a radiologist, a clinician, and a device-research reviewer) adjudicated a 1,200-pair sample and agreed with each other almost perfectly (98.5% raw, kappa 0.97 on the 400 triple-reviewed pairs), but the GPT-5-mini equivalence judge matched their consensus on just 72.3% of pairs (kappa 0.52). The judge and the clinicians agree on easy cases and diverge exactly on the adversarial and uncertain ones the sample was designed to over-represent.
Not model-specific · Multiple datasets · n = 1,200 · Cohen's kappa 0.52 (moderate); consensus labels 689 equivalent / 409 not equivalent / 102 uncertain
Source, denominator, and limits
Metric
raw agreement between clinician consensus and the GPT-5-mini equivalence judge