Binesh Sadanandan PhD Dissertation Companion
Read thesis

Questions

Questions, answered with the evidence

The questions a careful reader asks about this work. Each answer links to the result it rests on.

The equivalence audit uses a language-model judge. Why trust it?

It is validated. Three reviewers, a radiologist, a clinician, and a device-research reviewer, adjudicated a 1,200-pair sample. They agreed with each other on 98.5% of triple-reviewed pairs, kappa 0.97, so the task is well defined. The judge matched their consensus on 72.3%, and it disagrees most on exactly the adversarial cases the sample over-represents. That sample validates the judge. The flip rates are a separate matter, and the audit rejects roughly half of its candidate pairs, including 93.3% of generated negation-pattern paraphrases.

The adjudication result The audit result

Does the adapter make the model safer?

No. The adapter cuts flips from 8.5% to 3.5%, but text-only agreement rises with it, from 53.5% to 76.8%: the model returns its original answer with the image deleted 23 points more often. On the four-quadrant screen, the adapters concentrate predictions in the consistent-but-image-invariant cell, 84.1% of PadChest questions for the targeted adapter against 22.5% for the base model. The adapter buys consistency from the question text. Safer would require image dependence, which it does not buy.

The re-audit result The quadrant chart

Was the clinician study conducted?

Two different things carry that name, and they have different answers. The clinician adjudication of paraphrase equivalence was conducted: three reviewers, including a radiologist and a clinician, adjudicated 1,200 pairs, and that is what validates the audit judge. The clinician deployment consultation, on how the gate and escalation workflow would sit in practice, is designed but not conducted, and the dissertation's conclusion lists it as future work.

The adjudication result

The audit kept 61,761 pairs. Why is the benchmark 92,856?

Because the two numbers describe different stages, and they never combine. 122,778 candidate pairs went through the audit and 61,761, 50.3%, survived. The audited PadChest v1 portion was then superseded: PadChest was regenerated and separately filtered into a v2 set of 59,573 pairs. The shipped benchmark is audit-retained MIMIC-CXR, 8,938, plus audit-retained VinDr-CXR, 24,345, plus PadChest v2, 59,573. That sum is 92,856. The figure below draws the path.

The audit result The benchmark result

Are these results peer reviewed?

Five papers are accepted, and they carry Thrusts 1, 2, 3 and both results of Thrust 4: the four-quadrant screen, and the attention-faithfulness result, accepted at iMIMIC 2026, a MICCAI 2026 workshop. The entropy screen and the deployment audit of Thrust 5 are two separate manuscripts, one under review and one in revision, and the Chapter 2 scoping review is in open peer review at JMIR AI, which posts the named preprint itself.

The papers

Cite

Permanent link

Type to search. Press Escape to close.