The four-quadrant screen
A dummy example on 100 chest X-rays. An AI answers a yes/no question for each one, and every case gets two behavioral tests: reword the question, then change the image. Consistency crossed with image reliance sorts the 100 answers into four cells, and only one cell is safe. Run the screen step by step, and toggle the models to see why a low flip rate proves nothing by itself.
Toggle the models and watch the reversal. Model B posts the better flip rate, 4% against 12%, and the flip-rate card even turns green for it. Then the second test exposes the purchase: 81% of B's consistent answers never looked at the image, so its reliability is a costume. Model A flips more, yet most of its consistency is genuine. A low flip rate is where the audit starts, never where it ends, because consistency can be bought from the text prior, and the Dangerous cell is where that purchase hides.
Is the screen grading itself?
The cleanest check is information-theoretic. For every prediction, measure the Kullback-Leibler divergence between the output distribution with the image and without it: how much does removing the image move the whole distribution, not just the final answer? The quadrant labels are discrete and never see this number. If the labels are real, Dangerous cases should sit near zero and Ideal cases should sit high. Drag the threshold and watch how well one continuous signal recovers the behavioral labels on the 100 dummy cases.
Three independent signals, all agreeing with the behavioral labels, so the screen is a finding rather than a relabeling. In the study that introduced this screen, the divergence check separates Dangerous from Ideal at AUROC 0.76 on real predictions; this toy is cleaner than reality on purpose, so the mechanism is easy to see.