81% (range 45-100%)
model-dataset81% of consistent predictions are image-invariant
Averaged across ten model-dataset settings, 81% of each model's consistent predictions are unchanged when the image is removed (range 45-100%), so a low paraphrase-flip rate is not evidence of grounded visual reasoning.
Multiple models · Multiple datasets · n = 10 · unweighted mean over 10 settings; range 45-100%
Source, denominator, and limits
- Metric
- mean share of consistent predictions that are image-invariant (text-only agreement)
- Denominator
- model-dataset (per model-dataset setting)
- Sample
- n = 10
- Model
- Multiple models
- Dataset
- Multiple datasets
- Split
- eval
- Comparison
- range 45-100% across the ten model-dataset settings
- Uncertainty
- unweighted mean over 10 settings; range 45-100%
- Thesis
- Chapter 6, tab:quadrant_counts
- Source artifact
dissertation/tables/thrust4/table_quadrant_counts.tex- Last verified
- 2026-07-15