How to read a reader-study result
Tonight · ~25 min · read · energy: medium · setup: none
This lesson walks a real reader study — Herpe et al. — applying the design lens of lesson 3. It is honest about the evidence level: the full PDF was not retrievable at write time, so the facts are abstract-level, and that limitation is part of the lesson in reading critically.
The paper
Herpe et al. — AI assistance impact on radiologist interpretation of knee MRI, Eur Radiol 2026;36(2):1294-1305, PMID 40745051 (peer-reviewed; abstract inspected — full PDF not retrievable at write time, so the facts below are abstract-level).
The design (abstract-level)
- Design — retrospective two-session paired reader study (Jan–Apr 2024); 6 radiologists (2–10 years MSK experience); 165 knee MRIs; each reader read the same cases with and without AI, in two sessions 1 month apart (each reader is their own control, separated by a ~1-month interval — functioning as a washout).
- Independence of the AI — the model was trained on 23,074 studies separate from the study dataset, so the reader cohort did not leak into model training (the study itself is retrospective, not live deployment).
- Reference standard — consensus of three expert MSK radiologists (not a single reader).
- Lesions assessed — ACL, meniscus, cartilage, and MCL (multi-structure).
Note: these design facts are abstract-level; the abstract does not establish session ordering or counterbalancing, so I do not claim how order effects were handled (lesson 3).
The result (reader-alone → reader+AI)
sensitivity 81% → 86%, specificity 88% → 93%, accuracy 86% → 91% (all p<0.001); and the standout — inter-reader agreement (Fleiss’ κ) rose 0.54 → 0.78 (p<0.001).
What it demonstrates about radiologist + AI: on this retrospective cohort, AI assistance improved readers’ accuracy and made them more consistent with each other (κ 0.54→0.78) — i.e. AI acted as a harmonising/second-check aid, not just an accuracy boost. What it does not establish: it is retrospective (not a prospective live-workflow study), it does not prove a reading-time/efficiency benefit (the abstract names “efficiency” but does not report a reading-time metric), and it does not establish clinical outcomes. Read it as a strong retrospective reader+AI demonstration, not as deployment evidence — and note that even a clean reader+AI delta does not show how the tool behaves under workflow pressure (lessons 6–7) or drift (lesson 8).
Reading the result with the lens
Applying lesson 3’s checklist to Herpe et al.:
- MRMC? Multi-reader (6) × multi-case (165) — yes, the standard structure.
- Crossover with washout? Two-session paired design with a ~1-month interval — functions as a washout; counterbalancing/order not established from the abstract (a limitation to carry).
- Independent reference standard? Consensus of three expert MSK radiologists — stronger than a single reader (Ch. 7).
- Outcomes beyond AUROC? Sensitivity, specificity, accuracy, and inter-reader agreement reported; reading time not reported despite an “efficiency” framing — a gap.
- Leakage? Model trained on separate data — good.
The lesson: even a well-designed reader study has a scoped claim (retrospective, no efficiency metric, no clinical outcomes), and the honest reading keeps that scope. The κ improvement is the most interesting result — AI as a consistency aid is a different and arguably more deployable claim than AI as an accuracy aid.
Stop and think — then reveal
Herpe et al. report improved accuracy and a κ rise 0.54 → 0.78, but no reading-time metric. A hospital cites it to justify buying the tool “to speed up reporting”. What is the gap between the evidence and the justification?
The study shows accuracy and consistency gains on a retrospective cohort; it does not report a reading-time/efficiency metric (the abstract names efficiency without quantifying it). So “speed up reporting” is not supported by this evidence — speed is a workflow/efficiency claim that needs a time metric, ideally under live-workflow conditions (lesson 7). The justifiable claim from Herpe is “may improve accuracy and inter-reader consistency on retrospective knee MRI”; the efficiency claim needs a different study. This is exactly the overclaim discipline of Chapter 7 applied to deployment.
What to retain
- Herpe et al. (PMID 40745051): retrospective MRMC two-session reader study, 6 readers × 165 knee MRIs, consensus-of-3 reference standard, AI trained on separate data.
- Result: sensitivity 81→86%, specificity 88→93%, accuracy 86→91%, and inter-reader κ 0.54→0.78 — AI as a consistency aid as much as an accuracy aid.
- Scope: retrospective, no reading-time metric despite an “efficiency” framing, no clinical outcomes — a strong reader+AI demonstration, not deployment evidence.
- Even a clean reader+AI delta does not show behaviour under workflow pressure or drift; abstract-level facts (order/counterbalancing) carry an honest limitation.
Next: even with a positive reader study, why thresholds change when the workflow changes.