Recognise an overclaim
Tonight · ~15 min · read · energy: low · setup: none
Certain words in imaging-AI conclusions carry more weight than the evidence supports. This short lesson is a vocabulary of overclaim — the terms to flag, and the humbler statement that the evidence actually licenses. Once you can rewrite an overclaim into an honest claim, appraisal becomes almost mechanical.
The anchor: claims have a scope; evidence sets it
A study’s evidence sets the scope of what it licenses: an internal test licenses a within-distribution claim; an unvalidated feature licenses nothing about biology; a high AUROC licenses nothing about utility. Overclaim happens when the conclusion’s words are wider than that scope. The fix as a reader: take the strongest claim and ask whether the strongest metric supports it; if not, mentally narrow the claim to what the evidence actually shows.
The vocabulary to flag
- “Biomarker” — reserved for a feature validated against a biological/clinical reference (Ch. 3). An unvalidated texture feature is a feature, not a biomarker. Honest rewrite: “candidate feature associated with the outcome”.
- “Generalizable” / “generalisable” — licenses only what the testing type supports (lesson 4). Single-site internal data licenses “internally tested at one site”, not “generalisable”. Frozen external data licenses travel to that site; retrained external licenses transfer-after- adaptation, not raw generalisation.
- “Clinically useful” — licenses only what utility analysis shows (lesson 5). A high AUROC licenses “good discrimination”; clinical usefulness needs calibration + decision impact.
- “Validated” — ambiguous (tuning set or test set? internal or external?). Demand the precise testing type.
- “External validation” — check frozen-vs-retrained and whether the threshold was re-estimated (lesson 4).
The rewrite habit
For each flagged word, produce the honest version:
| Overclaim | Honest rewrite (given typical evidence) |
|---|---|
| “a new imaging biomarker” | “a candidate feature associated with outcome, pending validation” |
| “generalisable across centres” | “frozen-model performance held on one external site; wider generalisation untested” |
| “clinically useful” | “well-discriminating; clinical utility not assessed” |
| “externally validated” | “tested on an external site after fine-tuning, with per-dataset thresholds” |
This is not pedantry — it is the difference between reading a conclusion and reading the evidence. The Tran et al. dissection (lesson 8) is the worked example: “multi-continental external validation” becomes, after the frozen-vs-retrained check, a narrower and more honest statement.
Overclaim is not always dishonest
Often overclaim is optimism, not fraud: authors believe their model is great and their prose reflects that belief. Your job as a reader is not to impute motive but to narrow the claim to the evidence. Sometimes that leaves a solid, useful, narrower result — and sometimes it leaves almost nothing. Both outcomes are informative.
Stop and think — then reveal
A radiomics abstract says: “We identified a robust CT biomarker of early recurrence that generalises across scanners and is clinically useful for surveillance.” The study is single-site, internal test only, AUROC reported, no calibration or decision-curve. Rewrite the claim honestly.
“ robust CT biomarker“ → “a candidate feature associated with recurrence” (no biological/clinical validation reported, so it is a feature, not a biomarker; robustness also unproven without test–retest). “Generalises across scanners” → “tested at a single site” (no external or cross-scanner evidence). “Clinically useful for surveillance” → “discriminates recurrence internally” (no calibration or decision-utility analysis). Honest version: a candidate radiomic feature associated with recurrence in a single-site internal test (AUROC …); external validation, calibration, and clinical utility were not assessed. That is a much narrower — and far more common — claim than the abstract’s.
What to retain
- Overclaim = conclusion words wider than the evidence scope. Flag the vocabulary.
- “Biomarker” needs validation; “generalisable” needs external testing; “clinically useful” needs utility analysis; “validated” needs a precise testing type.
- Build the rewrite habit: narrow each flagged word to what the strongest metric supports.
- Overclaim is often optimism, not fraud — your job is to narrow the claim to the evidence, whatever the motive.
Next: the frameworks that force these questions into papers — the reporting and risk-of-bias frameworks.