Skip to content

Recognise an overclaim

Tonight · ~15 min · read · energy: low · setup: none

Certain words in imaging-AI conclusions carry more weight than the evidence supports. This short lesson is a vocabulary of overclaim — the terms to flag, and the humbler statement that the evidence actually licenses. Once you can rewrite an overclaim into an honest claim, appraisal becomes almost mechanical.

The anchor: claims have a scope; evidence sets it

A study’s evidence sets the scope of what it licenses: an internal test licenses a within-distribution claim; an unvalidated feature licenses nothing about biology; a high AUROC licenses nothing about utility. Overclaim happens when the conclusion’s words are wider than that scope. The fix as a reader: take the strongest claim and ask whether the strongest metric supports it; if not, mentally narrow the claim to what the evidence actually shows.

The vocabulary to flag

  • “Biomarker” — reserved for a feature validated against a biological/clinical reference (Ch. 3). An unvalidated texture feature is a feature, not a biomarker. Honest rewrite: “candidate feature associated with the outcome”.
  • “Generalizable” / “generalisable” — licenses only what the testing type supports (lesson 4). Single-site internal data licenses “internally tested at one site”, not “generalisable”. Frozen external data licenses travel to that site; retrained external licenses transfer-after- adaptation, not raw generalisation.
  • “Clinically useful” — licenses only what utility analysis shows (lesson 5). A high AUROC licenses “good discrimination”; clinical usefulness needs calibration + decision impact.
  • “Validated” — ambiguous (tuning set or test set? internal or external?). Demand the precise testing type.
  • “External validation” — check frozen-vs-retrained and whether the threshold was re-estimated (lesson 4).

The rewrite habit

For each flagged word, produce the honest version:

Overclaim Honest rewrite (given typical evidence)
“a new imaging biomarker” “a candidate feature associated with outcome, pending validation”
“generalisable across centres” “frozen-model performance held on one external site; wider generalisation untested”
“clinically useful” “well-discriminating; clinical utility not assessed”
“externally validated” “tested on an external site after fine-tuning, with per-dataset thresholds”

This is not pedantry — it is the difference between reading a conclusion and reading the evidence. The Tran et al. dissection (lesson 8) is the worked example: “multi-continental external validation” becomes, after the frozen-vs-retrained check, a narrower and more honest statement.

Overclaim is not always dishonest

Often overclaim is optimism, not fraud: authors believe their model is great and their prose reflects that belief. Your job as a reader is not to impute motive but to narrow the claim to the evidence. Sometimes that leaves a solid, useful, narrower result — and sometimes it leaves almost nothing. Both outcomes are informative.

Stop and think — then reveal

A radiomics abstract says: “We identified a robust CT biomarker of early recurrence that generalises across scanners and is clinically useful for surveillance.” The study is single-site, internal test only, AUROC reported, no calibration or decision-curve. Rewrite the claim honestly.

“ robust CT biomarker“ → “a candidate feature associated with recurrence” (no biological/clinical validation reported, so it is a feature, not a biomarker; robustness also unproven without test–retest). “Generalises across scanners” → “tested at a single site” (no external or cross-scanner evidence). “Clinically useful for surveillance” → “discriminates recurrence internally” (no calibration or decision-utility analysis). Honest version: a candidate radiomic feature associated with recurrence in a single-site internal test (AUROC …); external validation, calibration, and clinical utility were not assessed. That is a much narrower — and far more common — claim than the abstract’s.

What to retain

  1. Overclaim = conclusion words wider than the evidence scope. Flag the vocabulary.
  2. “Biomarker” needs validation; “generalisable” needs external testing; “clinically useful” needs utility analysis; “validated” needs a precise testing type.
  3. Build the rewrite habit: narrow each flagged word to what the strongest metric supports.
  4. Overclaim is often optimism, not fraud — your job is to narrow the claim to the evidence, whatever the motive.

Next: the frameworks that force these questions into papers — the reporting and risk-of-bias frameworks.