What evidence would make you trust this system?
Tonight · ~25 min · synthesis exercise · energy: medium · setup: none
This closing lesson is a synthesis: given everything in the chapter, what evidence would actually justify trusting a deployed imaging-AI system? Work the exercise before reading the model answer — the skill is assembling the case, not memorising the list.
The anchor: trust is a stack of scoped claims
Trusting a deployed system is not a single verdict; it is a stack of scoped claims, each backed by a different kind of evidence. A system you “trust” has positive answers across standalone accuracy, reader+AI delta, integration, threshold/calibration, human-factors safety, and ongoing monitoring. A gap in any layer is a gap in trust — and the honest reading names the gaps, not just the strengths.
The exercise
Take a concrete system — say, a TotalSegmentator-class assistant proposed for clinical volumetry — and assemble the evidence you would require before trusting it. For each layer, state what evidence you have and what is missing:
- Standalone correctness — internal + external testing, frozen model, with calibration (Ch. 4, Ch. 6).
- Reader+AI delta — an MRMC crossover reader study showing the assistant improves reader measurements/consistency, not just standalone Dice (lessons 3–4).
- Integration — evidence the result is retrieved, rendered, and linked to the source study in this site’s PACS (Ch. 5).
- Operating point — thresholds and calibration checked at this site’s prevalence and workflow role (lesson 5).
- Human-factors safety — evidence that automation bias/alert fatigue are monitored and mitigated (lesson 6).
- Deployment staging — shadow deployment results before live use (lesson 7).
- Lifecycle — monitoring for drift, versioning/provenance, rollback (lesson 8).
Model answer (assemble your own first)
A trustworthy deployment case reads roughly: the model was frozen-externally tested with calibration; an MRMC reader study showed reader+AI improved volumetry consistency without automation-bias harm; the result integrates and renders in our PACS; the operating threshold was set and calibrated for our prevalence and triage role; it ran shadow for N weeks with acceptable FP load; and there is monitoring with drift alerts, versioning, and rollback. Each clause is a scoped claim with its own evidence; absent clauses are absent trust. Most real “deployed AI” papers provide the first clause and gloss the rest — which is exactly why Chapter 7’s appraisal lens applies to deployment claims as much as to development claims.
The transferable skill
This is the same skill as designing a study (Ch. 4, lesson 8) and appraising a paper (Ch. 7): enumerate the claims, demand the evidence for each, and keep the scope honest. Trust is not granted by a single number; it is assembled, layer by layer, and it has to be re-earned as the world drifts.
Stop and think — then reveal
A vendor offers a deployed segmentation assistant with strong standalone Dice and a frozen external test. They have no reader study, no shadow-deployment data, and no monitoring plan. Which trust layers are present and which are missing?
Present: standalone correctness (frozen external test) — layer 1, partially (no calibration mentioned). Missing: the reader+AI delta (layer 2 — no reader study, so we do not know if it helps readers), integration (layer 3 — no evidence it renders in our PACS), operating-point/calibration (layer 4), human-factors safety (layer 5), deployment staging (layer 6 — no shadow data), and lifecycle monitoring (layer 7 — no plan). So this is a single-layer trust case dressed as a deployment: you would require the reader study, integration evidence, threshold/calibration at your site, shadow deployment, and a monitoring plan before trusting it. One good layer is not a stack.
What to retain
- Trust is a stack of scoped claims (standalone, reader+AI, integration, operating point, human-factors, staging, lifecycle), each with its own evidence.
- Assemble the case layer by layer; absent layers are absent trust.
- Most “deployed AI” claims provide the standalone layer and gloss the rest — apply the Ch. 7 appraisal lens to deployment claims.
- Trust must be re-earned as the world drifts; monitoring is not optional.
You have finished Chapter 9 — and the course. The deployment-path diagram, AIRA/AIW-I links, agentic-workflow risks, and further reading are in the deployment reference. Return to the start page for where to go next.