Skip to content

What evidence would make you trust this system?

Tonight · ~25 min · synthesis exercise · energy: medium · setup: none

This closing lesson is a synthesis: given everything in the chapter, what evidence would actually justify trusting a deployed imaging-AI system? Work the exercise before reading the model answer — the skill is assembling the case, not memorising the list.

The anchor: trust is a stack of scoped claims

Trusting a deployed system is not a single verdict; it is a stack of scoped claims, each backed by a different kind of evidence. A system you “trust” has positive answers across standalone accuracy, reader+AI delta, integration, threshold/calibration, human-factors safety, and ongoing monitoring. A gap in any layer is a gap in trust — and the honest reading names the gaps, not just the strengths.

The exercise

Take a concrete system — say, a TotalSegmentator-class assistant proposed for clinical volumetry — and assemble the evidence you would require before trusting it. For each layer, state what evidence you have and what is missing:

  1. Standalone correctness — internal + external testing, frozen model, with calibration (Ch. 4, Ch. 6).
  2. Reader+AI delta — an MRMC crossover reader study showing the assistant improves reader measurements/consistency, not just standalone Dice (lessons 3–4).
  3. Integration — evidence the result is retrieved, rendered, and linked to the source study in this site’s PACS (Ch. 5).
  4. Operating point — thresholds and calibration checked at this site’s prevalence and workflow role (lesson 5).
  5. Human-factors safety — evidence that automation bias/alert fatigue are monitored and mitigated (lesson 6).
  6. Deployment staging — shadow deployment results before live use (lesson 7).
  7. Lifecycle — monitoring for drift, versioning/provenance, rollback (lesson 8).

Model answer (assemble your own first)

A trustworthy deployment case reads roughly: the model was frozen-externally tested with calibration; an MRMC reader study showed reader+AI improved volumetry consistency without automation-bias harm; the result integrates and renders in our PACS; the operating threshold was set and calibrated for our prevalence and triage role; it ran shadow for N weeks with acceptable FP load; and there is monitoring with drift alerts, versioning, and rollback. Each clause is a scoped claim with its own evidence; absent clauses are absent trust. Most real “deployed AI” papers provide the first clause and gloss the rest — which is exactly why Chapter 7’s appraisal lens applies to deployment claims as much as to development claims.

The transferable skill

This is the same skill as designing a study (Ch. 4, lesson 8) and appraising a paper (Ch. 7): enumerate the claims, demand the evidence for each, and keep the scope honest. Trust is not granted by a single number; it is assembled, layer by layer, and it has to be re-earned as the world drifts.

Stop and think — then reveal

A vendor offers a deployed segmentation assistant with strong standalone Dice and a frozen external test. They have no reader study, no shadow-deployment data, and no monitoring plan. Which trust layers are present and which are missing?

Present: standalone correctness (frozen external test) — layer 1, partially (no calibration mentioned). Missing: the reader+AI delta (layer 2 — no reader study, so we do not know if it helps readers), integration (layer 3 — no evidence it renders in our PACS), operating-point/calibration (layer 4), human-factors safety (layer 5), deployment staging (layer 6 — no shadow data), and lifecycle monitoring (layer 7 — no plan). So this is a single-layer trust case dressed as a deployment: you would require the reader study, integration evidence, threshold/calibration at your site, shadow deployment, and a monitoring plan before trusting it. One good layer is not a stack.

What to retain

  1. Trust is a stack of scoped claims (standalone, reader+AI, integration, operating point, human-factors, staging, lifecycle), each with its own evidence.
  2. Assemble the case layer by layer; absent layers are absent trust.
  3. Most “deployed AI” claims provide the standalone layer and gloss the rest — apply the Ch. 7 appraisal lens to deployment claims.
  4. Trust must be re-earned as the world drifts; monitoring is not optional.

You have finished Chapter 9 — and the course. The deployment-path diagram, AIRA/AIW-I links, agentic-workflow risks, and further reading are in the deployment reference. Return to the start page for where to go next.