Chapter 9 — Clinical deployment, human-AI interaction, and lifecycle
This chapter · ~3 h total across 9 lessons · From a validated model to safe, monitored clinical use · Energy: low to medium · Setup: none
A model with great test-set numbers is not a deployed product. This chapter is the mental model for what happens between a validated model and safe clinical use — and for what must keep happening after deployment. It sits at the intersection of software, medicine, AI, and imaging informatics (Ch. 5). It is not a regulatory-compliance chapter; it is the clinical/technical model.
The thread this chapter pulls
Every previous chapter asked “is the model right?”. This one asks “does the model help, safely, in real practice, over time?” Those are different experiments. A model’s standalone AUROC does not establish the benefit of adding it to a radiologist; a correct result can be clinically invisible; a deployed operating point depends on prevalence and workflow; the human can become part of the failure mode; and the model starts ageing the day it is deployed. This chapter walks that gap and the lifecycle beyond it.
Lessons
- A good test-set model is not yet a clinical tool — bridging development to use.
- What role is the AI actually playing? — triage, concurrent, second reader, autonomous, measurement support.
- AI-alone and radiologist+AI are different experiments — reader studies, MRMC, crossover.
- How to read a reader-study result — Herpe et al.
- Why thresholds change when the workflow changes — prevalence, FP/FN cost, calibration.
- The human can become part of the failure mode — automation bias, alert fatigue, cognitive offloading.
- Shadow deployment before clinical deployment — external test, silent deployment, prospective evaluation.
- The model starts ageing the day it is deployed — drift, monitoring, versioning, rollback, provenance.
- What evidence would make you trust this system? — a synthesis exercise.
Dense reference — the deployment-path diagram, the AIRA/AIW-I links, the agentic- workflow risks, and the further reading — lives in the deployment reference.
How this chapter connects
- It assumes Chapter 4 (validation) and Chapter 5 (how results move and get retrieved).
- It applies Chapter 7’s appraisal lens to deployment claims.
If you only have an hour, lessons 3 (alone vs +AI), 6 (human failure modes) and 8 (drift/monitoring) are the core that “deployed AI” claims most often gloss over.