was successfully added to your cart.

Radiology AI Adoption: Why Pilots Stall

Avatar photo
Chief executive officer at John Snow Labs

Why radiology AI adoption stalls – and what health systems that scaled it did differently

By 2055, US imaging demand will rise 16.9%–26.9% above 2023 levels, while the radiologist workforce grows only 25.7% if residency positions stay flat, according to 2025 companion studies from the Harvey L. Neiman Health Policy Institute (Christensen et al., 2025). The FDA has cleared 723 AI-enabled radiology devices – more than any other specialty – yet fewer than 30% underwent clinical testing before market clearance (Sivakumar et al., 2025). Radiology does not lack AI tools. It lacks health systems that get one past the pilot. 

The pattern is consistent across the published evidence: a pilot runs cleanly on curated data, then stalls the moment a health system tries to run the same tool continuously on live studies. The barriers are rarely the detection algorithm itself. They typically sit in four places: the infrastructure that feeds the model data, the governance that clears it for use, the workflow that surfaces its output, and the financial case that funds it before any of that value shows up on a dashboard. 

Why a clean pilot rarely survives contact with production data

A radiology AI pilot typically runs on images selected to represent the use case, de-identified for that specific project, and formatted to the tool’s input specification – pulled from PACS without the volume or heterogeneity of daily practice. That demonstrates the algorithm works on data prepared for it. It says nothing about whether the health system can prepare data of that quality continuously, at the volume the department actually generates. 

The gap shows up in the literature. A 2025 review of 173 CE-certified radiological AI products found that peer-reviewed evidence grew from 36% of products in 2020 to 66% in 2023, but the added evidence remained concentrated at the diagnostic-accuracy level (Antonissen et al., 2025). A separate scoping review of 140 studies published between 2020 and 2025 identified high technical demand, unclear guidance, and limited expert engagement as the recurring adoption barriers – not detection performance (Lawrence et al., 2025). Tools are increasingly validated for accuracy. What stays thin is evidence that they hold up once deployed into a live operational environment. 

Production imaging infrastructure has to do five things at once: ingest DICOM at departmental volume; de-identify to the applicable regulatory standard across every metadata configuration the department’s scanners produce; normalize metadata across vendors and acquisition protocols; flag and route studies that fail quality control before they reach the model; and surface the model’s output inside the PACS and EHR systems radiologists already use. None of these is individually hard to build. Building them as one maintained system, instead of five one-off scripts solving immediate problems, is the part pilots skip. 

The clearest public illustration of that governance bar does not come from imaging – it comes from clinical text, where regulatory-grade de-identification has been validated at the largest scale published to date. Providence St. Joseph Health co-authored the largest de-identification validation published to date: 2 billion patient notes, 99% PHI obfuscation, HIPAA Expert Determination criteria met both in aggregate and per record, and an independent adversarial audit in which a red team working for three months could not re-identify any of 790 randomly selected patients (Kocaman et al., 2025). Those results are measured on the production data the system actually processes, not a curated holdout set. That is the same standard imaging pipelines have to meet: performance demonstrated on the full heterogeneity of what a department generates daily, not on the subset assembled for a demo. John Snow Labs’ de-identification technology, built on the approach validated in that study, extends that same validated standard across text, PDF, FHIR, and DICOM data rather than treating imaging de-identification as a separate problem to solve later. 

Why compliance uncertainty stalls systems that already have a working tool

Radiology operates in a liability environment that makes AI adoption risk-averse by design. When an AI system contributes to a clinical decision, someone has to answer, before deployment, who is accountable if that contribution leads to a poor outcome. Institutions that never answer the question in advance find out how unresolved it was only once an adverse event makes it urgent. 

The regulatory record shows how uneven the underlying evidence base still is. Of the 723 FDA-cleared radiology AI devices identified in a 2025 systematic review, fewer than 30% had undergone clinical testing, and a smaller share still had been evaluated prospectively or with a human-in-the-loop protocol (Sivakumar et al.,  2025). A separate 2025 review of the ethical, legal, and regulatory frameworks governing radiology AI concluded that oversight structures have not kept pace with the rate of device clearance (Aldhafeeri, 2025). Clearance is not the same as institutional readiness. The compliance, audit, and monitoring infrastructure a hospital’s risk management team requires is a separate build from the one the FDA requires. 

That build has to include documentation that the system was validated for its intended use, deployed with defined human oversight, monitored after deployment, and covered by a process for investigating adverse events. A 2024 review in Mayo Clinic Proceedings: Digital Health identified postproduction monitoring, not initial validation, as the persistent operational gap: institutions that lack a performance dashboard and an in-PACS correction workflow struggle to sustain both AI accuracy and radiologist adoption over time, regardless of how the model performed at go-live (Mayo Clin Proc Digit Health 2024). 

Dandelion Health’s approach to this problem, in the clinical-text domain, illustrates the standard institutional compliance teams are converging on: assess each data type by re-identification risk, adapt the de-identification process to that risk level, and validate continuously rather than once (Dandelion Health). Continuous QA is a governance instrument as much as a technical one: it is the evidence a compliance team can point to when a regulator or auditor asks whether the system’s performance has held up in production. Imaging governance needs the same kind of standing evidence, not a validation report filed once at go-live. 

Why a well-performing tool is not the same as a workflow that gets used

AI radiology tools are usually evaluated on task performance – detection accuracy, segmentation quality, report fluency – measured outside the clinical workflow. Deployment success depends on something different: whether the output reaches the radiologist in the format and timeframe that makes it useful, whether it adds steps to the PACS and RIS workflow or removes them, and whether a radiologist can correct the output without leaving their primary reading environment. 

A 2024 review in Radiology on standards-based AI integration found that variable adoption of interoperability standards across vendors leaves institutions managing AI results through custom, one-off integrations rather than a shared framework, which increases both the operational burden and the chance something breaks silently (Tejani et al., 2024). The Mayo Clinic Proceedings review made a related, narrower point: PACS viewers often lack a proper tool for radiologists to correct an AI-generated segmentation in place, so radiologists work around the AI output instead of through it. 

Alert fatigue compounds the problem. A 2026 narrative review spanning nine stages of the radiology workflow found that false-positive prioritization alerts measurably reduce radiologists’ responsiveness to genuine escalation signals over time (El Ray et al., 2026) – a tool that is right too rarely gets ignored even in the cases where it is right. The most common failure pattern documented across this literature is consistent: AI output presented as a separate interface radiologists must toggle to, rather than an integrated element of the environment they already work in. The productivity case that justified acquiring the tool disappears in the friction of using it. 

Intermountain Health’s approach to this problem, again in the clinical-text domain, shows what integration looks like when it works. Its Databricks Lakehouse infrastructure processes hundreds of millions of clinical documents and cut review time from roughly 10 minutes to 3 per document – a 70% efficiency gain – by surfacing structured output inside the tools clinicians already use, not a separate AI interface. The underlying principle transfers directly to imaging: the output has to arrive inside the PACS and RIS workflow, not alongside it. 

Why infrastructure has a harder ROI case than the tools it enables

Tool-level ROI is straightforward to model once a system is live in production. The harder – and more consequential – case is the infrastructure investment that has to happen first.

 

The financial case for a single AI tool is straightforward to build: it performs a task faster or more accurately, the improvement is measurable, and the savings can be estimated. A 2024 ROI calculator published in the Journal of the American College of Radiology modeled exactly this and found a 451% five-year return from an AI diagnostic imaging platform, rising to 791% once radiologist time savings were included (Bharadwaj et al., 2024). 

The infrastructure underneath that tool is a harder case to make, because its value is enabling rather than directly operational: the data pipelines, de-identification systems, and PACS integration do not move throughput on their own; they are what allows a tool that moves throughput to run in production at all. A 2025 systematic review of the economic value of AI in radiology, covering 21 studies with quantified economic outcomes, found that financial return is highly context-dependent – AI created value in high-volume, resource-intensive tasks under fixed-cost deployment, and in several cases increased cost when priced per use or when accuracy fell below the human baseline it was meant to improve on (Molwitz et al., 2025). Deployment model, in other words, drove ROI as much as the underlying algorithm’s accuracy did. 

That is the argument for building infrastructure before acquiring point tools, not after: a fixed-cost data and governance layer amortizes across every tool a health system adds to it, while a per-tool infrastructure build repeats the same compliance and integration cost with each new acquisition. Health systems that buy tools before the infrastructure exists to run them in production typically discover the gap only once the tool cannot go live – at which point the ROI case for the missing infrastructure is harder to make, because the money for the tool is already spent. 

The sequencing that separates health systems that scale radiology AI from those stuck at the pilot stage is consistent across the evidence: build the data infrastructure, compliance framework, and workflow integration first; deploy tools into that environment second; measure the operational result against a baseline set before the infrastructure existed. That sequencing means showing an infrastructure investment on the balance sheet before showing an AI productivity gain on the operational dashboard, and it takes institutional patience that a per-tool procurement process does not require. With a radiologist workforce that will not outgrow imaging demand by more than a few points through 2055, the health systems willing to make that sequencing decision now are the ones positioned to keep pace with volume without falling further behind on cost per interpretation. 

John Snow Labs’ de-identification technology applies this same validated approach – production-scale accuracy, continuous QA, on-premises deployment – across text, PDF, FHIR, and DICOM data, so imaging infrastructure does not have to be built as a separate project from the rest of a health system’s clinical data pipeline. 

Frequently asked questions

Why do radiology AI pilots so often fail to scale, even when the technology clearly works in the proof-of-concept phase? 

Because a pilot is designed to show what a tool can do under controlled conditions, not whether the health system can sustain those conditions at production volume. A curated, pre-formatted dataset tells you the model works; it does not tell you whether the organization can ingest, de-identify, normalize, and quality-control imaging data continuously across the full heterogeneity of a live radiology department. Published reviews consistently point to infrastructure, governance, and workflow integration – not detection accuracy – as the barriers that stall scale-up (Lawrence et al., 2025). The benchmark for that infrastructure is set by systems validated on production data at scale, not on holdout sets assembled for a demo. 

What compliance and governance requirements does radiology AI deployment actually need to satisfy before go-live? 

Documentation that the system was validated for its intended use, deployed with defined human oversight, monitored post-deployment, and covered by a process for investigating adverse events. Fewer than 30% of the 723 FDA-cleared radiology AI devices reviewed in a 2025 study had undergone clinical testing (Sivakumar et al., 2025), which is why institutional risk management teams typically require evidence beyond the FDA clearance itself before approving production deployment. 

What does regulatory-grade de-identification look like at production scale? 

Validation measured on the full production corpus rather than a curated sample, and repeated continuously rather than once at go-live. The largest published example comes from clinical text: a system de-identified 2 billion patient notes with 99% PHI obfuscation, satisfied HIPAA Expert Determination criteria both in aggregate and per record, and survived an independent adversarial audit in which a red team working for three months re-identified none of 790 randomly selected patients (Kocaman et al., 2025). Imaging pipelines have to clear the same bar across every scanner configuration and metadata variant a department produces, which is why de-identification is an infrastructure commitment rather than a per-project step. 

Why is workflow integration harder than acquiring a well-performing AI tool? 

Because tool performance is measured on isolated tasks, while deployment success depends on how the tool fits into the PACS, RIS, and EHR environment where radiology actually happens. Variable adoption of interoperability standards across vendors pushes institutions toward custom, one-off integrations (Tejani et al., 2024), and false-positive alerts measurably reduce radiologists’ responsiveness to real ones over time (El Ray et al., 2026). Radiologists stop using tools that perform well technically when the workflow the integration creates is worse than the one it replaced. 

Why is the ROI case for AI infrastructure harder to make than the ROI case for AI tools? 

Because infrastructure value is enabling rather than directly operational. A single tool’s financial benefit is traceable – one 2024 model found a 451% five-year ROI, rising to 791% with radiologist time savings included (Bharadwaj et al., 2024). But that calculation assumes the infrastructure to run the tool in production already exists. A 2025 review found that fixed-cost, locally deployed models delivered more consistent economic value than pay-per-use pricing across the reviewed studies (Molwitz et al., 2025), which is an argument for building the infrastructure layer once rather than re-justifying it with every new tool. 

What investment sequencing distinguishes health systems that successfully scale radiology AI from those that don’t? 

The successful ones build data infrastructure, compliance frameworks, and workflow integration capability first, then deploy AI tools into that environment, then measure operational benefit against a pre-infrastructure baseline. Organizations that acquire tools first consistently discover the infrastructure gap only once the tool cannot run in production – at which point the case for the missing infrastructure investment is harder to make, not easier. 

How large is the gap between radiologist supply and imaging demand, and why does it raise the stakes on solving these adoption barriers? 

US imaging demand is projected to rise 16.9% to 26.9% by 2055 relative to 2023, while the radiologist workforce is projected to grow 25.7% at current residency levels – a gap that Harvey L. Neiman Health Policy Institute researchers project will persist for decades absent deliberate action (Christensen et al., 2025). That gap is precisely the reason a stalled pilot is a costly outcome, not a neutral one: every year an AI deployment spends stuck at proof-of-concept is a year the underlying workforce shortage goes unaddressed.  

References

Christensen EW, Parikh JR, Drake AR, et al. “Projected US Radiologist Supply, 2025 to 2055.” Journal of the American College of Radiology, 2025;22(2). 

Christensen EW, et al. “Projected US Imaging Utilization, 2025 to 2055.” Journal of the American College of Radiology, 2025;22(2). 

Sivakumar R, Lue B, Kundu S. “FDA Approval of Artificial Intelligence and Machine Learning Devices in Radiology: A Systematic Review.” JAMA Network Open, 2025;8(11):e2542338. 

Aldhafeeri FM. “Governing Artificial Intelligence in Radiology: A Systematic Review of Ethical, Legal, and Regulatory Frameworks.” Diagnostics, 2025;15(18):2300. 

“Implementing Artificial Intelligence Algorithms in the Radiology Workflow: Challenges and Considerations.” Mayo Clinic Proceedings: Digital Health, 2024. 

Tejani AS, Cook TS, Hafezi-Nejad N, Sardana T, O’Donnell KP, et al. “Integrating and Adopting AI in the Radiology Workflow: A Primer for Standards and Integrating the Healthcare Enterprise (IHE) Profiles.” Radiology, 2024;311(3):e232653. 

“Artificial Intelligence Across the Radiology Workflow: A Nine-Stage Narrative Review.” Diagnostics, 2026. 

Bharadwaj P, Nicola L, Blankenburg M, et al. “Unlocking the Value: Quantifying the Return on Investment of Hospital Artificial Intelligence.” Journal of the American College of Radiology, 2024;21(10):1677-1685. 

Molwitz I, et al. “Economic Value of AI in Radiology: A Systematic Review.” Radiology: Artificial Intelligence, 2025;8(1). 

Antonissen N, Tryfonos O, Houben IB, Jacobs C, de Rooij M, van Leeuwen KG. “Artificial intelligence in radiology: 173 commercially available products and their scientific evidence.” European Radiology, 2025. 

Lawrence R, Dodsworth E, Massou E, et al. “Artificial intelligence for diagnostics in radiology practice: a rapid systematic scoping review.” eClinicalMedicine, 2025;83:103228. 

Kocaman V, Mico L, Kaya MA, Taiyab N, et al. “Automated De-Identification, Consistent Obfuscation, and Regulatory Grade Validation of 2 Billion Patient Notes.” Research Square, 2025 (preprint). doi:10.21203/rs.3.rs-6867162/v1. 

How useful was this post?

Healthcare LLM

Learn more
Avatar photo
Chief executive officer at John Snow Labs
Our additional expert:
David Talby is a chief executive officer at John Snow Labs, helping healthcare & life science companies put AI to good use. David is the creator of Spark NLP – the world’s most widely used natural language processing library in the enterprise. He has extensive experience building and running web-scale software platforms and teams – in startups, for Microsoft’s Bing in the US and Europe, and to scale Amazon’s financial systems in Seattle and the UK. David holds a PhD in computer science and master’s degrees in both computer science and business administration.

Reliable and verified information compiled by our editorial and professional team. John Snow Labs' Editorial Policy.

The AI-Ready Hospital: What Healthcare AI Implementation Requires

The AI-Ready Hospital: Architecture, Culture, Workflows, and Staffing for the Next Decade An AI-ready hospital is a health system whose data infrastructure,...
preloader