Clinical AI Has an Evidence Problem, and It Is Economic
Clinical AI adoption requires more than regulatory approval and accuracy. Without a standardized economic case and workflow integration, deployments will stall.
A persistent fallacy in clinical AI is that regulatory approval and high diagnostic accuracy are the finish lines. Teams celebrate when their model achieves parity with human specialists and secures an FDA clearance or CE mark. Yet, these models routinely stall before deployment. The friction is not clinical validity; it is the health-economic case. Accuracy and approval do not automatically produce adoption because hospital systems do not buy algorithmic performance, they buy outcomes that survive budgetary constraints.
Key takeaways
- The bottleneck has moved: Clinical AI is no longer bottlenecked by data science or regulatory clearance; the primary barrier to adoption is now health economics and workflow integration.
- Accuracy is just the baseline: High diagnostic accuracy gets a model into the conversation, but transparent economic evaluations are what secure hospital budgets.
- Usability is an economic variable: Poor interface design in complex workflows (like NGS) creates cognitive load and time penalties, which directly destroy the economic case for adoption.
- Transparency is now standardized: The CHEERS-AI standard forces vendors to detail the hidden costs of AI, from staff training and integration to the model’s learning curve over time.

The Shift from Data Science to Health Economics
For years, the development of clinical AI focused almost exclusively on improving diagnostic accuracy and securing regulatory approval. This created a generation of algorithms that perform exceptionally well in controlled environments but fail to scale in the real world. The bottleneck in clinical AI has shifted from data science to health economics. The Consolidated Health Economic Evaluation Reporting Standards for Interventions that use AI (CHEERS-AI) was introduced to ensure evaluations are reported transparently to support interpretation and comparison, as pilot exercises showed that important AI-specific details are frequently missing or inadequately reported.
Hospital procurement teams are increasingly skeptical of vendor claims regarding cost savings. When a vendor claims an AI intervention saves money, the underlying assumptions, such as the true cost of integration, the model’s learning curve, and the software’s maintenance overhead, are often opaque. CHEERS-AI forces a new standard of transparency, demanding that researchers detail the mechanism by which the intervention affects care pathways, costs, and health outcomes. Without this transparency, hospital systems cannot accurately assess the return on investment for deploying an AI tool, leading to stalled pilot programs and abandoned initiatives.
Why Transparency Fails in Practice
The CHEERS-AI framework is rigorous, demanding that developers provide detailed reporting on the model’s structure, the source data, and the sociotechnical context of deployment. However, the reality of clinical IT procurement means that these standards are often treated as compliance checkboxes rather than strategic planning tools. A hospital may demand to see a model’s cost-effectiveness analysis, but if that analysis assumes a frictionless integration into the electronic health record (EHR) system, an assumption that is almost universally false, the economic projections will collapse within the first month of pilot testing.
This disconnect between theoretical cost savings and practical implementation creates a credibility gap. Developers must recognize that the burden of proof is shifting. It is no longer sufficient to show that an AI tool could save money in a perfectly optimized workflow. Vendors must demonstrate that their tool will save money within the messy, fragmented, and legacy-bound reality of modern hospital IT infrastructure.

The Hidden Costs of Workflow Integration
The most significant of these hidden costs is rarely the software license itself; it is the sociotechnical cost of workflow integration. A model that requires a clinician to switch between multiple interfaces or manually input data introduces friction that can negate any theoretical efficiency gains. Research into AI-powered next-generation sequencing (NGS) workflows demonstrates that the transition of AI-driven tools into clinical practice is hindered by fragmented workflows, limited usability, and non-standardized data interfaces.
These integration challenges are particularly acute in specialized fields like genomics, where the volume of data is immense and the clinical decisions are highly consequential. When complex analytical pipelines fail to simplify into transparent, interactive systems, poor usability leads to underutilization, misinterpretation, or outright rejection. A model’s theoretical value is constrained by its usability. If a clinician cannot quickly and confidently interpret the AI’s output, they will revert to traditional methods, rendering the investment useless.
The NGS study underscores this by highlighting that the cognitive load on clinical staff is often ignored during the procurement process. When an AI tool forces a geneticist or pathologist to constantly switch context, re-enter data, or parse opaque algorithmic confidence scores, the time penalty accrues rapidly. If an AI tool saves ten minutes of diagnostic time but requires fifteen minutes of manual data wrangling to operate, the health-economic case is destroyed.
The researchers mapped where AI actually touches a sequencing workflow, and the answer is scattered across four roles and five stages rather than concentrated in the model.

Exhibit 1. Data requirements for AI in a next-generation sequencing workflow, mapped across four clinical roles and five workflow stages. Source: Plass, M., Holzinger, A., Reihs, R., & Müller, H., 2026, Integrating AI into clinical practice: Human-centered design requirements for next-generation sequencing workflows, Computerized Medical Imaging and Graphics, 130, 102758, p. 6.

Usability as a Core Economic Variable
As we have seen in the transformation of healthcare services with AI ecosystems, integrating these tools requires treating usability as a core economic variable. Poor human-computer interaction is not just an interface problem; it is a systemic workflow failure that incurs measurable financial penalties. Every extra click, every confusing visualization, and every disjointed workflow step adds to the cognitive load of the clinician and the operational cost of the hospital.
Designing for this reality means recognizing that requirements engineering for probabilistic software is entirely different from traditional IT procurement. Standard software provides deterministic outputs based on fixed rules; AI models provide probabilistic predictions that require human interpretation and oversight. The interface must be designed to support this nuanced interaction, providing clear explanations of the model’s confidence and reasoning.
Crossing the Chasm to Deployment
To cross the chasm from an approved algorithm to a deployed standard of care, clinical AI must prove its value through rigorous, standardized economic evaluation. The CHEERS-AI framework provides a roadmap for this evaluation, highlighting the need to account for AI-specific costs and operational impacts. Developers must move beyond demonstrating clinical equivalence and start demonstrating economic superiority in real-world settings.
Accuracy gets a model into the conversation. Workflow integration and hard economic evidence are what keep it in the hospital. The future of clinical AI will be defined not by the models with the highest area under the curve, but by the models that seamlessly integrate into clinical workflows and deliver clear, quantifiable economic value.
References
- Elvidge, J., Hawksworth, C., Avşar, T. S., Zemplenyi, A., Chalkidou, A., Petrou, S…. & Dawoud, D. (2024). Consolidated Health Economic Evaluation Reporting Standards for Interventions That Use Artificial Intelligence (CHEERS-AI). Value in Health, 27(9), 1196-1205. https://doi.org/10.1016/j.jval.2024.05.006
- Plass, M., Holzinger, A., Reihs, R., & Müller, H. (2026). Integrating AI into clinical practice: Human-centered design requirements for next-generation sequencing workflows. Computerized Medical Imaging and Graphics, 130, 102758. https://doi.org/10.1016/j.compmedimag.2026.102758
Frequently asked questions
Why is clinical accuracy not enough for AI adoption?
Hospital purchasing decisions require a clear health-economic case that demonstrates measurable value over the standard of care. Clinical accuracy proves a model works in a controlled setting, but it does not prove it is cost-effective to deploy at scale.
What is the CHEERS-AI standard?
CHEERS-AI is a reporting guideline extension designed to ensure that economic evaluations of AI-based health interventions are transparent and comparable. It requires authors to explicitly detail AI-specific costs, training data constraints, and integration requirements.
Where do the hidden costs of clinical AI typically emerge?
The hidden costs of clinical AI emerge during workflow integration and usability friction. Poor interface design, fragmented data pipelines, and a lack of role-specific decision support can slow clinicians down, erasing the efficiency gains the AI promised.
How does usability impact the economic value of an AI tool?
Usability directly dictates how efficiently clinical staff can operate a tool, which is a core economic variable. If an AI interface increases cognitive load or requires workflow workarounds, it incurs a time penalty that undermines its return on investment.
Evidence
The Consolidated Health Economic Evaluation Reporting Standards for Interventions that use AI (CHEERS-AI) was introduced to ensure evaluations are reported transparently to support interpretation and comparison, as pilot exercises showed that important AI-specific details are frequently missing or inadequately reported.
The CHEERS-AI checklist will ensure that important details relating to the AI nature of the intervention Introduction checklist, authors and implications for the analysis are reported in a transparent and should provide at least reproducible way.
Elvidge, J., Hawksworth, C., Avşar, T. S., Zemplenyi, A., Chalkidou, A., Petrou, S., ... & Dawoud, D. (2024). Consolidated Health Economic Evaluation Reporting Standards for Interventions That Use Artificial Intelligence (CHEERS-AI). Value in Health, 27(9), 1196-1205. (p. 1)
CHEERS-AI forces a new standard of transparency, demanding that researchers detail the mechanism by which the intervention affects care pathways, costs, and health outcomes.
and economic effects on resource use or system efficiency. Selection of outcomes 11 An outcome measure is used to quantify the extent to which an intervention has an effect. For example, a study may measure effectiveness (benefits and harms) in terms of changes in health outcomes, diagnostic outcomes (such as improved accuracy), process outcomes (eg, faster decision making), or several outcomes simultaneously. Measurement of 12 Assumptions regarding the effect of the AI intervention, such as the use of arbitrary or exploratory input outcomes values, should be transparently reported. Their theoretical or scientific basis should be explained. Measurement and 14 The purchase cost of an intervention with an AI component may include a purchase price and other valuation of resources and components, such as the developer implementing the AI into practice or maintaining it over time. There may costs be other relevant costs to the healthcare system relating to implementation of an AI intervention. Rationale and description 16 A model may be used in a health economic evaluation to estimate the cost-effectiveness of an intervention. of model Explain if a particular model structure or programming approach, such as individual patient simulation, has been chosen to characterize the AI intervention appropriately. Discussion Study findings, limitations, 26 There may be ethical and equity issues associated with AI that are relevant for decision makers alongside the generalizability, and cost-effectiveness results. Biases may include, for example, the AI intervention being developed using a current knowledge training data set that is not representative of the population of interest. AI indicates artificial intelligence; IMDRF, International Medical Device Regulators Forum; NICE, National Institute for Health and Care Excellence; SaMD, Software as a Medical Device. Reporting Trials–Artificial Intelligence [CONSORT-AI]9). Those specific nuances and implications are reported or cited in a checklists should enhance the reporting of the development and transparent and reproducible way. assessment of AI health interventions, and now we have devel- The importance of an AI extension to CHEERS 2022 was oped the CHEERS-AI reporting guideline extension to do the same highlighted during the development process, including through for EEs. Our reporting standards were developed using a very qualitative responses during the Delphi study. Respondents noted similar methodological approach to CHEERS 2022, including a as follows: first,
Elvidge, J., Hawksworth, C., Avşar, T. S., Zemplenyi, A., Chalkidou, A., Petrou, S., ... & Dawoud, D. (2024). Consolidated Health Economic Evaluation Reporting Standards for Interventions That Use Artificial Intelligence (CHEERS-AI). Value in Health, 27(9), 1196-1205. (p. 7)
Research into AI-powered next-generation sequencing (NGS) workflows demonstrates that the transition of AI-driven tools into clinical practice is hindered by fragmented workflows, limited usability, and non-standardized data interfaces.
While NGS enables Clinical genomics rapid and scalable analysis of complex genetic information—paving the way for precision diagnostics and AI usability engineering stratified treatment—the transition from potential to practice is hindered by fragmented workflows, limited Stakeholder workflows FHIR usability, and non-standardized data interfaces.
Plass, M., Holzinger, A., Reihs, R., & Müller, H. (2026). Integrating AI into clinical practice: Human-centered design requirements for next-generation sequencing workflows. Computerized Medical Imaging and Graphics, 130, 102758. (p. 1)
When complex analytical pipelines fail to simplify into transparent, interactive systems, poor usability leads to underutilization, misinterpretation, or outright rejection.
Poor usability can lead to underutilization, misinterpre tools often lack standardization, and their interfaces are seldom tation, or outright rejection of otherwise technically sound systems.
Plass, M., Holzinger, A., Reihs, R., & Müller, H. (2026). Integrating AI into clinical practice: Human-centered design requirements for next-generation sequencing workflows. Computerized Medical Imaging and Graphics, 130, 102758. (p. 2)