Last fall, I sat in a governance review where a health system was deciding whether to renew a clinical decision support contract. The model had performed well in its validation study. It had cleared the FDA. The procurement team loved the dashboard. But when the chief medical officer pulled up the override log, more than 40% of the tool’s recommendations had been manually reversed by physicians over the prior six months. Nobody had flagged it. The system was technically live and practically ignored.
That meeting changed how I think about healthcare AI. We spent two years persuading physicians to try these tools, and most of them have. Doximity’s 2026 State of AI in Medicine report puts daily usage at 63% among US physicians, up from 47% a year ago. The adoption debate is settled. What replaced it is whether the tools now running inside our health systems actually hold up under the conditions where patients depend on them.
I used to think the answer was better models. I don’t anymore. The answer is better accountability for what models do after deployment.
The gap between the demo and the ward
Controlled evaluations and live clinical environments are not the same thing. A Microsoft Research study found that top-performing clinical AI models still produce confident outputs when key inputs are withheld, reverse their positions under minor prompt variations and generate clinical reasoning that reads well but doesn’t track. These are not obscure research prototypes. They are the systems topping the benchmarks that health systems use to make purchasing decisions.
The Stanford-Harvard ARISE network’s NOHARM evaluation tested 31 large language models across 100 real primary care cases annotated by 29 specialist physicians. Even the best-performing models produced severe clinical errors in 12 to 15 out of every 100 cases. Across all models tested, the severe harm rate reached 22%. That would not survive a single morbidity and mortality conference if a physician were responsible.
We have seen this play out with deployed systems, not just benchmarks. Researchers externally validated Epic’s sepsis prediction model and found an area under the curve of 0.63, barely better than chance, with sensitivity of just 33%. This was a tool already running in hundreds of US hospitals. It was eventually deactivated at Michigan Medicine when a dataset shift during the COVID-19 pandemic caused it to generate spurious alerts. In my own experience evaluating AI tools across radiology, pathology and clinical documentation, I have watched systems that score above 90% on curated test sets struggle with edge cases that any second-year resident would catch, because the training data never included them.
The industry keeps treating this as a model architecture problem. It is a data coverage problem and a validation design problem, and until we take it as seriously as we take accuracy on leaderboards, the gap between the demo and the ward will keep widening.
Cleared is not the same as ready
The FDA has authorized more than 1,000 AI-enabled medical devices. Each clearance is a point-in-time assessment: The system met a performance threshold under defined conditions on a specific date. It says nothing about what happens when patient demographics shift, clinical workflows evolve or the model sits untouched for 18 months while the standard of care moves on.
Too many organizations treat regulatory clearance as a finish line. It is a starting gun. The AlignInsight framework for post-deployment monitoring is one of the more practical approaches I have seen for closing that gap, because it treats monitoring as a continuous obligation rather than an annual checkbox. But frameworks only work if someone owns them operationally, and in most health systems I work with, that ownership is still undefined.
Fairness as a baseline, not an upgrade
A system that performs well on average but poorly for specific populations accelerates existing inequity and makes the source harder to trace. A 2019 study found that a widely used algorithm systematically underestimated the health needs of Black patients because it used cost as a proxy for illness. The fix required retraining on different labels, not a better architecture. That finding should have permanently changed how every health system evaluates AI procurement. For many, it has not.
Stratified validation across race, age, sex and socioeconomic status before deployment should be table stakes for any tool that touches clinical decisions. Most procurement processes still do not require it.
Someone needs to own this
When AI contributes to a diagnostic error today, accountability is genuinely unclear. The model vendor, the deploying institution, the clinician who acted on the output: The contracts and bylaws most health systems operate under were not written for this scenario. A randomized trial found that physicians receiving erroneous AI recommendations saw their diagnostic accuracy drop by 18 percentage points, even after completing 20 hours of AI literacy training. Automation bias is real, measurable and resistant to education alone. That ambiguity in who is responsible will become expensive for someone soon.
The organizations getting ahead of it are treating AI governance the way they treat credentialing and informed consent: as a non-negotiable operating condition. Every AI-assisted decision that touches a patient should leave an auditable trail — what was observed, what was recommended, what the clinician chose, what the outcome was. Without that trail, you cannot improve the system, you cannot explain a failure and you cannot defend a decision. The institutions building that infrastructure now are accumulating institutional knowledge that will compound. The ones waiting are accumulating risk.
The work that matters now
We got AI through the door. The harder work is determining what it is actually doing now that it is inside. The good news is that most of what is needed already exists: better real-world evaluation, continuous monitoring, clear ownership, data-layer accountability. It is governance, not invention. The only question is whether we build it before the first high-profile failure forces the issue, or after.
Healthcare has rarely been rewarded for waiting.
Opinions expressed by SmartBrief contributors are their own.
_______________
Subscribe to AI Impact: Healthcare powered by SmartBrief and SmartBrief for Healthcare Leaders, one of our more than 30 healthcare publications.
