Ask a conference panel what AI is doing in clinical governance and you will hear about predictive risk stratification, automated root cause analysis and real-time safety intelligence. Ask a governance lead what they actually use it for and the answer is usually: writing things up faster.
The gap between those two answers is the subject of this article.
What is genuinely deployed
Ambient documentation is the only category with real scale. A survey of 598 UK GPs published in npj Digital Medicine in May 2026 found 40% currently using AI scribes, with a further 23% having used them previously. Among users, they were used in anywhere from 5% to 100% of consultations, averaging 60%.
Adoption was higher among men, in private practice, among more experienced GPs, and among GP trainers.
The authors’ own framing of that number is the important bit. They describe adoption as “relatively high despite regulatory issues and recent official cease communication.”
Read that again. Two-fifths of GPs adopted a technology while national bodies were telling people to stop using it. That is the single most useful fact in this article, and it tells you something about how governance is actually working in 2026: deployment is running ahead of assurance, and the gap is being closed by frontline pragmatism rather than by policy.
Everything else is smaller and patchier. Evidence mapping, policy drafting, regulatory horizon-scanning, turning walkround notes into tracked actions, drafting incident summaries and board papers. These are real and they are spreading, but there is no equivalent adoption survey — largely because they are features inside compliance platforms rather than distinct products anyone counts.
What is genuinely rare: anything making or grading a decision without a person. Severity grading, regulatory interpretation, causal analysis. Vendors demo these. Very few organisations have signed them off.
What the evidence actually supports
Here is the uncomfortable structure of the evidence base: almost all measured accuracy data is about clinical documentation from live speech. There is essentially no measured evidence on AI applied to healthcare-administrative tasks — incident reports, complaints responses, board papers, regulatory returns, governance summaries.
That is exactly the work governance teams are using it for.
The nearest proxies bracket the problem usefully:
- Summarising a document you supplied is a low-single-digit-error task. Research in npj Digital Medicine hand-labelled 450 transcript–note pairs from UK primary care and found hallucinations in 1.47% of generated sentences, omissions in 3.45%.
- Producing content referencing things not in front of it is a roughly 20%-fabrication task. A JMIR Mental Health study found 19.9% of GPT-4o’s citations entirely fabricated.
Most administrative AI use sits somewhere between those two poles, and nobody has measured where. When you are deciding what to trust a tool with, the honest question is: is this task closer to summarising what I gave it, or closer to generating what it thinks I want?
Two details that change how you should review output:
Omission dominates. Omissions ran about 2.3 times more common than hallucinations. And omission is the error humans are worst at catching — reading for a wrong statement is straightforward; noticing that a fluent, professional summary is silently missing the thing the family said in the second meeting is much harder.
Design changes matter enormously and unpredictably. In the same study, one architectural change eliminated major omissions entirely; a template-driven variant increased major hallucinations. Changes intended to help can hurt, and you only find out by measuring. Which means a vendor who has never measured cannot tell you whether their last update made things worse.
Where the hype has outrun the evidence
The GOSH/TORTUS pilot deserves scrutiny. It is the most-cited NHS ambient AI evaluation: more than 17,000 patient encounters across nine London NHS sites, reporting a 23.5% increase in direct patient interaction time, an 8.2% reduction in appointment length, and economic modelling extrapolating to hundreds of millions of pounds of annual benefit.
Three problems. There is no peer-reviewed publication, no control arm, and — most importantly — no accuracy or safety outcomes at all. Every reported outcome is efficiency or satisfaction. The economic extrapolation rests on assuming one extra patient per shift per clinician generalises nationally.
And the study was NHS England-sponsored, with NHS England then citing it to support its own instruction to deploy at pace. That is not fraud, but it is not independent evidence either, and it is currently doing a lot of load-bearing work in national policy.
Time savings are real but much smaller than people think. A 2026 study in JMIR Medical Informatics measured documentation time objectively using Epic Signal data. Time in notes fell 21% — a genuine, statistically significant improvement. In absolute terms that was 13.6 minutes per day.
The same providers self-reported saving “hours per week.” The authors flagged the gap explicitly as a possible “error in judgment among providers.”
There was no significant change in burnout. And while clinicians rated their own attentiveness 17% higher, patients’ ratings of whether their provider listened did not move significantly.
If you are building a business case on self-reported time savings, expect to overstate the benefit by roughly an order of magnitude.
“91% of AI models degrade over time” is circulating widely. We could trace it only to an MLOps vendor blog. Don’t use it. What is defensible is that only around 9% of FDA-registered AI healthcare tools have a post-deployment surveillance plan — meaning drift would be undetectable regardless of its true rate. That is better sourced and more alarming.
The governance gap nobody is closing
The most important finding for anyone running an incident process is this.
Researchers reviewed 429 FDA medical device reports relating to AI-enabled devices, each assessed by a physician safety leader and a human factors expert. They found 34.5% contained insufficient information to determine whether AI had contributed at all. (npj Digital Medicine, 2024.)
The authors’ explanation is the crux: reporters “may not have insight on whether AI/ML are contributing to the safety issue they are observing given that these algorithms are at work ‘behind the scenes’.”
The implication for UK healthcare is direct. If your incident reporting has surfaced no AI-related events, that is not evidence AI is causing no harm in your organisation. It is at least as likely that your reporting system — designed long before any of this — cannot see it. Neither LFPSE nor your local system has a field for “the summary I relied on was missing something.”
If you take one action from this article, make it adding an explicit prompt to your incident process asking whether an AI tool was involved in producing, summarising or triaging any information the decision relied on. It costs nothing and it is the only way you will ever find out.
What good practice looks like right now
Know which task you’ve delegated. Grounded summarisation of your own documents: relatively safe, verify for omission. Anything requiring facts the model wasn’t shown: high risk, verify every reference.
Complete the paperwork properly. DCB0160 with a real safety case, hazard log and monitoring framework; a DPIA; and your supplier’s DCB0129. NHS England’s ambient scribing guidance sets this out clearly, and it was developed with MHRA and CQC input.
Audit by sampling, not just by review. NHS England’s guidance asks for “ongoing audits of clinical documentation” alongside per-case review — which is a quiet admission that per-case review alone isn’t trusted. It shouldn’t be. Pull a random sample monthly, check against source material, record the error rate, and watch it over time.
Check the medical device question honestly. NHS England’s position is that generative summarisation “would be treated as high functionality and likely would qualify as a medical device.” If a supplier is summarising clinical content and says otherwise, get the reasoning in writing.
Assume automation bias, because NHS England does. It is named as a hazard in the national guidance, and the evidence from radiology is stark: when an AI suggestion was wrong, experienced radiologists’ accuracy fell from 82.3% to 45.5%. Reviewers do not reliably catch confident errors. Design for that rather than hoping.
The summary, if you want one line for your board: the tools that are genuinely deployed are documentation tools, the measured benefit is real but modest, the assurance infrastructure is behind the deployment curve, and the reporting systems that would tell you if something went wrong probably cannot see it yet.
For the wider picture see our guide to AI in healthcare compliance. If you are procuring, start with the twelve questions to ask any vendor. And for the regulatory backdrop — including what is happening to the Single Assessment Framework — see PSIRF, SAF and the new bar for assurance.
