Most AI procurement checklists are written from first principles by people who have never seen one of these systems fail. This one is written backwards, from the failure modes that have actually been documented in the literature — which means every question below exists because something went wrong somewhere and somebody published it.
We sell software in this category. Read the questions with that in mind, and note that several of them are ones we would rather you didn’t ask us either.
First, the parable
In 2021, researchers at Michigan Medicine ran an external validation of the Epic Sepsis Model — a proprietary sepsis prediction tool deployed at hundreds of US hospitals. They studied 27,697 patients across 38,455 hospitalisations.
Epic’s internally reported performance was an AUC of 0.76 to 0.83. The independent result was 0.63.
At the threshold actually in clinical use, sensitivity was 33% and positive predictive value 12%. The model alerted on 18% of all hospitalisations, missed 67% of sepsis patients, and identified only 7% of cases that clinicians hadn’t already caught. Depending on how you counted the alerts, staff were evaluating between 8 and 109 patients for every true case.
The authors’ conclusion is worth reading twice: “the widespread adoption of the ESM despite its poor performance raises fundamental concerns about sepsis management on a national level.”
Nothing about that story is American. Every structural feature reproduces in an NHS procurement: a proprietary model, performance figures supplied by the vendor, no independent validation until academics did it unilaterally, a threshold chosen by a local committee, and an alert burden nobody modelled in advance.
The questions below are designed so you don’t buy that.
The twelve questions
1. What is your measured hallucination rate and your measured omission rate — separately?
Not “our accuracy is 95%.” Two numbers, because they are different failure modes with different detectability.
Research published in npj Digital Medicine hand-labelled 450 transcript–note pairs from UK primary care consultations. Hallucinated content appeared in 1.47% of generated sentences. Omitted source content ran at 3.45% — roughly 2.3 times more common.
Omission is the one that matters more, because a reviewer reading a fluent, complete-looking summary has no cue that something is missing. If a vendor only quotes you one number, it will be the flattering one.
2. What proportion of errors were classified as major, and on what severity framework?
In the same study, 44% of hallucinations were rated major — capable of affecting diagnosis or management. So were 16.7% of omissions. A low overall error rate with a high major-error share is worse than the reverse, and averages hide it.
3. Who adjudicated the accuracy — and was it the people using the product?
This one matters more than it sounds. A 2026 study in JMIR Medical Informatics had providers self-audit AI-generated notes; they rated 86.6% accurate. The same study measured their time savings objectively at 13.6 minutes per day, while those providers self-reported saving “hours per week.”
The authors flagged the gap as a possible “error in judgment among providers.” If clinicians overestimate their time savings by an order of magnitude, their self-certified accuracy figures deserve the same scepticism. Ask for independent adjudication against ground truth, or treat the number as marketing.
4. What is the error rate in multi-party, complex, accented, non-English and speech-impaired consultations?
Not the average — this specific breakdown.
A study presented at ACM FAccT found around 1% of Whisper transcriptions contained entirely hallucinated phrases, with 38% of those including explicit harms. Critically, hallucinations concentrated on speakers with aphasia — 1.7% of segments versus 1.2% for controls.
That is an equality-of-access problem, not just a safety one. Your patients with communication difficulties are the ones the technology serves worst, and they are also the ones least able to spot and correct a wrong note.
5. Is transcription error reported separately from summarisation error, and how much propagates?
Ambient tools chain speech recognition into a language model. A perfectly faithful summariser summarising a hallucinated transcript produces a confident, internally consistent, entirely wrong note. If your vendor reports one blended figure, they are hiding where the error enters.
6. Has your performance been externally validated by anyone who does not sell the product?
See Epic. If the answer is no, you are buying self-reported performance, and the honest thing is to price that risk rather than pretend it isn’t there.
7. At the threshold we will actually deploy at, what is the positive predictive value and the review burden per true positive?
Vendors quote performance at the threshold that flatters the model. You will deploy at a threshold your governance committee picks. Those are rarely the same number, and the alert burden between them can differ by an order of magnitude.
8. Is this a medical device — and if you say no, on what basis?
This is the question most likely to be fudged, and the most consequential.
NHS England’s information governance guidance on ambient scribing is explicit: products that only produce easily-verified transcriptions are “likely not medical devices”, but “the use of Generative AI for further processing… summarisation, would be treated as high functionality and likely would qualify as a medical device.”
If your vendor is summarising clinical content and telling you they’re not a medical device, ask them to put the reasoning in writing. Then ask for the MHRA registration and the UKCA or CE certificate.
Note also that MHRA Class I registration is self-declared, not independently assessed. It is a statement by the manufacturer, not a verdict on the product.
9. Provide your DCB0129.
DCB0129 is the clinical risk management standard for health IT manufacturers. Its counterpart, DCB0160, is yours — you complete the safety case, hazard log and monitoring framework for the deployment. Along with a DPIA, this is the documented baseline NHS England expects.
A vendor who cannot produce a DCB0129 has not done the work, and you will be completing your DCB0160 on sand.
10. What is your post-deployment monitoring plan — metrics, frequency, and what triggers rollback?
Work presented at NeurIPS 2025 reported that only 9% of FDA-registered AI-based healthcare tools include a post-deployment surveillance plan. (It’s a position paper citing a secondary review, so attribute it as reported rather than established — but the direction is not seriously contested.)
Model drift is well understood as a mechanism. Its real-world magnitude in deployed clinical AI is poorly quantified — largely because almost nobody is monitoring, which means drift would be undetectable regardless of how bad it is. That is the more alarming version of the problem, and it is better sourced than the drift statistics circulating on vendor blogs.
11. What is your model change-control process — will we be told before the model changes?
NICE’s Evidence Standards Framework anticipates this, requiring agreement on post-deployment reporting of performance change for adaptive technologies. In practice, the underlying model can be swapped by the supplier with no notice and no re-validation. Your safety case was written about a system that no longer exists.
12. Does the interface present information or a recommendation — and does it flag low confidence?
This is the one question on the list with a solid human-factors evidence base behind the answer.
A systematic review in JAMIA covering 74 studies found that automation bias is worsened by exactly the conditions of NHS work — high workload, task complexity, time pressure “which pressurized cognitive resources.” Among the few mitigations with evidence behind them: presenting information rather than a recommendation.
Why it matters so much: in a study in Radiology, 27 radiologists read mammograms with AI assistance. When the AI was wrong, accuracy among the very experienced fell from 82.3% to 45.5%. Among the inexperienced, from 79.7% to 19.8%. A confident wrong recommendation didn’t just fail to help — it destroyed clinical capability that was already there.
NHS England has, to its credit, named “overreliance or automation bias from users” as a hazard in its own guidance. Ask your vendor what they have designed to counter it. Most will not have considered the question.
The thing nobody will tell you
There is a national tension running through all of this that you should factor into your risk appetite.
NHS England’s Medium-Term Planning Framework directs organisations to deploy ambient voice technology “at pace.” Meanwhile the AVT Supplier Registry, launched in January 2026, is self-certified, NHS England performs “only preliminary completion checks”, explicitly “does not endorse any of the suppliers”, and is expressly “not a commercial framework.” The hazard-log and safety-case templates that would let trusts do the assurance work properly are promised across 2026 and 2027.
Registry inclusion is not assurance. The responsibility remains entirely yours, and the tooling to discharge it is still being written.
Meanwhile 40% of UK GPs are already using AI scribes, according to a survey of 598 GPs published in npj Digital Medicine in May 2026 — adoption the authors describe as “relatively high despite regulatory issues and recent official cease communication.”
Deployment is running ahead of assurance. That is the environment you are procuring in, and the twelve questions above are what standing still in it looks like.
For the wider picture on where AI genuinely helps in healthcare compliance and where it doesn’t, see our state-of-the-nation guide to AI in healthcare compliance. If you’re being asked to evidence continuous assurance rather than inspection-week readiness, our guide to PSIRF, SAF and the new bar for assurance covers the regulatory side.
