Safe Workplace
Products
CalmERCompliantCare
Pricing
CalmER pricingCompliantCare pricing
More
CustomersGuidesBlogBook a demo
Employee Relations

Keeping AI-assisted casework tribunal-ready

Safe Workplace29 July 20267 min read

The question HR teams keep asking about AI in casework is “are we allowed to?” It is the wrong question. Nothing prohibits you from using AI to draft an investigation report.

The right question is what happens when the other side asks how the report was produced — and whether your answer survives contact with a tribunal.

We have a useful preview of that, because a different profession has already been through it in public. Since mid-2025, UK courts have dealt with a run of cases where lawyers put AI-generated material in front of a judge without checking it. A running tally kept by barrister Matthew Lee of Doughty Street Chambers recorded 67 UK cases of hallucinated or false citations by July 2026. It is his count rather than an official statistic — but the judgments are real, published, and quoted verbatim.

The reasoning in those cases transfers to HR almost without modification.

What the courts actually decided

The anchor authority is R (Ayinde) v Haringey and Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin), decided by the Divisional Court in June 2025. Two referrals concerning fictitious case law and misstated statute. The court directed judges to take “a robust approach”, and said lawyers citing fictitious cases “must face serious consequences.”

The case that should make HR sit up is Rodney v Gee’z Micro Bar and Pitstop, heard in March 2026. A paralegal had used research-assistant tools. Three filed documents contained mis-cited authorities. The supervising solicitor, in the judge’s words, “reviewed the said documents but did not verify the authenticity of the cases or the citations.”

The judge’s assessment:

“Misleading material was placed before the court… when even the most simple of checks would have shown that not to be the case… That is inexcusable on the part of a professionally qualified lawyer.”

Contempt wasn’t made out — negligence alone isn’t enough for that. But both solicitors were referred to the SRA, and the judgment was published at public expense.

In Rafique v HMRC [2026] UKFTT 673 (TC), the tribunal put it more bluntly still: “it is a contempt of court for fabricated authorities to be cited to a court or tribunal as if they are genuine.”

Three principles emerge, and all three apply to you:

  1. Using AI is not itself prohibited. No court has said otherwise.
  2. The duty to verify is unaffected and non-delegable. You cannot subcontract it to the tool, or to a junior, or to the fact that it looked right.
  3. Reviewing is not verifying. This is the distinction that caught the solicitor in Rodney, and it is the one most likely to catch you.

Why HR is more exposed than it thinks

The mechanism in these cases is not exotic. A professional under time pressure delegates drafting, doesn’t check the output, and signs a document carrying a formal attestation.

An investigation report is exactly that document. So is a disciplinary outcome letter, a grievance response, and an ET3. They are relied on by people making decisions about someone’s livelihood, they are disclosable, and their author will be cross-examined on them.

Two features of AI error make this harder than it sounds.

The failure mode is designed to pass a quick check. In a study of GPT-4o generating literature reviews, 19.9% of citations were entirely fabricated — but among fabricated citations that carried a DOI, 64% of those DOIs were valid and resolved to a real article, just an unrelated one. Among the citations that were genuine, DOIs were wrong in 36% of cases. Fabrication is not obviously fake. It is plausible metadata that survives the check most people actually perform.

That study tested a bare model with no retrieval, and it was in a clinical-academic domain, so don’t read the 19.9% across to your policy assistant. Read the shape of the failure across, because that generalises.

Omission is more common than fabrication, and much harder to spot. In research on clinical summarisation published in npj Digital Medicine, hallucinated content appeared in 1.47% of generated sentences — but 3.45% of source content was omitted. Omissions ran roughly 2.3 times more common than fabrications. And 44% of the hallucinations were rated major, meaning they could change a decision.

Think about what that means for a case summary. Reading for a wrong statement is a task humans are reasonably good at. Reading a fluent, coherent, professional-looking summary and noticing that the mitigating circumstance the employee mentioned in hearing two is simply not there — that is a much harder cognitive task, and it is the more likely failure.

The uncomfortable evidence on human review

“Human in the loop” is doing enormous work in this market as a reassurance. The evidence for it is thinner than the confidence with which it is asserted.

The most direct test comes from radiology. In a study published in Radiology, 27 radiologists read mammograms, some accompanied by incorrect AI assessments. When the AI was right, accuracy sat around 80%. When the AI was wrong:

  • Inexperienced readers dropped from 79.7% to 19.8%
  • Very experienced readers dropped from 82.3% to 45.5%

Experts, reading images they are trained to read, lost roughly half their accuracy when the machine was confidently wrong. The AI didn’t merely fail to help — it destroyed capability they already had.

That was an experiment, using a simulated AI, with a wrong-suggestion rate far above real life. It generalises as a human-factors principle, not as a rate. But the principle is the point, and it is corroborated by a systematic review in JAMIA which found the conditions that worsen automation bias are high workload, task complexity and time pressure. Which is a fair description of an ER team’s Tuesday.

And then there is Rodney, which is a natural experiment in exactly this. Review happened. Verification did not. Across 67 UK cases, the review step existed and failed repeatedly — among professionals with formal duties, regulatory exposure, and a statement of truth attached to the document.

The honest conclusion is not that human review is theatre. It is that review as usually designed — per-case, in-workflow, under time pressure, performed by the person who benefits from the output being fine — is the weakest possible implementation.

What tribunal-ready actually looks like

Know which task you gave it. Summarising a document you supplied is a fundamentally more constrained task than asking a model to produce content about things it wasn’t shown. The error rates differ by roughly an order of magnitude. “Summarise these four statements” is low-risk. “What does the ACAS Code say about this?” invites fabrication. Keep the tool grounded in your own evidence.

Verify the specific things AI gets wrong. Not a general read-through. Dates, names, quotations, and any reference to a policy clause or legal authority — checked against the source document, every time. In Rodney, “even the most simple of checks” would have caught it.

Check for what’s missing, deliberately. Since omission is the dominant failure and the hardest to see, it needs its own step. Read the source and ask what should be in the summary, rather than reading the summary and asking whether it looks right.

Keep a record of the process, not just the output. If you are asked in cross-examination how a report was produced, “our system logs which documents the summary drew on, and the investigating officer signed off against the source material” is an answer. “I think someone used the AI feature” is not.

Make the sign-off mean something. The Amsterdam Court of Appeal held in 2023 that nominal human review of an algorithmic decision was “not much more than a purely symbolic act” — and therefore did not stop the decision being solely automated. Apply that test to yourself. If your reviewer has no time, no access to the underlying evidence and no realistic route to disagree, you have a symbolic act, and a tribunal may see it the same way.

Sample and audit independently. Per-case review by the person producing the case cannot be your only control. Pull a random sample each month, check it against source material, and record the error rate. This is the control that would actually have caught Rodney — and it is the one nobody budgets for.

Be able to say what the tool is. Not the marketing name. What model, what data it was given, what it does with it, whether your case material trains anything, and where it is processed.


None of this is an argument against using AI in casework. The administrative load around a case is real, and the tools genuinely help with it. The efficiency case is not in dispute.

The argument is narrower: the moment you use AI to produce something that will be relied on by a decision-maker and scrutinised by a tribunal, you inherit a verification duty that the tool cannot discharge for you. The legal profession learned that in public, expensively, across 67 reported cases. HR gets to learn it by reading.

For the broader picture, see our guide to AI in employee relations, or the companion piece on bias, black boxes and the law.

Ready to see it on your data?

Book a demo of the platform behind these insights — CalmER for employee relations, CompliantCare for healthcare compliance.