Safe Workplace
Products
CalmERCompliantCare
Pricing
CalmER pricingCompliantCare pricing
More
CustomersGuidesBlogBook a demo
Employee Relations

Bias, black boxes and the law: a plain-English guide

Safe Workplace29 July 20269 min read

Most writing about AI bias in HR runs on the same two examples: Amazon’s scrapped recruiting tool, and a vague sense that facial recognition doesn’t work on darker skin. Both are weaker than their reputation. Meanwhile the genuinely important developments — a Dutch appeal court ruling on what counts as human review, a Nature paper on what safety training actually does to bias, and a quiet change to UK data law — get almost no coverage.

This is the plain-English version. No legal advice, and we’re not lawyers — but the sources are all linked, so you can check us.

Start with the finding that should worry you most

In 2024, researchers published a study in Nature using a technique called matched guise probing. They took identical content, wrote it once in African American English and once in Standardised American English, made no mention of race whatsoever, and asked language models to make decisions about the speaker.

The models were more likely to assign speakers of African American English less prestigious jobs, to convict them, and to sentence them to death. The authors describe the models as embodying “covert stereotypes that are more negative than any human stereotypes about African Americans ever experimentally recorded.”

Then comes the part that matters for anyone buying this software:

“Existing methods for alleviating racial bias… such as human feedback training do not mitigate the dialect prejudice, but can exacerbate the discrepancy between covert and overt stereotypes, by teaching language models to superficially conceal the racism that they maintain on a deeper level.”

Read that again. The standard safety training used across the industry made the models’ stated views more positive while leaving the underlying prejudice intact. (Hofmann et al., Nature 633, 2024.)

Why this matters more in employee relations than in recruitment: the structure of the task in that study — take a person’s own words, judge their character, decide a sanction — is the structure of a disciplinary process. This is a hypothetical-vignette study on US dialect, not a study of real HR cases, and we should be honest about that extrapolation rather than smuggle it past you. But it means the reassurance most vendors offer — “we tested it and it produced no biased outputs” — is close to worthless. Testing the outputs is exactly the check this failure mode is built to pass.

The evidence on screening bias, and how to read it

Names. The strongest study we found ran 120 names across more than 550 real CVs and 500 real job listings, across nine occupations and three open language models — over three million comparisons. White-associated names were preferred 85% of the time against 9% for Black-associated names. Male-associated names 52% against 11% for female. The systems never preferred Black male-associated names over white male-associated ones. (Wilson & Caliskan, AIES 2024.)

Two details deserve more attention than the headline. First, Black female names were preferred 67% of the time against 15% for Black male names — meaning the harm concentrated on Black men and would be invisible if you analysed race and gender separately, which is how most bias audits are done. Second, bias got worse when CVs were shorter: less information, so the demographic signal carries more weight.

The caveat: these were open-weight models, not the commercial ATS products actually in use. As one of the authors put it, proprietary systems can only be studied by approximation.

A complication worth sitting with. Bloomberg ran a comparable test on GPT models and found names associated with Asian women ranked top most often. The direction of bias was not the same as the Washington study. That instability is itself the finding: these systems are not reliably fair in any particular direction, which makes “we checked for bias against group X” a much weaker assurance than it sounds.

Faces. Two studies get conflated constantly. Gender Shades (2018) found error rates up to 34.7% for darker-skinned women against 0.8% for lighter-skinned men — but that was gender classification, guessing whether a face is male or female, not identity verification. The relevant study for workplace ID checks is NIST’s FRVT Part 3 (2019): 189 algorithms, 18.27 million images. It found false positives for Asian and African American faces elevated by “a factor of 10 to 100 times” in one-to-one matching.

But NIST’s own framing is the bit everyone drops: different algorithms perform very differently, and the most equitable were also among the most accurate. Algorithms developed in Asian countries showed no Asian/Caucasian gap at all — strong evidence that training data, not the task, is the problem. “Facial recognition is racist” misstates NIST. “Facial recognition is racist unless you check, and most buyers don’t check” is accurate.

Voice. Five commercial speech recognition systems tested by Koenecke et al. in PNAS showed word error rates for Black speakers roughly double those for white speakers. If you use AI notetaking in disciplinary or grievance hearings, that is a fairness problem before any model reasons about the content.

Two examples to handle carefully

Amazon’s recruiting tool. Reported by Reuters in October 2018, sourced to five people speaking anonymously, with Amazon declining to comment. There is no document, no dataset, no audit, no effect size. The tool was a prototype built on 2014-era keyword matching, recruiters “never relied solely” on its rankings, and it was killed before deployment. It is arguably a governance success story — bias found internally, mitigation attempted, mitigation judged unreliable, project cancelled. Cite it as the origin of corporate awareness that proxies defeat naive de-biasing. Don’t cite it as evidence that AI rejected women.

NYC’s bias audit law. Local Law 144 requires an annual independent bias audit, a public summary, and candidate notice. Researchers checked 391 employers and found 18 with published audit reports. That gets quoted as a 4.6% compliance rate. The authors coined the term “null compliance” specifically to stop that reading: the law lets employers decide whether they’re in scope, so silence cannot be read as breach. The real finding is that the law is unfalsifiable from the outside — which is a more interesting problem than mass non-compliance, and a warning about transparency-based regulation generally.

The ruling that matters most

If you remember one legal authority from this article, make it this one.

In April 2023 the Amsterdam Court of Appeal ruled on drivers’ challenges to automated decisions by Uber and Ola. Uber’s defence was that a human reviewed the decisions, so they weren’t automated. The court rejected it: on the facts, the reviews were “not… much more than a purely symbolic act.”

Because the review was symbolic, the decisions were solely automated, which engaged GDPR Article 22 and the explanation rights in Articles 13–15. The trade-secrets defence also failed — withholding information about the fraud-detection algorithms was disproportionate “compared to the negative effects of unexplained automated dismissals and the disciplining of workers.” In October 2023 Uber was ordered to pay €584,000, with €4,000 for every further day of non-compliance. (Fountain Court analysis.)

This is a Dutch judgment, not binding here. But it is the clearest judicial statement anywhere that a rubber stamp is not human review — and “we have a human in the loop” is the single most common compliance claim in this market.

A second European case is worth knowing. In December 2020 the Court of Bologna held Deliveroo’s “Frank” ranking algorithm indirectly discriminatory because it penalised riders for cancelling sessions without distinguishing why — including strike action, illness and caring responsibilities. The platform “does not know and does not want to know the reasons why the riders cancel.” The read-across to conventional employment is direct: facially neutral attendance or reliability scoring is indirectly discriminatory if it is blind to protected reasons for absence. If you are scoring absence patterns, that is your case.

The UK position, honestly

There is no reported UK employment tribunal or appellate judgment squarely deciding an algorithmic management discrimination claim. Anyone telling you otherwise is overselling.

The case that should have produced one was Manjang v Uber Eats. Pa Edrissa Manjang, a Black courier, was removed from the platform after repeated failures of Uber’s facial verification checks. Worker Info Exchange obtained every selfie he had submitted through a subject access request; all of them were of him. He filed in October 2021 with backing from the Equality and Human Rights Commission and the App Drivers & Couriers Union.

It settled in March 2024 for an undisclosed sum. No judgment, no findings of fact, no disclosure of what went wrong. A seventeen-day hearing listed for November 2024 never happened. Uber maintains its process “includes robust human review” and that facial verification was not the reason for his loss of access. EHRC chair Baroness Falkner noted they were “particularly concerned that Mr Manjang was not made aware that his account was in the process of deactivation, nor provided any clear and effective route to challenge the technology.” The ICO did not investigate, despite complaints dating to 2021.

Two and a half years, regulator-funded, and it still ended without a ruling. That tells you something about how well the current system adjudicates algorithmic discrimination.

And UK law just moved in the opposite direction. Section 80 of the Data (Use and Access) Act 2025 narrows UK GDPR Article 22. The general prohibition on solely automated decision-making now applies only where the processing relies on special category data. An automated decision based on performance metrics or attendance data falls outside the prohibition and into a lighter safeguards regime. Whatever your view of that policy, it means the Article 22 protection HR teams have relied on covers less than it used to.

Meanwhile the Equality Act 2010 is untouched and does all the same work it always did. Indirect discrimination doesn’t care whether the disadvantage came from a manager or a model.

So what should you actually do

Assume the audit you can afford won’t find covert bias. The Nature dialect finding means output testing catches the bias that was already visible. Design for containment instead: keep the model away from the decision, not just away from the protected characteristic.

Audit intersectionally or don’t bother. The Washington study’s harm was concentrated on Black men and would have been invisible to separate race and gender analyses.

Make human review real, or stop calling it review. The Amsterdam test is a good one to apply to yourself: if your reviewer sees the recommendation, has no practical route to disagree, and has thirty seconds, you have a symbolic act. Real review means the reviewer sees the underlying evidence, has time, and — critically — sometimes says no. If nobody in your organisation has ever overturned the model, that is data.

Check what your scoring is blind to. The Bologna principle. If your system counts absences without knowing which were disability-related, pregnancy-related, or protected industrial action, you have built the thing that court struck down.

Keep the record. Every case that has gone badly for an employer has turned on an inability to explain what happened and why. Which is the subject of our next piece, on keeping AI-assisted casework tribunal-ready.

For the wider context on where AI belongs in employee relations, see our state-of-the-nation guide.

Ready to see it on your data?

Book a demo of the platform behind these insights — CalmER for employee relations, CompliantCare for healthcare compliance.