Dark teal cover with a magnifying glass over a line of text, for an article on checking AI output against its sources.
AI riskQuality controlProfessional servicesLeadership

How to stop AI making things up in business use

Almost everyone checks that a source is real. That check misses AI's worst mistakes, because those mistakes cite real sources and then describe what the source says incorrectly.

Good Transformer12 min read

AI tools do more than draft this year. They act, pulling a figure into a document or updating a record directly, often with nobody retyping what they produced first.

So the check most firms run for AI mistakes is the wrong one. Nearly everyone checks that a source is real. That check catches invented sources. It misses the mistakes that use real ones: a genuine document, correctly cited, summarised as saying something it does not say.

Only reading the document catches those, and nobody can read everything. So the useful thing a firm can do this week is write down a threshold. If someone outside the firm will act on a claim, one named person checks it against the source before it goes out.

This piece is general information, not legal advice.

The failure everyone catches

In October 2025 a Leeds employment tribunal heard an application against Mr D May, trading as Leeds Gymnastics Academy, who had been defending an employment claim without a lawyer. He had filed written representations. The tribunal recorded what happened next at paragraph 49 of its judgment:

"At least half of those references were non-existent and Mr May admitted that he had just used Chat GPT to produce his representations without checking any of the results. He said he was reasonably entitled to conclude that everything that Chat GPT said was reliable, and in fact, he said there was no reason to fact check the internet at all."

The tribunal ordered him to pay £2,178 for the other side's preparation time, calculated as 49 hours and 30 minutes at £44 an hour. Payment is stayed pending his appeal to the Employment Appeal Tribunal, so the money has not changed hands. But the finding stands.

That is what everyone pictures when they hear the word hallucination, and it is the easy one. A case that does not exist fails the cheapest test there is. Someone types the name into a search box, finds nothing, and nothing else in the document can be trusted either.

The failure that survives the check

In January 2026 the First-tier Tribunal's tax chamber gave judgment in Gary Elden v HMRC. HMRC had applied to strike out the appeal. The written argument filed in response summarised five cases in support.

The citations were correct. The cases were real. There was nothing to catch by looking.

Judge Allatt read them anyway. The written argument offered Atlantic Electronics Ltd v HMRC as confirming that delay alone is not enough for a strike-out. Judge Allatt read all nine pages of it and recorded this:

"This case concerned HMRC's appeal against a First-tier Tribunal decision refusing permission to admit additional evidence from two witnesses in an MTIC (Missing Trader Intra-Community) VAT fraud appeal. It is not a case about strike out. It says nothing about what powers Tribunals should consider using before strike out."

She did the same with the next one, an eleven-page decision the written argument had described as being about refusing a full hearing. It turned out to concern a costs application filed four working days late. A third case, offered on the same point, concerns PAYE penalties, and passages attributed to it came from a different case altogether.

The defence given on the appellant's behalf is worth reading closely, because it sets out the assumption behind the whole problem:

"The suggestion that citing a published authority amounts to providing false material is misconceived. A court decision is a matter of public record. Whether a case applies is a matter of legal argument and opinion, not misrepresentation."

The citations themselves were accurate. What the document said those cases had decided was not. The tribunal found the summaries had been produced with AI and had not been verified with sufficient care. It ordered that future submissions in the case state whether AI had been used.

Checking that a source exists and checking that it says what you claim are two different jobs, and most working processes only do the first. Closing that gap took a judge nine pages of reading on one case and eleven on the next.

How common this failure is, and who it happens to

There is a public database of court cases involving AI-generated false material, maintained by the researcher Damien Charlotin. On 17 August 2026 it held 1,922 cases, and it had last been updated the day before. Sixty-two of them are British.

Its published data shows between 115 and 178 new cases in every complete month of 2026 so far. Treat those as minimum counts rather than final ones. The database fills in older decisions as they surface, so the recent months are the least complete.

One more figure needs care. Of the cases dated 2026, 850 involve something invented and 424 involve a real source described wrongly. The database applies both labels where both apply, so those two counts overlap.

Misdescription matters because it is what gets through: the check most people run cannot detect it.

Then there is the breakdown nobody quotes. Thirty-five of the 62 British cases were filed by someone with no lawyer, and 24 by a lawyer. The remaining three were filed by a government lawyer, a judge, and one jointly by a litigant and a lawyer. Worldwide the pattern runs the same way, with 1,110 cases filed by people without lawyers against 754 filed by lawyers.

So this is not only a lawyers' problem. It is best documented in law because courts write down what they find.

What it describes is what happens when a capable person does serious work with AI and nobody checks behind them. Courts are simply the one place where a stranger reads every source and publishes the result. Your firm has no such person.

What a real check looks like

The clearest published description of a proper check comes from the Divisional Court in June 2025, in the joined cases of Ayinde and Al-Haroun. The court set out the duty in these words:

"Those who use artificial intelligence to conduct legal research notwithstanding these risks have a professional duty therefore to check the accuracy of such research by reference to authoritative sources, before using it in the course of their professional work"

Check the accuracy against the source.

Two qualifications apply. The first is that this duty is owed to the court by lawyers, and does not bind an architecture practice, a recruitment agency or a design studio at all. It is quoted here as the sharpest available statement of what a real check is.

The second is that the same court opened its judgment by calling artificial intelligence a powerful technology and a useful tool in litigation. The judges who have seen the worst of this are not arguing against the tool, and neither are we. The problem is sending what the tool produced to a client without anyone reading the sources.

If your firm reviews work it did not produce, the same problem shows up in ordinary review: the reviewer checks that a document has the right sections rather than reading what it says.

You cannot check everything, so decide what you check

The instruction to verify everything before it goes to a client is the one nobody follows, because at any real volume it cannot be followed. A firm that tries either stops trying within a fortnight or slows down so much that the AI stops paying for itself.

So the question is which claims someone has to open the source for.

The obvious ways of dividing the work are the wrong ones. Do not divide it by the value of the job, because a small job can carry a claim that ends up in a contract. Do not divide it by subject matter, because this failure reaches well beyond law. Do not divide it by how confident the output sounded, because it always sounds confident, and that is the whole problem.

The test that works is whether you can still fix the mistake cheaply. A wrong claim in an internal draft gets corrected by the next person who reads it, and costs nothing to put right. A wrong claim that has already reached a client gets corrected in front of the person who relied on it, if it gets corrected at all.

Draw the line at the point where cheap correction ends: the moment someone outside the firm is going to act on the claim.

The threshold, written out

Take this and put it in your AI policy. It is written to be copied.

Source-checking threshold. Where AI has helped produce a claim of fact, and someone outside this firm will act on that claim, one named person checks it against the underlying source before the work goes out. That person confirms the source supports the claim as stated. Confirming that a cited source exists does not satisfy this. Where the source cannot be reached, or does not support the claim, the claim is removed, or restated in general terms without the specific fact. The work does not go out until that has been done.

Three tiers, and what each one gets

The work What it gets Who does it
Internal draft, thinking, notes Nothing. Use it freely. The person writing it
Informs a decision inside the firm Spot-check the claims the decision depends on The person taking the decision
Someone outside the firm will act on it Every claim of fact checked against its source, in full One named person, not the author

The middle row is where most firms sit and where most of the argument happens. Keep it cheap: a decision that stays inside the firm can be revisited, so spot-checking is proportionate.

The five checks that catch what an existence check misses

A claim of fact, for this purpose, is anything with a number in it, anything attributed to a named person or organisation, and anything carrying a date. Run these five on each of them, before the work goes to anyone outside the firm.

  1. Open the source and find the sentence. Find the actual passage that supports the claim, not the abstract, the summary, or the search result. This is the check that would have caught Elden, where every citation was genuine and no citation said what it was offered for. If you cannot find the passage in the document, the claim does not go out.
  2. Check the date and the edition. Many organisations publish the same survey or guidance every year with different numbers. A figure attached to the wrong year is a false claim even when the figure is real somewhere.
  3. Check that you are at the originating source. A statistic quoted in a blog post that quotes a press release that quotes a report has been retold three times, and each retelling could have changed it. Go to whoever produced the number.
  4. Check the scope as well as the number. A finding about firms with more than 250 staff does not describe a firm of twelve people. This is where correct figures become misleading claims.
  5. Read one paragraph either side of the quotation. A quotation lifted out of its setting can carry the opposite of what its author meant, and the sentences on either side are where you find that out.

Write down that the check was done

If nobody records that the check was done, it stops being done. One line in the file note or the document properties is enough:

Sources checked against originals by [name], [date]. Claims verified: [list]. Claims removed as unverifiable: [list].

The last line matters most. It gives the checker somewhere to put a claim they could not stand up, which makes removing it routine rather than awkward.

What you can control

You cannot stop a model producing a confident false claim. That is how these systems work, and no amount of prompting removes it.

What you can decide is which work gets checked against its sources. Write the threshold down. Name the person.

Tell people that confirming a source exists is not the same as confirming it supports the claim. Accept that the small share of work leaving your control carries a real cost. That is a rule a firm of twelve can actually follow.

If something has already gone out with a false claim in it, that is a different job and the first hours matter: we have written a plan for that. We have also written more broadly about keeping AI mistakes away from clients.

If you want help setting the threshold for your own work, and training people to work to it rather than sending round a policy nobody reads, book a conversation with us.

FAQ

Does a better model fix this? A better model produces fewer false claims. It does not stop producing them. A false claim from a better model is also harder to spot, because the work around it is more plausible. Model choice is worth getting right, and it stands alongside the threshold rather than replacing it.

We give it the source documents. Are we covered? You are in a better position, but you are not covered. Supplying the material means the model has no gap to fill with an invented source, which is real progress. It does not stop the model summarising a supplied document as saying more than it says. That is the Elden failure exactly, and the documents there were real and in front of everyone.

The assistant gives citations now. Is that not the check? It is the beginning of one. A live link tells you a document exists, which was never really the problem. Whether the document supports the sentence attached to it is a separate question, and answering it means opening the link.

We do not have time to read every source. Nobody does, which is why the threshold exists. For most firms, work that someone outside will act on is a small fraction of output. Measure it for a week before deciding it is unaffordable.

Sources


The duties described in Ayinde apply to lawyers conducting litigation and are quoted here as a standard worth borrowing, not as an obligation on your firm. Take advice on your own position.

Work with Good Transformer

Turn this thinking into working practice.

Explore team advisory

Newsletter

Get new Insights by email

Practical notes on using AI with judgement, and the AI news leaders actually need. No hype, no spam, unsubscribe anytime.

Choose how often you want the digest

Keep reading

AI risk6 min read

The AI that tidies your writing has a slant of its own

A new Oxford study finds AI editing tools introduce a consistent slant into the text they tidy, even when told to preserve it. For firms that run client writing through AI, here is the check worth keeping.

6 July 2026