
Patients have caught errors in AI-written NHS notes that clinicians missed
Healthwatch England reported cases of AI scribes producing the wrong drug name and an incorrect diagnosis in clinical documentation, spotted by patients rather than clinicians. The rule that follows applies to any firm taking meeting notes.
An AI-generated hospital summary recorded a woman as having serious nerve damage she did not have. The scan had found none. One word had gone missing between the result and the summary: "null", the word the report used to say there was none. She spotted it herself. Her clinicians had not.
Your firm runs the same kind of tool over client meetings, and it can fail the same way. So treat anything an AI notetaker produces as a draft rather than a file note: nothing goes onto the file or out to a client until one named person who was in the meeting has checked it and put their name to it.
The check that works is not a careful read. It is a comparison, against your own account of the meeting written before you open the note, and against the papers. The errors reported on 31 August included an instruction left out and a negating word that vanished, and both of those survive a read.
What the watchdog reported, and what a doctor described
The Guardian reported on 31 August that AI scribes running in NHS consulting rooms are putting errors into patient records. The finding comes from Healthwatch England, the statutory patient watchdog. Its write-up of that research, published in July 2026, draws on two things: a YouGov poll of 4,039 adults in England, fielded between 16 and 27 April 2026, and 44 written accounts from members of the public.
Be clear about what that evidence can carry. The Healthwatch accounts show the kinds of failure that can occur, not how often they occur: the 44 written accounts were self-selected, and the poll measured public attitudes rather than error rates.
Four failures are on the record. They are worth taking one at a time, because a check that catches one will not catch the others.
A negation lost. A woman's summary recorded demyelination, serious nerve damage. She queried it, and the hospital corrected the record to what it should have said all along: "null demyelination".
A drug name replaced by another. A scribe recorded a different drug from the one the doctor had prescribed. The patient, not the doctor, spotted it.
An instruction left out. A summary letter did not carry the consultant's instruction to get a repeat prescription from the patient's own doctor. Without it, the patient could have been left without their medication.
Something recorded that never happened. Separately from the Healthwatch work, a London doctor, Shier Ziser Dawood, wrote in BJGP Life in August 2025 that a scribe was "adamant" Dawood had told a patient to continue a drug she had never prescribed. It had not been mentioned in the consultation, and it did not appear in the transcript the tool itself produced.
Healthwatch summarises the pattern like this: "We have heard multiple stories from patients who have noticed these errors when a health professional hasn't. These inaccuracies may persist in their records if the patient doesn't catch them."
How common is this, and what do people make of it
None of that means AI notes are generally worse than human ones, and there is no good comparative error-rate evidence either way. A survey of 598 UK GPs published in npj Digital Medicine in 2026 found 240 of them already using AI scribes, and found the benefits widely accepted: over 65% agreed there are safety benefits such as reducing errors and omissions. The same GPs saw the other side, with over 60% agreeing there are risks of "inaccuracies, errors, misinterpretations and associated medicolegal threats". Reviewing the wider literature, the authors note clinicians reporting inaccuracies "such as mistaking 'did' with 'did not'". The claim here is the narrow one: a generally useful tool can still produce errors that fluent prose makes unusually hard to see.
The public is divided about the tools themselves. Healthwatch found support and opposition close to level, at 38% and 37%, with almost twice as many strongly opposed (21%) as strongly supportive (11%). What shifts that balance is a visible safeguard. Over two-thirds, 69%, said they would be more comfortable if there were "a clear commitment from the healthcare professional to check the accuracy of the content the AI scribe produces", and 81% wanted to be told the tool was being used and asked before it was. Of those who had had an appointment in the previous year, nearly 90% said they had not been aware of AI scribing being used, which is not evidence it was used on them unannounced; in most of those appointments it may simply not have been used.
The useful part for a firm is the shape of that: what reassures people is not the technology getting better, it is being told the tool is running and knowing a named person checked what it wrote.
Why an NHS story is your story
Strip out the medicine. What is left is a machine producing a fluent, confident, well-formed account of a professional conversation, and that account can become the official record of what was agreed if nobody compares it with what was actually said.
That is your file note. Your attendance note, your meeting minute, your call summary, your client letter. The tools are already in the room: Microsoft 365 Copilot, Otter, Fireflies, Granola, Zoom's AI Companion.
The read-across is not automatic, and the limits are worth stating. A patient reading their own medication list has a reason to read every line that a client skimming a meeting note does not, and the cost of an error in a consulting room is not the cost of an error in an attendance note. What carries over is the failure mode: a fluent summary of a conversation, produced by a system that bears no responsibility for the consequences, entering a record that later decisions will rest on.
What the regulator actually said, and what it did not
On 29 July 2026 the Medicines and Healthcare products Regulatory Agency announced guidance, developed with NHS England, on how medical device law applies to ambient voice technology, or AVT, the NHS term for these tools. It confirms that products "intended solely for transcription, summarising of clinical conversations, drafting letters, or suggesting clinical codes for a clinician to review are not regulated as medical devices under the current framework". Products intended to support diagnosis, treatment or prevention, or to take automated clinical action such as placing orders without clinician review, are regulated as medical devices.
The test is what the product is for, not whether a human happens to read the output afterwards.
The obligation to check sits alongside that classification rather than underneath it, and it is stated plainly on both sides. The MHRA: "Clinicians remain responsible for reviewing and verifying AI generated transcripts, summaries and other outputs before they are used in patient care. This responsibility is unchanged by the guidance." NHS England, in its guidance on ambient scribing, tells patients what that means in practice: staff "must check the notes or letters that have been made by the ambient scribe to make sure they are accurate", and they do this "before adding any information to your record".
So no medical-device approval regime stands behind these tools for their administrative functions. Assurance comes from elsewhere: NHS clinical risk management standards, compliance with which "is mandatory under the Health and Social care Act 2012", together with procurement, local governance, and a person checking the output.
None of that regime reaches your practice. The MHRA does not regulate you, and no equivalent body has said any of this about your file notes. Nothing is lost by that. If you are a solicitor, accountant or other professional adviser, your existing professional, evidential and often regulatory reasons for keeping an accurate record of what you advised do not disappear because software drafted it. What the NHS material gives you is not an obligation but a rule worth borrowing, already written and already tested at scale: the output is a draft, somebody checks it, and then it becomes the record.
Why a read-through is not enough, and what to check instead
"Just check it" is weaker advice than it sounds.
Of the four failures, only the swapped drug name was caught by someone noticing the line itself, and even then only because the patient knew independently what the correct drug should have been. Take that knowledge away and it reads like any other line.
The other three survive a read, each for its own reason. It is hard to notice an instruction that is not there by reading what is there, because an omission leaves no mark on the page. An invented line reads much like a real one. And a dropped negative produces a sentence that is fluent, confident and says the opposite of what was said.
That last one is worth seeing in your own work. The client says they will not be claiming research and development relief this year. The note says they will be. Nothing about that sentence looks wrong on the page, and a year later it is the only account of what was decided.
There is also an order problem. Reading the AI version first makes its wording harder to separate from your own memory of the meeting, so the note you are checking quietly becomes the standard you are checking it against. That is the reason the check starts before you open the document.
So start before you open it. Write your own three bullets of what was agreed, from memory, while the memory is still yours. Then open the note and work through four things in this order, comparing against those bullets and against the papers:
- Every name and figure. Parties, entities, statutes, amounts, percentages, dates, deadlines, notice periods, reference numbers. Check them against a document rather than against your recollection: the engagement letter, the papers in front of you, the transcript if the tool keeps one. That independent reference point is what made the drug name catchable.
- Every negative, in both directions. Check every "not", "no", "won't", "isn't", "unless", "except" on the page against your bullets, then check your bullets for a negative the note has lost. Each is one short word that reverses its sentence, and the missing ones are not on the page to be found.
- Every action. Who agreed to do what, by when. Then ask separately what was agreed that is not there, which is the only way to catch an omission.
- Anything that surprises you. This is the weakest of the four, because an invented line looks like every other line and a general read may slide past it. Use it as a prompt rather than a net: where a substantive statement surprises you, verify it against the transcript, the recording or contemporaneous material rather than against your impression of it.
On a short note this may take only a few minutes. On a long one, where figures have to be read against papers or a transcript, it takes considerably longer. Know that before you set the expectation with your team. It is not a proofread; the prose does not need improving.
Write the rule down, as a rule about the document
Define the document's status first. Everything else follows from it.
Four lines, which you can put in a policy or a team channel this week.
- Anything an AI notetaker produces is a draft. It is not a file note, an attendance note or a minute until line 3 has happened.
- A draft does not go to a client or a third party.
- One named person who was in the meeting checks it and puts their name on it. One person, not the team, because a note everybody owns is a note nobody has read. Anything going to a client or a third party gets the three bullets and all four checks. Purely internal notes get the bullets written and compared against the note, which is less, but is still a comparison rather than a read.
- If it is still unchecked after two working days, it is marked unverified on the file. Handle it from there under your own retention and record-keeping policy rather than letting it sit and eventually pass as a checked note. Whether an unchecked draft should be kept, corrected or removed depends on your professional obligations and on what other record of the meeting exists, so that is a decision for your policy, not for this article.
Line 3 is where the work happens, and line 4 is what stops the whole thing becoming decoration. "Unverified" has to be visible on the document, or the two-day rule means nothing.
The honest risk in writing a rule you do not keep
A control you have written down and routinely fail to perform can be worse than having no such control. If people initial notes without reading them, you have manufactured evidence that a check happened, and that evidence will be produced at the exact moment it hurts.
So keep it small enough to be done. Three bullets, four checks, one name, and the lighter internal tier in line 3 for the notes that never leave the building.
Be wary of the narrowing that sounds more sensible than it is. "Only where the note carries advice, money or a deadline" seems like a reasonable limit until you notice that in an accountancy or law practice it covers most substantive client calls, which means it is not a limit at all. Divide the work by where the note ends up, which is a line you can actually draw, rather than by how important it felt at the time. A rule that quietly lapses is the one that will be produced against you.
FAQ
Does this mean we should stop using AI notetakers? No. In the npj survey above, GPs widely endorsed the efficiency benefits, including reducing administrative burden and time spent writing notes, though how much time a firm actually saves varies with how much checking and editing the output needs. The NHS is expanding their use rather than pulling back: NHS England's ambient scribing service page says organisations are "directed to deploy AVT at pace". The point is that the output is a draft, and the net saving is smaller than the gross, because the check is part of the job.
Will a better model fix this? Better models should reduce some kinds of error. They will not remove the check, because even a low error rate matters when the error is hard to detect, and because more consistently fluent output invites more trust in the occasional mistake.
Who should do the checking? Someone who was in the meeting, named individually rather than as a team, and ideally the person whose advice is recorded in it. Someone who was not there is less well placed to spot omissions or distortions that are not apparent from the text itself.
Is it enough to send the note to the client and ask them to confirm? It is a useful second check and it is not a substitute. It should not transfer responsibility for your own record to the client, and a confirmation you do not trust is worse than none, because it goes on the file. Check it first, then send it.
What if we find an error after the note is on the file? Correct it, date the correction, and keep the original readable. A correction nobody can see is not much of a correction.
What about the recordings and transcripts themselves? That is a separate problem with a separate answer, and for a small firm it may be the larger exposure: where recordings are stored, who can reach them, and how long you keep them. We have covered it in our piece on notetakers in client meetings, which is also where the question of whether to let these tools into the room at all belongs.
If you want a second pair of eyes on how AI is writing things down inside your firm, book a conversation.
This is general information about record-keeping practice, not legal advice.