
Patients have caught errors in AI-written NHS notes that clinicians missed
Healthwatch England reported cases of AI scribes producing the wrong drug name and an incorrect diagnosis in clinical documentation, spotted by patients rather than clinicians. The rule that follows applies to any firm taking meeting notes.
An AI-generated hospital summary turned "null demyelination" into "demyelination", and a woman was recorded as having serious nerve damage she did not have. She caught it herself. Her clinicians had not.
So treat anything an AI notetaker produces as a draft, not a file note. Nothing goes into a client file or out to a client until one named person who was in the meeting has checked it and put their name on it.
Check it against the names, the figures, every negative and anything that is missing. The errors reported on 31 August included an instruction left out and a negating word that disappeared, and both of those can survive a skim.
What the watchdog found, and what a doctor reported
The Guardian reported on 31 August that AI scribes running in NHS consulting rooms are putting errors into patient records. The finding comes from Healthwatch England, the statutory patient watchdog, whose write-up of its research, published in July 2026, reports a YouGov poll of 4,039 adults in England fielded between 16 and 27 April 2026, alongside 44 written accounts from members of the public.
Four failures are on the record. Take them one at a time, because they are not the same kind of mistake and they do not fail in the same way.
One caveat first, because it governs everything that follows. These cases show the kinds of failure that can occur, not how often they occur. The 44 written accounts were self-selected, and the 4,039-person poll measured public attitudes rather than error rates.
A negation lost. A woman's summary said she had demyelination, serious nerve damage. The record was later corrected to what it should have said, "null demyelination". The word that reversed the meaning was gone. She queried it herself and the hospital corrected it.
A drug name replaced by a similar one. A scribe recorded a different drug from the one the doctor had prescribed. The patient, not the doctor, spotted it.
An instruction left out. A summary letter did not carry the consultant's instruction to get a repeat prescription from the patient's own doctor, which could have left them without their medication.
Something recorded that never happened. Separately from the Healthwatch work, a London doctor, Shier Ziser Dawood, wrote in BJGP Life in August 2025 that a scribe was "adamant" she had told a patient to continue a drug she had not prescribed, which had not been mentioned in the consultation and did not appear in the transcript the tool itself produced.
These cases do not show that AI notes are generally worse than human ones, and we do not have good comparative error-rate evidence either way. A survey of 598 UK GPs published in npj Digital Medicine in 2026 found use already relatively high, with 240 of them current users, and found GPs endorsing the benefits: over 65% agreed there are safety benefits such as reducing errors and omissions. The same GPs were alert to the other side, with over 60% agreeing there are risks of "inaccuracies, errors, misinterpretations and associated medicolegal threats". Reviewing the wider literature, its authors also note clinicians reporting inaccuracies "such as mistaking 'did' with 'did not'", which is the failure this article is about. The argument here is the narrow one: a generally useful system can still produce errors that fluent prose makes unusually hard to notice.
Healthwatch summarises the pattern like this: "We have heard multiple stories from patients who have noticed these errors when a health professional hasn't. These inaccuracies may persist in their records if the patient doesn't catch them."
Why an NHS story is your story
Strip out the drug names. What is left is a machine producing a fluent, confident, well-formed account of a professional conversation, which can become the official record of what was agreed if nobody meaningfully compares it with what was actually said.
That is your file note. Your attendance note, your meeting minute, your call summary, your client letter. The tools are already in the room: Microsoft 365 Copilot, Otter, Fireflies, Granola, Zoom's AI Companion. Whether to let them in at all is a separate question, covered in our piece on notetakers in client meetings. This one is about whether what they write is true.
The transfer is not automatic, and it is worth being honest about the limits. A patient reading their own medication list has a reason to read every line that a client skimming a meeting note does not. The failure costs are different. But the relevant failure mode is the same, and it is the failure mode that matters: a fluent summary of a conversation, produced by a system that bears no responsibility for the consequences, entering a record that other decisions will rest on.
The MHRA has already decided who checks
The regulatory position settles where the duty sits.
On 29 July 2026 the Medicines and Healthcare products Regulatory Agency announced guidance, developed with NHS England, on how medical device law applies to ambient voice technology, the NHS term for these tools. It confirms that products "intended solely for transcription, summarising of clinical conversations, drafting letters, or suggesting clinical codes for a clinician to review are not regulated as medical devices under the current framework". Products intended to support diagnosis, treatment or prevention, or to take automated clinical action such as placing orders without clinician review, are regulated as medical devices.
Read the reasoning rather than the headline. A tool limited to transcription or summarisation can fall outside medical-device regulation because its intended purpose is administrative rather than diagnostic or therapeutic. That is the legal test: what the product is for, not whether a human happens to read it afterwards.
The checking requirement sits alongside that classification rather than underneath it, and the regulator states it directly: "Clinicians remain responsible for reviewing and verifying AI generated transcripts, summaries and other outputs before they are used in patient care. This responsibility is unchanged by the guidance."
So for these administrative functions there is no medical-device approval regime providing the assurance. Assurance comes from somewhere else: NHS clinical risk management standards, compliance with which "is mandatory under the Health and Social care Act 2012", along with procurement, local governance and human verification of the output. Human verification remains a required safeguard even where the tool itself falls outside medical-device regulation.
The MHRA does not regulate your practice, and no equivalent body has said any of this about your file notes. But you do not need it to. If you are a solicitor, accountant or other professional adviser, your existing professional, evidential and often regulatory reasons for keeping an accurate record of what you advised do not disappear because software drafted it. The tool has not moved them.
NHS England's rule, which you can borrow today
NHS England has already written the rule. Its patient-facing explanation of how ambient scribing is governed puts it plainly, and it gives professional firms a useful rule to borrow:
"Notes, summaries and letters that have been created by ambient scribes might sometimes contain errors. People using ambient scribes must check the notes or letters that have been made by the ambient scribe to make sure they are accurate. They do this before adding any information to your record and are responsible for making sure the information in your record remains correct."
Its letter to trust and integrated care board chief executives, published the same day as the MHRA guidance, says it in one line: "Users remain responsible for reviewing, validating and approving any information generated by AVT before it is relied upon for patient care." AVT is ambient voice technology, the NHS term for these tools.
That is the whole control. The output is a draft. Somebody checks it. Then it becomes the record.
Why a read-through is not enough, and what to check instead
"Just check it" is weaker advice than it sounds.
Look again at the four failures. The swapped drug name was catchable, but only because the patient had independent knowledge of what the correct drug should have been. Take that knowledge away and it reads like any other line.
The other three can survive a read. It is hard to notice an instruction that is not there by reading what is there, because an omission leaves no mark on the page. An invented line reads much like a real one.
And a dropped negative produces a sentence that is fluent, confident and says the opposite of what was said. The client says they will not be claiming research and development relief this year. The note says they will be. Nothing about that sentence looks wrong on the page, and a year later it is the only account of what was decided.
There is a second problem with reading, and it is the reason the check below starts before you open the document. There is a cognitive risk here: a later summary can influence what you subsequently recall, so reading the AI version first may make it harder to separate its wording from your own memory of the meeting.
So the check has to be targeted at the failure modes rather than general.
Start before you open the note. Write your own three bullets of what was agreed, from memory, while it is still yours. Then read the note against those. Without that step, most of the checks below are asking you to consult a memory the note has already shaped.
Then four things, in this order:
- Every name and figure. Parties, entities, statutes, amounts, percentages, dates, deadlines, notice periods, reference numbers. The swapped drug name was caught only because the reader knew independently what the right one was, so check these against a document, not against your recollection: the engagement letter, the papers in front of you, the transcript if the tool keeps one.
- Every negative, in both directions. Check every "not", "no", "won't", "isn't", "unless", "except" on the page against your bullets, then check your bullets for a negative the note has lost. Each is one short word that reverses the meaning of its sentence, and the missing ones are not on the page to be found.
- Every action. Who agreed to do what, by when. Then ask separately what was agreed that is missing, which is the only way to catch an omission.
- Anything that surprises you. A general read may not catch an invented line, because it looks like every other line. If a substantive statement surprises you, verify it against the transcript, the recording or contemporaneous material where you have them, rather than against your impression of it.
On a short note this may take only a few minutes. On a long one, where the figures have to be read against papers or a transcript, it takes longer, and it is worth knowing that before you promise it to the team. It is not a proofread; the prose does not need improving.
Write the rule down, as a rule about the document
Start by defining the document's status. If every AI note stays a draft until it is checked, the need for one named reviewer follows naturally.
Four lines, which you can put in a policy or in a team channel this week.
- Anything an AI notetaker produces is a draft. It is not a file note, an attendance note or a minute until line 3 has happened.
- A draft does not go to a client, onto the file, or into a system another decision reads from.
- One named person who was in the meeting writes the three bullets, runs the four checks, and puts their name on it. One person, not the team. If three people were in the meeting, one of them owns the note.
- If it is still unchecked after two working days, it is marked unverified and does not go to the client. Handle it from there under your own retention and record-keeping policy rather than letting it sit and eventually pass as a verified note. Whether an unchecked draft should be kept, corrected or removed depends on your professional obligations and what other record of the meeting exists, so that is a decision for your policy, not for this article.
That fourth line is what stops the rule becoming decoration, and the word "unverified" is the part doing the work.
The honest risk in writing a rule you do not keep
A control you have written down and routinely fail to perform can be worse than having no such control. If people initial notes without reading them, you have manufactured evidence that a check happened, and that evidence will be produced at the exact moment it hurts.
So keep it small enough to actually be done. Three bullets, four targets, one name.
Be careful with the obvious escape hatch. "Only where the note carries advice, money or a deadline" sounds like a sensible narrowing until you notice that, in an accountancy or law practice, it may cover most substantive client calls. If you cannot sustain the check everywhere, narrow it by something that actually divides your work: notes that go to a client or a third party get the full check, purely internal notes get the three bullets and nothing else. Write the narrower rule down. A rule that quietly lapses is the one that will be produced against you.
There is one more thing this article deliberately does not cover, and it may be the larger exposure for a small firm: the recordings and transcripts themselves, where they are stored, who can reach them and how long you keep them. That is a different problem with a different answer, and it is in the piece on letting notetakers into the meeting in the first place.
What the public wants from NHS scribes, and what firms can take from it
Healthwatch asked. Over two-thirds of people, 69%, said they would be more comfortable with these tools if there were "a clear commitment from the healthcare professional to check the accuracy of the content the AI scribe produces". Four-fifths, 81%, wanted to be told the tool was being used and asked before it was. Of those who had had an appointment in the previous year, nearly 90% said they had not been aware of AI scribing being used. That is not evidence that it was used on them without their knowledge; in most of those appointments it may simply not have been used.
Read that as a client-relations finding rather than a health one, and read it accurately: the public is divided. Healthwatch found support and opposition close to level, at 38% and 37%, with almost twice as many strongly opposed (21%) as strongly supportive (11%). What moves the number is visible safeguards. People want to be told the tool is running, and they want to know a person is still accountable for what it writes. Disclosure costs almost nothing. Human review costs time, and it is part of using the tool responsibly.
FAQ
Does this mean we should stop using AI notetakers? No. There is good evidence that they can save time, although the size of the saving varies by setting and by how much checking and editing the output needs. The NHS is expanding their use rather than pulling back: NHS England's ambient scribing service page says organisations are "directed to deploy AVT at pace". The point is that the output is a draft, and the saving is smaller than the raw time it appears to save, because the check is part of the job.
Will a better model fix this? Better models should reduce some kinds of error. They will not remove the check, because NHS guidance still requires verification, and because even a low error rate matters when the error is hard to detect. There is also a risk that more consistently fluent output encourages more trust in the occasional mistake.
Who should do the checking? Someone who was in the meeting, named individually rather than as a team, and ideally the person whose advice is recorded in it. Someone who was not there is less well placed to spot omissions or distortions that are not apparent from the text itself.
Is it enough to send the note to the client and ask them to confirm? It is a useful second check and it is not a substitute. It should not transfer responsibility for your own record to the client, and a confirmation you do not trust is worse than none, because it goes on the file. Check it first, then send it.
What if we find an error after the note is on the file? Correct it, date the correction, and keep the original readable. A visible correction trail is ordinary good practice, and a correction nobody can see is not much of a correction.
If you want a second pair of eyes on how AI is writing things down inside your firm, book a conversation.
This is general information about record-keeping practice, not legal advice.