Dark teal cover with a single-page checklist motif and the Good Transformer wordmark, for an article on what a firm does in the first hour after an AI tool gets something wrong.
AI governanceRisk managementComplianceLeadership

What to do in the first hour after AI gets something wrong

Your AI policy tells people how AI should be used. A one-page incident plan tells them what to do when something has already gone wrong.

Good Transformer14 min read

AI mistakes used to mean bad text. Increasingly they mean actions taken on email, files, records and live systems.

That changes what a firm has to be ready to do. Sometimes you have hours to correct an answer. Sometimes you have minutes to stop a tool.

Government security agencies published guidance on this in spring 2026. On 1 May the NCSC and five partner agencies published joint guidance on adopting agentic AI telling organisations to "develop and test incident response procedures to detect, contain and recover from agent compromise".

Two weeks later the NCSC told firms to be clear about "who owns an agentic system, who approves its access, who monitors its behaviour, who reviews incidents, and ultimately who can stop it".

This article is general information, not legal advice.

A law firm had a policy. It did not help.

On 22 May 2026 the High Court handed down a judgment that every firm using AI on client work should read. A junior associate at a large law firm used an AI tool to research whether a court could release liquidators on a particular kind of insolvency application. The tool produced a passage set out as a direct quotation from the Insolvency Rules.

In the court's words, "the Purported Text was offered as a quotation from statute. That quotation did not exist and nor did anything like it exist."

The firm already had a written AI policy, and the judgment quotes it. It warned that AI can produce output that appears believable but is "highly inaccurate, outdated or entirely fabricated", it required users to apply critical thought and to "fact and sense check" all outputs, and it said AI systems "must be supervised by humans".

So the first failure was not a missing policy. It was a failure to follow one.

Two details are worth more than the error itself.

The tool warned the associate. It said plainly that it could not verify the wording of the current rule from a primary source, and it recommended checking the rule as published on legislation.gov.uk before filing. Nobody acted on that warning. The supervising solicitors, on the court's findings, did not know AI had been used at all, so had no reason to check anything.

What happened next is the part a policy cannot reach. Once the court pointed out that the quoted rule did not exist, the firm needed someone to stop, establish what had happened, check the primary source and correct the record. Instead AI was used again to help construct a reply. The judge found that "an opportunity to set the record straight became a further instance of misleading information being put before the court", and that once the court had identified the problem the supervising solicitor "should have established what had happened". That did not happen.

The firm paid its former clients' costs and referred itself to the Solicitors Regulation Authority.

That gap is what an incident plan is for. A policy tells people how AI should be used. An incident plan tells them who takes control once those rules have already failed.

The two shapes of an AI incident

An AI incident comes in two shapes, and the first move is different for each.

The first is a wrong answer that has already gone to a client: a figure in a client report, a citation in a letter, a summary that got the instruction backwards. How fast you have to move depends on who has seen it, and the job is to correct the record.

The second shape is newer. A tool is still running. In April 2026 an AI coding agent at a small software business deleted its production database and its backups in about nine seconds.

On the founder's own account, reported at the time, the agent was supposed to be working on a staging problem. When it hit a credential mismatch it searched elsewhere and found an unrelated access token in another file. That token carried much broader permissions than its owner realised, including the ability to delete production data. The agent used it. The business was down for over thirty hours.

That is a software business, and the same shape in a professional-services firm rarely involves a database. It looks like an AI assistant with access to your mailbox sending a reply to a client before anyone reads it. It looks like an agent filing documents into the wrong client's folder, for an hour, while nobody is watching. The common feature is that the tool is acting on real systems and it has not finished.

Here you have minutes rather than hours, and what matters is what the tool can still reach. That is why the 1 May guidance asks firms to sort agent actions by "potential impact, likelihood and reversibility". It also asks them to keep a way of reverting a system to "known-good agent behaviours when unpredictability is observed".

A wrong answer has gone out A tool is still running
First move Preserve the original, stop it going further, correct the record Stop or isolate the tool, then find out what it did
The clock Hours, set by who reads it next Minutes, set by what the tool can still reach
Who acts Whoever owns the client relationship Whoever can stop or isolate the tool
The fix afterwards The check that should have caught it The permission or credential that allowed it

The page itself, ready to copy

Here is the whole thing. Copy it, overwrite the square brackets, and keep it alongside your AI policy (we have a simple one here).

AI INCIDENT PAGE                     Firm: [name]    Last agreed: [date]

WHAT THIS COVERS

Use this page when an AI mistake has escaped normal review and could
matter to a client, or when an AI tool has taken an unintended action
on live data or systems.

Two cases:
  A. A wrong answer or other bad output has already gone out.
  B. A tool or agent is still acting on files, email, records or systems.


START A RECORD

Time found: [ ]
Reported by: [ ]
AI tool or agent: [ ]

Preserve the original output, the prompt or chat if available, the
affected files and the relevant logs, before they are overwritten.
Record the actions taken from here.


1. WHO TAKES CONTROL

First: [name]. If away: [name].
Timescale: immediately, or same day depending on impact.

The report says which AI tool was involved, what appears to have
happened, and who or what may be affected.

Nobody is penalised for reporting an AI mistake. [Say this and mean it.]


2. WHO CAN STOP IT

People who can stop or isolate a tool or account: [name], [name].
Our route: [saved admin link / internal procedure / IT provider]
Out-of-hours contact: [number]
Emergency-admin route: [location or process]    Last tested: [date]

For tools that can send, edit, delete or move real information, we do
not run them unattended unless we can contain them out of hours.


3. WHO DECIDES WHO ELSE NEEDS TO KNOW

Decision maker: [name].
Client contact: [name].
Broker or insurer: [details].
Other escalation: [legal / data protection / regulator where relevant].

If incorrect client work has gone out and matters, correct it promptly.
Decide separately whether the circumstances, including the use of AI,
also need to be explained.

If personal data may have been exposed, altered, lost or made
unavailable without permission, escalate it immediately as a suspected
data breach. If it is reportable, the ICO deadline is 72 hours from
awareness.

Check any insurer or broker notification requirements promptly and in
parallel. Do not delay urgent containment, correction or a legal
reporting deadline while waiting for the broker.


4. WHAT WE CHANGE AFTERWARDS

Owner: [name].    Reviewed within: [five working days].
Recorded here: [location].

Fix the control that failed: a permission, a check, an approval step, a
credential or a logging rule. A better prompt can support the fix, but
it should not be the only response to a material incident.


CASE A FIRST MOVE
Preserve the original. Stop further use or distribution. Work out who
received it, and correct or escalate it.

CASE B FIRST MOVE
Stop or isolate the tool first. Preserve the relevant logs. Then
establish what it accessed or changed.

The four questions, explained

1. Who takes control, and does the report say AI was involved?

Name one person who hears about it first, and a second person to take over when the first is away. Then add the line the law firm above needed: whoever reports it says whether AI was used, and which tool. Supervisors who do not know AI was involved have no reason to check anything.

2. Who can stop it, and can they do it on a Saturday?

This is the question most firms cannot answer. Two people should be able to stop or isolate an AI tool and cut its access to your files and email, and both should have found the relevant screen once while nothing was wrong. Finding it is the point, not pressing it. If you want to test a stop control properly, test it on something nobody depends on.

As of July 2026, in Microsoft 365 the list of agents sits under Agents, then All agents, then Registry, and the quickest way to stop one is the Block button. That same documentation rates an agent with no owner on record as a critical risk, which is a fair indication of how much the ownership question matters. In Google Workspace, one screen holds both the list of apps that can reach your mail and files and the button that blocks one, under Security, then Access and data control, then API controls. Google notes that the list can be a day or two out of date, so block the app first and check the list afterwards. Admin interfaces move, so check the current path when you fill the page in.

If your IT is outsourced, this line is a contract question rather than a technical one. Do not assume your support contract includes weekend action. Check it.

Work out the emergency route now: named internal administrators, a properly secured break-glass process, or a contracted out-of-hours route through your provider. Microsoft treats top-level administrator access as highly privileged and recommends tightly controlled emergency accounts, so this is worth setting up properly rather than handing two people permanent keys.

One thing not to write down is that you will wait until Monday. For a tool that can send, edit, delete or move real information, that is not a containment plan. The NCSC is unusually blunt about it: "If you cannot understand, monitor or contain an agent's actions, it is not ready for deployment." Either create an out-of-hours route, or restrict the tool so it cannot keep acting unattended.

3. Who decides who else needs to know?

This is the decision firms put off, so write down who makes it. There are two questions and they are separate: correcting the work, and explaining the circumstances.

If incorrect work has reached a client and matters, correct it promptly. Then decide, separately, whether the circumstances including the use of AI also need explaining. The Institute of Chartered Accountants of Scotland, writing about this case on 2 June 2026, is clear on the general posture: if an error is identified "it should be investigated promptly and addressed openly", and trying to rationalise or minimise it risks making matters worse.

On personal data, be careful in both directions. Personal data being somewhere in an incident does not by itself make the incident a reportable breach. What matters is whether personal data may have been exposed, altered, lost or made unavailable without permission. If it may have been, escalate it immediately as a suspected data breach, and if it is reportable the deadline is 72 hours from when you became aware.

On insurance, check now rather than later whether your professional indemnity and cyber cover respond to an error caused by AI. Aon reported in May 2026 that insurers are handling this in three different ways, one of which is wording that narrows the cover a firm believed it had. Ask your broker what the policy covers and what it requires you to tell them, while nothing is wrong. During an incident, make that call in parallel with everything else, and never let it delay a correction or a legal deadline.

4. What do you change afterwards?

Fix the control that failed. That is usually a permission, a check, an approval step, a credential or a logging rule.

The law firm needed a check: verify any quoted authority against the source before it goes out, and say when AI has been used. The software business needed a permission: no broadly privileged production credential sitting within an agent's reach. A better prompt or instruction can support the fix, but it should not be the only response, and the April case shows why. The agent reportedly had explicit instructions against destructive behaviour and went ahead anyway.

What your tools will not tell you afterwards

The 2026 guidance is good on ownership, permissions and logging, but it does not give a small firm a step-by-step order for the first hour. That is why the page above is four questions rather than a framework.

It is also worth knowing what your own tools record. OpenAI's admin documentation for ChatGPT Work says its compliance log covers prompts and responses, and "doesn't track files, actions, or tool calls". It keeps those prompts and responses for thirty days.

So the record of what an agent actually did to your documents comes from the system it touched, not from the AI tool. That is exactly why the page starts with preserving evidence rather than fixing things.

Many firms already have part of this plan

The Cyber Security Breaches Survey 2025/2026, published on 30 April 2026 by the Department for Science, Innovation and Technology and the Home Office, found that 25% of UK businesses have a formal incident response plan. Among micro businesses the figure is 21%, against 76% of large businesses.

The next numbers are more useful. 39% of businesses have already assigned incident roles to individuals, and 34% already have written guidance on who to notify. So more firms hold a piece of the plan than hold the plan itself. For those firms the answers exist already, in different people's heads rather than written down.

What to do this week

Open the two screens while nothing is wrong: your agent or app list, and the place you stop a tool or revoke an account's access. Copy the page above and fill in the brackets with real names. Then send it to whoever holds admin access, including your IT provider if that is someone else.

Filling it in takes about an hour. Be honest that two of the lines may take longer, because they are other people's decisions: the out-of-hours answer from a provider, and the question to your broker. Do not wait for either before you write the rest down. A page with two lines still open beats no page, and it tells you exactly which two calls to chase.

Read it alongside our guide to stopping mistakes reaching a client at all, and our guide to what happens when a tool you rely on disappears.

If you would rather work through it with someone, book a short call and we will help you map where AI now touches your client work and the checks that belong around it.

Frequently asked questions

Is an AI mistake something you have to report to a regulator?

Usually not by itself. There is no general duty to tell anyone that an AI tool produced a bad answer.

A duty is triggered by the consequence, not by the AI. If personal data may have been exposed, altered, lost or made unavailable without permission, that is a suspected data breach and the ordinary rules and timetable apply. Professional obligations sit on top of that, and some regulators expect openness where client work was affected. Take advice on your own position.

What if the AI tool is the one that spotted the mistake?

Treat it as a report like any other, and record that it came from the tool. In the case above the tool flagged its own uncertainty and recommended checking the source, and nobody acted on it. A plan that only starts when a person notices the mistake will miss the warnings the tool itself gives you.

How is the one-page plan different from an AI policy?

A policy sets the rules for using AI: what people may put into which tool, what needs checking, what is off limits. It is written for a normal Tuesday.

The incident page is written for the day those rules did not hold, and it answers a different question: not what should have happened, but who takes control now. The law firm above had the first and not the second. Our guidance on letting an agent act on its own sits between the two.


This is general information, not legal advice. Where an incident involves personal data, a regulated activity or a professional obligation, take advice on your own position.

Work with Good Transformer

Turn this thinking into working practice.

Explore team advisory

Newsletter

Get new Insights by email

Practical notes on using AI with judgement, and the AI news leaders actually need. No hype, no spam, unsubscribe anytime.

Choose how often you want the digest

Keep reading