Dark teal cover with a node-and-edge motif and the Good Transformer wordmark, marking an article on checking whether AI tools train on your data.
AI data privacyData protectionAI toolsSmall businessAI policy

Is your AI tool training on your data? How to check

On a business plan, the major providers do not train on what your staff type in, but that promise does not cover a free account or a bundled assistant.

Good Transformer15 min read

On the main business plans, none of the major AI providers trains its models on what your staff type in. Two of them put that in a contract rather than a settings page.

That promise is narrower than it sounds. It does not cover the free account a member of staff signed up for, and it does not cover the assistant that came bundled with Microsoft 365 or Google Workspace, or the notetaker that joins meetings and writes them up.

It does not even cover every paid plan. Some, including Google Workspace Individual, are sold to solo businesses on consumer-style terms.

Those have their own data settings, on their own screens, and some are switched on by default. The table below gives the wording each provider publishes and the exact menu path.

Why this question changed this year

Two years ago the question was simpler. Somebody signed up for ChatGPT, typed a client's material into it, and a leader wanted to know whether that material was feeding the model.

That question still matters, and the answer to it has improved. What changed is where AI now sits, which is inside software the firm already pays for.

Microsoft 365 and Google Workspace now come with their own assistants built in. Teams, Zoom and Google Meet will write up a meeting without anyone installing anything. Agents read files and act on them, instead of only responding to what someone types.

None of that was chosen tool by tool. None of it has been checked tool by tool either.

Where to find the setting for each tool

Find the row for the plan your firm actually pays for, not the one with the product name you recognise. The business rows all say the same thing. The consumer rows do not.

Every row below was read off the provider's own page on 4 August 2026. Providers revise this wording without announcing it, so treat that date as part of the information.

Product and plan What the provider says about training Where the setting is
ChatGPT Free, Plus, Pro ChatGPT "improves by further training on the conversations people have with it, unless you opt out" Profile icon, then Settings, then Data Controls, then "Improve the model for everyone". Turning it off "applies to your entire account"
ChatGPT Business, Enterprise, Edu, and the API Data from these "isn't used for training our models, unless you have explicitly opted in" Nothing for an individual to switch. It is set by the plan
Claude Free, Pro, Max The Consumer Terms say Anthropic may use your material to improve the services, "including training our models, unless you opt out of training through your account settings" Settings, then Privacy, then "Help Improve our AI models"
Claude Team, Enterprise, and the API A term of the contract: "Anthropic may not train models on Customer Content from Services" Nothing to switch. An owner can turn off "Rate chats" under Organization settings, then Data and Privacy
Gemini in a personal Google account Activity is used "to provide, develop and improve its services (including training generative AI models)", and "Keep activity is on by default" for anyone aged 18 or over gemini.google.com, then Activity
Gemini on a qualifying Workspace edition Content "is not human reviewed or otherwise used for Generative AI model training outside your domain without permission" Administrator-controlled. Users cannot change it
Microsoft 365 Copilot "Prompts, responses, and data accessed through Microsoft Graph aren't used to train foundation LLMs" No training toggle for a user. The setting that does matter is web search, below

Anthropic publishes no default state for its consumer toggle at all. It makes you choose during signup, and the contract says training happens unless you opt out. So the honest answer to "is it on?" for a member of staff on a personal Claude account is that nobody outside that account can tell you, including us. Somebody has to open the setting and look.

Google is the opposite. It publishes the default, and the default is on.

What paying for the business plan changes

You stop relying on a promise that anybody in the firm can undo.

On the consumer tiers, the no-training position is only a setting. A member of staff can turn it off, and a member of staff can turn it back on, and nobody in the firm gets told either way.

On Claude Team, Claude Enterprise and the Anthropic API it is a term of the commercial contract. On a qualifying Google Workspace edition it sits in the service terms, which say Google "will not use Customer Data to train or fine-tune any AI/ML models without Customer's prior permission or instruction". A settings page can be changed by whoever is holding the laptop. A contract cannot.

Paying for a business plan does not always buy the contract terms, and Google is the one to check. Google splits its editions by whether Gemini counts as a core service. On Business Starter, Business Standard, Business Plus and the Enterprise editions it does. Google's published wording for those is that chats and uploaded files "won't be reviewed by human reviewers or otherwise used to improve generative AI models".

On Workspace Individual, Essentials Starter and several others it does not. The wording for those editions is that chats "may be reviewed by human reviewers and used to improve" Google's products. A sole trader paying every month for Workspace Individual is on consumer-style terms and has no reason to suspect it.

Moving client work onto a paid plan is the most valuable single change a firm can make here, and the only cost is the subscription.

A paid plan also does nothing about the personal account someone set up without telling anyone. We set out a one-page rule for that in whether staff can put client data in ChatGPT.

Two Microsoft settings the training answer does not cover

Microsoft 365 Copilot does not train on your content, and Microsoft says so plainly on its privacy page. Two other settings sit outside that answer.

Web search. When it is enabled, Copilot reads the prompt, works out which parts need public information, and sends a short generated query to the Bing search service. It does not send the whole prompt or the document.

In Microsoft's worked example, a prompt about an internal clean energy strategy paper produces the query "Fabrikam clean energy policy announcements". In another of its examples, the query includes a person's name.

Microsoft acts as your data processor on prompts and responses. On those generated search queries it says that it "acts as a data controller", and that its Data Protection Addendum "doesn't apply" to them. The default runs the other way. If an administrator has not configured the "Allow web search in Copilot" policy, web search is available to users.

The obvious switch does not help. Microsoft says that the privacy setting for optional connected experiences, the one a user can see inside Word or Excel, "has no effect on the availability of web search".

Recording consent in Teams. There is a policy called "Require participant agreement for recording and transcription". When it is on, participants are asked to agree before they are recorded or transcribed. Microsoft's settings table records the default as off, so participants are not asked unless an administrator turns it on.

Both defaults are set for easy adoption rather than for a firm holding other people's confidential material. Both take an administrator a few minutes to change.

Meeting notetakers record your clients too

A chat box holds what your staff chose to type. A notetaker holds what a client said, in the client's own words, often without the client thinking about it.

So consent matters more here than the training question does.

Notetaker Does it train its own models on your content?
Teams intelligent recap No. "Customer data isn't logged or used for any AI model training or testing"
Zoom AI Companion No. Zoom "does not use any customer audio, video, chat, screen sharing, attachments, or other communications-like customer content to train Zoom's or its third-party artificial intelligence models"
Google Meet automatic notes No. Gemini in Meet "doesn't use your content to train or improve Gemini or other generative AI models"
Otter.ai Yes, on de-identified content. Otter "uses a proprietary method to de-identify user data before training our models"
Fathom Yes, on de-identified content. Fathom "may use and create de-identified data generated from Meeting Content Information to improve our Services by training, improving, and customizing our in-house artificial intelligence models"

Neither Otter nor Fathom is doing anything hidden. Both say so in their own published policies, and de-identification is a real safeguard. It is simply not the same answer the suite vendors give, and a firm that assumed all notetakers behave the same way is wrong.

Fathom also puts the consent duty back on the customer, in terms a professional reader will recognise: "Please make sure you have the necessary permissions and consents from other meeting participants before using our Services". Our piece on AI notetakers in client meetings covers what to say and who checks the written record.

What "we do not train on it" does not mean

It does not mean nobody reads it. Google tells consumer users that "a subset of chats is reviewed by human reviewers (including Google's trained service providers)". OpenAI tells consumer users that a limited number of authorised personnel and trusted service providers may access their conversations for security investigations, for support, for legal matters and to improve model performance.

Both then give the same advice, in almost the same words. Google: "Please don't enter confidential information that you wouldn't want a reviewer to see." OpenAI: "Please do not enter sensitive information that you would not want reviewed or used."

It does not mean that deleting a chat removes it. Google is explicit that chats which have been through human review "are not deleted when you delete your activity", and are kept for up to three years. Anthropic keeps consumer material in its training pipelines for up to five years where the setting was on, and thirty days where it was not.

Microsoft publishes no retention period for Microsoft 365 Copilot. Prompts and responses are stored in a hidden folder in the user's own mailbox, and the firm's own retention rules apply to them.

Microsoft warns: "Messages visible in your AI apps are not an accurate reflection of whether they are retained or permanently deleted for compliance requirements." Under a one-day deletion rule, Microsoft says a message "could take 16 days" to be deleted, and permanent deletion "is always suspended" where a legal hold applies.

A court can also override a provider's deletion practice. In May 2025 a United States court directed OpenAI to preserve output logs that would otherwise have been deleted, including material deleted at a user's own request. That direction covered consumer and Team accounts and did not cover ChatGPT Enterprise.

It was terminated as of 26 September 2025. OpenAI kept the logs it had already set aside before that date, apart from those originating in the United Kingdom, the European Economic Area and Switzerland, which the same order excluded.

That is not United Kingdom law and it is not in force now. But for about four months, a vendor's ordinary deletion practice was suspended by an order the vendor did not control, and its customers found out afterwards.

This is background on why the training answer carries legal weight, not legal advice on your own position.

There is no United Kingdom statute governing artificial intelligence. The obligations that apply are the ones a firm already has under UK GDPR and the Data Protection Act 2018.

Under those rules, a supplier that only handles your material on your instructions is your processor. Article 28(10) says that a processor which decides the purposes and means of processing "shall be considered to be a controller in respect of that processing".

The Information Commissioner's Office has addressed this directly. In the outcomes report from its consultation series on generative AI, it rejected the argument that a developer improving its own models on a customer's data is still a processor: "We do not accept this argument. It contradicts our established position that if a developer is processing data for their own purposes, they are a controller for that processing."

The answer to the training question therefore affects more than how comfortable a firm feels. Where the answer is yes, the provider is deciding what to do with that material for its own reasons.

Two limits are worth stating, because this is easy to over-apply. The Information Commissioner's Office is stating its interpretation and expectation, not the statute. The ICO's own guidance on AI and data protection still carries the line "This guidance was updated on 15 March 2023". A note says it is under review following the Data (Use and Access) Act.

Your professional regulator may ask more of you than data protection law does. In November 2023 the Solicitors Regulation Authority described the risk of "a staff member using an online AI, such as ChatGPT, to answer a question on a client's case".

The Institute of Chartered Accountants in England and Wales has told members that where information going into an AI tool includes client data, "it's very likely that this will breach client confidentiality".

For what the law asks of a firm more broadly, our piece on GDPR and AI is the fuller answer. This one is about finding the settings.

Three questions to ask about any tool you are about to add

Settings move. Products get renamed and menus get rearranged, so the table above will be out of date before long. These three questions will still work, including on a tool nobody has written a guide for yet.

  1. Does the no-training position come from the contract or from a settings page? A term you signed still applies after someone changes a setting. A default does not.
  2. How long is your material kept after you delete it, and what suspends that deletion? Ask for the number, then ask what overrides the number. Legal holds, safety review and human review can all keep material longer than the published figure.
  3. Who else handles it? Every major provider publishes a sub-processor list naming the other companies that touch your material. Almost nobody reads it.

If the tool is one you are about to buy rather than one already running, our fuller version is the four things to check before you buy an AI tool.

What to do this week

Make a list of every AI tool anyone in the firm uses. Include the ones nobody chose. Think of the assistant inside the suite, the notetaker that joins meetings, and the AI features Microsoft and Google have added to software the firm already pays for.

Against each one, write which plan it is on, what its data setting is today, and the date you looked.

That list does two jobs. Clients increasingly ask how you handle their information, and the list is the answer. It is also the thing nobody can reconstruct from memory a year later. When a client does ask, what to say when a client asks how your firm uses AI sets out the answer.

Where anyone is doing client work on a free account, move them onto the paid plan. With Anthropic and Google, it turns a setting into a contract term.

If you would rather go through your own tools with someone, that is what our AI Lessons for Leaders sessions are: an hour on the tools your firm actually uses and the settings inside them.

Book a discovery call if that would help.

Common questions

How do you find out whether a tool trains on your data?

Go to the provider's own policy page rather than a summary of it, and check the plan you are on rather than the product name. The wording differs between a free account and a business account on the same product, and the business wording is usually on a separate page aimed at administrators.

Does turning the training setting off remove what is already there?

Not necessarily. Anthropic says the five-year retention in its training pipelines applies to new or resumed chats after the setting is switched on. Turning it off stops new material going in. It does not delete what is already there.

Is a paid plan enough on its own?

Usually it is, for the training question. Google is the exception to check, because some paid Workspace editions treat Gemini as an additional service on consumer-style terms.

What about the assistant that came bundled with software the firm already pays for?

Check it separately. Copilot's training answer is no, but two settings sit outside it. Web search is available unless an administrator has turned it off. The privacy switch a user can see in Word or Excel does not control web search.

Do you need to tell clients which AI tools you use?

That depends on your engagement terms and your professional obligations rather than on any general rule. What helps in every case is being able to answer accurately when asked, which is the reason for writing the list down.


This article is general information, not legal advice. Take proper advice on how these rules apply to your firm.

Work with Good Transformer

Turn this thinking into working practice.

Explore team advisory

Newsletter

Get new Insights by email

Practical notes on using AI with judgement, and the AI news leaders actually need. No hype, no spam, unsubscribe anytime.

Choose how often you want the digest

Keep reading

Client confidentiality22 min read

Keeping client confidentiality when you use AI

Client information may only leave the firm if the law allows it or the client agrees. An AI tool is run by another company, so that agreement belongs in your engagement letter.

30 July 2026

AI procurement15 min read

The four things to check before you buy an AI tool

Before you sign up to any AI tool, check four things: whether it trains on your data, where it is stored, what it can do once connected, and how you leave. The catch nobody mentions: at the tiers small firms buy, no one answers your email, so the answers are in the vendor's own terms.

22 July 2026