Dark teal cover with a node-and-edge motif and the Good Transformer wordmark, marking an article on the four things to check before you buy an AI tool.
AI procurementAI governanceData protectionSmall business

The four things to check before you buy an AI tool

Before you sign up to any AI tool, check four things: whether it trains on your data, where it is stored, what it can do once connected, and how you leave. The catch nobody mentions: at the tiers small firms buy, no one answers your email, so the answers are in the vendor's own terms.

Good Transformer15 min read

Before you sign up to any AI tool, pin down four things: whether it trains on your data, where that data lives, what the tool can do once it is connected to your systems, and what happens to your data when you leave.

Those four cover the main risks to your data. They are not the whole of buying well (a note below the checklist covers what they leave out), but they are the four that catch small firms out.

Here is the part most guides skip, though.

You are not going to ask them. Not really. A giant AI lab does not put a person on the end of the line to answer a small firm's questions. You sign up on a web page, click to accept the terms, and that is the deal.

So the job is not to interrogate the vendor. It is to find the four answers in what they have already published, save them, and decline the ones that do not add up.

One of the four now matters more than it used to. In 2026 an AI assistant does not just draft text. Connected to your inbox and your shared drive, it can send, move and change things on your behalf. So the thing to check hardest is less about what the tool reads and more about what it is allowed to do.

You are not going to get a phone call

It helps to be honest about how buying these tools works, because most procurement advice pretends otherwise.

For a firm of three to fifty people, adopting ChatGPT, Claude, Copilot or Gemini means going to a website, picking a plan and clicking to accept. There is no account manager, and no number to ring where someone talks you through the data terms. The contract is the standard one on the page, take it or leave it, the same for you as for everyone else on that plan.

That is not a scandal, it is how these products work. They are self-serve, sold to millions, and answering bespoke questions from every buyer would not scale.

But it changes what "get it in writing" can mean. It cannot mean a reply to your email, because you will not get one. It means the vendor's own published terms: the trust page, the privacy notice, the sub-processor list. Saving a dated copy of the relevant page is your record.

You do genuinely get to ask a human in two situations. One is an enterprise deal, where the spend justifies a salesperson and a signed contract with its own data terms. The other is buying through a reseller or IT provider, who answers to you even when the lab will not. If you are neither, you are reading the published pages, and this piece is about reading them well.

The four things to check, and where to find each answer

Here is what to look for, why each matters, and where each vendor tends to publish it. Ask a human only in the two cases where there is one to ask.

1. Do they train on your data?

This is the question everyone asks, and on a paid business plan it is usually already answered, in your favour. The business and enterprise tiers of the mainstream tools state that they do not train on your data by default.

OpenAI says it does not use business, enterprise or API data, inputs or outputs, to train its models. Microsoft says the prompts and data its Copilot reaches across Microsoft 365 are not used to train the foundation models. Google says Workspace does not use customer data to train its models without your instruction. Anthropic says the same for its commercial plans: on Team, Enterprise and the API, your prompts and results are not used to train its models by default.

The catch is the tier, and Claude is the sharpest example. On a paid business plan, Anthropic does not train on your work. On the consumer Free, Pro and Max plans, since a change in August 2025, you are asked to choose at sign-up, with the sharing switch already set to on. Leave it on and your chats can be used for training and kept for up to five years instead of thirty days. Same brand, opposite default.

So the thing to check is never "does this company train on data" in the abstract. It is "on the exact plan we are buying, and where do their terms say so". The plan is the contract.

2. Where does the data live, and who else can reach it?

Check where the data is stored and who the vendor passes it to. This one has more variation than firms expect, and it is where a good-looking tool can still be wrong for you.

Some vendors let eligible business customers keep data in the UK or Europe. OpenAI, for example, now offers UK and European data residency on its enterprise and API tiers, though it applies to newly created workspaces and covers where data is stored, not every step of processing. Others do not. Claude is a useful warning here: the first-party Claude enterprise product does not offer UK or European data residency of its own. You get regional residency only by running Claude through a separate cloud arrangement, such as AWS or Google Cloud in a European region.

That may be fine for your firm, or it may not, depending on what you handle and what you have told your own clients. The point is to know before you sign, not after.

The second half is sub-processors, the other companies a vendor relies on to run the service. The major vendors publish these, Anthropic's list sits at anthropic.com/subprocessors, and a vendor is meant to bring new ones in only with your authorisation and to tell you when the list changes. Find the list, and check it against the promises you have made your own clients.

3. What is the tool allowed to do once it is connected?

This is the 2026 question, and the one to spend most time on. Once an assistant is wired into your systems, what can it actually do? Can it only read, or can it also send, delete and change? Does it ask before it acts, or act and tell you afterwards?

The better tools now let you set this, and the setting is worth finding before you switch anything on. In ChatGPT's business and enterprise apps, for example, an administrator can hold a connected tool to read-only, or to a chosen set of actions, and make it ask before doing anything with real effect. Set these tight to start with: read-only where you can, ask-before-acting for anything that goes to a client or outside your firm. You can always loosen them later.

In the security world this maps to two entries on a well-known list, OWASP's Top 10 for large language model applications: excessive agency, meaning a tool that can do more than it should, and sensitive information disclosure. More on that list below.

4. What happens to your data when you leave?

The last thing to check is the one nobody looks at and the one that catches firms out. You can leave a vendor, but your data may not. Look for what happens when you stop: is your data deleted or returned, how quickly, and what about copies in backups or anything used to tune the service.

Here the published terms often do answer well. Anthropic, for instance, offers a zero-retention option, where inputs and outputs are not stored at all beyond what safety and law require, though only for eligible customers on its API and Claude Code enterprise products, not the ordinary paid plans. UK law backs this up: when a vendor handles personal data for you, the contract is meant to give you the choice to have it deleted or returned at the end.

If the published terms cannot tell you what leaving looks like, treat the silence as the answer. Our note on what to do if a vendor fails or is bought covers the same ground for the day you have to move in a hurry.

The four-question check

Because you are mostly reading rather than asking, here is the check as a table you can fill in from the vendor's own pages before you commit. One tool, ten minutes, one saved and dated screenshot per row.

Check What a good answer looks like Where to find it Walk away if
Training "Not used to train, by default" on the exact plan you are buying Trust or privacy page; business-data or enterprise-data page The no-training promise only applies to a tier above the one you are buying
Storage You know the country, and it suits what you handle; sub-processors are listed Trust centre; data-residency page; the sub-processor list You cannot find where the data is stored at all
What it can do You can hold it to read-only or ask-before-acting when connected Admin or security docs; the app's permission settings It connects to your systems with no way to limit what it does
Exit Data deleted or returned on cancellation, in a stated time Privacy or data-retention page; the DPA if there is one No statement on deletion, and no one to ask

The four questions are deliberately narrow. They do not tell you whether the vendor is reliable or financially sound, what its security track record is, who owns what the tool produces for you, or whether your industry adds rules of its own. A firm regulated in finance or health has extra homework. And if the tool will help make decisions about people, do a proper data-protection impact assessment before you start. Treat the four as the minimum, not the whole job.

If you are the exception who does have a salesperson or a reseller, the same four become an email worth sending: on our plan, do you train on our data and where do your terms say so; where is it stored and who are your sub-processors; can we hold the tool to read-only; and on cancellation, is our data deleted or returned and how fast. A written reply, kept on file, is the enterprise version of the check above.

The security list worth borrowing

If you want a second opinion that is not the vendor's own, the OWASP Top 10 for Large Language Model Applications is a free, widely used list of the main ways AI tools go wrong. The current version is the 2025 edition; some sites relabel it 2026, but it is the same list.

Most of it is written for the people who build these tools. A handful is for you, the buyer. The two rows to weigh for any tool are the first two.

Risk (OWASP, 2025) In plain terms Yours to weigh?
Sensitive information disclosure The tool reveals data it should not Yes
Excessive agency The tool can do more than it should Yes
Supply chain A weakness reaches you through the vendor's own suppliers Yes, check who they rely on
System prompt leakage The tool's hidden instructions leak out Worth noting
Poisoning, output handling, embeddings How the model is built and wired together Mostly the vendor's; yours is to check outputs before relying on them

There is a reason this list matters more than it used to. The security researcher Simon Willison calls the dangerous combination the lethal trifecta: a tool that has access to your private data, is exposed to content from outside your control, and has a way to send information out. Put those three together and a booby-trapped email or web page can, in theory, steer the tool into doing something you never asked. That is the risk the read-only setting in check three is there to contain.

Where the law actually sits

None of this needs a statute to justify it, but it helps to know the law is behind you. In the UK there is no dedicated AI Act. The rules that apply are the ones you already have: UK data-protection law, the Data Protection Act 2018, and the guidance of the Information Commissioner's Office. Those rules are being updated rather than replaced. The Data (Use and Access) Act 2025 amended them, and the ICO has been directed to produce a statutory code of practice on AI and automated decision-making. None of that changes the four checks; it makes them more clearly the baseline.

When a vendor handles personal data on your behalf, the law treats them as your processor and expects a written contract that covers, among other things, security, the use of sub-processors, and deleting or returning your data at the end. For a self-serve tool that written contract is the vendor's standard data-processing terms, which is exactly why you save a copy of them. The major vendors build these in: OpenAI publishes a data-processing addendum, and Anthropic's commercial terms include one with the EU's standard contractual clauses attached, so "no one to ask" does not mean "no contract". The ICO's own guidance for AI warns that a verbal-only arrangement leaves you with no recourse if something goes wrong.

The EU's AI Act, which reaches its main application date this summer, is a separate matter: it can touch a UK firm through EU clients or operations, but it does not create these duties for you. Your obligations run through UK data-protection law, and the check above is how you meet them in practice.

This connects to a related topic, the AI wording in your own client contracts: one check protects you as a buyer, the other protects you as a supplier.

Where this fits

Getting this right is mostly about knowing what to read and being willing to walk away, not about legal expertise or negotiating skill you do not have. The reason it matters is that most AI tools still arrive without any of this being checked.

One recent industry report from the security firm Netskope found that most people using generative AI at work were doing so through personal, unmanaged accounts, outside any procurement check at all. A ten-minute read of the vendor's terms, saved and dated, is what moves a tool from that pile into one you can stand behind.

Sometimes you cannot walk away, because the tool is the one your clients already expect you to use. The check still earns its keep. Buy the safest tier, turn training off, keep your most sensitive data out of it, and keep a note of what you decided and why, so there is a record if a client or a regulator ever asks.

If you would like help turning this into a simple step your team runs before adopting any tool, that is the practical work our AI Lessons for Leaders sessions cover. The next time a tool lands with a free trial and a friendly-looking sign-up button, run the four-question check first, then decide.

What should you check before signing up to an AI tool?

Check four things: whether the tool trains on your data on the exact plan you are buying, where that data is stored and who else can reach it, what the tool is allowed to do once it is connected to your systems, and what happens to your data when you leave. For a small firm on a self-serve plan you will not get a human to answer these, so the answers come from the vendor's own trust and privacy pages, which are worth saving with a date.

Do AI tools like ChatGPT, Claude or Copilot train on your business data?

On the paid business and enterprise tiers, the mainstream tools state that they do not use your data to train their models by default. The consumer tiers of the same brands often lean the other way; Claude's consumer plans, for example, ask you to choose at sign-up with the sharing switch already set to on. This is why the plan you buy matters as much as the brand, and why it is worth reading where the vendor's terms say so.

Can a small business negotiate AI vendor terms?

Usually not. At the self-serve plans most small firms buy, the terms are standard and fixed, and there is no one to negotiate with. Your only real say is whether to buy and which tier to buy, not the wording. You get to ask questions and agree bespoke terms on an enterprise deal, or when you buy through a reseller who answers to you.

What is the most important thing to check in 2026?

What the tool is allowed to do once it is connected. Older AI tools mostly read and drafted; this year's assistants can act inside your systems, sending and changing things. Finding the controls that hold a tool to read-only or ask-before-acting, and setting them tight to start with, is the check that has grown most important.

Sources


This is general information, not legal advice. Where an AI tool will handle personal data or client information, take advice on your own situation before you rely on it.

Work with Good Transformer

Turn this thinking into working practice.

Explore team advisory

Newsletter

Get new Insights by email

Practical notes on using AI with judgement, and the AI news leaders actually need. No hype, no spam, unsubscribe anytime.

Choose how often you want the digest

Keep reading

AI policy8 min read

The AI policy clauses most templates miss

A workable AI policy is short. But most templates skip the clauses that decide whether it protects you: what your tools are allowed to do, firm limits on what data can go in, who owns the output, when you tell clients, and a standing training line. Each comes with copy-paste wording.

21 July 2026