Dark teal cover with the Good Transformer wordmark, for an article on AI agents that keep working on their own.
AI agentsFuture of workLeadershipAI strategy

Always-on AI agents and agent swarms are about to rebuild work around AI

Since August, OpenAI, Meta and SpaceXAI have launched AI agents that each have their own computer in the cloud, so they keep working after you close the app. Firms will get the most from them by handing an agent a whole job and setting its rules in writing.

Good Transformer14 min read

Since August, the biggest AI companies have launched agents that each have their own computer in the cloud, so they keep working after you close the app. In the business versions now arriving, each agent also gets its own login and a defined job inside the firm.

These agents are built to take a whole job off your hands, so firms will have to change how work is organised to get the most from them. You set the agent's rules and check its work.

The place to start is one recurring job, such as chasing unpaid invoices. Write down what the agent may see, what it can do alone and what it must bring to you. Later in this article there is a one-page sheet for writing those rules down.

What launched in the last eight weeks

On 11 August, SpaceXAI, the maker of the Grok assistant, launched Grok Bot. Its announcement describes bots that "have their own computer", sign into the tools you already use, and "finish jobs end to end, and only come back when something needs your approval" (SpaceXAI).

Meta followed on 8 September with Muse, a personal agent that runs in the cloud. Meta says Muse "keeps working after people close the app, and comes back when something changes or when it needs approval, like before it sends an email or makes a purchase" (Meta).

On 29 September, OpenAI launched agents it calls dots. They have "their own cloud computer", can work on your goals around the clock, and connect to more than 4,000 apps.

OpenAI gives one small example from an early tester. He had forgotten to invoice a publication. His dot noticed, prepared the invoice and sent it once he approved it (OpenAI, via the Internet Archive).

Microsoft is close behind. Its Autopilot agent, announced on 25 September, runs inside your firm's Microsoft 365 account "with its own identity, memory, computer and workspace" and keeps working "while you sleep". Microsoft said Autopilot would move into private preview at the end of September (Microsoft).

These products share a design. Each agent runs on a computer in the cloud, so it can keep working when you are not watching. It remembers what you have told it and works inside the apps you connect. And it stops to ask before anything sensitive, such as sending an email, spending money or publishing.

Not all of them can be used in the UK yet.

Agent Launched Can a UK firm use it today?
OpenAI dots 29 Sep 2026 Yes, on ChatGPT Business Premium. The individual Pro plan excludes the UK for now (OpenAI Help Centre, via the Internet Archive).
Meta Muse 8 Sep 2026 No. Meta says it is available in the US and Canada.
SpaceXAI Grok Bot 11 Aug 2026 SpaceXAI's announcements do not say; ask its sales team.
Microsoft Autopilot Announced 25 Sep 2026 Private preview only.

The business versions give each agent its own login and a defined job

The personal agents act for one person. The business versions are designed to work for the firm, more like a member of staff with a job description.

OpenAI is previewing what it calls specialist dots. "Your company sets up each dot with its own identity, credentials, and access to the systems it needs to complete its tasks." OpenAI says it has tested them internally on procurement, invoice processing, email marketing, customer support and commercial contracting, and is starting with "focused enterprise pilots".

SpaceXAI launched Grok Bot for Enterprise on 3 September. The company says "a Bot has no access by default and reaches only the accounts you sign it into" (SpaceXAI). Its customers include the legal technology company Legora.

Meta has also released Muse for small businesses. On 29 September it added connectors for tools such as QuickBooks, Shopify, Slack and Stripe, with the promise that "nothing publishes, sends, or spends without your approval" (Meta). It is not yet available in the UK.

With all of them, you do not sit with the agent and drive it. You give it a login, a job and rules about what needs your approval, and it gets on with the work.

Where firms will manage many agents

Once a firm has several agents, someone has to keep track of what each one can reach and what it has done. The large AI and software companies are building platforms for exactly that.

  • OpenAI Frontier, launched on 5 February, helps large companies "build, deploy, and manage AI agents". Each agent "has its own identity, with explicit permissions and guardrails". It launched with a limited set of customers, including HP, Intuit and Uber (OpenAI, via the Internet Archive).
  • Microsoft Agent 365, generally available since 1 May, is described by Microsoft as a "control plane to observe, govern, and secure agents", including agents from partner companies (Microsoft). OpenAI says it is working with Microsoft so that specialist dots can be managed there.
  • Anthropic's Claude Managed Agents, in public beta since 8 April, is a set of tools that developers use to build and run agents in the cloud (Anthropic). It needs a developer to set up.
  • Google's Gemini Enterprise lets staff create and manage agents in one place. Its Business edition, for teams of up to 300, costs from $21 per user per month, and Google says it needs no IT setup (Google Cloud).

Frontier and Agent 365 are aimed at large companies, Managed Agents at developers, and Gemini Enterprise's Business edition at teams of up to 300.

Our view is that a firm of 5 to 50 people will most likely manage its agents through the office software it already pays for, such as Microsoft 365. It will probably not buy a separate platform. Either way, agents will be managed much like staff accounts, each with a login, a list of permissions and a record of what it did.

Agents now work in teams, and in swarms

SpaceXAI says its own staff run "a chief of staff" bot that directs specialist bots for the inbox, expenses and recruiting. The bots "can independently message each other". Anthropic's Managed Agents lets "a lead agent break the job into pieces and delegate each one to a specialist" (Anthropic).

In September, OpenAI ran about 10,000 agents together to produce a proof of a famous maths problem. People still steered that run. OpenAI's Noam Brown gave most of the credit to the underlying model rather than to how the agents worked together. We covered it in our piece on what the Navier-Stokes result means.

Earlier this year, OpenAI agents being tested on a security exercise built their own message board without being told to. In July they used it to coordinate a real break-in to the computer systems of Hugging Face, a company that hosts much of the world's freely available AI software. Hugging Face detected the intrusion and disclosed it on 16 July. We set out what happened in our piece on the 2026 agent incidents.

On 29 September, OpenAI said it would not release its next agent model, GPT-6.1 Astra. Its head of safety systems, Saachi Jain, said the model fell short on "staying within scope and authorisation" (BBC News).

For a small firm, two things follow. First, agent teams work well when a job divides into separate pieces, such as research across many sources. They help little where each step depends on the one before. Second, the more agents you run, the more it matters that each has limits on what it can reach.

Why firms will have to change how work is organised

In 1990 the economist Paul David studied why electric power took decades to raise factory productivity. At first, factories replaced the steam engine with an electric motor and kept everything else, including the long shafts and belts that ran every machine from one power source (Paul David, American Economic Review, 1990).

The gains came when factories put a small motor on each machine. Then, David wrote, "factory structures could be radically redesigned", with lighter buildings and single-storey layouts.

Most firms that use AI are at the first stage: they use it to do their existing work faster. A March 2026 survey of 668 UK businesses, by the British Chambers of Commerce and the University of Essex, found that 54% of small and medium firms now use AI. In 2024 the figure was 25%.

The survey also found that general AI tools "tend to augment existing work, helping employees complete tasks more efficiently with minimal organisational change". The researchers say restructuring is more likely as firms move to "embedding more advanced AI systems in their core operations" (British Chambers of Commerce).

Agents that run a whole job are a step in that direction. They only pay off if someone has decided what the job is, what the agent may touch and where it must stop.

EY's audit business already works this way. It first released agents that answered research questions in 2025, then project management and administrative agents in spring 2026, then agents that draft audit work and support reviews over the summer.

Richard Harrison, an EY UK partner, told the accountancy body ICAEW: "Our agents are focused on very specific tasks, and operate within very clearly defined boundaries." Each auditor remains the "mandatory 'human in the loop'" with responsibility for reviewing the agents' work (ICAEW).

Ethan Mollick of the Wharton School writes that "work is increasingly about assigning work to agents, rather than working together with chatbots". He adds that "the best way to use agents is to think of yourself as a manager" (One Useful Thing). He is equally clear on the limits. Agents "should not decide by themselves to spend money, contact outsiders, access sensitive material", or take actions their managers did not authorise (One Useful Thing).

How a small firm could redesign invoice chasing around an agent

Take chasing unpaid invoices in a 15-person accountancy or architecture practice.

Today, someone checks the list of overdue invoices each week, decides who to chase, writes the emails and tells a partner about anything awkward. Much of what they know is in their head: which client always pays late but always pays, which one is in a dispute, which client contact is on holiday.

Fitting an AI assistant into that process means the same person asks a chatbot to draft the reminder emails. It saves a few minutes.

Reorganising the job around an agent works differently. The agent checks the accounts software every day. It sends the first polite reminder on its own, using wording the firm has approved.

The agent drafts the second reminder for a person to approve. Anything disputed, or over an amount the firm sets, goes straight to the partner with a short summary. A person reviews a weekly list of what the agent did.

To make that work, the firm has to write down the knowledge that used to sit in one person's head. It has to give the agent its own login, rather than lending it a member of staff's. And it has to decide, in advance, which steps need a human. The job is redesigned so an agent can run it, and the person who did the job now sets the agent's rules and checks its work.

A one-page handover sheet for your first agent job

Copy this, fill it in for one recurring job, and save it where the person who checks the agent's work can find it. If you cannot fill in a line, the job is not ready to hand over yet. Because the agent may read client emails, agree the "may see" line with whoever is responsible for client confidentiality in your firm.

AGENT HANDOVER SHEET

Job:                      (one recurring job, start to finish)
Done well means:          (what a good result looks like, in one sentence)
Named owner:              (the person who checks the agent's work)

The agent may see:        (which systems, folders and mailboxes)
The agent must not see:   (anything outside the job)

The agent may do alone:   (actions with no approval needed)
The agent must ask first: (actions that need a person's approval)
The agent must hand back: (situations a person deals with entirely)

Rules a new colleague would need to know:
  1.
  2.
  3.

How often the owner reviews its work:  (daily / weekly)
How we would switch it off:            (who, and how)

Here is the sheet filled in for the invoice example.

Line Invoice chasing
Job Chase overdue client invoices until paid or passed to a partner
Done well means Every invoice over 30 days old has had a polite reminder within a week
Named owner The practice manager
May see The accounts software (read only) and the shared accounts mailbox
Must not see Partners' inboxes, client project files
May do alone Send the first reminder using the approved wording
Must ask first Any second reminder, and any other message to a client
Must hand back Disputes, payment-plan requests, anything over the firm's set amount
Rules Client A pays on the 28th; never chase during an open complaint; copy the account partner on second reminders
Review Weekly list of every message sent
Switch off The practice manager removes its access to the mailbox

Filling in the sheet moves the unwritten rules onto paper. It also gives the agent a boundary that matches the approval settings in products such as dots, which let you allow an action, require approval or block it. We go further into writing an agent's remit in our guide to an AI agent job description, and into choosing a small first job in our piece on the minimum viable agent.

What to do this month

  1. Pick one recurring job with a clear end. Good candidates are chasing invoices, booking meetings, preparing a weekly client update or filing incoming documents. Avoid work that depends on professional judgement, such as a design or an advice letter.
  2. Fill in the handover sheet with the person who does the job today. They know the exceptions the agent will need.
  3. Try it on a tool you already have, or can get. A UK firm on ChatGPT Business Premium can set up a dot now. That dot works for the person who sets it up, using the apps they connect. Dots with their own login are still in enterprise pilots, so use this first trial to test the rules on your sheet. Firms on Microsoft 365 can watch for Autopilot, which is in private preview.
  4. Start with strict rules. Run the agent with approval required for every outside message for the first fortnight, then loosen the rules where the agent has proved reliable.

We help leaders choose that first job and write its rules with their team in AI Lessons for Leaders.

FAQ

What is an always-on AI agent? It is an AI system with its own computer in the cloud, so it keeps working after you close the app. It can use the apps you connect, remember your preferences and come back to you when it needs approval. OpenAI's dots, Meta's Muse and SpaceXAI's Grok Bot all work this way.

Can a UK firm use OpenAI's dots? Yes, on the ChatGPT Business Premium plan, which OpenAI says includes dots across all supported regions. The individual Pro plan excludes the UK, the European Economic Area and Switzerland for now. OpenAI prices a premium seat on ChatGPT Business at $100 per user per month billed annually, or $125 billed monthly (OpenAI Help Centre, via the Internet Archive).

How is an agent different from a chatbot? A chatbot answers when you ask and stops when you close it. An agent takes a job, works through the steps on its own and only comes back for approval or when something changes.

Does a firm need a developer to use agents? A firm does not need one for ready-made agents, such as dots or the agents built into office software. Platforms such as Anthropic's Claude Managed Agents are built for developers.

What should a firm never let an agent decide alone? Ethan Mollick's list is a good place to start: spending money, contacting people outside the firm and opening sensitive material. A firm can still approve one specific action in advance, as the invoice example does with the first reminder. Set everything else on that list as "ask first" or "hand back" on the handover sheet.


If you want to work out which job in your firm to hand to an agent first, book a conversation.

Work with Good Transformer

Turn this thinking into working practice.

Explore team advisory

Newsletter

Get new Insights by email

Practical notes on using AI with judgement, and the AI news leaders actually need. No hype, no spam, unsubscribe anytime.

Choose how often you want the digest

Keep reading

AI strategy10 min read

Is this the AGI era? What actually changed

Nobody agrees what AGI means, so the argument cannot be settled. What can be measured is how long a job an AI now finishes on its own, and that has been doubling every few months.

7 September 2026