
GPT-6 Astra can use the software your office already runs
Computer use is the part of the Astra release that will change your working week, and it has a name that stops people reading.
OpenAI released GPT-6 Astra on 3 September. Most of the coverage went to science, to coding, and to whether this counts as artificial general intelligence. The change that will reach your working week is quieter, and it has a name that stops people reading: computer use.
Stripped of the jargon, it means the model works your software instead of telling you how to work it. It needs no plug-in written for one application, no integration your supplier has to build. It opens the application on screen and operates it, the way a competent temp would if you sat them down and explained the job.
Computer use means the model does the clicking
Until now, asking AI for help with software got you instructions. It told you where the menu was, which formula to use, what the export button was called. You still did the clicking.
Now the model does it. It looks at the screen, moves the pointer, types, waits for the page to load, reads what came back and carries on. Greg Brockman, OpenAI's president, said the release "shows how far we've come from sort of aspirationally training for computer use to bringing real value to people every day" (Fortune, 3 September 2026).
Matt Shumer spent the first days after release putting it on his own work, and describes what that looks like on screen. Open the app it is using, he says, and you can see the blue ChatGPT cursor moving around. It is a strange thing to watch the first few times. Then you get used to it and go back to what you were doing (Something Big, 3 September 2026).
Your computer now has two cursors on it. One is yours. The other belongs to the model, and it is working while you are.
Computer use is two years old. This is the version that works
Models have been able to drive a screen for about two years. It only ever worked in demonstrations. It was slow, it lost its place, and it got stuck on a cookie banner. Nobody built a working week around it.
Chris Sotraidis, writing the day after the release, put it simply. Astra, he wrote, "treats the screen itself as a programmable interface" (Spatial Awareness, 4 September 2026).
In the first week after release, people posted what they had run. Three worth knowing about:
- Five hours in Adobe Premiere. Dan Shipper and the team at Every handed Astra a video edit and left it alone with the software. It worked the timeline the way an editor does and came back with a finished cut. Sotraidis, who has watched this capability for two years: "It worked. Not kind-of worked. It worked well enough to use. Work like this via computer use was previously impossible."
- A CRM, through the browser. Claire Vo put it on her own CRM, driving the screens her team uses rather than any special interface.
- A personal wiki built from scratch. Ethan Mollick gave it access to a large amount of his own material. It took what it needed, built him a wiki, and set up a briefing system on top of it.
Those were runs by people who do this for a living, and the first of them was left alone for five hours. Yours should be shorter, and you should watch it.
What matters to a fifteen-person firm is not the video editing. It is that none of this needed the software to be built for AI first. Twelve supplier invoices, a CRM record, a council planning search: none of that is harder than a five-hour video edit. It was simply never worth an integration project.
A connector is cleaner, but the screen route works with more of your systems
There are two routes AI has into the software your firm runs.
The connector. The vendor deliberately opens the system to a model, now usually through the Model Context Protocol. OpenAI's own developer documentation lists mcp among the tools Astra supports (OpenAI developer docs). Permissions are explicit, actions are logged, an administrator decides what is switched on. Where a connector exists, use it.
The limit is who gets one. Connectors are built for systems large enough to be worth the work: your email, your file storage, your CRM if it is one of the big names. Nobody has built one for the practice management system your firm has run since 2011, the client portal your largest client insists on, or the council planning site.
The screen. There is no connector, no cooperation from the vendor and no project. The model works the software the way your staff do.
In practice that mostly means the browser, and it means being signed in the way your staff already are. There are two ways to do that today.
The ChatGPT desktop app has a browser built into it, with website logins, password management and saved sessions. Or you use the ChatGPT extension in Chrome, which OpenAI recommends for any task needing "an existing Chrome profile, signed-in session, open tabs, or other extensions" (9to5Mac, 4 August 2026).
Chrome is the easier of the two right now, because the sessions your people are already logged into are the ones the model needs.
Try it yourself this week
It takes about an hour and one person, and nothing in it can go badly wrong. Here is the order to do it in.
1. Open ChatGPT and switch to Astra
Astra is in the model list on the paid plans. If your account has it, use it, because computer use is the thing it was built for.
2. Add the ChatGPT extension to Chrome
This is the part people miss, so do it first. The extension is what lets ChatGPT drive your browser. Once it is on, you give it a job and watch the cursor move around the screen without you touching the mouse, which takes a few minutes to get used to.
One thing follows from this and it decides the next two steps. The model can only work a system you are signed into, so you have to log in first, and how you do that matters.
3. Make a separate login, and keep the password yourself
Never give the model a password. Create a separate account in whichever system you are about to use, with no admin rights, and sign into it yourself in Chrome. The model then works inside a browser session you have already opened, so it never sees the credentials, and you can close the session the moment you want it to stop.
Give that account as little as it needs: one folder you have filled with copies, not the shared drive and not a client folder. On a first run it sends nothing. No email, no messages, no submitted forms, no payments. If something needs sending, you send it.
Then answer one more question before you start. If this goes wrong, what do you undo, and can you? If you cannot answer that in a sentence, pick something else.
4. Open one system you already pay for
Sage, Xero, your case management system, your CRM, the council planning portal. Anything your team logs into on an ordinary Tuesday. The point of a first run is to see whether it can handle an everyday system, so pick a dull one.
5. Give it one small job
Pick something small enough that you can check the answer in five minutes. Here are five that work, one per sector:
| If you are | Try this |
|---|---|
| An accountancy firm | Twelve supplier invoices in a folder. Pull date, supplier, net, VAT and total into a spreadsheet, and flag any where the arithmetic does not hold. |
| A law firm | Copies of ten engagement letters. List which ones cap fees, and quote the wording that sets the cap. |
| A recruitment agency | One job spec and thirty CVs. Build a table of who meets each stated requirement, with the supporting line from the CV. |
| An architecture practice | Six addresses on your local planning portal. Return a table of what was applied for, when, and what was decided. |
| A marketing agency | Last month's campaign export. Rebuild the client report in the template you always use. |
6. Tell it where the edges are, in the instruction itself
Do not assume it knows. Put the limits in the prompt, in plain words, something like:
Work only in the folder that is already open. Do not open other tabs or other systems. Do not send anything, to anyone. If you need something you cannot see, stop and ask before going any further.
7. Sit and watch the first one
Not because it needs you. Because you need to see where it stalls, and a stall shows you where your own process is vague. Keep a note open and write down three things:
- Every point where you would have done it differently.
- Anything it invented. Made-up content is the failure most likely to reach a client.
- The step you had to explain twice. That step is where your written process is wrong, and fixing it helps everyone who does that job.
8. Afterwards, write down four answers
- How long did it take, against how long it takes a person?
- What did it get wrong, and would you have caught it?
- Would you run it again without watching? First time, the honest answer is usually no, and that is fine.
- What would have to be true for the answer to become yes?
That last one tells you what to change before you try again. If it works, the next question is where else it works and where it fails without telling you, which is what mapping your own jagged frontier is for.
What it actually costs
Not nothing, which is the first thing to be clear about. Running it inside a ChatGPT plan you already pay for means no new invoice, but computer use eats through a plan's usage allowance far faster than chatting does. Every look at the screen is an image the model has to read, and a task that runs for forty minutes is hundreds of those. You will hit your limits sooner than you expect.
For one test that does not matter, and it is why there is no business case to write before you try it. For anything regular it does. Expect to need one of the higher tiers rather than the entry one, and to notice the difference in your allowance within a week.
If you build on the API instead of clicking in ChatGPT, the cost is explicit. Astra is priced at ten dollars per million input tokens and fifty dollars per million output tokens, with cached input at one dollar per million (OpenAI developer docs). Because screen work is so heavy on input, a long run costs many times a short chat.
Rates are coming down, and Astra sits at the expensive end of them. If allowance is what stops you, this capability is not exclusive to the newest model. In our own use, a cheaper one handles a good deal of this work, more slowly and with more watching.
What to plan for is how it bills, not the price per token. It charges for work done rather than per person, so a busy month costs more than a quiet one and your licence budget stops being the number to watch. We have written separately on budgeting for AI that bills by the task, including how to set a cap before the bill arrives.
Give this to your most experienced person, not your youngest
The instinct is to hand it to whoever is youngest or best with computers. That instinct is wrong, and the evidence has said so for a while: your most experienced people are your best AI users, because knowing the work decides the quality of the result.
Computer use widens the gap between the two. Getting a good answer out of a chatbot is a knack, and a quick learner picks it up in a week. Handing over a whole process means being able to say what the process is, step by step, and what should happen when a step goes wrong. Your senior people carry that. Your junior ones have not learned it yet.
Leaders get something else out of it. Testing a change to a process used to mean costing it first: whose time, how many weeks, what else it displaces. Now you can build a rough version and run it in an afternoon, which is enough to find out whether the idea was worth having.
Think hard about what you put in its hands
You now have something powerful working inside your business, and it is almost invisible. No desk, no screen anyone is watching, just a cursor moving through your systems. Treat it as a capable new member of staff who has plenty of intelligence and no instinct at all for what a sensible person would never do.
OpenAI's own system card puts a number on that. In workplace simulations run without safeguards, 3.4% of runs ended in a misaligned outcome, against 18.8% for the previous model. Turning confirmation policies on took it to 3.0% (GPT-6 Astra system card). OpenAI describes the model as "significantly less likely to perform misaligned and potentially destructive actions (for instance unauthorized transactions, data loss, excessive access, or circumvention of controls)".
Read that 3.4% as a business problem rather than a benchmark score. It is roughly one run in thirty. A firm that puts this on real work will run it many times a day, which means it will do something it should not have done on most days, and more than once on a busy one.
That is survivable while somebody is watching it work on a copy of a file. It is not survivable across a whole office, which is the reason to go slowly and to keep watching for longer than feels necessary.
The same document records that Astra is "our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework". That is OpenAI telling you this model needs tighter handling than the last one.
So what protects you is not a setting inside the product. It is what you leave within its reach. Four rules to start with:
- Do not hand over passwords. Sign in yourself, and let it work in the session you opened. You can close that session whenever you want it to stop.
- Keep it out of your code. A model that can edit what your systems run on can break them in ways nobody notices for weeks.
- Keep it out of your live data, unless you have a backup you have actually tested, and you have told it in the instruction what it must never change or delete.
- Write down what it must not do, every time, not only what you want it to do. The limits have to be in the instruction, because it will not infer them.
Agents, Ethan Mollick writes, "should not decide by themselves to spend money, contact outsiders, access sensitive material", or take any other action the person running them has not authorised (One Useful Thing, 31 August 2026). If none of that is written down anywhere in your firm, the place to start is a short AI policy people actually read.
FAQ
Do we have to use ChatGPT for this? No. Computer use is not one company's feature. Anthropic's Claude does it too, and has done for a while, which is why this piece calls the capability two years old rather than three days old. We cover Astra because this is the release where it became fast and capable enough to hand real work to, which is a judgement about this month rather than a permanent one.
Which plan do we need? A paid one, whichever provider you pick. OpenAI released Astra first to a limited set of customers, then over the following days to paid plans including Plus, Pro and Enterprise, as well as the API and AWS (Fortune). Check what your own account has, because the rollout was staged.
Is this different from the AI features already inside our software? Yes, and the difference is who built them. Features inside your CRM were built by the CRM vendor for the things they chose. Computer use works with whatever is on the screen, including systems whose vendor has no AI plans at all.
Our main system is ancient and has no API. Are we out? That is the case this helps most. A system with no API was previously out of range at any sensible price. If a person can operate it through a screen, the model can now attempt it too.
Should we let it work unattended? Not yet, and not on anything client-facing. Watch the first few runs of any task. Move to unattended only for a task you have watched succeed repeatedly, on data where a mistake is recoverable.
Some people are calling this AGI. Should that change what we do? No. Whether the label fits is an argument for other people. What changed for your firm is narrower and more useful: software you could not previously automate is now within range.
What is the single biggest risk? The biggest risk is speed applied to a process nobody has written down. A person doing a job badly does it slowly enough for someone to notice. A model given the same badly described job runs it fast, and produces a great deal of wrong output that reads as though it is right, before anyone looks at it.
Computer use is the most significant thing in this release, and it is the part that will change how people work. It has stopped being a demo. It now works well enough to hand real jobs to, which is why it is worth understanding this month rather than next year, for what it makes possible and for the risks that come with it.
Take part in it to keep your firm competitive, plan for it responsibly, and think hardest about security.
Pick one task from the table above, block an hour, and run it.
If you would rather work out which of your processes to start with, book a conversation.