
Is this the AGI era? What actually changed
Nobody agrees what AGI means, so the argument cannot be settled. What can be measured is how long a job an AI now finishes on its own, and that has been doubling every few months.
When OpenAI released GPT-6 Astra on 3 September, its president Greg Brockman said it was "not unreasonable to feel that we are now in the AGI era" (Fortune, 3 September 2026).
Artificial general intelligence is meant to mean a system that can do more or less any thinking job a person can. People are arguing about the claim because nobody agrees what that means in practice. OpenAI's own charter defines it one way and its contract with Microsoft defines it by profit.
That argument cannot be settled. A different thing underneath it can be measured, and it matters more to you: the length of job one of these systems finishes on its own has been doubling every few months for three years. That is what decides which of your work can be handed over, and it has moved since you last looked.
There is no UK statute on AI to wait for, so none of this arrives as a compliance deadline. It arrives as a management decision, which is why it is worth ten minutes now.
Nobody agrees what AGI means
Ethan Mollick of the Wharton School, writing yesterday about the Astra release, put the problem in one line: "The problem is that AGI is not a well-defined term" (LinkedIn, 6 September 2026).
Four definitions, from the people best placed to set one (Fortune, 30 March 2026):
- OpenAI's 2018 charter calls it "highly autonomous systems that outperform humans at most economically valuable work".
- OpenAI's 2024 contract with Microsoft defines it by money rather than capability, at a threshold of $100 billion in profit.
- Google DeepMind proposes ten cognitive faculties, judged against a well-educated adult.
- François Chollet argues the test is how efficiently a system learns something new, not how much it already knows.
So the word settles nothing. Jensen Huang of Nvidia told Lex Fridman in late March that AGI had already been achieved. Plenty of researchers say the opposite with equal confidence. If your team is arguing about whether this counts, the argument cannot resolve.
Jagged, not general, is the answer that helps
Mollick offers a version a manager can use. He says he would accept that we are in an AGI era for what he calls "jagged AGI", which he defines as "better than average human in many areas, worse in others". On the harder test he is clear: "If you mean better than a human expert at most human tasks, we aren't there (yet?)."
He also makes the point that gets lost in the row about labels: "it is still astonishing that AI is where it is & got there so fast, especially less than a decade after the invention of the transformer, and less than four years since ChatGPT."
Jagged is the word to take into a management meeting. It says the capability is real and uneven, that it will be superhuman on some of your tasks and poor at others, and that you cannot predict which from the outside. That is a testing problem, not a philosophy problem, and we have written about how to map that unevenness in your own firm.
How long a job a model can finish on its own
METR, which evaluates frontier models, asks how long a task a model can complete without help. They take jobs of known human duration and find the length at which a model succeeds half the time. That length is its time horizon.
In plain terms, a five-hour time horizon means the system can finish a job that would take a competent person about five hours, and it will get it right roughly one attempt in two.
A model fails about half the time at its stated horizon. That is not a rounding error, it is the definition, and it is the reason the work still needs checking. Shorter jobs succeed far more often. It is the number vendor claims often leave out, and it explains why a demo dazzles and a rollout disappoints.
The horizon is doubling quickly. METR measured a doubling roughly every 197 days across 2019 to 2025. For models released from 2023 onwards, about 131 days. For 2024 onwards, about 89 days. So the length of job a model can finish alone, half the time, has been doubling every three to six months, and the three-month figure comes from the most recent models.
Work you ruled out eighteen months ago as too long to hand over may no longer be too long. Your judgement was correct when you made it and is now out of date.
The models are now speeding up the building of the next models
On 6 September OpenAI published two pieces on using its own models to accelerate its own research, one of them by chief scientist Jakub Pachocki. Simon Willison, reading them, noted that the company had made recursive self-improvement the theme of the day, and that 2026 has been the year agentic engineering took off inside OpenAI (Simon Willison, 6 September 2026).
The figure he highlights is the daily spend on AI per OpenAI researcher. It sat near zero in February 2026, reached roughly $150 by June, and climbed to about $600 by late August. Willison's reading of the late-July jump is that internal staff had gained access to the model later released as Astra.
Treat the causal story as his inference rather than a fact, because it is. What is not in doubt is that a leading lab is now spending heavily on having its models do the work of building its next models.
If that works, the doubling figures above are a floor and the next few years are faster than the last three. If it does not, the trend carries on as it was. Nobody outside those labs can tell you which, and anyone who says otherwise is guessing.
What this changes in an ordinary week
A written process becomes the thing you hand to a system. When a process can be given to a model, the quality of your written process sets the limit on what you can delegate. Firms that never wrote theirs down have nothing to hand over. Firms that did can hand them over now.
The model works inside your applications, not alongside them. Computer use means it operates the software your people already use, at the same time they do. We covered that separately in what Astra's computer use means for your office. Work nobody is watching keystroke by keystroke is now happening on your systems.
Speaking starts to replace typing. Speaking an instruction to a system that can see your screen takes less time than typing one, and it removes prompt-writing as the skill barrier. We have no reliable measurements on how fast this arrives, so expect it, but do not plan around a timetable.
Per-seat licensing and hourly fees both come under pressure
Software vendors stop billing per person. When a system does whole jobs rather than helping a person do them, a per-seat licence stops matching what is being sold, and pricing moves to how much the system runs. A busy month then costs more than a quiet one, and your finance lead will need a variable line in the budget. We have written on budgeting for AI that bills by the task.
Firms that bill for hours will bill fewer of them. The client is buying an outcome and the hours are the proxy. That does not mean fees must fall, but the basis of the fee now has to be defended, and the firms that move first will choose their own new basis rather than accept a client's.
The first of these is already happening in software pricing. The second has not arrived yet in professional services.
What can honestly be said
Nobody can tell you whether this is AGI, because that depends on whose definition you accept, and the definitions genuinely conflict. Jagged is the honest description: real, uneven, and superhuman in places.
Three things are sourced and checkable, and none of them requires you to take a position on the word. The length of job these systems finish unaided, half the time, has been doubling every few months for three years. They can now operate ordinary business software directly. And the labs building them are spending heavily on having the models help build the next ones.
Four things follow for an employer, and all of them are cheap.
Write down how your work actually gets done, because that description is what becomes delegable. Re-run the test you last ran eighteen months ago. Decide now which of your systems an agent may access, and write that down too. And give this work to the people who understand your processes, not the people most comfortable with software.
The label will be argued over for years. The capability is already changing what your staff can hand over.
FAQ
Is AGI just a marketing term? Partly, and the incentives are worth naming. OpenAI's contract with Microsoft ties the term to a $100 billion profit threshold, so the word carries commercial weight for at least one company using it. That does not make the progress underneath it fake. Assess the label and the capability separately.
Should we change our plans because of this? Not your plans. Your assumptions. Anything you decided AI could not do, decided more than eighteen months ago, should be re-tested rather than trusted.
Does this mean job losses at our firm? Nobody can tell you that honestly, and anyone offering a percentage is guessing. What is visible now is a change in the unit of work: fewer tasks needing a person for the whole job, more needing a person to specify them and check them. That shifts what people spend their days on before it shows up in headcount.
How do we judge claims like this ourselves? Ask what was measured, over what length of task, and at what success rate. Published claims often leave the success rate out, and a system that succeeds half the time is a very different proposition from one that succeeds nine times in ten.
Does UK law treat any of this differently? Nothing about the AGI question changes your existing duties. UK GDPR and the Data Protection Act 2018, the Equality Act 2010 and ICO guidance apply exactly as they did, and there is no separate UK statute on AI. Treat it as an ordinary question about data, fairness and record-keeping.
If you would like to work out which of your own assumptions are now out of date, book a conversation.