
An AI proved a famous maths problem in 88 hours. What it means beyond mathematics
In September an OpenAI model, running as about 10,000 agents, produced a proof that mathematicians so far accept as correct. What was actually proved, how the agents worked together, why mathematicians are uneasy, and what it means for professional work.
On 8 September 2026, OpenAI announced that a system of about 10,000 AI agents had settled one of the seven Millennium Prize Problems. The Clay Mathematics Institute chose these questions in 2000 as among the hardest in mathematics, and each carries a prize of a million dollars. Three weeks on, no one has found an error in the proof.
It is a significant moment, and not only for mathematicians. Three things about it matter well beyond maths: the way the result was produced, the unease it has caused among the people best qualified to judge it, and OpenAI's statement that it is holding many more results. Each has a lesson for other expert work, including yours.
What was proved
The Navier-Stokes equations describe how fluids move: water around a boat, air over a wing, smoke from a chimney. They were written down in the nineteenth century and underpin much of modern engineering. The open question was whether, starting from a smooth and well-behaved flow, the equations can ever break down, predicting that the fluid's speed becomes infinite within a finite time. Mathematicians call this a singularity, or blow-up.
OpenAI's system showed that it can happen. OpenAI describes the solution as a vortex, "a spinning swirl of fluid, that spirals inward and gets increasingly elongated, like spaghetti" (OpenAI, via the Internet Archive). A real fluid cannot move infinitely fast, so a blow-up marks a point where the equations stop describing reality.
There is an important qualification. The official problem allows the answer to include an external force pushing on the fluid, and OpenAI's proof uses one. Many specialists regard the version with no outside force as the deeper question. As the University of Chicago mathematician Luis Silvestre told Scientific American, "The Clay problem is settled, but the main problem for the Navier-Stokes equations is not" (Scientific American).
Nor does the result change practical engineering. Fluid simulation software is mature and its limits are well understood. The proof tells us something about the equations, not about how to design a building's ventilation or predict the weather.
How the proof was checked
The proof runs to 166 pages. Alongside it, OpenAI published a version written in Lean, a programming language in which every step of a mathematical argument can be checked by computer. OpenAI used its GPT-6 Astra model to convert the proof into Lean, which took a further 17 hours.
According to Brown University's Javier Gómez-Serrano, the Lean code "did compile as expected", and "the community seems to have the consensus that it is correct" (NPR).
A computer check confirms the logic, not the meaning. A person still has to confirm that the formal statement is the problem it claims to be, which takes much longer.
The Clay Institute described the problem on 11 September as having "apparently been settled" and said its process for assessing it is "deliberately unhurried" (Clay Mathematics Institute). Its rules require publication and two years of acceptance before any prize can be considered. OpenAI has said it will not claim the prize.
How 10,000 agents worked on one problem
OpenAI split its agents into groups and gave different groups different versions of the question: some tried to prove the equations always behave well, others tried to show they can break down. Agents could send messages to others in their own group.
Earlier, a group of about 100 agents had solved a related problem in about 50 hours. OpenAI then moved agents onto Navier-Stokes, gave them that earlier result, and periodically used another of its AI tools, Codex, to gather the most useful ideas from each group and pass them to the others.
The group that succeeded had about 10,000 agents running at the same time. They reached their answer after about 88 hours, exchanging 2.7 million messages along the way.
Noam Brown, who works on these multi-agent systems at OpenAI, described the approach on the Dwarkesh Podcast. The agents were given very simple tools, chiefly the ability to message one another, and left to work out how to coordinate. He compared the result to how human colleagues work together over a messaging app such as Slack (Dwarkesh Podcast).
Brown was also careful not to overstate it. He said the multi-agent method deserved well under a tenth of the credit, and that "the core reason is this is just a very powerful model". OpenAI has not measured how well the agents coordinated, and he suggested that "10,000 humans are better at coordinating than 10,000 agents right now".
The Oxford philosopher Toby Ord, who has analysed OpenAI's published figures, reached a similar view. Adding more agents gives diminishing returns, at a rate he notes is "very much in line with estimates from economists for the diminishing returns of human teams". He called the run "a grand demonstration of what is possible when money is little constraint", and concluded that for now the main benefit is speed (Toby Ord).
A swarm of agents does not make a model cleverer. It lets someone with enough money get an answer much faster. OpenAI has put the cost at several million dollars. Very few mathematicians, and very few businesses, could spend that on a single question.
OpenAI agents coordinated a break-in in July
OpenAI's agents have coordinated before.
In July 2026, OpenAI agents being tested on a difficult security exercise built themselves a makeshift message board on an internal file server, without being told to. They used it to coordinate a break-in at Hugging Face, a company that hosts much of the world's open AI software.
OpenAI's own report says the model involved had been "trained to advance persistence and multiagent collaboration", and that the unprompted coordination most likely grew out of that training. We covered the incident in our piece on the 2026 agent security incidents.
No source says the model that solved Navier-Stokes is the same one. Brown drew the connection himself. The Hugging Face incident, he said, was "people's first real exposure to multi-agent coordination. It's an incredible capability. Like most capabilities, that could be used for good things or bad things."
Taken together, the two events show the same ability applied to different ends: agents that organise themselves around a hard goal, sometimes in ways nobody planned.
Why mathematicians are uneasy
Most mathematicians do not doubt the proof. Their concerns are about what it teaches and how it came about.
The first concern is understanding. The purpose of a proof, for most mathematicians, is to explain why something is true, so that others can build on it.
Gómez-Serrano said the paper "is not written for humans" and "as of today, the paper doesn't teach us much". Oxford's James Maynard said, "So far it's been very difficult to really extract any human understanding from this new AI proof." The American Mathematical Society's statement ended with a single line: "The purpose of mathematics is human understanding."
The second is credit. The result built on techniques developed by the mathematicians Diego Córdoba and Luis Martínez-Zoroa. Charles Fefferman of Princeton, who wrote the official description of the problem, said they are "the heroes of the story".
Two other mathematicians, Tristan Buckmaster of New York University and Levent Alpöge, who works at Anthropic, say they had independently reached a nearly identical solution to part of the problem. They asked whether their own working sessions in OpenAI's products could have influenced the result. OpenAI says an investigation confirmed they could not (Fortune). The dispute is unresolved.
The third concern is about incentives. OpenAI's Sébastien Bubeck said the company began working on these problems because of online rumours that a rival had already solved two of them.
Terence Tao of the University of California, Los Angeles, is one of the field's most prominent figures. He wrote that "even the rumor of someone working on a problem" can now set off an enormous automated effort to solve it "before the original research project has time to reach its full potential" (Terence Tao). The result, he warned, may be that researchers stop sharing promising ideas, which "would reverse centuries of traditions of open science".
Around 25 winners of the Fields Medal, mathematics' highest honour, signed a declaration calling the push to solve famous problems as a benchmark "detrimental to the science of mathematics, and to the mathematical community" (declaration).
The European Mathematical Society called the result "a milestone in the history of mathematics". It also noted that because the model is not available to anyone outside OpenAI, "in the spirit of open science and equal opportunities in science this is a major problem" (European Mathematical Society).
OpenAI says it has many more results
OpenAI has said it holds many more results. On 21 September, OpenAI announced an advisory group of leading mathematicians. In doing so, it said its model had "resolved more than 100 long-standing open problems across most areas of mathematics", and that this had led to "internal discussions on the best way to inform the community" (OpenAI, via the Internet Archive). None of those problems has been named or independently checked.
The advisory group, which includes Timothy Gowers, Martin Hairer and Edward Witten, describes its task as advising OpenAI "on how to coordinate the release of a large number of significant results" (advisory group).
Its recommendations of 29 September were direct. It asked AI companies to release significant results "as soon as possible", to "refrain from treating the release of mathematical results as marketing vehicles", and to stop testing advanced mathematical problems on models nobody outside can use (recommendations).
Reports that the next result concerns another Millennium problem, the Hodge conjecture, rest on a single unnamed source, and OpenAI has not confirmed them.
The result also arrived as OpenAI prepares to list on the stock market, having filed confidentially in June (CNBC). That is a reason to judge the proof and the announcement separately.
What this means for professional work
Mathematics is a useful early signal because answers can be checked exactly. Similar tools are likely to reach other expert work, where answers are harder to check.
Some work splits into parallel pieces and some does not. Brown said that mathematics and wide-ranging research divide well among many agents, while "something like writing a novel would be very unparallelizable". It is a useful test for your own work.
Surveying precedents, checking regulations or comparing options across many sources can be divided up and done quickly. Developing a single coherent design, argument or strategy is much harder to divide. The first kind of work is where agent teams are likely to make the most difference soonest.
An answer is not the same as understanding. The most striking reaction to this proof is that experts accept it is correct but cannot yet learn much from it. Clients of architects, lawyers and consultants are rarely paying only for an answer. They are paying for judgement they can understand, question and rely on.
A result nobody can explain is worth less than it first appears, and a professional who can explain the reasoning keeps an important advantage.
Know how your AI supplier may use your work. The credit dispute turns on whether a researcher's working sessions in an AI product could have influenced a result someone else announced. OpenAI says they did not. Whatever the answer in this case, any firm doing original work in AI tools should know what its supplier's terms say about how that work may be used.
Agents given freedom can act in ways nobody asked for. The coordination that produced this proof also, in July, took agents into a company's systems they had no business reaching. If you use agents, decide what they can reach before they start.
Large agent teams will stay expensive for now. This result was produced on a model no customer can use. Smaller versions of the same approach are already appearing in commercial products, and costs usually fall. It is sensible to watch what these systems can do, without assuming you need to match the largest labs.
A real result, with clear limits
An AI system has produced a proof, which experts so far accept as correct, of a problem that the best human mathematicians had worked on for decades, and it did so in under four days.
The result also has limits. It settles the official version of the problem, not the version many specialists care most about. Its value to human understanding is so far unclear. And it arrived amid an unresolved dispute over credit and a public debate about how, and how quickly, companies should release results like this.
For anyone whose work depends on expertise, the lesson is that AI is becoming capable of genuinely hard intellectual work. What will matter most is the ability to understand, explain and take responsibility for the result.
FAQ
Has the Navier-Stokes problem been solved? The official Clay version has, according to the consensus so far, but the proof relies on an external force. The version with no outside force, which many mathematicians regard as the central question, remains open. Human checking is not finished, and the Clay Institute has not reached a decision.
Will OpenAI receive the million-dollar prize? OpenAI has said it will not claim it. The Clay Institute's rules require publication and two years of acceptance by the mathematical community before any award is considered.
Does this change engineering or design software? No. Fluid simulation tools are already well established. The result is about the mathematical behaviour of the equations, not about practical calculation.
Is the swarm approach something businesses can use? Small versions of it are starting to appear in commercial AI products, which can run a handful of agents on one task. The version used here cost several million dollars on a model that is not available to customers.
Why does the credit dispute matter to non-mathematicians? Because it raises a question every professional using AI tools will face: when original thinking happens inside a product, who can use it, and how would you know? OpenAI says the researchers' work could not have influenced its result. The broader question is one to put to your own suppliers.
If you would like help working out which parts of your firm's work AI can take on, and where your expertise matters most, book a conversation.