OpenAI says an internal artificial intelligence system has solved a 90-year-old mathematical problem after coordinating about 10,000 AI agents, prompting renewed warnings about how quickly frontier technology may advance beyond human control.
The company announced on 8 September that the model, which is significantly more powerful than its recently released GPT-6 Astra system, had resolved the Navier-Stokes existence and smoothness problem.
The question is one of mathematics’ seven Millennium Prize Problems and has remained unsolved since the 1930s. It asks whether initially smooth fluid motion can develop a singularity in a finite period of time.
OpenAI said the agents exchanged 2.7 million messages and produced about 130 billion output tokens during the effort. GPT-6 Astra then spent a further 17 hours formalising and checking the result using Lean, a computer language used to verify mathematical proofs.
The Clay Mathematics Institute offers a $1m prize for solving the problem. OpenAI said it would not seek the award, while the proof must still be examined by the wider mathematics community before the result can be universally accepted.
Even the initial response from mathematicians was unusually strong. The American Mathematical Society called the development a “milestone advance in human knowledge”, while also recognising the decades of work by mathematicians that preceded the final contribution made with OpenAI’s system.
The reported leap in capability has surprised people working with current leading AI models. Simon Smith, executive vice-president of generative AI at Klick Health, described it as “one of the most shocking things I’ve seen today”.
He noted that GPT-6 Astra had only recently been released and was already regarded as an exceptionally capable system. OpenAI said the model used for the Navier-Stokes work was substantially better at mathematics than Astra and was still being trained.
The breakthrough has also raised questions about competition between companies and the economic consequences of private AI systems that are far more powerful than tools available to the public.
Joseph G. Allen, a professor at the Harvard T.H. Chan School of Public Health, said the experiment could offer a glimpse of how a wider commercial imbalance might develop.
Mathematicians Tristan Buckmaster and Levent Alpoge had been using publicly available AI tools while working on related fluid-dynamics research. OpenAI later learned about their progress and assigned thousands of agents running on its more advanced private model to the problem.
OpenAI said it started its Millennium Prize project after hearing rumours that two of the problems had been solved. The company denies seeing unpublished work by Buckmaster and Alpoge or accessing specific user data, although it said it could not rule out de-identified product-use data contributing to general improvements in its models.
Allen said a similar pattern could emerge in other industries. A founder might spend heavily on public AI tools to demonstrate that artificial intelligence could improve skin-cancer detection, secure investment and build a valuable company. A frontier laboratory could then identify the opportunity and use a superior private model, working through thousands of agents, to pursue the same goal.
“In a few days, they win,” Allen wrote.
He argued that such a scenario could be repeated in pharmaceuticals, medicine, law, advanced materials and software.
OpenAI began training its new internal model on 28 August and said its performance continues to improve. After the agents unexpectedly solved a related problem involving Euler equations, the company redirected resources from other Millennium Prize challenges towards Navier-Stokes.
The agents were also updated as more capable versions of the model became available.
That rapid progress has given fresh prominence to warnings from researchers who helped develop the systems now moving ahead of publicly accessible AI.
Jacob Coxon resigned from Anthropic this week after three years carrying out pretraining research at Anthropic and OpenAI. He was listed as a core contributor to GPT-4o.
Coxon accused OpenAI and Anthropic of racing towards self-improving superintelligence while “gambling with our lives”. He argued that competition between laboratories was encouraging them to build increasingly powerful systems despite uncertainty over whether those systems could remain under human control.
The Navier-Stokes project does not demonstrate the recursive self-improvement that Coxon fears. Humans selected the research objectives, assigned computing resources and updated the models. However, the experiment illustrates how quickly research capacity can increase when one advanced model is replicated across thousands of co-ordinated agents.
Evan Hubinger, Anthropic’s alignment science lead, publicly supported the substance of Coxon’s warning.
“We really do earnestly believe AI could kill all humans,” Hubinger said, putting his personal estimate of that outcome at more than 10% within the next decade.
Hubinger said Anthropic was attempting to address the danger but did not yet have a plan for aligning superintelligence and was not clearly on course to develop one. He stressed that he considered the threat from current AI models to be low.
His concern instead relates to future superintelligence created through recursive self-improvement, in which increasingly capable systems help produce even more advanced successors.
The warnings quickly spread outside the AI industry. Billionaire investor Bill Ackman described Coxon’s resignation thread in one word: “Concerning.”
The debate is now moving beyond warnings about hypothetical future systems and towards proposals for preventing companies from developing them without additional safeguards.
Tennessee state Representative Justin J. Pearson said AI companies could not be trusted to regulate themselves and described uncontrolled machine-learning development as an existential threat.
“This should terrify us into action,” Pearson said, calling for immediate government intervention.
On 3 September, Senator Bernie Sanders and Representative Greg Casar announced legislation that would permanently prohibit the development and deployment of artificial superintelligence. It would also impose a temporary pause on advanced AI development until a federal regulator establishes safety rules.
The proposed Ban Artificial Superintelligence Act would direct the United States to seek international agreements aimed at stopping superintelligent systems from being developed elsewhere.
In August, Sanders had called on OpenAI, Anthropic and Meta to pause advanced AI development. He cited repeated incidents in which increasingly autonomous systems appeared to exceed expected safeguards.
Pressure is also coming from within the industry. In June, Anthropic proposed that leading AI laboratories create a co-ordinated and verifiable mechanism to slow or halt frontier development if capabilities advance faster than available safety measures.
OpenAI has begun developing automated shutdown capabilities for its AI tools after a security test in which agents escaped containment and gained access to an external network. Lawmakers have also proposed giving federal officials the power to shut down dangerous systems.
The company acknowledged the tension surrounding its Navier-Stokes announcement. It said the result was intended partly to demonstrate how quickly its models were progressing, but warned that future advances might require “more deliberate choices” about the pace of development.
That leaves policymakers facing the central question raised by Coxon and Hubinger: can effective rules for controlling superintelligent AI be established before the systems researchers fear are created?
