IA · 9 September 2026 · 5 min read

OpenAI solves Navier-Stokes with ten thousand agents, but academia cries foul

In brief: OpenAI announced a formal proof for the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize Problems, utilizing an autonomous swarm of 10,000 agents and millions of dollars in compute. The landmark achievement has sparked fierce academic blowback, as NYU professor Tristan Buckmaster accused the lab of racing ahead after learning of his joint research with an Anthropic scientist, raising serious questions about the use of proprietary Codex session logs.

by Team Mocchi's

OpenAI solves Navier-Stokes with ten thousand agents, but academia cries foul

Theoretical mathematics finds itself facing an unprecedented turning point: the resolution of one of the world's most elusive mathematical problems risks being defined not by pure intellectual discovery, but by the fierce ethical and industrial controversy running alongside it. OpenAI announced it has formulated a complete proof for the Navier-Stokes existence and smoothness problem, a foundational equation describing fluid dynamics and one of the seven Millennium Prize Problems established by the Clay Mathematics Institute, each carrying a one-million-dollar bounty.

While the scientific implications are profound, the academic community's reaction has been deeply fractured by accusations of intellectual preemption and improper competitive conduct leveled by university researchers.

Brute force: ten thousand agents on Lean

To achieve this milestone, OpenAI relied not merely on heuristic guidance, but on an unprecedented allocation of computational resources dedicated to pure research. As reported by Wired, the San Francisco lab trained an internal frontier model exhibiting exceptional mathematical reasoning, deploying up to 10,000 concurrent agents operating across more than fifty hours. Crucially, the resulting proof was fully formalized in Lean, the proof-assistant language that mathematically verifies logical integrity step by step without human error.

The cost of this operation illustrates the immense divide between traditional academic departments and frontier AI labs. According to details published by TechCrunch, the computational run consumed roughly 300 billion output tokens—equivalent to approximately $22.5 million worth of compute under current enterprise rates. Mark Chen, head of research at OpenAI, confirmed that operational compute costs exceeded millions of dollars, while noting that OpenAI does not intend to claim the Clay Institute's cash prize.

Academic outcry: a scientific race or intellectual preemption?

The timing of the release ignited immediate friction. Just hours before OpenAI's announcement, NYU mathematics professor Tristan Buckmaster published three intermediate proofs moving toward the same solution, developed in direct collaboration with Levent Alpöge, a mathematician and researcher at Anthropic.

Buckmaster subsequently published a detailed statement asserting that OpenAI learned of their imminent breakthroughs and rapidly redirected vast infrastructure to beat them to the finish line. The central issue centers on cloud data hygiene: Buckmaster and Alpöge had relied on both Claude and OpenAI Codex to write and test theoretical drafts throughout the project. When Buckmaster approached OpenAI executives to confirm whether their private Codex logs had influenced the model's trajectory, the company failed to provide an explicit denial regarding training pipelines.

Furthermore, Buckmaster claimed OpenAI proposed terms to publish a joint paper recognizing the AI-generated proof provided that Alpöge's name was removed due to his corporate ties with Anthropic, an accusation that OpenAI leadership strongly denied.

OpenAI's defense and the research privacy dilemma

OpenAI has mounted a firm defense, arguing that its mathematical proof takes a fundamentally distinct theoretical trajectory from the NYU-Anthropic work. As covered by The Verge, Sébastien Bubeck from OpenAI stated that the team intensified its effort following industry whispers of external progress, but maintained that no private drafts had been inspected.

Yet, a specific clause in OpenAI's official disclosure has raised red flags among corporate counsel and researchers alike: while OpenAI stated that no specific user data was viewed manually, the company admitted that "we cannot rule out that de-identified data derived from their usage of our products helped improve our models." That statement strikes at the heart of modern enterprise concerns, blurring the boundaries between telemetry, model refinement, and the potential assimilation of proprietary user innovation.

Mocchi's take

This dispute represents a defining case study for organizations conducting R&D and core engineering on closed-source AI platforms. For software agencies and enterprise research teams, the warning is unambiguous: feeding early-stage logic, sensitive codebases, or mathematical formulations into public cloud interfaces without strict zero-retention and non-training guarantees exposes intellectual property to systemic absorption. While pairing frontier reasoning agents with formal verification frameworks like Lean heralds a transformative leap forward for software reliability, securing proprietary advantage will increasingly require air-gapped models and private enterprise deployments.

Further reading

All articles on the Mocchi's blog