Paris, 1900, the International Congress of Mathematicians. The published text of David Hilbert's lecture opens with a question: "Who of us would not be glad to lift the veil behind which the future lies hidden; to cast a glance at the next advances of our science and at the secrets of its development during future centuries?"
What follows is a numbered list of problems for the mathematicians of the coming century. The numbering stops at 23. The eighth, on prime numbers, contains Riemann's statement that the zeros of the zeta function, apart from the well-known negative integer ones, all have real part one half.
A hundred and twenty-six years later, at 6:58 a.m. Japan time on October 7, 2026 (21:58 UTC on October 6), OpenAI published 722 mathematical manuscripts written by an unreleased internal model to a GitHub repository called openai/math. The accompanying README explains that the company widened its evaluations to open research problems "after performance on our existing mathematical evaluations saturated."
The same README names, among its results, work on a zero-free region for the Riemann zeta function, a direct descendant of Hilbert's eighth problem. Lists of problems that have pointed mathematics forward for more than a century are now serving as exam questions for a machine.
Faced with 722 answer scripts, we at Ai believe the question that weighs as much as whether they are correct is who will set the questions for the next hundred years.

The lists that steered mathematics
Before listing a single problem, Hilbert explains why problems matter. "As long as a branch of science offers an abundance of problems, so long is it alive; a lack of problems foreshadows extinction or the cessation of independent development."
He also says where problems come from. The first ones, in his account, spring from experience of the outside world.
As a field matures, the human mind begins to generate new problems on its own and "appears then itself as the real questioner." The 23 problems are that questioner's handover to the next century.
Twentieth-century mathematics took this list as one of its signposts. In 2000 the Clay Mathematics Institute named seven Millennium Prize Problems and allocated $1 million to each. Under its rules, the problems were selected by the institute's Scientific Advisory Board, "focusing on important classic questions that have resisted solution over the years."
The problems Paul Erdős left behind were gathered on a website that the mathematician Thomas Bloom launched in 2023 with just over 200 problems; it now hosts 1,221. A declaration issued on September 11 by 28 Fields medalists records the same lineage: "Famous problems have often served as landmarks and lighthouses against which one can measure an improved understanding of this landscape."
A list, in other words, has been a tool for deciding what to think about next. Hilbert wrote, for each of his 23 problems, why the question was interesting and what a solution would open up. The seven Clay problems came with a committee that chose them and a stated reason for the choice.
Writing the list, we would argue, has meant pointing the direction of mathematics.
An exam of about 4,000 problems
Using problems as a test is an old mathematical habit. In the same lecture Hilbert recalls how Johann Bernoulli publicly posed the problem of the curve of quickest descent so that the analysts of his day could use it "as a touchstone" to "test the value of their methods and measure their strength." It is by solving problems, Hilbert adds, that "the investigator tests the temper of his steel."
Bernoulli's problem was posed where everyone could read it.
According to OpenAI's README, the model was posed approximately 4,000 problems over the course of the evaluation. Grouping the output and "requiring an appropriate level of significance" produced the catalogue: 722 manuscripts in 372 result families across 17 fields.
Every manuscript is a preprint on GitHub and stands at the stage of a claimed proof. For 162 of the 722, the main result has been formalized in the Lean proof language.
What has been disclosed about the exam itself amounts to the number of problems and the compute: an average of roughly three hours of ChatGPT Pro thinking per result. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), an independent panel of mathematicians at the Institute for Advanced Study, had asked in its September 29 guidelines that any large release come with a document explaining "how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen." The same guidelines open by asking AI labs "to stop testing advanced mathematical problems on proprietary models."
OpenAI has its reasons. Scientific American reports that the company "maintains that it can't slow down because these math problems are an indispensable test to show that their AI really is getting smarter." OpenAI's own blog post says it is important "to continue to evaluate our internal frontier models on mathematics and other sciences."
The conversion of a problem list into an exam had already happened on the Erdős problems site. Writing on October 6, Bloom said of the site: "This was not at all my intention when I created the site -- I wanted to promote these problems to a human audience." A large pool of questions that are simple to state, he observed, "provided the ideal showcase for AI."
There was a contrasting scene. On the day of the release, before the results appeared, the number theorist Frank Calegari posted his own exam on his blog: open problems from his field, "Each question is worth 10 points."
Lehmer's problem, Artin's conjecture, the Birch and Swinnerton-Dyer conjecture in arbitrary rank, the Riemann hypothesis. He explained what made them worth asking: each one "is only a very special case of a much larger problem, and even solving that larger problem would not be the final goal."
His update after the release reads: "So that's on the one hand 5/100 but on the other hand still [...] Jaw-droppingly amazing."
In parentheses he added: "who knows how well it would have done had it been given the list of questions."
The human exam and OpenAI's exam were different things. One was published with a reason attached to every question; for the other, what has been published is a count and a single line describing the cut.
The 722 manuscripts are also a record of the judgment of whoever chose the questions. We want to know what was set, and what was left out, as much as we want to know what was solved.
Harvested questions
Hilbert's lecture contains a passage of pure optimism: "The supply of problems in mathematics is inexhaustible, and as soon as one problem is solved numerous others come forth in its place."
Just before it comes his famous line that "in mathematics there is no ignorabimus." The person who solves a problem finds the next one, and questions beget questions. That cycle is the mathematics Hilbert described.
About four hours before the release, Terence Tao posted a four-part thread on Mastodon about AI solving open problems. "Solutions to open problems are now being harvested at large scale in an unsustainable fashion," he wrote, "leaving entire fields of mathematics much less fertile than when such problems were solved in the traditional 'Math 1.0' fashion."
A problem that has been "solved" cannot be reverted to "unsolved," he pointed out, and the mere knowledge that a solution exists "contaminates" efforts by humans and AI alike to find other routes that might reveal new insight.
The Fields medalists' declaration describes problems as material for raising people, too: "For students we often suggest problems with the core intention of developing skills making them well-positioned for advances in research and elsewhere."
The series' Gaussian moat installment deals with a manuscript claiming to reach the experts' expected conclusion from grounds different from the ones that supported their intuition. As it puts it: "A conjecture turning out to be right and the reasons behind it being right are two different things." The arrival of an answer and the birth of the next question from it are two different events as well.
Hilbert's "as soon as one problem is solved numerous others come forth" assumed that whoever solved a problem understood it and went on to ask what lay beyond. When solutions arrive first and understanding follows behind, whether the next questions appear depends on how fast understanding can catch up.
The 722 manuscripts put that assumption to the test on hundreds of questions at once. Questions, as we see it, are an inexhaustible spring and also a field that can be farmed into exhaustion.
The people who asked
Each of the eight questions this series examines has a person and a scene behind it. In 1859 Riemann wrote the paper tying the zeros of the zeta function to the distribution of primes. In 1917 a paper by Sōichi Kakeya and a coauthor appeared in the Tôhoku Mathematical Journal and became the starting point for questions about the smallest area needed to turn a figure around.
In 1936 Erdős and Paul Turán wrote a joint paper on sequences of integers that avoid arithmetic progressions. In July 1944 Turán, called up for labor service and working at a brick factory near Budapest, watched loaded trucks jump the rails at the crossings and wondered: "But what is the minimum number of crossings?"
In his address to the 1950 International Congress of Mathematicians, W. V. D. Hodge set out the cases he could prove and wrote: "Beyond this, the problem is an unsolved one." In the early 1960s Birch and Swinnerton-Dyer ran experiments on one of the early EDSAC computers at Cambridge and found the form of what became the Birch and Swinnerton-Dyer conjecture.
In 1962 Basil Gordon asked whether one could walk to infinity on the Gaussian primes with steps of bounded length. On June 27, 1983, Alexander Grothendieck set out the section conjecture in a letter to Gerd Faltings.
The series' Turán installment follows where that question went: "Eighty-two years after the trucks jumped the rails, the reply arrived, written by a machine, in a GitHub repository rather than the specialist journal Turán had hoped for."
The Grothendieck installment traces how an AI manuscript uses theorems by Japanese researchers as its components, and sums it up: "The dream Grothendieck sketched was turned into tools over thirty years by researchers in Japan, and in Kyoto above all, and an AI opened that toolbox and used it."
Every one of these questions began where a particular person stopped work in a particular place: a derailed truck, a printout from a computer, a letter to a colleague.
All eight were given their shape by human hands long before any machine was asked to answer them. The ability to set questions, we believe, has grown in a different place from the ability to produce answers.
Work chosen by the metric
What is happening in mathematics applies to work outside it, and the Fields medalists' declaration says so itself: "In many fields and activities, years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas."
When companies bring in AI, automation starts with the tasks that are easiest to measure, and the work that appears on a metric gets done fastest. Summarizing meeting minutes, a first pass over a contract, routine research. Each of these used to be a practice problem handed to junior staff.
Once the answering moves to AI, output speeds up and the training ground for asking questions shrinks. The series' Erdős installment draws a lesson for organizations from a website's decision to stop displaying how many problems had been solved: "A metric is only a stand-in for what you actually want to measure."
The scarcest work left in an organization, we believe, will be deciding what to ask next. Who writes the "problem list" handed to AI, and on what grounds are the items chosen? Leadership should be able to answer that in its own words.
The next 23
In its October 6 statement AGMAI called the release "the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge." The same statement says that "the future of mathematical research cannot consist only of understanding results produced by AI labs," and that "Mathematicians must be able to formulate their own questions, develop their own approaches, and explore directions that have not been selected as examples of an AI system's capabilities."
The full picture of the 722 is the job of the next piece, "A Map of the 722". From there, the eight questions await.
Hilbert's lecture, once the list is done, closes with a wish: "may the new century bring it gifted masters and many zealous and enthusiastic disciples." What he asked for at the end was people more than answers.
Who will write the next list of questions to set beyond the veil? That seat is still open.
References
- Mathematical Problems (David Hilbert, Bulletin of the American Mathematical Society, 1902)
- openai/math README and CONTENTS (OpenAI, GitHub)
- Sharing AI progress in mathematics (OpenAI)
- On OpenAI's Release of Mathematical Results (AGMAI)
- Responsible Release of AI-Generated Mathematics (AGMAI)
- A Severe Misalignment of AI in Mathematics (Math and AI)
- Thread on Math 1.0 and Math 2.0, part 3/4 (Terence Tao, Mathstodon)
- OpenAI unleashes hundreds more math results upon a field already in shock (Scientific American)
- The OpenAI problem dump (Frank Calegari, Persiflage)
- Changes (Thomas Bloom, erdosproblems.com)
- Millennium Prize Description and Rules (Clay Mathematics Institute)
- On some problems of maxima and minima for the curve of constant breadth and the in-revolvable curve of the equilateral triangle (Tôhoku Mathematical Journal, J-STAGE)
- On some sequences of integers (P. Erdős and P. Turán, 1936)
- A Note of Welcome (Paul Turán, Journal of Graph Theory, 1977)
- Proceedings of the ICM 1950, Vol. 1 (International Mathematical Union)
- The Birch and Swinnerton-Dyer Conjecture, official problem description (Clay Mathematics Institute)
- A Theorem about Gaussian Moats (Theorem of the Day)
- Letter to G. Faltings (A. Grothendieck)