# openai/math: 372 claimed theorems, 26 million lines of Lean, and what the green tick means

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/openai-math
> date: 2026-10-06
> tags: formal-methods, math, reasoning, verifiers

The repository went up on October 6 with a one-line description and a catalogue that reads like the open-problems section of every graduate textbook at once. A zero-free half-plane for the Riemann zeta function. The Hodge conjecture for CM abelian varieties. The Unique Games Conjecture. All free group factors isomorphic. Matrix multiplication with exponent at most 9/4. The irrationality exponent of $\pi$ is exactly 2. Each of these alone would be the mathematical story of a year. `openai/math` claims 372 of them, give or take how you count "of them".

The link reached me with no comment, which I read as "is any of this real?". I can't referee 34,769 pages of mathematics, and neither can anyone else this week. What I can do is the part an engineer is good at: read the artefacts. How was this produced? What exactly does the Lean library check, and what does it leave to a human? Which of the headline claims would I have to take on trust, and which come with a certificate that a laptop could, in principle, re-check overnight?

The answer surprised me in both directions. The Lean side is far bigger and cleaner than I expected: about 26 million lines, not one `sorry`, every statement written against plain mathlib, checked by a tool designed for adversarial proof submissions. For a real fraction of the headline results the question stops being "is this 199-page proof right?" and becomes "does this nine-line statement say what the abstract says?", which is a question a careful reader can answer. Then there is everything else: 137 families with no Lean at all, Lean pages that formalize a weaker corollary than the headline, and a review field in the project's own metadata that says `unchecked`.

This page is the map. First how the release was made and what its verification actually covers, then the results in order of importance, then every family by discipline, then a searchable table of all 372.

## What landed on GitHub

The repository is a release of "mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model". The README is short and worth reading in full, because it is the only methodological statement there is:

> The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems.

So: roughly 4,000 problems posed, 372 result families kept "requiring an appropriate level of significance", about one in eleven. The README frames this as evaluation, not a research programme: "As part of model development, we evaluate our models on open research problems. We expanded these evaluations after performance on our existing mathematical evaluations saturated." It also says, plainly, that "some of the unformalized results could have issues".

Two exceptions to the fixed procedure are named: the work on a zero-free region for the Riemann zeta function, and the proof of the Hodge conjecture for CM abelian varieties. The 11/12 zeta write-up was additionally "human edited for readability". Only one of the 722 manuscript READMEs carries a human-assistance note, and it is that one (`preprints/The-Quasi-Riemann-Hypothesis-October-5-2026/README.md`: "This paper was written with human assistance."). The Hodge paper's README does not mention it. The README does not say what "exception" means in either case: more compute, more human steering, or several model runs stitched together. That is the first thing I would want to know about the two results most likely to be scrutinised.

I counted the rest. The 722 manuscripts are PDFs plus LaTeX sources, 606 MB in `preprints/`, 34,769 pages in total. The median manuscript is 39 pages; the longest is a 262-page paper on spacetime Penrose inequalities, and the 7/8 quasi-Riemann paper is 199. The directory names carry dates, and they cluster hard: 563 of the 722 (78%) are dated September 23 to 27, another 136 October 4 and 5. Two predate September 11. The overview PDF is dated October 6.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/fig1.png"
  alt="The first page of OpenAI's overview PDF: the title OpenAI Research Catalog, the line 372 result families in 722 manuscripts, dated October 6, 2026, and a contents list of seventeen subjects from number theory to partial differential equations."
  caption="The catalogue's own table of contents. Entry numbers 'follow the catalog order and do not indicate a ranking' (OpenAI Research Catalog, overview.pdf, page 1)."
/>

The 372 families are grouped into 17 disciplines in `overview.tex`. A family is a principal result plus companions: alternative proofs, consequences, special cases. Family 003, the quasi-Riemann hypothesis, has three manuscripts: the 199-page 7/8 proof, the 49-page human-edited 11/12 proof, and a 9-page paper on Landau–Siegel zeros. Family 032, Hodge, has eight.

<ReleaseMap />

The bars show the first thing that matters for trust. Theoretical computer science and combinatorics, the two largest groups, have Lean pages for 80% and 89% of their families. Algebraic and complex geometry has 7 of 36, topology 3 of 18. That is not OpenAI hiding the hard ones. Mathlib has little of the machinery a Hodge-theoretic or 4-manifold proof stands on, so formalizing those results would mean formalizing a field first. It does mean the claims I would rate most consequential in algebraic geometry are exactly the ones with the least machine support.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/fig2.png"
  alt="A page from the overview PDF listing catalogue entries 011 to 019, each a bold title followed by a one-paragraph summary and links to its manuscripts, including the Ford–Konyagin–Luca conjecture, Ostmann's inverse Goldbach conjecture, restricted geometric Langlands, Zilber–Pink, the irrationality exponent of pi and the p-adic section conjecture."
  caption="One page of the catalogue: nine families, among them Ostmann's inverse Goldbach conjecture, abelian Zilber–Pink, the irrationality exponent of π and the p-adic section conjecture, each in a paragraph (OpenAI Research Catalog, overview.pdf, page 3)."
/>

## What the Lean library checks

`lean/` is a single Lake project. I measured it with `find`, `wc` and ripgrep. I did not build it: the README warns that compiling the whole thing can exhaust Linux's `vm.max_map_count`, and on a shared machine I was not going to try.

- 121,734 `.lean` files under `lean/OAI`, 25.9 million lines, 1.7 GB on disk, against Lean `v4.34.1` and a pinned mathlib.
- 30 pinned dependencies in `lakefile.lean`, from mathlib and the PrimeNumberTheoremAnd project to Carleson's theorem and sphere eversion, several carried with local compatibility patches. Mathlib accounts for 56,196 of the import lines; PrimeNumberTheoremAnd for 244.
- No `sorry`, no `admit`, no `axiom` declaration and no `native_decide` anywhere in `lean/OAI`. The three hits ripgrep found for the word "axiom" are in doc comments.
- 405 Comparator challenges in `lean/ComparatorChallenges/`, naming 507 theorems in all. Every one of the 405 configurations permits exactly the three standard axioms, `propext`, `Quot.sound` and `Classical.choice`.
- `formalization.yaml` lists 162 source papers and 185 "main results", and it is honest about its own status: `scope: "Partial progress."`, `automation: method: agent`, `review: status: unchecked`.

The last point frames everything else. The Lean was written by agents and nobody has reviewed it. That is less alarming than it sounds, because of how the checking is set up.

### Comparator, and why it matters here

[Comparator](https://github.com/leanprover/comparator) is the Lean FRO's judge for proofs from an untrusted party, originally built for the AIMO competitions. You give it two modules. The challenge contains the theorem statement with `sorry` as its proof; you write or audit it, and it is trusted. The solution contains the same theorem with a real proof; it is untrusted. Comparator builds both in a sandbox, exports them with `lean4export`, checks that every declaration the challenge's statement uses is identical in the solution's environment, checks that the proof uses no axiom outside the permitted list, and replays the solution into the Lean kernel. If it succeeds, its README says, the theorem in the solution is guaranteed to "prove the same statement as provided in `Challenge`", use only the permitted axioms, and be accepted by the kernel.

That design defeats the classic ways an AI-written proof can cheat: an extra `axiom`, a redefinition of a familiar name, a tactic that quietly elaborates to something else, a `native_decide` that trusts the compiler. The statement file is the contract, and every one of the 405 challenge files imports nothing but mathlib. None of the 30 dependencies, none of the 26 million lines, can change what the statement means.

Here is the whole challenge file for the quasi-Riemann hypothesis, `lean/ComparatorChallenges/QuasiRiemannHypothesis.lean`:

```lean
import Mathlib

namespace OAI

theorem riemannZeta_ne_zero_of_seven_eighths_lt_re
    {s : ℂ} (hs : (7 / 8 : ℝ) < s.re) : riemannZeta s ≠ 0 := by
  sorry

end OAI
```

`riemannZeta` is mathlib's. There is nothing else to read. If Comparator passes on this file, then, modulo the Lean kernel, ζ has no zeros with real part above 7/8. The one convention worth knowing: mathlib defines `riemannZeta 1` as a junk value, $(\gamma - \log 4\pi)/2$, which is not zero, so including the pole in the half-plane neither trivialises nor falsifies the claim.

And here is the solution side of another one, `lean/OAI/NumberTheory/PiExponent/Main.lean:17`, the theorem that the irrationality exponent of $\pi$ is 2:

```lean
theorem main :
  (∀ ν : ℝ, 2 < ν → ∃ Q : ℤ, 2 ≤ Q ∧
    ∀ p q : ℤ, Q ≤ q →
      (q : ℝ) ^ (-ν) ≤ |Real.pi - (p : ℝ) / (q : ℝ)|) ∧
  sSup {ν : ℝ | 0 < ν ∧
    Set.Infinite {r : ℚ | 2 ≤ r.den ∧
      0 < |Real.pi - (r : ℝ)| ∧
      |Real.pi - (r : ℝ)| < (r.den : ℝ) ^ (-ν)}} = 2 := by
  refine ⟨pi_integer_eventual_lower_bound, ?_⟩
  change irrationalityExponent Real.pi = 2
  exact pi_irrationalityExponent_eq_two
```

The first conjunct is the classical statement: for every $\nu > 2$, all large enough denominators $q$ satisfy $|\pi - p/q| \ge q^{-\nu}$. It does not lean on any definition the model wrote. The proof behind those three lines lives in 869 files and about 95,000 lines under `OAI/NumberTheory/PiExponent/`. The same file also proves the Flint–Hills series $\sum 1/(n^3 \sin^2 n)$ converges, though that theorem is not one of the Comparator challenges.

### What the green tick means, and what it doesn't

The guarantee has five layers. The first three are mechanical. The last two are the reader's job, and they are where every interesting doubt about this release now lives.

<TrustStepper />

The kernel layer is as strong as a guarantee in mathematics gets. One of the 405 configs (`ArtinParabolicIntersections`) also asks for nanoda, an independent kernel; the rest rely on Lean's own. I did not run Comparator, and as of writing I have not found anyone who has published a run. That is the cheapest missing piece in the whole story and I expect it to be filled within days.

The fourth layer, statement fidelity, depends heavily on the statement. Across the 405 challenge files the median is 79 lines; ten are 20 lines or fewer, and 75 run past 200. The longest, `EditApproximation.lean`, is 19,169 lines, because the statement includes a concrete binary algorithm and its exact resource accounting. Nine lines about `riemannZeta` can be checked by anyone who knows what the zeta function is. 225 lines that build group von Neumann algebras from scratch, because mathlib has none, need an operator algebraist with an afternoon. A 19,000-line statement is itself a formalization project to review.

Two traps recur in the long statements, and they are worth knowing before you trust a green tick. The first is Lean's junk values. `MatrixMultiplication.lean` defines $\omega(\mathbb{C})$ as the `sInf` of the admissible exponents, and in Lean the infimum of an empty set of reals, or of one that is unbounded below, is 0. So `omega ℂ ≤ 9/4` is a meaningful theorem only because the set is nonempty (schoolbook multiplication is admissible at 3) and bounded below (a correct program needs on the order of $n^2$ gates). Both facts are true and easy. Neither is part of the statement, so a reviewer has to know to check them. The second is `Classical.epsilon`: the free-group-factor statement picks "a projection of trace $(r-1)^{-1/2}$" with `Classical.epsilon`, which returns an arbitrary element if no such projection exists. One does, so the definition is fine, but that is a fact about von Neumann algebras the reader brings, not something Comparator checks.

Nine of the configurations use `definition_names`, Comparator's mechanism for definitions the statement pins or leaves open. Comparator's own README warns that definition holes can be gamed and "must always be checked with an additional (potentially human) verifier". I read the one that is a genuine hole (`ElementaryPositivity`): its type is itself the claim, a witness structure whose fields carry the proof obligation, so there is nothing to game. The other eight pin definitions the statement provides.

The fifth layer, scope, is the one the summaries most often blur. Every family with Lean has a page in `lean/docs/` that says what was selected, and 85 of the 235 pages explicitly carve something out ("is outside this statement", "are not included"). Some carve-outs are harmless; the Flint–Hills corollary sits outside the π challenge but is proved in the same file. Some are not:

- Family 159 headlines a quasipolynomial bound $C_k N \exp[-c_k (\log N)^{\varepsilon_k}]$ for sets without $k$-term progressions. The Lean page: "The paper's quantitative upper bound for the largest progression-free subset of $\{1,\ldots,N\}$ is outside this statement." What is checked is the qualitative consequence, Erdős's conjecture that a set whose reciprocals diverge contains arbitrarily long progressions. That is still a famous open problem, now with a 36-line statement, but it is not the headline.
- Family 130 headlines a deterministic DFT in $O(n(\log n)^{1-\delta})$ operations for every $n$, with $\delta = 10^{-13}$. The Lean statement is subsequential: for every $c > 0$ there are infinitely many lengths $n$ with a circuit cheaper than $c\, n \log_2 n$, in a model where multiplying by any fixed complex constant costs one gate. The Lean page says so: "no all-length, bounded-coefficient, conditioning, or bit-complexity claim."
- Family 197's Lean page constructs the Kaplansky counterexample but puts "the further conclusion that the group is nonsofic" outside. Here the gap closes by a known theorem: Elek and Szabó proved in 2004 that group algebras of sofic groups are directly finite, so a checked counterexample is, by their result, a checked nonsofic group. Whether a non-sofic group exists at all has been open since soficity was defined.

## How I'd read an extraordinary claim in this release

The catalogue makes claims at a rate no human field has ever absorbed. Terence Tao has called the pace "insane", in Scientific American's paraphrase. The useful question for any single claim is not "do I believe OpenAI?" but "what would it take to believe this one?". The release sorts itself into four tiers, and the tier matters more than the fame of the problem.

### Tier 1: a short, mathlib-only statement, checked

The 7/8 zero-free half-plane, the Kaplansky counterexample (17 lines, every term from mathlib), the irrationality exponent of π, Erdős's reciprocal-sum conjecture. What it takes: one independent Comparator run, and ten minutes reading the statement. If both come back clean, I would treat these as settled to the standard of any refereed theorem, and higher than most. My worry here is not the proof. It is that nobody has published the run yet.

### Tier 2: a long, bespoke statement, checked

Matrix multiplication, free group factors, the Unique Games statement (239 lines), anything in the 75 statements over 200 lines. What it takes: a specialist reading the statement's own definitions against the field's. This is days of expert work per result, not months, and it is the bottleneck I would fund first.

### Tier 3: checked, but weaker than the headline

Families 159 and 130 above, and the 85 Lean pages with explicit carve-outs. Believe the checked part on the tier-1 standard. Treat the rest as manuscript-only.

### Tier 4: manuscript only

137 of the 372 families have no Lean page at all, among them the Hodge conjecture for CM abelian varieties, the full Birch–Swinnerton-Dyer formula in Selmer corank at most one, Goldfeld's conjecture for quadratic twists, and the Kakeya maximal conjecture in three dimensions. These are ordinary preprints from an anonymous author with an unusually prolific output. What it takes is the usual: experts reading 50 to 200 pages each, and the README already concedes that "some of the unformalized results could have issues". I would hold every claim in this tier at the level of an unrefereed arXiv posting, and the extraordinary ones lower than that until someone in the field has read them.

The Hodge paper is the sharpest example of the last tier: one of the two results the README says were not produced by the standard procedure, 53 pages, no Lean, and a claim that would also settle the Tate conjecture for abelian varieties over finite fields. Its section below says what came before it and what I would check first.

## What the reasoning summaries show

For ten families the release includes "abridged summaries of the model's reasoning" in `reasoning_traces/`. They are summaries written about the model ("The assistant first seeks…"), with short boxed `VERBATIM EXCERPT`s of its own notes. I extracted all ten with `pdftotext` and counted.

<TraceAnatomy />

The first thing they show is how much failure there is. The free-group-factor summary spends six of its seven sections on invariants that do not work: free entropy, $L^2$-homology, rigidity, cost, each abandoned with a reason ("Difference P(V)-U small but derivative of difference can be huge; no closability. Same old."). The constructive turn comes from a classical analogy the model writes down itself: "Aha! For classical vertex labels in a group with non-uniform distribution, edge differences determine absolute labels almost surely". It closes with a hedge the catalogue summary drops: "This conclusion depends on the trace-preserving flow and limiting-generation steps described above."

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/fig6.png"
  alt="The first page of the free group factor reasoning summary: the problem statement in monospace, asking whether L(F_2) and L(F_3) are isomorphic, followed by section 1, From factor rigidity to the missing width bound, describing attempts using Dykema–Rădulescu rescaling, Boutonnet–Drimbe–Ioana–Popa and Popa–Vaes, and a verbatim excerpt from the model."
  caption="The problem as posed to the model, and the start of six sections of failed invariants before the constructive turn (Summarized chain of thought: Isomorphism of free group factors, page 1)."
/>

The second thing is that the model audits. In the Kaplansky summary, sections 13 to 22, ten in a row, are titled as audits and stress tests of one candidate counterexample. The π summary is the most striking. Part I set out to prove $\mu(\pi) < 5/2$, which Meiburg showed in 2022 is enough for the Flint–Hills series to converge (Alekseyev had shown in 2011 that $\mu(\pi) > 5/2$ would make it diverge), starting from the best known bound of about 7.10 (Zeilberger and Zudilin's 7.103205334137 from 2020, and a September 2026 preprint by Yufei Bai at 7.101862832357, both in the summary's references). After some thirty sections of failed routes it reached a bound of 62/25 = 2.48. Then, in section 36, it noticed that its own parameter inequality $\mu > 1 + 1/\sqrt{a}$ tends to 2 as $a \to 1$. The section titles say what happened next: "Stress-testing the interpolation proof and its unexpectedly strong exponent", "Rebuilding the geometry and testing an unexpectedly strong consequence", "Repeated audits of the determinant and its geometric foundations". Part II is a separate run that restarts from the 62/25 technique and claims exactly 2.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/fig5.png"
  alt="Page 33 of the pi reasoning summary: section 36, Stress-testing the interpolation proof and its unexpectedly strong exponent, showing the inequality mu(b minus a) greater than 1 minus a, b less than root a, implies mu greater than 1 plus 1 over root a, and the remark that taking a toward one suggested exponents approaching two."
  caption="The moment the π proof outgrows its target: the model's own inequality allows any exponent above 2, and the next sections audit that instead of celebrating it (Summarized chain of thought: The irrationality exponent of π is 2, page 33)."
/>

The third thing is that the model builds on itself. The π summary cites "OpenAI. Catalan's constant is irrational. 2026." as its main source of technique, more often than any human paper; that is family 005 in this release. The free-group summary and the catalogue lean on each other the same way. The README says "some outputs build upon earlier results produced by the models". It means a flaw in one family can propagate into others, and that the families are not independent draws you can average over.

What the summaries do not show is the compute, the number of attempts per problem, or what was thrown away. "Abridged" is doing a lot of work. They are curated narratives of ten successes, and the ten were chosen by OpenAI.

## What mathematicians have said so far

The release is a day old, and the public reaction has been about process more than proofs. Scientific American's report on October 6 ([Howlett](https://www.scientificamerican.com/article/openai-unleashes-hundreds-more-math-results-upon-a-field-already-in-shock/)) quotes Andrew Sutherland (MIT): "Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified. We should ask for receipts." Daniel Litt (Toronto) takes the other side: "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us. To me, it's going to be a good thing for mathematics." The same report notes that an Advisory Group on Mathematics and AI was set up on September 21 with recommendations on transparency; news coverage elsewhere says OpenAI's posting to its own GitHub went against its advice to use an independent repository, which I could not confirm from a primary source.

Gil Kalai, a few hours after the release, on [his blog](https://gilkalai.wordpress.com/2026/10/07/updates-sharing-ai-progress-on-mathematics-amazing-and-my-lecture-plans/): "Certainly this is an amazing milestone for mathematics, and the results will needs to be verified and digested by human mathematicians in the months to come." That sentence is the right frame for everything that follows.

I found no published expert verdict on any individual result yet, positive or negative, and no published Comparator run. When those appear they will matter more than anything on this page.

## The results that matter most

Ranking 372 claimed theorems is presumptuous, so here is the rule I used. A result ranks high if it resolves a problem people outside its subfield know by name, and if it would change what other mathematicians or engineers do the day it is confirmed. How much of it a machine has checked is a separate axis, and I say it for each one rather than letting it move the order.

By that rule two results lead, and both are from theoretical computer science: the matrix multiplication exponent and the discrete Fourier transform. They are not the most famous problems in the release. The Hodge conjecture and the Riemann zeta function are older and better known. But they are the two claims every engineer who reads this site has a stake in, both have a Lean statement (the Fourier one weaker than its headline), and both are the kind of claim I expected to fall apart on a close reading and did not. After them come the Unique Games Conjecture, derandomized logspace, the free group factors, the irrationality exponent of π, the quasi-Riemann hypothesis, Hodge for CM abelian varieties, the elliptic-curve results around Birch and Swinnerton-Dyer, the Mahler conjectures, critical percolation, Donaldson's tamed-to-compatible question, Kadison's similarity problem, the crossing number of the complete graph, Nagata's conjecture, Hilbert's tenth problem over the rationals, the Kaplansky counterexamples, Kakeya, Hadwiger and Erdős's progression problem. Each one says what is checked in Lean and what is not, next to the claim. Everything else follows by discipline.

### Matrix multiplication: ω ≤ 9/4 in thirteen pages

This is the claim I went in expecting to break. The matrix multiplication exponent $\omega$ has been the slowest-moving number in algorithms: since 1990 it has gone from 2.3755 to 2.3712, in steps that by the end were measured in the fourth decimal place, each one a long paper and a large optimisation. Family 107 says $\omega \le 9/4 = 2.25$ over the complex numbers. That is a bigger drop than the previous thirty-six years combined, and the paper that proves it is thirteen pages long.

I read all three manuscripts in the family, the Lean statements and the Lean source tree around them, and I re-ran every finite construction in the 9/4 paper in my own Python. Every finite step checks. Every step of the main argument that I could follow by hand checks too. The statement is formalized in an explicit, honest cost model with only the standard axioms. I can't tell you it is right; that takes experts and time. I can tell you where it could be wrong, and that I couldn't find the spot.

#### What ω measures, and why it moved so slowly

$\omega$ is the smallest exponent such that two $n \times n$ matrices can be multiplied with $O(n^{\omega+\varepsilon})$ additions, subtractions and multiplications, for every $\varepsilon > 0$. Writing down the answer already costs $n^2$, so $\omega \ge 2$, and the schoolbook method gives $\omega \le 3$. The conjecture most people believe is $\omega = 2$.

Strassen opened the gap in 1969. Seven products of sums of blocks are enough for a $2\times 2$ block product instead of eight, and since nothing in the identity uses commutativity the blocks can themselves be matrices, so recursion gives $7^k$ multiplications at size $2^k$, which is $\omega \le \log_2 7 \approx 2.807$. The widget below runs his identity on numbers you can edit, then shows the general rule: a scheme that multiplies $d\times d$ blocks with $r$ products gives $\omega \le \log_d r$.

<StrassenLadder />

Small schemes stall quickly (Laderman's 23-product $3\times3$ scheme gives only 2.854), so the field moved to asymptotic tools. Bini and coauthors showed that *approximate* schemes, exact only in the limit of a parameter $\varepsilon\to 0$, can be converted to exact ones at no cost in the exponent. Schönhage's asymptotic sum inequality (1981) showed that computing many *independent* small products at once is as good as one big one. Strassen's laser method (1986) supplied a way to carve independent products out of high tensor powers of a small tensor. Coppersmith and Winograd (1990) found the right small tensor, now called $\mathrm{CW}_q$, and used Salem–Spencer progression-free sets to separate the pieces, reaching 2.375477.

Everything after that is the same machine, tuned. Stothers analysed the fourth power of the CW tensor, Vassilevska Williams the eighth, Le Gall the thirty-second ([arXiv:1401.7714](https://arxiv.org/abs/1401.7714), 2.3728639). Alman and Vassilevska Williams reduced the loss when the marginal distributions allow several joint ones ([arXiv:2010.05846](https://arxiv.org/abs/2010.05846), 2.3728596). Duan, Wu and Zhou attacked the loss from requiring whole variable blocks to be disjoint ([arXiv:2210.10173](https://arxiv.org/abs/2210.10173), 2.371866). Vassilevska Williams, Xu, Xu and Zhou ([arXiv:2307.07970](https://arxiv.org/abs/2307.07970), 2.371552) and then Alman, Duan, Vassilevska Williams, Xu, Xu and Zhou ([arXiv:2404.16349](https://arxiv.org/abs/2404.16349), 2.371339) made the analysis more asymmetric. The last published step before this release was Dupont et al., who re-optimised that framework on the eighth power with AlphaEvolve in the loop and got 2.371177 ([arXiv:2608.16884](https://arxiv.org/abs/2608.16884), August 2026).

<OmegaTimeline />

Switch the chart to the last 0.003 and the shape of the problem is obvious. Between 2010 and 2026 eight papers moved the bound by 0.0026. Then the release puts three points on the chart in one go.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/matrix-multiplication-fig4.png" alt="Table 1 of More Asymmetry Yields Faster Matrix Multiplication: upper bounds on omega(1,k,1) for k from 0.33 to 2, compared with the previous bounds; at k = 1 the bound is 2.371177 from Dupont et al., previously 2.371552; at k = 0.7 it is 2.152770." caption="The state of the art the release is measured against: rectangular and square bounds before September 2026. At k = 1 the square bound is 2.371177; at k = 0.70 it is 2.152770 (More Asymmetry Yields Faster Matrix Multiplication, Table 1)." />

There were also known reasons this machine had a floor. Ambainis, Filmus and Le Gall proved that a wide class of laser-method analyses on CW powers cannot go below 2.3078 ([arXiv:1411.5414](https://arxiv.org/abs/1411.5414)). Alman proved that the far more general "universal method" applied to any $\mathrm{CW}_q$ cannot beat 2.16805 ([arXiv:1812.08731](https://arxiv.org/abs/1812.08731)), and Christandl, Vrana and Zuiddam showed that any route through an intermediate tensor $T$ is stuck above twice its *irreversibility*, $\log \tilde R(T)/\log \tilde Q(T)$ ([arXiv:1812.06952](https://arxiv.org/abs/1812.06952)). Any claim near 2.25 has to say which of these doors it walks through. I come back to that below.

#### Three papers, three claims of very different weight

The family is three manuscripts, and they are not three proofs of one thing.

| Paper | Claim | Field | Method | Length | Lean |
|---|---|---|---|---|---|
| [An Upper Bound of 9/4 for the Matrix Multiplication Exponent](https://github.com/openai/math/blob/main/preprints/Matrix-Multiplication-Nine-Fourths-October-2-2026/paper.pdf) (2 Oct) | $\omega \le 9/4$ | $\mathbb{C}$ | tensor characters and polynomial multiplication | 13 pp. | yes |
| [Complex Matrix Multiplication Below 2.258 and Rectangular Bounds](https://github.com/openai/math/blob/main/preprints/Complex-Matrix-Multiplication-Below-2.258-and-Rectangular-Bounds-September-24-2026/Complex-Matrix-Multiplication-Below-2.258-and-Rectangular-Bounds-September-24-2026.pdf) (24 Sep) | $\omega < 2.258$, $\alpha > 0.465$, $\omega(1,0.709,1) < 2.092$ | char. 0; square and 0.709 also in all but finitely many $p$ | CW₁ plus a group tensor that separates labels | 83 pp. | rectangular parts over $\mathbb{C}$ |
| [Staggered extraction for exact matrix multiplication over every field](https://github.com/openai/math/blob/main/preprints/Staggered-extraction-for-exact-matrix-multiplication-over-every-field-September-24-2026/Staggered-extraction-for-exact-matrix-multiplication-over-every-field-September-24-2026.pdf) (24 Sep) | $\omega < 2.371054886006746$ | every field | refined laser method on CW₅ | 35 pp. | yes |

The dates matter. The two September papers are the older work; the 9/4 paper came eight days later and makes the 2.258 square bound redundant over $\mathbb{C}$. The every-field paper is a different kind of object: an incremental laser-method improvement, 0.000122 below Dupont et al., which is exactly the size of step this literature has been taking. That one is the conservative claim of the three, and the only one that covers positive characteristic.

#### How the 9/4 proof works

The proof never builds a fast algorithm. It bounds $\omega$ from the dual side, and it is worth slowing down here because this is where the novelty is.

Think of a matrix multiplication algorithm as a tensor. The bilinear map $(A,B)\mapsto AB$ is encoded by the trilinear form $T_n = \sum_{i,j,k} x_{ij}\,y_{jk}\,z_{ki}$, its three variable groups are called legs, and the number of products an algorithm needs is the tensor's rank. Strassen showed in the late 1980s that asymptotic questions about tensors are governed by a set of numerical invariants he called the asymptotic spectrum. The paper calls them *characters*: maps $\lambda$ from tensors to non-negative reals that add on direct sums, multiply on tensor products, never increase under restriction, and send the one-term tensor $xyz$ to 1. The three flattening ranks are examples. Strassen's duality theorem says, roughly, that the asymptotic rank of a tensor is the largest value any character gives it. The paper reproves the special case it needs in an appendix (its Lemma 2.2: if $k < d^\omega$ some character has $\lambda(T_d)\ge k$). So to prove $\omega \le 9/4$ it is enough to show that **every** character has $\lambda(T_d) \le d^{9/4}$.

Each character has a fingerprint. On the "dot product" tensor $B_X(m) = x\sum_{i=1}^m y_i z_i$ its value must be $m^{p_X}$ for some $p_X\in[0,1]$, and likewise $p_Y, p_Z$ for the other two orientations. Matrix multiplication is the product of three dot products, one per pair of legs, so $\lambda(T_m) = m^{p_X+p_Y+p_Z} = m^{3t}$ with $t$ the average. The goal becomes $t \le 3/4$.

The paper then studies a different tensor, polynomial multiplication: $C(a,b) = \sum_{i<a,\,j<b} x_i y_j z_{i+j}$, which multiplies a polynomial with $a$ coefficients by one with $b$ coefficients. Its rank is exactly $a+b-1$ (evaluate at $a+b-1$ points, interpolate), so every character gives it at most $a+b-1$. Averaging a character over the six ways to permute the legs gives a symmetric profile $P(a,b)$ with $P(1,b) = b$ and, from the rank bound, $P(a,a) \le (2a-1)^{1/t}$.

The work is in two lower-bound inequalities on $P$, and both run on one new lemma about sharing. Suppose a tensor is a sum of blocks that share the first leg but have their own, disjoint variables on the other two legs. A full direct sum needs disjoint variables on all three legs, so these blocks are not independent, and in the laser method this is exactly the kind of sharing that costs you. Proposition 3.1 separates them. Take $L = 5M$ copies of the tensor ($M$ blocks). In copy $r$, substitute each shared variable by a sum over tentative labels $g$ with phase $\zeta^{2rg}$, and the block-$h$ variables on the other legs with phases $\zeta^{r(u-h)}$ and $\zeta^{r(-v-h)}$, where $\zeta$ is a primitive $L$-th root of unity. Summing over the copies is a discrete Fourier projection: only terms with $u - v + 2(g-h) = 0$ survive. Then give the variables weights $g^2$, $hu - h^2$ and $-hv$ and keep only the lowest power of a formal parameter. On every surviving term the total weight is exactly $(g-h)^2$, so the weight-zero part forces the tentative label to be correct, $g = h$, and then $u = v$. What is left is the $M$ blocks, now fully separated, each tensored with a dot product of length $M$. You paid $5M$ copies and got back $M$ separated blocks plus a free $M$-dimensional dot product, and when $M$ is exponentially large (it is a multinomial coefficient in the application) the factor 5 is noise. Corollary 3.2 turns this into an entropy inequality: for any character and any probability vector $q$ over the blocks, $\lambda(T) \ge e^{p_X H(q)} \prod_i \lambda(T_i)^{q_i}$. Sharing a leg, it turns out, is something a character can be made to pay you for.

The first use is concavity. Tensoring $C(a,b)$ with $B_X(2)$ and changing bases with the exact sequence $0 \to D V_{e-1} \to V_e \otimes W \to V_{e+1}\to 0$ (with $D = uw - vs$, a determinant) makes the tensor block-triangular with diagonal blocks $C(a,b+1)$ and $C(a,b-1)$; a degeneration kills the off-diagonal block. Feeding that into the separation inequality and optimising $q$ gives $2P(a,b) \ge P(a,b+1) + P(a,b-1)$, so the profile is discretely concave.

The second use is a shifted tripling. Split the second and third legs of $C(a, 3h+a-1)$ into left, middle and right ranges and weight the middle ones $+1$ and $-1$. The terms that cross ranges get positive weight and vanish; what survives is two copies of $C(a,h)$ and one copy of $C(a,h)$ with two legs swapped, sharing the first leg. The paper's own Figure 1 is the smallest case.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/matrix-multiplication-fig1.png" alt="A 4 by 5 grid for the polynomial multiplication tensor C(2,4): rows y = 0 to 3, columns Z-index 0 to 4. Three shaded blocks along the diagonal survive the degeneration, two crossed cells holding x0 and x1 have weight plus one and disappear, blank cells have no term." caption="The shifted tripling on its smallest case, C(2,4): the three shaded blocks have disjoint Y and Z coordinates but share x0 and x1; the two crossed terms vanish in the degeneration (An Upper Bound of 9/4 for the Matrix Multiplication Exponent, Figure 1)." />

The separation inequality with $q$ uniform turns this into $P(a, 3h+a-1) \ge 3P(a,h)$.

The last page is pure discrete analysis. Concavity says increments of $P$ shrink; tripling at $h = a$ says $P$ must triple over a stretch of length $3a-1$. Together they force the diagonal increments to stay large, and a short recurrence gives $P(a,a) \ge a^{4/3}$. Compare with the upper bound $(2a-1)^{1/t}$: if $t > 3/4$ the lower bound grows faster and eventually overtakes the upper one, which is impossible. So $t \le 3/4$ for every character, $\lambda(T_d) \le d^{9/4}$, and $\omega \le 9/4$.

<NineFourthsSqueeze />

Slide $t$ past 0.75 and the curve crosses zero; at $t = 0.80$ the contradiction arrives at $a \approx 63$, at 0.76 it needs $a$ beyond ten million. It doesn't matter how large: it only has to exist. Note what $t\le 3/4$ does not contradict. Every character anyone has written down, including the quantum functionals of Christandl, Vrana and Zuiddam ([arXiv:1709.07851](https://arxiv.org/abs/1709.07851)), has $t = 2/3$ on matrix multiplication, which is what $\omega = 2$ would need.

Converting the bound back into an algorithm is standard but not explicit. The definition of the exponent hands you, for each $\varepsilon$, some fixed size $u$ and some exact rank decomposition of $T_u$ with fewer than $u^{9/4+\delta}$ products, and recursion does the rest. The paper does not say what $u$ is. Its own words: the proof "does not specify a competitive finite matrix size". Remark 5.2 adds that the coefficients can be taken algebraic, in one number field.

#### Rerunning the finite constructions

I wrote my own Python (numpy, small sizes) for every finite construction in the 9/4 paper, and for the explicit pieces of the other two. The core of the Proposition 3.1 check tracks each coefficient of the substituted tensor as a polynomial in the formal parameter, by its power:

```python
# my check of Proposition 3.1: substitute L = 5M Fourier copies, collect by weight
for (h, a, b, cc), coef in c.items():
    for g, u, v in itertools.product(G, G, G):
        phase = sum(zeta ** (2 * r * g) * zeta ** (r * (u - h)) * zeta ** (r * (-v - h)) for r in range(L)) / L
        if abs(phase) < 1e-9:
            continue
        w = g * g + (h * u - h * h) + (-h * v)
        out[((a, g), (h, b, u), (h, cc, v))][w] += coef * phase
```

With random complex blocks for $M = 2, 3, 4$, no negative weight survives, and the weight-zero tensor equals the block direct sum tensored with $B_X(M)$ coefficient for coefficient. The concavity degeneration I built literally, with the paper's quotient lifts and kernel basis for $D = uw - vs$: for every $1\le a\le 6$ and $2\le b\le 7$ the input-kernel-to-output-quotient block is zero and the two diagonal blocks are exactly $C(a,b+1)$ and $C(a,b-1)$; the off-diagonal block the degeneration removes is genuinely non-zero (two entries at $a=3, b=4$), so the degeneration is doing real work. The tripling weights leave no negative term for any $a, h \le 7$, the three surviving blocks are $C(a,h)$, $C(a,h)$ and the swapped, reversed $C(a,h)$, and $C(2,4)$ loses exactly the two terms crossed out in Figure 1. The recurrence in Lemma 5.1 gives $D_a/a^{4/3} \ge 1.32$ for every $a$ up to 2,000.

I also asked whether the argument has slack. I set up a linear program over the profile values themselves: minimise $P(A,A)$ subject only to symmetry, $P(1,b) = b$, concavity and tripling. The minimum grows like $a^{1.354}$ between $A = 16$ and 32 and $a^{1.343}$ between 32 and 64, drifting down toward $4/3$. Adding five- and seven-sector versions of the tripling (which the same weighting trick allows) changed nothing. So with these ingredients 9/4 is where the method stops; a better bound would need a new inequality, not a better optimiser.

Checking the steps one at a time is not the same as checking the proof. What I can say is that each of the four constructions is a finite identity that my code confirms, that the logic connecting them is short enough to follow line by line, and that I checked the separation inequality against characters I know: it holds with equality for the flattening ranks, and with equality for matrix multiplication itself split along one index. The one ingredient I cannot check numerically is Lemma 2.2, the existence of a detecting character. It is Strassen's theory, reproved in a three-page appendix via a Farkas argument and a Schauder–Tychonoff fixed point, and it is the part the Lean development spends real effort on.

#### Why the barriers don't apply

The barrier results above all constrain one shape of proof: start from a fixed intermediate tensor $T$ (in practice a CW tensor), take powers, and degenerate them into many independent matrix products. The 9/4 proof has no intermediate tensor that gets turned into matrix multiplication. It proves an inequality about every point of the asymptotic spectrum directly, and Strassen's duality is exact, so that route is not limited in principle; it could prove $\omega = 2$ if the right inequalities were true. Polynomial multiplication appears only as a measuring instrument, never as something degenerated into $T_n$.

The 2.258 paper does use the CW route, so it has to respect the barrier. It starts from the small tensor $\mathrm{CW}_1$, whose border rank is 3, and adds a group tensor; over $\mathbb{C}$ an abelian group tensor is a diagonal tensor in disguise, which does not change the irreversibility accounting. $\mathrm{CW}_1$ is tight, so its asymptotic subrank is $\max_P \min_i e^{H(P_i)}$ over distributions on its six support points. I computed that: 2.7551, which makes the Christandl–Vrana–Zuiddam bound $2\ln 3/\ln 2.7551 = 2.16805$, the same number as Alman's limit for all CW tensors. 2.258 sits above it. The Ambainis–Filmus–Le Gall 2.3078 limit covers only a specific family of laser-method analyses; the group-tensor separation is outside that family, so going below 2.3078 is not a contradiction either.

#### The 2.258 paper: separating labels with a group tensor

The older paper carries a different idea, and it is the one I'd expect people to reuse even if the 9/4 argument holds. Its Lemma 3.3 handles the same situation as the separation lemma, pieces that two legs can tell apart while the third shares variables, by tensoring with the multiplication tensor of a finite abelian group $(\mathbb{Z}/4Q)^d$ and zeroing out coordinates. Tags for the $K$ pieces are points on one sphere; each piece keeps auxiliary points on one affine slice orthogonal to its tag. If anything but the intended triple survived, a non-zero displacement perpendicular to the tag would land on another point of the same sphere, which Pythagoras forbids.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/matrix-multiplication-fig2.png" alt="A circle centred at 0 with a tag vector t_s to its rim; a displacement d = u minus u-prime perpendicular to t_s points out of the circle to c = t_s + d. Text: the affine slice gives t_s dot d = 0; the group equation gives c = t_s + d; if c is another tag on the same sphere, then norm of c squared equals norm of t_s squared plus norm of d squared, so d = 0 and c = t_s." caption="Why the group-tensor separation leaves no cross terms: a displacement orthogonal to a tag leaves the tag sphere (Complex Matrix Multiplication Below 2.258 and Rectangular Bounds, Figure 1)." />

The group tensor costs about $K$ in rank, but the separation hands back $K$ direct pieces and a length-$K$ pairing, so the charge cancels and the pairing is pure volume. I built a small instance ($K = 4$ pieces, $m = 4$, $G = (\mathbb{Z}/16)^3$) and the zeroed-out tensor is exactly the direct sum the lemma promises, with the shared variables now carrying distinct tags. With that tool, the paper's warm-up already beats the published record: degenerate $\mathrm{CW}_1$ to five terms, complete three copies together, and 75 scalar pieces survive, so $\omega \le 9\ln 3/\ln 75 = 2.2901$. I recounted the pieces by enumeration: 17, 75 and 305 for blocks of two, three and four copies, matching the paper's $4^b + 3^b - 2^{b+1}$, and the three-copy block is the best of them (the ladder widget above lets you slide $b$). The headline 2.258 then needs much heavier machinery: banked constructions, a comparison argument modelled on viscosity solutions, and a GMP/C++ certificate with 128-bit integers that the README says checks "finite arithmetic interfaces", with "the tensor constructions and analytic allowances" left to the manuscript.

The rectangular results are the part of this paper that 9/4 does not supersede. A dual exponent $\alpha > 0.465$ means an $n\times n^{0.465}$ by $n^{0.465}\times n$ product costs $n^{2+o(1)}$; the previous best was 0.321334. And $\omega(1, 0.709, 1) < 2.092$ compares with 2.152770 at $k = 0.70$ in the table above, a large drop for a quantity that sits inside the running times of several graph algorithms. Both are formalized over $\mathbb{C}$. The extension of the square bound to "every field except possibly in one finite set of positive characteristics" works by realising two fixed rank schemes over a number field and reducing mod $p$; the exceptional primes are not computed, and that transfer is not among the Lean statements.

#### Every field, and why the field matters

$\omega$ depends only on the characteristic of the field (Schönhage), so "every field" really means "characteristic 0 and every prime $p$". The new complex results need roots of unity and division by group orders: the group tensor's rank equals its size only when the characteristic does not divide it, and the 9/4 separation step divides by $L$. That is why the release states 9/4 over $\mathbb{C}$ only and keeps a separate all-fields claim.

That claim, $\omega_F < 2.371054886006746$, is the staggered-extraction paper. It stays inside the Alman et al. asymmetric framework on $\mathrm{CW}_5$, with strands of length 8 split into lengths 4, 2 and 1. The new move is to run several "lots" offset in time, so that at each tick one lot is at each splitting stage, and to pool their capacities in one joint extraction: a shortage on one side at one depth can be paid for by surplus at another depth.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/matrix-multiplication-fig3.png" alt="Panel (a): a schedule table with rows A: 8 to 4, B: 4 to 2, C: 2 to 1 and columns tick 1, tick 2, tick tau, tick K+1, tick K+2; lot labels advance diagonally. Panel (b): in one full tick, lots tau, tau minus 1 and tau minus 2 feed a factor for each of six physical orders, giving one joint extraction with limiting log-yield C-star over 6 per factor." caption="Staggered lots: each tick processes three different lots at three different depths so their capacities can be pooled in one extraction (Staggered extraction for exact matrix multiplication over every field, Figure 1)." />

The rank bound for powers of $\mathrm{CW}_5$ in this paper avoids interpolation entirely: it reads the constant coefficient of a Laurent-polynomial identity with integer coefficients, which works over any field, finite ones included. I expanded that identity (its equation 8) for $q = 1$ to 6; the negative powers cancel and the constant term is exactly $\mathrm{CW}_q$. The final number is $3(8\ln 7 - H_0 - C_*)/S_*$, and from the paper's own enclosures for $H_0 + C_*$ and $S_*$ I get 2.3710548860067453. Its tables already certify 2.371056 with a margin of about $4\times 10^{-6}$. This is a credible, careful, small improvement, and it is the one I'd bet on first. It is also, if the 9/4 paper holds, mostly of interest for positive characteristic.

I looked for the step in the 9/4 argument that needs characteristic zero and did not find an essential one: the Fourier step only needs some modulus larger than $3(M-1)$, which can be chosen prime to $p$, with roots of unity in an extension field. The release doesn't claim 9/4 in positive characteristic, though, and I wouldn't either without someone going through the spectral appendix over finite fields.

#### What Lean actually pins down

The Comparator challenge file restates everything from scratch, on top of Mathlib only. The cost model is a straight-line program: gates are constants, inputs, and binary add, subtract and multiply; constants are free and each arithmetic gate costs one.

```lean
-- lean/ComparatorChallenges/MatrixMultiplication.lean:89-95
/-- Positive slack, one uniform constant, and a correct program at every size. -/
def AdmissibleExponent (F : Type*) [Field F] (τ : ℝ) : Prop :=
  ∀ ε : ℝ, 0 < ε → ∃ C : ℝ, 0 < C ∧
    ∀ n : ℕ, 1 ≤ n → ∃ P : MatrixAlgorithm F n n n,
      P.Correct ∧ (P.cost : ℝ) ≤ C * (n : ℝ) ^ (τ + ε)

def omega (F : Type*) [Field F] : ℝ := sInf {τ : ℝ | AdmissibleExponent F τ}

-- lines 118-120
theorem complex_omega_le_nine_quarters :
    Arithmetic.omega ℂ ≤ (9 : ℝ) / 4 := by
  sorry
```

The `sorry` is the challenge; the solution module `OAI.LinearAlgebra.MatrixMultiplication.Main` proves it as `AuxiliarySeparation.omega_le_nine_quarters` (`Main.lean:14-16`), and that subtree's docstring cites "Theorem 1.1" of the 9/4 paper. The same file pins $\alpha(\mathbb{C}) > 93/200$ and $\omega(\mathbb{C}; 1, 0.709, 1) < 523/250$. A second challenge, `MatrixFields.lean:109-112`, states `Arithmetic.omega F < 2371054886006746 / 10 ^ 15` for every field `F`, with every numerical condition discharged inside Lean rather than assumed. Both configurations permit only `propext`, `Quot.sound` and `Classical.choice`.

I checked the usual trap in statements like this. `sInf` of an empty or unbounded-below set of reals is 0 in Mathlib, which would make "≤ 9/4" vacuous. Here the set is non-empty (the schoolbook method makes $\tau = 3$ admissible) and bounded below (a correct program needs at least one multiplication, so no negative exponent works), so the infimum is the real thing. The model is the standard arithmetic one. It says nothing about bit complexity or numerical stability.

The scale is large. The `MatrixMultiplication` tree is 509 files and 83,302 lines, of which the 9/4 proof (`AuxiliarySeparation/`) is about 10,900; the all-fields tree is 354 files and 181,751 lines. `grep` finds no `sorry`, `axiom` or `native_decide` in either. The fixed-point step leans on a patched copy of an external Brouwer library (harfe/fixed-point-theorems-lean4), which Comparator's axiom check covers transitively. I did not build any of it; the release's own instructions require a full Mathlib toolchain. If it builds and Comparator passes, the 9/4 bound is machine-checked in a model I'd accept.

#### How plausible is it?

Put plainly, the 9/4 result needs three things to be true. Strassen's duality, in the special form of Lemma 2.2, which is established mathematics. The separation inequality, which is new but elementary, and which I verified as an identity on small cases and as an inequality on every character I know. And the bookkeeping that applies it under all six leg permutations and multiplies the results, which is exactly the kind of bookkeeping where a human proof goes wrong and a Lean proof doesn't. The weight of the whole result sits on Proposition 3.1 and Corollary 3.2. If there is a mistake, I'd look there first, in how the per-character exponent $p_X^{(\pi)}$ is matched to the permuted tensor; I traced it and it holds.

What makes me take it seriously: the proof is short, uses no numerical certificate, routes around every known barrier for a principled reason, saturates at its own stated constant when I optimise over its ingredients, and is formalized against a self-contained statement. What keeps me from saying more: it is a few days old, no specialist has published a reading of it, and "every character satisfies these inequalities" is a strong statement about an object (the asymptotic spectrum over $\mathbb{C}$) that is still not fully known.

Ranked by how readily I'd believe them today: the every-field 2.371054886 (incremental, certified, formalized), then 9/4 over $\mathbb{C}$ (short, checkable, formalized, unrefereed), then the rectangular bounds (formalized, but resting on the long paper), then the 2.258 square bound in positive characteristic (unformalized transfer with uncomputed exceptions), which is the weakest claim in the family and also the least important.

#### What it means if you write GEMM kernels

Nothing, directly. Every bound since 1990 is a *galactic* algorithm: correct, asymptotically faster, and slower than the cubic method at every size that fits in memory. The 9/4 result is further from practice than most, because it is non-constructive; it proves a fast base case exists without exhibiting it. The algorithms that matter for hardware are the blocked $O(n^3)$ schedule that cuBLAS, CUTLASS and every BLAS implement, and occasionally Strassen–Winograd for very large dense products, where one level of recursion saves 1/8 of the multiplications at the cost of extra additions, workspace and weaker error bounds. The site's [DeepGEMM-Ascend teardown](/articles/deepgemm-ascend) found a dense BF16 matrix multiply running at 431 of 432 peak TFLOPS. At that point the hardware's fixed-shape tensor-core instructions and memory traffic set the cost, and an exponent says nothing about either.

Where it matters is on paper. Every algorithm stated as $O(n^\omega)$ inherits the new exponent: determinant, inverse and linear systems (the 2.258 paper says so for its own bound), triangle detection, transitive closure, Valiant's context-free parsing, and the many graph algorithms built on rectangular products, where $\alpha$ and $\omega(1,k,1)$ appear directly. If 9/4 survives review, a lot of exponents in a lot of papers change in the second decimal place.

#### Reactions so far

The release went public on 6 October; the reactions I could find are hours old and none is a technical reading of this family. Gil Kalai's [post](https://gilkalai.wordpress.com/2026/10/07/updates-sharing-ai-progress-on-mathematics-amazing-and-my-lecture-plans/) is about the release as a whole: "the results will needs to be verified and digested by human mathematicians in the months to come." On the Hacker News thread about [OpenAI's announcement](https://openai.com/index/sharing-ai-progress-in-mathematics/), one commenter wrote "I find the claimed matrix multiply result (w&lt;= 2.25) shocking. I hope it holds up." Scott Aaronson's most recent post predates the release, and I couldn't load the Computational Complexity blog from here. When people who work on the asymptotic spectrum write about Proposition 3.1, that will be the reaction worth reading.

### The FFT is not optimal (in the right model, by a hair)

The second claim I expected to be wrong is about the FFT. Family 130 says you can compute a discrete Fourier transform of length $n$ in $O(n(\log n)^{1-\delta})$ operations, with $\delta = 10^{-13}$, for every $n$. Anyone who has written an FFT has absorbed $n\log n$ as a law of nature. It is not one. It is an upper bound that nobody has beaten and a lower bound that nobody has proved, except in models that tie the algorithm's hands. Once I read the two papers behind this family, the claim stopped looking crazy and started looking precise: it beats $n\log n$ in exactly the model where no lower bound was ever known, by an amount so small that it will never touch a real FFT, and the part that matters for the open question is checked in Lean.

The family is two manuscripts. [An explicit power saving for the exact discrete Fourier transform](https://github.com/openai/math/blob/main/preprints/An-explicit-power-saving-for-the-exact-discrete-Fourier-transform-September-25-2026/main.pdf) (31 pages) gives a single uniform algorithm with the explicit exponent. [Finite tensor savings and exact Fourier circuits](https://github.com/openai/math/blob/main/preprints/Finite-tensor-savings-and-exact-Fourier-circuits-September-25-2026/main.pdf) (49 pages) proves a weaker, non-constructive statement by a completely different route, and that weaker statement is the one in Lean.

#### What was actually open

Cooley and Tukey's 1965 [radix-2 algorithm](https://doi.org/10.1090/S0025-5718-1965-0178586-1) computes $F_n x$, with $F_n = (\zeta_n^{jk})$ and $\zeta_n = e^{2\pi i/n}$, in $O(n\log n)$ operations at highly composite lengths. Good's coprime factorization (1958) and Bluestein's chirp trick ([1970](https://doi.org/10.1109/TAU.1970.1162132)) extend that to every $n$. In sixty years the constant moved and the shape did not. Split radix brought the real-operation count from $5n\log_2 n$ to $4n\log_2 n$; Alman and Rao ([STOC 2023](https://arxiv.org/abs/2211.06459)) got it to $\tfrac{15}{4}n\log_2 n$ by way of Walsh–Hadamard non-rigidity. A better constant does not answer the question anyone cares about, which is whether $cn\log n$ is required for some fixed $c$.

The lower-bound side is where the modelling choices bite, so it is worth being exact about it.

- **Morgenstern (1973).** In a [two-page JACM note](https://doi.org/10.1145/321752.321761), he showed that a linear algorithm whose constants are bounded in modulus needs order $n\log n$ steps for the *unnormalized* $F_n$. The argument is a potential function: $|\det F_n| = n^{n/2}$, and a gate with bounded coefficients can only multiply the determinant by a bounded factor, so you need about $\tfrac{1}{2}n\log_2 n$ gates to get there. With unbounded constants the argument says nothing, since one multiplication by $10^{100}$ buys you as much determinant as you like.
- **Ailon (2013).** [arXiv:1305.4745](https://arxiv.org/abs/1305.4745) took the *normalized* transform, whose determinant has modulus one, so Morgenstern's potential starts at zero. He proved $\Omega(n\log n)$ for layered circuits of unitary $2\times 2$ gates on exactly $n$ registers, using matrix entropy as the potential. The FFT lives in this model; almost nothing else does.
- **Ailon (2014).** [arXiv:1403.1307](https://arxiv.org/abs/1403.1307) relaxed that to any scaling of $F_n$ and allowed extra memory, at the price of requiring every intermediate composition of gates to be $R$-well-conditioned, and got $\Omega(n\log n/R)$. Read the other way, this is a speed-versus-conditioning trade-off: any algorithm that beats $n\log n$ by much must pass through badly conditioned intermediate maps.
- **Papadimitriou (1979).** His [JACM paper](https://doi.org/10.1145/322108.322118) proved the FFT optimal inside a restricted, graph-structured class of algorithms. I couldn't read the full text (the ACM copy is paywalled from here), so I'll only say that, like the others, it constrains the form of the algorithm.

Ailon's 2014 abstract describes a super-linear lower bound in the plain linear circuit model as "a long standing open problem for over 40 years". That plain model is: inputs, gates computing $u+v$, $u-v$ or $\lambda u$ for any fixed complex $\lambda$, each gate costing one, unlimited reuse, no bound on coefficients, conditioning or memory. That is the model of the second paper, and it is word for word the model in the Lean statement below. So yes, this is the standard open problem. The answer is the surprising one: the conjectured $\Omega(n\log n)$ is false there.

The integer-multiplication story is the useful comparison. Harvey and van der Hoeven ([Annals, 2021](https://hal.science/hal-02070778)) reached $O(n\log n)$ for multiplying $n$-bit integers, and $n\log n$ was conjectured optimal, with a lower bound known only conditionally ([Afshani et al., 2019](https://arxiv.org/abs/1902.10935), assuming a network-coding conjecture). The same release's family 109 claims $O(n(\lg n)^{1-\kappa})$ with $\kappa = 2^{-182}$ on a multitape Turing machine, and the explicit FFT paper takes its key finite gadget from that manuscript ("Proposition 8 and Section 3.5"). The two results share an engine.

<FourierModelToggles />

Flip the toggles and the pattern is plain: every proved bound needs at least one restriction that the new construction breaks. It uses coefficients of unbounded size (Morgenstern is out), shears and multi-wire gates that are not unitary (Ailon 2013 is out), and it makes no promise about conditioning (Ailon 2014 is out). Nothing here contradicts a theorem. It walks through the one door those theorems left open.

#### The model, assumption by assumption

The explicit paper states its model in a paragraph, and every clause is there for a reason.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/fft-fig0.png" alt="Page 2 of the explicit power saving paper: the absolute constants m equals ten to the six, W star equals two to the seventy-one, Delta equals 6871402692000000, lambda and theta, the decimal value of theta, Theorem 1.1 with running time n log n to the theta times log log n to the four minus theta, and Corollary 1.2 with exponent one minus ten to the minus thirteen." caption="The whole claim on one page: the five constants that fix the exponent, Theorem 1.1 and Corollary 1.2 (An explicit power saving for the exact discrete Fourier transform, page 2)." />

*Exact complex arithmetic.* Registers hold exact complex numbers and every field operation costs one. That is the algebraic-complexity convention, and it is what Cooley–Tukey's $n\log n$ counts too. It is not bit complexity. The paper is blunt: "no finite-precision or numerical-stability assertion".

*Unrestricted coefficients.* The input is only ever added, subtracted or multiplied by input-independent prepared scalars, but those scalars can be as large as needed and intermediate values can blow up. This is the clause that steps around Morgenstern. I'll show below what it costs in floating point.

*A supplied root of unity.* The algorithm is given one specified root $\zeta_{D_*}$ with $D_* \lt 1024n^3$ and builds every other constant from it by powering and field operations. This sounds like cheating until you see the alternative: field operations on rationals cannot produce $e^{2\pi i/n}$ at all, so every FFT in this model is handed its roots somehow. The paper charges the preparation of everything else, including conjugates (it re-evaluates each coefficient's arithmetic program with roots inverted, rather than conjugating data).

*Charged preparation and indexing.* Integer operations on $O(\log n)$-bit words cost one, and the schedule, the address permutations and all scalar preparation are counted. This is the difference between the two papers. The second paper is non-uniform: coefficients are free and the circuit for each $n$ simply exists. The first paper pays for everything, which is why it needs a whole section on traversing mixed-radix addresses in linear time ("scanning all the axes anew at each array entry would lose the saving").

So paper two answers the textbook circuit question. Paper one goes further and turns it into an honest algorithm with a number attached. Both stay in exact arithmetic.

#### How you beat $n\log n$: save on one tiny matrix, then amplify

The core idea is in the second paper's abstract and it is genuinely elegant. Take a fixed invertible $q\times q$ matrix $A$ that is not monomial (not just a permutation times a diagonal). The obvious way to apply $A^{\otimes b}$ to a vector of length $q^b$ is axis by axis: $b$ passes, each making $q^{b-1}$ calls to $A$, so $bq^{b-1}$ calls. Call it a *finite win* if some cleverer word, with free monomial maps between calls, does it in fewer. One finite win, at one fixed size, is enough to break $n\log n$.

Why? Amplification. If $A^{\otimes b}$ costs $g \lt bq^{b-1}$ calls, then for $A^{\otimes k}$ you group the $k$ axes into blocks of $b$, run the winning word on the blocks, and notice that each call slot of the tensored word decomposes into a direct sum of smaller tensor powers $A^{\otimes r}$, which you handle recursively. The normalized cost obeys a recurrence of the form $f(k) \le C_0 + g\,\mathbb{E}f(K)$ with $K$ binomial, and because $g$ is strictly below the break-even point you get $A^{\otimes k}$ in $O(q^k (k+1)^\alpha)$ for some $\alpha \lt 1$, against $kq^k$ for the axis-by-axis method. A fixed constant-size saving turns into a saving in the *exponent* of $k$.

The explicit paper makes this concrete with one specific matrix:

$$C = \tfrac12\begin{pmatrix}1+i & 1-i\\ 1-i & 1+i\end{pmatrix} = H\,\mathrm{diag}(1,i)\,H,\qquad H = \tfrac{1}{\sqrt2}\begin{pmatrix}1&1\\1&-1\end{pmatrix}.$$

It is unitary, $\det C = i$, and $C^2$ is the swap. So $C^{\otimes k}$ is a Walsh–Hadamard conjugate of a diagonal of fourth roots of unity. Its determinant has modulus one, which is why Morgenstern's argument cannot even see it.

The saving for $C$ comes from a scalar network inherited from the integer-multiplication paper. Take $h = 100$, the $v = \binom{100}{3} = 161{,}700$ three-element subsets of $\lbrace 1..100\rbrace$, and call two triples neighbours when they intersect in 0 or 2 elements. Eight rows of linear updates, built from maps $V, G, J, R$ on side wires and central wires, add $x$ into $y$ while restoring every auxiliary wire to its *arbitrary* starting value; the proof is the identity $RG + JV = I$. Three stages of that give a signed exchange of two banks. Then the trick: replace each scalar wire by an array indexed by $\mathbb{F}_2^{m}$ with $m = h^3 = 10^6$, give every gate a binary subspace label, and insert a change of frame along each edge. Moving between nested labels costs one directional kernel $C_z = aI + bR_z$ per dimension of the gap (Lemma 2.2), and because frames commute with the scalar gates, the whole network collapses to $C^{\otimes m}$ applied to every one of its $W$ physical arrays. The total dimension of all frame changes comes to $s = Wm - \Delta$ with

$$W = 1{,}873{,}807{,}244{,}643{,}542{,}670{,}000,\qquad \Delta = 2v^2\bigl(v - 3h(h+1)\bigr) = 6{,}871{,}402{,}692{,}000{,}000.$$

That is the finite win: applying $C^{\otimes m}$ to $W$ arrays the obvious way takes $Wm$ directional steps, and the network takes $\Delta$ fewer. Pad the roles to $W_* = 2^{71}$ arrays (each padding array costs its plain $m$ steps, so the absolute saving survives), recurse on the fibres, and you get Theorem 2.6: $C^{\otimes k}$ in $O(2^k(k+1)^\theta)$, with

$$\lambda = m - \frac{\Delta}{W_*},\qquad \theta = \log_m \lambda = 0.99999999999978935615699598\ldots$$

One detail I liked: the network has to restore auxiliary wires that hold *arbitrary* data, not zeros. In the recursion, every one of the $2^{71}$ roles is a live fibre of somebody else's transform, so a gadget that only worked on clean scratch space would be useless (the paper's Remark 2.7). The same "borrow dirty memory and give it back" idea, which the authors trace to Bennett's reversible computation and to catalytic space, shows up again in the Fourier compiler.

#### From $C^{\otimes k}$ to an actual Fourier transform

A saving on tensor powers of one $2\times 2$ matrix is not yet a DFT. Three classical tools, each pushed to exact width, close the gap.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/fft-fig1.png" alt="Box diagram of the proof: on the left, the fixed network on arbitrary roles giving C tensor k in time two to the k times k plus one to the theta (Theorem 2.6); on the right, exact-width Fourier words with O of one plus log to the fourth r common slots on exactly r coordinates (Proposition 3.1); both feed sector packing and synchronization (Proposition 4.2), which feeds the small-prime working length, CRT and chirp convolution of Theorem 1.1." caption="How the pieces connect: the tensor-power algorithm and the exact-width Fourier words are separate inputs to synchronization (An explicit power saving for the exact discrete Fourier transform, Figure 1)." />

First, Good–Thomas. Pick a working length $L = r_1 r_2\cdots r_\ell \cdot 2^e$ between $2n$ and $4n$, where the $r_j$ are the first $\ell$ odd primes, so $\ell = \Theta(\log n/\log\log n)$ and every $r_j$ is $O(\ell^2)$. Chinese-remainder reindexing turns $F_L$ into $F_{r_1}\otimes\cdots\otimes F_{r_\ell}$ (times the power-of-two part) with no twiddle factors. I checked the index map on $L = 15$; it is the textbook one.

Second, each small $F_r$ is compiled into a short word of layers on *exactly its own* $r$ coordinates. This is the part that surprised me most. The route is a Newton-basis factorization $F_r = N D_N^{-1} N^{T}$, where $N$ is two diagonals around a triangular Toeplitz matrix (the paper derives it; it matches Kuznetsov's LDLᵀ of geometric Vandermonde matrices). The Toeplitz factor is applied by divide and conquer, with each off-diagonal block written through displacement rank 3 as a few convolutions, and every convolution's temporary storage is borrowed from coordinates that already hold other data and restored by a compute–copy–uncompute replay. The result: $O(1 + \log^4 r)$ layers, each either a monomial or disjoint copies of $C$ on pairs. Any shear $y \leftarrow y + tx$ becomes exactly three copies of $C$ between diagonal scalings, through the identity $H'\,\mathrm{diag}(2,1)\,H'\,\mathrm{diag}(-3,1)\,H' = \begin{pmatrix}-8&-10\\0&-6\end{pmatrix}$. Width matters because these $r$-coordinate sets are tensor factors; padding each one would multiply the total size.

Third, synchronization. Pad all the axes' words to a common sequence of slots. A monomial slot is cheap. A pair slot, tensored across $\ell$ axes, splits into sectors, and a sector with $k$ active pairs is precisely $C^{\otimes k}$ on $2^k$ contiguous entries:

<Figure src="https://ai.thesatyajit.com/articles/openai-math/fft-fig2.png" alt="Two axes each split into a pair and a singleton give a two by two grid of sectors labelled C tensor C, C, C and 1, which are packed into contiguous blocks of widths 4, 2, 2 and 1." caption="A simultaneous pair layer on two axes becomes four independent tensor-power calls, packed into contiguous sectors of widths 2^k (An explicit power saving for the exact discrete Fourier transform, Figure 2)." />

That is where the saving lands: each pair slot costs $O(L(\ell+1)^\theta)$ instead of $O(L\ell)$. Multiply by the $O((\log\log n)^4)$ slots and you get Theorem 1.1, $O(n(\log n)^\theta(\log\log n)^{4-\theta})$. A Bluestein chirp convolution (three transforms of length $L$, one of them the fixed kernel, charged as preparation) gives every $n$. The corollary to $\delta = 10^{-13}$ is just $1-\theta \approx 2.1\times10^{-13}$, which leaves room to swallow the $(\log\log n)^3$.

The non-constructive paper reaches the same circuits from the other side. It never names a winning matrix. It assumes there is no finite win for *any* matrix, builds from that assumption a "price" on all invertible matrices with exact tensor and direct-sum laws (finite comparisons, convex separation, then Tychonoff compactness over all matrices), and drives the prices into a contradiction using signed Jordan–Wigner-style operators: one description prices a reflection at $dw - O(w)$, another at $\tfrac34 dw + O(w)$, impossible for $d \gt 30$. A packing step that puts many comparison programs on disjoint tensor sectors with random start times is the combinatorial heart of it:

<Figure src="https://ai.thesatyajit.com/articles/openai-math/fft-fig3.png" alt="A chain of monomial maps M0, M1 up to Mt interleaved with call slots B1 up to Bt; each Bt expands into gather Pt, one call to J holding k_i copies of each A_i, and scatter Pt inverse, with identity spectators carrying arbitrary data." caption="Packing comparisons: each simultaneous call slot is one embedded call to a block-diagonal J, with gather and scatter permutations (Finite tensor savings and exact Fourier circuits, Figure 1)." />

#### How small is $10^{-13}$?

The headline says "below $n\log n$". It is worth putting a number on "below". I recomputed every constant in (1.1) and (2.4)–(2.10) with exact integers and they all match the paper, including $W$, $s$, $\Delta = 2v^2(v-3h(h+1))$, the fraction $\Delta/(mW_*) = 1717850673/590295810358705651712$ and $\theta$ to every printed digit.

Each level of the recursion saves a fraction $\varepsilon = \Delta/(mW_*) \approx 2.9\times10^{-12}$ of the work, about three parts in a trillion. Each level also shrinks the tensor order by a factor of a million. To save 1% in total you need $(1-\varepsilon)^d = 0.99$, so $d \approx 3.45\times10^{9}$ levels, which means a tensor order around $10^{6\times 3.45\times10^9}$. Said through the bound itself: the factor $(\log n)^{-(1-\theta)}$ equals 0.99 only when $\ln\ln n \approx 4.77\times10^{10}$. The *logarithm* of such an $n$ has about $2.07\times10^{10}$ decimal digits.

That is still flattering, because it ignores two things that only make it worse.

- The bound carries $(\log\log n)^{4-\theta}$, roughly $(\log\log n)^3$ more than $n\log n$. The bound itself drops below $n\log n$ only when $(\log n)^{2.1\times10^{-13}}$ outgrows $(\ln\log n)^3$, which, with every constant set to one, first happens near $\ln\ln n \approx 4.8\times10^{14}$.
- The constants. The recursion only starts at $k \ge K = 72\times10^6$ tensor axes, and the tensor order in the FFT is at most $\ell$, the number of primes. Reaching $\ell = 7.2\times10^7$ needs $n$ beyond the product of the first 72 million odd primes, about $e^{1.44\times10^9}$. And a single call to $C^{\otimes k}$ runs all $W_* = 2^{71}$ roles at its top level, with $2^{71}-1$ of them zero-filled (proof of Theorem 2.6), a constant factor of about $2.4\times10^{21}$ that the $(\log n)^{-\delta}$ factor has to repay. One level of the network also pays a fixed pointwise overhead $A$ per entry against a saving of $\varepsilon k$, so it breaks even only when $k \gt A/\varepsilon \approx 3.4\times10^{11}\cdot A$.

The paper does not hide any of this ("The constants in the algorithm are enormous, so these are asymptotic existence and explicitness results, with no useful crossover estimate"). The interactive below makes the scale concrete. Drag $n$ through the sizes anyone actually transforms and watch the best-case saving sit at $10^{-13}$:

<FftCostExplorer />

One curiosity. $\delta$ is a function of the gadget size, not a barrier. The only inequality the dimension count needs is $v \gt 3h(h+1)$, and $h = 100$ looks chosen for comfort. Plugging smaller $h$ into the same formulas, $h = 25$ (with $W_* = 2^{46}$) would give $1-\theta \approx 3.5\times10^{-10}$, about 1,700 times larger. I haven't checked that the rest of the construction goes through unchanged at $h = 25$, and it would not change any conclusion: $10^{-10}$ is just as galactic. The Lean library also proves a quadratic layer bound for the local Fourier words, $8{,}388{,}608(\lfloor\log_2 r\rfloor+2)^2+5$ layers (`ToeplitzCross.lean:274`), stronger than the paper's quartic one, which would shave the $\log\log$ exponent if carried over. The paper doesn't use it.

#### FFTW and cuFFT are unaffected

There are three separate reasons, and any one of them is enough.

First, scale, as above. Second, the claim lives in exact arithmetic with unbounded coefficients, and those coefficients are where floating point dies. I ran the paper's own local ingredient in float64: compute $F_r x$ as $N\,(N^T x)/\mathrm{diag}(N)$ from Lemma 3.2's exact formula. The Newton denominators $H_j = \prod_{s\le j}(1-u^s)$ grow exponentially, and so does the error:

| r | max $\lvert H_j\rvert$ | cond(N) | relative error in float64 |
|---|---|---|---|
| 16 | 54 | $1.5\times10^{4}$ | $1.0\times10^{-13}$ |
| 32 | 999 | $4.5\times10^{8}$ | $1.3\times10^{-8}$ |
| 48 | $1.6\times10^{4}$ | $1.4\times10^{13}$ | $4.2\times10^{-5}$ |
| 64 | $2.5\times10^{5}$ | $4.2\times10^{17}$ | 0.71 |

At $r = 64$ the answer is noise; numpy's FFT on the same input is right to $5\times10^{-14}$. The real construction compiles primes up to $O(\log^2 n)$, far past 64. The zero-test-free shear split in (3.14), which writes a coefficient $\mu$ as $(1+|\mu|^2) + (\mu - 1 - |\mu|^2)$, loses $2\log_{10}|\mu|$ digits to cancellation in the same way. None of this is a flaw in the papers: Ailon's 2014 theorem says any big speed-up must go through ill-conditioned intermediate maps, and here they are.

Third, production FFTs are not optimizing this cost model. FFTW and cuFFT live and die by memory traffic, cache blocking and vector width; a $2^{71}$-role batch structure is not a memory layout. Practical speed comes from the constant and the data movement, which is why the Alman–Rao $\tfrac{15}{4}$ result, not this one, is the closer cousin to anything you would implement.

#### What is checked, and what isn't

The Lean side is unusually strong for this release and unusually narrow. The Comparator challenge states the textbook circuit model in a few dozen self-contained lines, with nothing from the solution library in its definitions:

```lean
-- lean/ComparatorChallenges/ExactFourier.lean:13-16, 45-58
inductive Gate (w : ℕ) where
  | add (i j : Fin w)
  | sub (i j : Fin w)
  | scale (c : ℂ) (i : Fin w)

noncomputable def fourierMatrix (n : ℕ) : Matrix (Fin n) (Fin n) ℂ :=
  fun j k => zeta n ^ (j.val * k.val)

def MainStatement : Prop :=
  ∀ c : ℝ, 0 < c → ∀ N₀ : ℕ, 2 ≤ N₀ → ∃ n : ℕ, N₀ ≤ n ∧
    ∃ C : Circuit n, C.Computes (fourierMatrix n) ∧
      (C.size : ℝ) < c * (n : ℝ) * Real.logb 2 (n : ℝ)

theorem main_theorem : MainStatement := by
  sorry
```

The `sorry` is the challenge stub, as with every Comparator file; the solution is `OAI.ExactFourier.main_theorem` at `lean/OAI/Computability/FourierCircuit/Main.lean:277`, which is `win_to_fourier finite_win`. The library behind it is 51 files and 19,103 lines; I grepped it for `sorry`, `axiom`, `native_decide` and `implemented_by` and found none, and the Comparator config permits only `propext`, `Quot.sound` and `Classical.choice`. Note what the gate type says: `scale (c : ℂ)` takes any complex constant, outputs may name any available value, and there is no cost for preparing `c`. That is exactly the unrestricted non-uniform model, with every scalar multiplication charged. The library also proves the liminf form, `liminf L(n)/(n log₂ n) = 0` for the literal minimum circuit size (`Main.lean:314`). I did not build it or run Comparator, and `formalization.yaml` lists the review status as "unchecked".

What Lean does not cover is the headline number. The uniform, every-$n$ algorithm with $\delta = 10^{-13}$, the charged preparation and the address arithmetic exist only on paper. The repository's [own scope note](https://github.com/openai/math/blob/main/lean/docs/130.md) is careful about this ("The result is subsequential, with no all-length, bounded-coefficient, conditioning, or bit-complexity claim"), but the family's one-line summary puts the $\delta = 10^{-13}$ sentence right next to a "Lean" link, which a quick reader will take as covering it. It doesn't. Also, the formal proof finds its finite win by contradiction, so the Lean development never exhibits the $10^{6}$-bit network; the explicit paper's Lemma A.1 is what connects that network to the same finite-win definition.

The finite pieces I could check by hand all hold. Besides the constants above: $RG + JV = I$ exactly, for every $h$ from 5 to 8 (the identity does not depend on $h$); Lemma 2.2's frame-change formula, by brute force over nested subspaces of $\mathbb{F}_2^4$ and $\mathbb{F}_2^5$ (ten cases, including $\Phi_D = C^{\otimes m}$); the three-$C$ shear identity for real, complex and large $t$; and $F_r = N D_N^{-1} N^T$ for $r = 2, 3, 5, 7, 8$. The core of that check fits in a few lines:

```python
# constants of (1.1) and (2.4)-(2.10), exact integers
from math import comb
from fractions import Fraction
h = 100; v = comb(h, 3); N0 = v**3; m = h**3; I0 = 3*v**2
dnbr = comb(h-3, 3) + 3*(h-3)                 # |S ∩ T| in {0, 2}
W = 2*N0 + I0*(v*dnbr + 101)                  # 1873807244643542670000
Delta = 2*v*v*(v - 3*h*(h+1))                 # 6871402692000000
assert W*m - (W*m - 2*N0 + 2*I0*101*100) == Delta
eps = Fraction(Delta, m * 2**71)              # 1717850673/590295810358705651712
# theta = log_m(m*(1-eps)) = 0.999999999999789356156995981891...
```

For contrast, the thing it is all measured against, the radix-2 FFT, is ten lines:

```python
import numpy as np
def fft(x):                      # len(x) a power of two
    n = len(x)
    if n == 1:
        return x
    even, odd = fft(x[0::2]), fft(x[1::2])
    w = np.exp(-2j * np.pi * np.arange(n // 2) / n) * odd   # n/2 twiddles
    return np.concatenate([even + w, even - w])             # n adds/subs
```

On reception: as of the day after the release I found no published expert analysis of either Fourier paper. The public reaction has gathered around the integer-multiplication sibling, posted on X by [@AcerFur](https://x.com/AcerFur/status/2107606747972309163) ("I was definitely very surprised when this one came in"), and the replies split between disbelief ("this would mean that FFT can be computed in less than O(n log n), which I think has been proved impossible") and jokes about $\kappa = 2^{-182}$. The first reply has it exactly backwards, which is the most useful thing to take from this section: it was never proved impossible. It was proved impossible for algorithms with bounded constants, unitary gates or good conditioning, and these papers show that the unrestricted version was genuinely different.

My read: as mathematics this is the most clarifying result in the release's theory-of-computing group, because it settles a question people stated carefully for decades, and the part that settles it is formally verified. The explicit $\delta$ is a real but unverified strengthening. As engineering it changes nothing, and the paper says so first.

### The Unique Games Conjecture, claimed proved

<Figure src="https://ai.thesatyajit.com/articles/openai-math/theoretical-cs-fig4.png" alt="Page 12 of the release's overview catalogue: the start of the Theoretical computer science section, listing family 102 (The Unique Games Conjecture and optimal approximation thresholds), 103 (L = RL = BPL), 104 (quasipolynomial mean-payoff games), 105 (2-to-1 games with perfect completeness), 106 (hardness of coloring three-colorable graphs) and 107 (matrix multiplication with exponent at most 9/4)." caption="How the release introduces this shelf: the first six theoretical-CS entries in the catalogue (openai/math overview.pdf, page 12)." />

Subhash Khot posed the Unique Games Conjecture in 2002 ("On the power of unique 2-prover 1-round games", STOC 2002), and it became the most productive unproved statement in complexity theory. The object is simple. You have a graph, a set of labels, and on every edge a permutation that says "if the left endpoint gets label $i$, the right endpoint must get $\pi(i)$". Finding a labeling that satisfies every edge is easy, by propagation. The conjecture says that once you only promise 99% of edges can be satisfied, telling that case apart from one where no labeling satisfies even 1% is NP-hard.

Why anyone outside complexity theory should care is the reason I put it first. Assuming the conjecture, a long list of algorithms that engineers already use turn out to be exactly optimal. Khot, Kindler, Mossel and O'Donnell (SICOMP 2007) showed that Goemans and Williamson's random-hyperplane rounding for Max-Cut, with its odd constant $\alpha_{GW} \approx 0.87856$, cannot be beaten. Khot and Regev (JCSS 2008) showed that the textbook factor-2 vertex cover (take both endpoints of a maximal matching) cannot be beaten. Raghavendra (STOC 2008) showed something stranger and more general: for every constraint satisfaction problem whatsoever, one fixed semidefinite program, the "basic SDP", plus a generic rounding scheme achieves the best ratio any polynomial-time algorithm can. That last result is the one the release's reasoning trace starts from. The prompt OpenAI published asks the model to prove NP-hardness at the basic-SDP threshold for every finite Max-CSP *without* the conjecture, and the trace shows the model deciding, early, that the only route runs through proving the conjecture itself.

The conjecture was never a safe bet. Arora, Barak and Steurer (FOCS 2010) found a subexponential-time algorithm for Unique Games, something no NP-hard problem is believed to have, and for years a respectable minority expected the conjecture to be false. Then Khot, Minzer and Safra, building on Dinur, Khot, Kindler, Minzer and Safra and on Barak, Kothari and Steurer, proved the 2-to-2 Games Theorem (FOCS 2018). That gave Unique Games hardness with completeness about one half: NP-hard to tell "half the edges satisfiable" from "almost none". The missing piece was always completeness close to one.

#### What the paper claims

[The Unique Games Theorem](https://github.com/openai/math/blob/main/preprints/The-Unique-Games-Theorem-September-23-2026/paper.pdf) (58 pages) states it in the strongest ordinary form. For every fixed $\varepsilon, \delta \in (0, 1/2)$ there is an integer $s$ and a deterministic polynomial-time reduction from 3SAT to explicit, unweighted, simple bipartite Unique Games over the alphabet $\mathbb F_2^s$, in which every constraint is a translation $a(v) = a(u) + c$, with value at least $1-\varepsilon$ on satisfiable formulas and at most $\delta$ on unsatisfiable ones. Translations over $\mathbb F_2^s$ are a known equivalent form of the conjecture (Khot, Kindler, Mossel, O'Donnell), so this is the full thing, not a variant.

Section 8 then cashes it in. It re-derives Raghavendra's consequence for every fixed CSP (Corollary 8.1: NP-hard to beat the basic SDP on its own integrality-gap instances), optimal hardness for every ordering CSP, so Maximum Acyclic Subgraph at $1/2$ and Betweenness at $1/3$ (Corollary 8.2), and any-constant hardness for Multicut, nonuniform Sparsest Cut, Min-2CNF deletion and Correlation Clustering (Corollary 8.3). Four companion papers in the same family go further and prove the most-cited consequences *directly* from standard Label Cover, without the Unique Games paper: [Max-Cut beyond 0.878 on unweighted graphs](https://github.com/openai/math/blob/main/preprints/A-Direct-Proof-of-Optimal-Max-Cut-Hardness-September-23-2026/paper.pdf), [Vertex Cover below 2](https://github.com/openai/math/blob/main/preprints/The-Factor-Two-Hardness-Threshold-for-Vertex-Cover-September-23-2026/paper.pdf) (minimum cover below $(1/2 + 1/m)n$ versus above $(1 - 1/m)n$), and any-constant hardness for [Min-UnCut](https://github.com/openai/math/blob/main/preprints/Constant-factor-hardness-of-Min-UnCut-September-23-2026/paper.pdf) and [directed feedback vertex set](https://github.com/openai/math/blob/main/preprints/Constant-factor-hardness-of-directed-feedback-vertex-set-September-23-2026/paper.pdf). I like that design choice. If the main proof had a hole, the four most practically relevant corollaries would still stand on separate arguments.

#### How the proof goes

<Figure src="https://ai.thesatyajit.com/articles/openai-math/theoretical-cs-fig1.png" alt="Flow diagram with two columns. Left, 'The reduction': 3SAT to weighted parity gap instance, to latent matrix test with a stable nonlinear map (Lemma 3.1), then rounding, subdivision and parallel repetition, to unweighted simple bipartite Unique Games (Theorem 1.1). Right, 'Soundness on a NO input': test acceptance at least 0.99 gives decoded advice strategies with agreement at least gamma (Proposition 5.3); a NO parity instance gives every pair of strategies agreement at most exp(-c k^(1/3)) (Lemma 6.2); the two contradict for large k." caption="The whole argument on one diagram: the reduction on the left, and the two incompatible bounds that give soundness on the right (The Unique Games Theorem, Figure 1)." />

The obstacle is easy to see, and the paper shows it in a two-line calculation that I turned into the widget below. The 2-to-2 machinery tests a table on binary matrices: an honest prover answers with $f_z(M) = Mz$ for its true answer $z$, and the verifier perturbs $M$ by a random rank-one matrix $a\,l^\top$ and asks for the same answer. But $f_z(M + a l^\top) - f_z(M) = a\,(l^\top z)$, and $l^\top z$ is a fair coin. An honest prover passes with probability $(1 + 2^{-\ell})/2$, about one half no matter how you tune things, which is exactly the "completeness one half" that 2-to-2 left behind.

<RankOneCompleteness />

The new idea is to stop perturbing the answer directly. Lemma 3.1 builds a "latent alphabet gadget": a bigger binary space $V$ containing the answer space $K = \mathbb F_2^\ell$, a nonlinear map $C: V \to K$ that commutes with translations by $K$, and a noise distribution $\mu$ on $V$ such that $C(x + a) = C(x)$ except with probability $p$, for any small $p$ you like. Honest provers now read their answer through $C$, so the perturbation is absorbed and completeness climbs to $1 - p/2$. The other half of the lemma is what keeps soundness: any family of *linear* observations whose restriction to $K$ has rank above a threshold $r^\ast$ (fixed before $\ell$) still detects the noise with probability at least $1/8$. So cheaters who answer with low-degree, linear-looking strategies still get caught by exactly the Fourier argument the 2-to-2 proof used, and the paper reuses the Khot–Minzer–Safra inverse shortcode theorem unchanged. The gadget itself is built from recursive quadratic blocks over $\mathbb F_{2^d}$, where the nonlinear output gets more stable with each level while a harmonic rank potential limits how much linear information is lost.

Soundness is the right-hand column of the figure. If some labeling passes the test with probability 0.99, a decoding argument (Section 5) extracts "advice strategies" for a repeated parity game that agree with probability at least some fixed $\gamma \gt 0$. If the source formula is unsatisfiable, parallel repetition (the Dinur–Steurer bound) caps every pair of strategies at $\exp(-ck^{1/3})$. For large tuple length $k$ those two numbers cross, so the 0.99 labeling cannot exist.

#### What is checked, and what that buys you

All five results in the family are Comparator challenges, and the main statement is short enough to quote. The `BinaryGapReduction` structure it refers to requires a Mathlib `TM2ComputableInPolyTime` machine from binary-encoded 3SAT to explicit games, a fixed alphabet $\mathbb F_2^s$, simple bipartite output, translation constraints, and the two value bounds.

```lean
-- lean/ComparatorChallenges/UniqueGamesTheorem.lean:227-235
/-- For every fixed pair of errors in `(0, 1/2)`, binary 3SAT reduces in
polynomial time to translation Unique Games with completeness at least `1 - ε`
and soundness at most `δ`. The alphabet and machine depend only on the errors. -/
theorem theorem11 (ε δ : ℝ)
    (hε : 0 < ε) (hεhalf : ε < 1 / 2)
    (hδ : 0 < δ) (hδhalf : δ < 1 / 2) :
    Nonempty (Explicit.MachineOutputContract.BinaryGapReduction ε δ) := by
  sorry
```

The solution module is `OAI.Computability.UniqueGames.Main`. What surprised me is what lives under it: a `PCP` directory of about 40,000 lines whose file names (`AmplificationIteration`, `AssignmentTester`, expander tables) read like a Dinur-style proof of the PCP theorem, a `Repetition` directory for parallel repetition, and an `Inverse` directory of about 18,700 lines full of `KMS...` files, which is the Khot–Minzer–Safra Grassmann expansion theorem. The paper cites those as external results. The Lean development, if it compiles with only the three standard axioms, has to prove them too. I did not read those files; I am going by names and line counts. If that holds up, it is the most ambitious formalization in the release and a stronger certificate than the paper.

Two caveats are worth keeping in mind. The Section 8 corollaries that import other authors' reductions (Raghavendra's, the ordering-CSP and cut reductions) are paper arguments, not Lean. And nothing here is fast. The degree of the reduction's polynomial grows as $\varepsilon$ and $\delta$ shrink, which is forced: Arora–Barak–Steurer solve Unique Games in time $\exp(n^{\mathrm{poly}(\varepsilon)})$, so a reduction from 3SAT with a fixed polynomial blow-up would refute the Exponential Time Hypothesis. A proof of the conjecture had to look like this, and it does.

#### What changes for engineers

Less than the headline suggests, and in a useful direction. Nothing gets slower. What the theorem does is tell you when to stop looking for a better polynomial-time approximation algorithm. If your problem is Max-Cut, the Goemans–Williamson SDP with hyperplane rounding is the end of the road for worst-case guarantees; if it is vertex cover, the matching trick is; if it is any CSP you can write down, the basic SDP is. The map above shows how much of the "unknown" grey disappears. Worst-case hardness says nothing about your instances, so the practical consequence is to spend effort on exact solvers, local search and instance structure, where all the real progress on these problems has come from anyway. It also retires a generation of "assuming UGC" footnotes in papers, if the proof survives review. No expert I could find has said anything specific about this proof yet; the coverage I found (officechai, startupfortune) only repeats OpenAI's framing and Sam Altman's caveat that the headline claims are not yet confirmed by outside mathematicians.

### L = BPL: randomness doesn't help small memory

[Exact derandomization of logarithmic space](https://github.com/openai/math/blob/main/preprints/Exact-Derandomization-of-Logarithmic-Space-L-equals-RL-equals-BPL-September-23-2026/paper.pdf) (family 103) is the claim I would most like a referee to read, and the one with the least independent support. It states $\mathsf L = \mathsf{RL} = \mathsf{BPL}$: any language decided by a polynomial-time, logarithmic-space machine that flips coins and errs with probability at most $1/3$ can be decided by a deterministic logarithmic-space machine. Quantitatively (Theorem 1.2) it approximates the acceptance probability of such a machine to within $2^{-q}$ in $O(\log n + q)$ space and $\mathrm{poly}(n)\,2^{O(q)}$ time.

This is the space-bounded twin of P = BPP, and among theorists it is the more believed of the two. Aleliunas, Karp, Lipton, Lovász and Rackoff asked for it in 1979. The record had barely moved since Saks and Zhou put BPL in deterministic space $O(\log^{3/2} n)$ in 1999; Hoza's 2021 work shaved a sub-logarithmic factor; Nisan's generator gives polynomial time with $O(\log^2 n)$ space; Reingold's SL = L (2005, JACM 2008) settled the undirected-connectivity special case. A full proof would be the biggest result in derandomization this century.

The approach, as far as I can follow the 108 pages, does not build a pseudorandom generator. It writes the acceptance probability as $p_0 = (I - S)^{-1}e$ for the strictly forward substochastic transition matrix $S$ of the configuration graph, and replaces $S$ step by step with a hierarchy of "corrected" matrices on sets of configuration copies, using the identity $I - D = (I + E)(I - C)$ to move removed transition mass into accumulated rewards. Each level's entries are estimated by sample-dependent tables with bounded row support, entries are compared by short fingerprints, and a shared "catalytic" bit vector and a telescoping allowance scheme keep every simultaneously live register within $O(\log n)$ bits. More than three quarters of the $2^{O(\log n)}$ random environments give an accurate estimate, so a deterministic machine enumerates them all and takes the median. An external input from group theory, Shalom's property (T) results, supplies a uniform mixing bound.

There is no Lean statement (the `Logspace` directory holds 585 lines of machine definitions and nothing else), no companion paper, and no reasoning trace. Space-bounded simulations are exactly the kind of proof where a single register that silently grows to $\log^2 n$ bits kills the result, and that is the accounting this paper spends most of its length on. I rate it a landmark claim with no independent verification. Practically, it would change nothing you run: polynomial time with unspecified exponents, and log-space is a theorist's resource. Its importance is that it would close the question of whether randomness is ever essential for small-memory computation.

### $L(\mathbb F_2) \cong L(\mathbb F_3)$: the free group factors are all the same

If I had to bet on one result among the operator algebras, it would be this one, and it goes the direction almost nobody expected.

Take the free group $\mathbb F_n$ on $n$ letters, let it act on $\ell^2(\mathbb F_n)$ by left translation, and close up the operators it generates. The result, $L(\mathbb F_n) = \lbrace \lambda(g) : g \in \mathbb F_n\rbrace''$, is a II₁ factor: an infinite-dimensional von Neumann algebra with a trace and trivial centre. Murray and von Neumann could tell these factors apart from the hyperfinite one in the 1940s, but nobody could tell $L(\mathbb F_2)$ from $L(\mathbb F_3)$. Does the algebra remember how many generators the group had? That question, credited to Kadison, became the central open problem in the classification of II₁ factors.

There was real structure around it. Voiculescu's free probability showed that the fundamental group of $L(\mathbb F_\infty)$ contains every positive rational, and Rădulescu extended this to every positive real ([JAMS 1992](https://doi.org/10.1090/S0894-0347-1992-1142260-1)). Dykema ([Pacific J. Math. 1994](https://doi.org/10.2140/pjm.1994.163.123)) and Rădulescu ([Invent. Math. 1994](https://doi.org/10.1007/BF01231764)) then built interpolated free group factors $L(\mathbb F_r)$ for every real $r > 1$, with an amplification rule

$$
L(\mathbb F_r)_t \;\cong\; L\big(\mathbb F_{1 + (r-1)/t^2}\big),
$$

and proved a dichotomy: either every $L(\mathbb F_r)$, $1 < r \le \infty$, is the same factor, or they are pairwise different. Voiculescu's free entropy dimension was built largely in the hope of proving "pairwise different", since a free semicircular $n$-tuple generating $L(\mathbb F_n)$ has entropy dimension $n$.

The paper ([family 287](https://github.com/openai/math/tree/main/preprints/An-isomorphism-of-the-free-group-factors-September-23-2026)) picks the other branch. Its Theorem 1.1 gives a trace-preserving isomorphism $L(\mathbb F_n) \cong L(\mathbb F_{n+1})$ for every $n \ge 3$, and the dichotomy does the rest. The widget shows the deduction. Theorem 1.1 makes every $L(\mathbb F_n)$ with $n \ge 3$ one algebra; amplify that one algebra by $t$ and each $L(\mathbb F_n)$ lands on rank $1 + (n-1)/t^2$, so all those image ranks are identified too. At $t = \sqrt 2$ the image of $L(\mathbb F_3)$ is $L(\mathbb F_2)$ and the image of $L(\mathbb F_5)$ is $L(\mathbb F_3)$.

<FreeGroupCompression />

The new input is the consecutive-rank isomorphism, and the construction is short and concrete. Start with free Haar unitaries $(A_1, \ldots, A_n, C)$ generating $L(\mathbb F_{n+1})$ and take a bounded logarithm $S = \arg C$. For each group word $g$, the conjugates $S_g = A_g S A_g^*$ are free from one another. Pick finitely supported real coefficient vectors $h_j$ and extend them to a cocycle on the free group by Fox's rule, $D_{gb} = D_g + \lambda(g) D_b$. These drive a polynomial differential equation that moves each $A_j$ by a sum of the $S_g$, and a formal trace identity plus analyticity shows the flow preserves every moment, so its time-one map is a trace-preserving automorphism. A free-sum norm bound (constant $3\pi$, independent of how many terms) then gives two estimates: each $A_j$ moves by at most $3\pi\lVert h_j\rVert_{\ell^2}$, and a chosen word $A_w$ lands within $3\pi\lVert D_w - \delta_e\rVert_{\ell^2}$ of $C A_w$.

So the whole problem reduces to finding coefficients that are small but whose cocycle value at some word is close to $\delta_e$. Lemma 4.1 does that with words having $m$ free prefixes and a Catalan moment count: the operators $T_m = 1 + \sum_{j \le m} \lambda(p_j)$ have $T_m T_m^*/m$ converging to a law with no atom at zero, so a truncated inverse produces a small vector mapped near $\delta_e$. Each step therefore moves the first $n$ generators a tiny amount while making a word in them approximate the extra generator $C$. Iterate with error budgets chosen after each witness word is fixed, and the first $n$ coordinates converge in norm to a free Haar tuple. A separate $L^2$-density argument with the trace-preserving conditional expectation shows that the limit still generates everything. That last step is where I would look hardest, because norm convergence of generators says nothing by itself about generation, and the paper treats it separately for exactly that reason.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/physics-operator-algebras-fig1.png"
  alt="Flow chart of six boxes: approximate targets in the current tuple, choose tolerance epsilon_k, perturb the tuple so a witness word approximates the old extra generator, substitute the witness word, choose a Lipschitz bound and future budget, and tail control that preserves earlier approximations."
  caption="The bookkeeping that makes the free-group-factor iteration converge: the witness word is chosen only after the tolerance is fixed, and its length only affects the future error budget (An isomorphism of the free group factors, Figure 1)."
/>

Two side results come for free. The fundamental group of the common factor is $\mathbb R_{>0}$, and Corollary 7.1 shows that none of Voiculescu's four free entropy dimensions ($\delta$, $\delta_0$, $\delta^*$, $\delta^\star$) is an invariant of the generated algebra: each takes every integer value $n \ge 2$ on generating tuples of $L(\mathbb F_2)$. That kills the program most people were betting on.

The Lean is the strongest part. The Comparator statement defines the group von Neumann algebra from scratch as the double commutant of left translations, builds the interpolated factors as corners of $L(\mathbb F_2) \otimes B(\ell^2)$ cut by a projection of trace $(r-1)^{-1/2}$, and asks for a normal trace-preserving isomorphism for every pair of parameters, the infinite one included (`lean/ComparatorChallenges/InterpolatedFactors.lean`, lines 216–219):

```lean
theorem allInterpolatedIsomorphic (r s : ℝ≥0∞) (hr : 1<r) (hs : 1<s) :
    Nonempty (NormalTracialEquiv (interpolatedTrace r) (interpolatedTrace s)
      (interpolatedTopology r) (interpolatedTopology s)) := by
  sorry
```

The `sorry` is by design; the proof lives in `OAI.Analysis.InterpolatedFactors.Interpolation`, about 17,600 lines plus a `FreeGroupFactor` directory, permitted axioms `propext`, `Quot.sound` and `Classical.choice`. I read the definitions and they are the standard ones. What I cannot tell you is that the build passes; I did not run it.

OpenAI also released an abridged reasoning summary for this result, and it is worth reading for one detail: the model spent most of its effort trying to prove the factors are *different*. It tried rigidity of free products, $L^2$-homology, Jung's strong 1-boundedness, width dimension of matrix microstates, Gaboriau-style cost arguments and more, and recorded each failure, before turning to constructive changes of generators.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/physics-operator-algebras-fig5.png"
  alt="First page of OpenAI's summarized chain of thought for the free group factor problem: the problem statement in monospace, then a section describing the model first searching for an obstruction to isomorphism using rigidity, free entropy and 1-bounded entropy."
  caption="The released reasoning summary shows the model first hunting for a proof that the factors are different, the direction most experts expected (Summarized chain of thought: Isomorphism of free group factors, page 1)."
/>

One piece of context. On 10 September 2026 Dima Shlyakhtenko posted [arXiv:2609.11074](https://arxiv.org/abs/2609.11074), identifying the factors of Fuchsian groups with interpolated free group factors, and the abstract says the results were obtained with ChatGPT Pro. The OpenAI paper cites it, notes that both constructions need separate care for the limiting distribution and for generation, and says it uses no result from it.

### The irrationality exponent of π is 2 (family 017)

How well can fractions approximate $\pi$? Every irrational $x$ has infinitely many $p/q$ with $|x - p/q| < 1/q^2$, and for almost every real number you cannot do much better: the irrationality exponent is 2. For $\pi$ the best known upper bound went from Mahler's 42 in 1953 to Hata's 8.0161, Salikhov's 7.6063 and Zeilberger and Zudilin's 7.103205334137 in 2020. The paper notes a September 2026 preprint by Bai claiming 7.101862832357. All of these come from cleverer and cleverer integrals. The claim here jumps straight to 2.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig3.png" alt="First page of 'The irrationality exponent of pi is 2': abstract, the definition of the irrationality exponent, Theorem 1.1 stating |pi - p/q| >= q^(-nu) for every nu > 2 and q >= Q(nu), and Corollary 1.2 on the Flint-Hills series." caption="The whole claim fits on page one; the proof is 22 more pages (The irrationality exponent of pi is 2, page 1)." />

The method is not an integral at all. It is a Roth-style argument. Suppose some exponent $\nu > 2$ allowed infinitely many exceptionally good approximations. Pick several of them, successively and far apart, and use the points $c_{ji} = 2ij\,p_i/q_i$, which sit next to the periods $2\pi i j$ of the exponential. A weighted multi-variable interpolation theorem (the paper's Theorem 2.1, with weights chosen before the points are known, which is what avoids a circular choice) produces a nonzero determinant over $\mathbb Q(i)$. Clearing denominators gives a lower bound. Translating the rows to the exact periods and expanding gives an upper bound with a quadratic saving, because rows testing the same entire function at repeated Taylor degrees cancel. For large enough dimension the two bounds collide. The threshold $Q(\nu)$ is ineffective, as in Roth's theorem.

The corollary is fun: the Flint–Hills series $\sum 1/(n^3 \sin^2 n)$ converges, because Alekseyev showed convergence needs exponent at most $5/2$ and Meiburg showed strictly below $5/2$ suffices. The paper also settles the general family: $\sum 1/(n^a |\sin n|^b)$ converges exactly when $a > \max(1, b)$. It even flags a sign error in a competing claim.

The Lean statement, quoted earlier with its proof's top level, is elementary and uses `Real.pi`, so there is very little room for a definitional trick. The solution directory, `OAI/NumberTheory/PiExponent`, is about 95,000 lines. The Flint–Hills corollary is outside the formal statement. The model's reasoning summary for this result is the one I walked through earlier, where the target was 5/2 and the argument overshot. One question a referee will ask, which the paper does not address: the argument only uses that $2\pi i j$ are periods of $\exp$, so does it also give exponent 2 for $\log 2$ and friends? If it does, that is a sign the method is real; if it is somehow specific to $\pi$, I would want to know why.

### A fixed gap next to the line Re s = 1 (family 003)

If I had to pick one claim in the whole release to check first, it would be this one. The paper says the Riemann zeta function, every Dirichlet $L$-function and every finite-order Hecke $L$-function over $\mathbb Q(\sqrt{-3})$ have no zeros with real part above $7/8$.

Some background. Riemann tied the primes to the zeros of $\zeta(s)$: a zero with real part $\sigma$ lets the error in the prime count be about $x^\sigma$. The Riemann hypothesis puts every nontrivial zero on $\Re s = 1/2$. What is actually known is far weaker. Hadamard and de la Vallée Poussin showed in 1896 that there are no zeros on the line $\Re s = 1$, and every improvement since has been a region that hugs that line more tightly as the height $t$ grows. The best explicit version of the classical region, from Mossinghoff, Trudgian and Yang, says there is no zero with $\sigma \ge 1 - 1/(5.558691 \log t)$. Ford's explicit Vinogradov–Korobov region, $1/(57.54 (\log t)^{2/3} (\log\log t)^{1/3})$, only overtakes that once $\log t$ is around 10,000. Both widths go to zero. Nobody could rule out zeros at real part $0.999999$ far enough up. The "quasi-Riemann hypothesis" is the statement that some fixed $\sigma_0 < 1$ works at every height.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig2.png" alt="Theorem 1.1 of the quasi-Riemann hypothesis paper: every finite-order Hecke L-function over Q(sqrt(-3)) has no zero in Re s > 7/8, the same holds for every Dirichlet L-function including zeta, and the paragraph noting this does not prove the Riemann hypothesis." caption="The headline statement, as the paper puts it, including its own note that this is not the Riemann hypothesis (The Quasi-Riemann Hypothesis: A Zero-Free Half-Plane Re s > 7/8, Theorem 1.1)." />

The widget below is the cleanest way I found to see what changes. It plots the width of the known zero-free strip against the height. The classical curves fall away to nothing; the claim is a flat line at $1/8$. The shaded band is a reminder that Platt and Trudgian have already checked numerically that every zero up to height $3 \times 10^{12}$ sits on the critical line, so the claim only says something new above that.

<ZeroFreeMargin />

The proof does not touch $\zeta$ directly. It works over the Eisenstein integers, where Kubota and Patterson's cubic theta function lives, and treats every Dirichlet $L$-function as a Hecke $L$-function there. For each target character it forms one character-weighted sum of cubic-theta Fourier coefficients and computes it two ways: a reflection formula for the theta coefficients, and Poisson summation in the averaging variable. The Poisson side expresses one of the rows through $1/L_F(s,\eta)$. If the target had a zero too far right, that reciprocal would have to continue further left than the two exact expansions allow. Part I of the 199-page paper runs this with balanced scales and gets $11/12$; Part II adds prime compensation and two more moment estimates to reach $7/8$. A separate 49-page manuscript gives an independent proof of $11/12$, and the README says that write-up was human-edited for readability and that this family, with the CM Hodge result, was an exception to the uniform three-hours-per-problem procedure.

The 11/12 manuscript spells out what a fixed half-plane buys: an effective $x^{11/12}\log x$ error in the prime number theorem for progressions uniformly in the modulus, Vinogradov's conjecture that the least quadratic nonresidue is $O(p^\varepsilon)$, and deterministic polynomial-time square roots modulo a prime. A third, nine-page paper proves the uniform Landau–Siegel statement $(1-\beta)\log q \ge c$; for large $q$ that is already implied by $7/8$, but its proof is independent and short.

The Lean is the reason I take this more seriously than its size would otherwise allow. The target statement is about Mathlib's own `riemannZeta`, not a custom definition, so it cannot be a quietly weakened version; I quoted the whole nine-line file in the section on Comparator. There are matching challenges for every Dirichlet $L$-function (`DirichletSevenEighths`), for the Hecke family over $\mathbb Q(\sqrt{-3})$ and for the Siegel-zero gap, all allowed only `propext`, `Quot.sound` and `Classical.choice`. The `sorry` is the challenge; the solution lives in `OAI/NumberTheory/DirichletL`. What I cannot tell you is whether it compiles. If it does, this is the largest result in analytic number theory in a very long time, and it was posted on GitHub with no referee. I found enthusiasm on X on release day and process complaints in the news, but no technical comment from an analytic number theorist by October 7.

### The Hodge conjecture for CM abelian varieties (family 032)

The Hodge conjecture says that on a smooth projective complex variety, every rational cohomology class of type $(p,p)$ is a rational combination of classes of subvarieties. Abelian varieties with complex multiplication are the most symmetric test case. Their Hodge classes are completely understood as linear algebra, and Deligne proved in 1982 that they are "absolute Hodge", which is everything a cycle class would be except an actual cycle. Building the cycles was the open part. Until 2025 the best general results were on Weil classes, and Markman's 2025 preprint (arXiv:2502.03415) made all abelian fourfolds work by algebraizing Weil classes on certain sixfolds.

Family 032's main paper claims the whole CM case. Theorem 1.1 says the cycle class map $\mathrm{CH}^p(A)_{\mathbb Q} \to H^{2p}(A,\mathbb Q)\cap H^{p,p}(A)$ is surjective for every CM abelian variety $A$ and every $p$. Section 8 then cashes in two of Milne's reductions. Milne showed in 1999 (Compositio 117) that Hodge for CM abelian varieties implies the Tate conjecture for abelian varieties over finite fields, and in 2002 (Annals 155) that it implies Grothendieck's Hodge standard conjecture for abelian varieties in characteristic $p$. Those are decades-old arguments, so if Theorem 1.1 holds, both conjectures follow. The catalogue pairs this family with family 001, Milne's rationality conjecture, which is filed under number theory.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig2.png"
  alt="The first page of the manuscript 'The rational Hodge conjecture for CM abelian varieties', dated September 30, 2026, showing the abstract, the definition of CM type, and Theorem 1.1 stating that the Betti cycle-class map from CH^p(A) tensor Q to rational (p,p) classes is surjective."
  caption="The whole claim fits on one page; the proof takes 52 more (The rational Hodge conjecture for CM abelian varieties, page 1)."
/>

The proof has a shape I could follow, which I did not expect. A CM type is a sign function $v$ on the embeddings of a CM field, odd under complex conjugation. The paper reduces every CM Hodge class, using auxiliary CM varieties, to one building block: a line in $H^4$ of a product of four CM abelian varieties $B(v_1)\times B(v_2)\times B(v_3)\times B(v_4)$, where $v_1+v_2=v_3+v_4$. That relation is exactly what makes the line type $(2,2)$ under every conjugate. You can check the bookkeeping below.

<CmFourFactorSwitch />

Then comes the geometry. To show the four-factor line is algebraic, it is enough to find one smooth projective surface $S$ with maps to the four abelian varieties such that the four pulled-back one-forms have a nonzero integral $\int_S h_1(\alpha_1)\wedge h_2(\alpha_2)\wedge h_3(\alpha_3)\wedge h_4(\alpha_4)$. The image of $S$ is then the cycle. The paper takes $S$ to be a compact quotient of the complex two-ball for a unitary group of signature $(2,1)$. It builds theta one-forms in the Kudla–Millson tradition and uses weak approximation to get a mixed period with two holomorphic and two antiholomorphic factors that is still nonzero. A Hecke and Frobenius slope computation then shows the one-forms come from the right CM Hodge structures. I can't judge Sections 4 to 6. They are where a specialist in Shimura varieties should start.

The companions go further. One claims rational Hodge for every product of projective K3 surfaces, another that the Kuga–Satake correspondence is algebraic for every projective K3, another all Weil classes on split abelian eightfolds. Those three run through immersed Lagrangian Floer theory and Fukaya categories of a mirror quartic. That is an unusual road to algebraic cycles, and its analytic foundations are hard to referee. One companion (032_7) is explicitly conditional, and the K3-products paper cites four other release preprints.

Three caveats. Nothing in 032 is in Lean, and the release says so. The README says the CM Hodge result was one of two exceptions to the release's fixed procedure of about three hours of compute per problem, without saying what was different. And this is a special case of the Hodge conjecture, not the Millennium problem; press coverage already points that out.

### The elliptic-curve stack: BSD in rank at most one, Goldfeld, and Fontaine–Mazur at 2 (families 002, 006, 010)

These three families are best read together, because each uses the others.

Family 006 is the base. Its 2-converse says that if the 2-power Selmer group of $E/\mathbb Q$ has corank 0 or 1, the analytic rank and the Mordell–Weil rank equal that corank and Sha is finite, with no hypothesis on reduction, torsion, isogenies or the residual representation. Alexander Smith proved in 2022 that, among quadratic twists of any fixed curve, the 2-Selmer corank is 0 or 1 half the time each. Put the two together and you get Goldfeld's 1979 conjecture for analytic rank: half the twists have analytic rank 0, half have rank 1. A companion adds a tail estimate so the mean analytic rank is exactly $1/2$. One honest caveat, which the papers state: twists are indexed by signed squarefree $d$ ordered by $|d|$, not by Goldfeld's fundamental discriminants.

Family 002 then claims the full Birch–Swinnerton-Dyer leading-term formula, every prime factor of it, for every curve whose $q$-power Selmer corank is 0 or 1 at some prime $q$. Gross–Zagier and Kolyvagin already give rank and finiteness of Sha in analytic rank at most one, so this amounts to: the exact BSD formula holds for every elliptic curve over $\mathbb Q$ of analytic rank 0 or 1. Before this, the $p$-parts of that formula were known for many primes under hypotheses, and the whole formula only for curves checked one at a time. The proof is a long sequence of integral determinant comparisons with Heegner points and theta series, prime by prime. The Selmer-converse companion has a pleasant aside: every prime $\ell \equiv 4, 7, 8 \pmod 9$ is a sum of two rational cubes, a case of Sylvester's problem that the paper says Yin and Burungale–Tian had also reached by Heegner-point methods.

Family 010 proves the odd, regular two-dimensional Fontaine–Mazur conjecture over $\mathbb Q$ at the prime 2 with no residual hypothesis, so scalar and reducible residual images are included. At odd primes this was essentially done (Kisin, Emerton, Pan, and a last case at 3); at 2, Allen and Tung had the solvable and nonsolvable residual cases. It is also the modularity input to the Hilbert's-tenth paper.

None of the roughly 620 pages here is formalized. These are exactly the claims where I trust the release least, not because anything looks wrong but because there is no independent check of any kind yet and they are the foundation of family 004.

### The Mahler conjectures, twice over, and a symplectic ball that proves one of them

Of the analysis families, this is my bet, and it is also the one I would most like a geometer to look at first.

Take a convex body $K \subset \mathbb R^n$ that is symmetric about the origin, and its polar $K^\circ = \lbrace y : \langle x, y\rangle \le 1 \text{ for all } x \in K\rbrace$. The volume product $|K|\,|K^\circ|$ does not change under linear maps. Blaschke and Santaló showed its maximum is the Euclidean ball's. Mahler asked in 1939 for the minimum, and conjectured it is the cube's value,

$$
|K|\,|K^\circ| \;\ge\; \frac{4^n}{n!}.
$$

For a body that is not symmetric, polarity is taken about the Santaló point, and the conjectured floor is the simplex's, $(n+1)^{n+1}/(n!)^2$. The symmetric floor has many minimizers, not just the cube and its dual the cross-polytope. Every Hanner polytope attains it: build from intervals by Cartesian products and by taking the convex hull of two bodies in complementary subspaces. That is part of why the problem is hard. No single "nice" extremizer exists that a symmetrization argument could flow towards.

The widget below makes the shape of the problem concrete along one curve through the space of bodies. The unit $\ell_p$ ball has the $\ell_q$ ball as its polar, with $1/p + 1/q = 1$, and both volumes have closed forms. Along that curve the product touches $4^n/n!$ only at $p = 1$ and $p = \infty$, the cross-polytope and the cube, and it peaks at the Euclidean ball. In the plane the ball sits $\pi^2/8 \approx 1.234$ times above the floor; in $\mathbb R^3$ it sits $\pi^2/6 \approx 1.645$ times above. The widget proves nothing. Every other body is the hard part.

<MahlerVolumeProduct />

Before this release the conjecture was settled in dimensions two and three only. Mahler did the plane himself. Iriyeh and Shibata did the symmetric case in $\mathbb R^3$ ([Duke 2020](https://arxiv.org/abs/1706.01749)). Chen, Li, Xi and Xu did the general case in $\mathbb R^3$ only this June ([arXiv:2605.09334](https://arxiv.org/abs/2605.09334)). Beyond that, people had special classes and asymptotics. Saint-Raymond did unconditional bodies, Reisner zonoids, and Nazarov, Petrov, Ryabogin and Zvavitch showed local minimality at the cube. Bourgain and Milman's 1987 theorem gives $c^n/n!$ for some constant $c$, and Kuperberg's 2008 bound gives $(\pi/4)^n$ times the conjectured floor. The release claims the whole thing: both inequalities in every dimension, with the equality cases classified exactly.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig1.png"
  alt="Entry 087 of the release overview: the Mahler conjectures and symplectic width, resolving the symmetric and nonsymmetric Mahler conjectures in every dimension with Hanner polytopes and simplices as minimizers, plus Gromov width 4 for symmetric polar products."
  caption="How the release itself summarises family 087, under the Convex and metric geometry heading (OpenAI Research Catalog overview, page 10)."
/>

The family is three papers, and they prove the symmetric inequality twice by unrelated routes. That is the detail that most raised my confidence.

The symplectic route is the one I would show a colleague. Put $K$ in position space and $K^\circ$ in momentum space, and look at $U_K = \operatorname{int}K \times \operatorname{int}K^\circ \subset \mathbb R^{2n}$ with the standard symplectic form $\omega_0 = \sum dq_i \wedge dp_i$. Symplectic maps preserve volume. A ball of "capacity" $c$ (meaning $\pi|z|^2 < c$) has volume $c^n/n!$. So if you can fit a symplectic ball of every capacity below 4 inside $U_K$, you get $|K|\,|K^\circ| \ge 4^n/n!$ immediately. Artstein-Avidan, Karasev and Ostrover saw in [2014](https://arxiv.org/abs/1303.4197) that Viterbo's volume–capacity conjecture would imply symmetric Mahler this way. They computed that the Hofer–Zehnder capacity of $K \times K^\circ$ is exactly 4. But Viterbo's conjecture is false: Haim-Kislev and Ostrover found a counterexample in [2024](https://arxiv.org/abs/2405.16513). So the bridge looked closed. Their counterexample is a Lagrangian product of a non-symmetric pentagon with a rotated copy, though, which leaves symmetric polar products untouched. The paper ["Symplectic Balls in Symmetric Polar Products"](https://github.com/openai/math/blob/main/preprints/Symplectic-Balls-in-Symmetric-Polar-Products-September-22-2026/paper.pdf) (19 pages) walks straight through that gap. It claims that for $n \ge 2$ and every origin-symmetric $K$, the Gromov width of $U_K$ is exactly 4.

The upper bound is the easy half. Pick $q_0 \in K$ and $p_0 \in K^\circ$ with $\langle q_0, p_0\rangle = 1$. A linear symplectic change of coordinates puts $U_K$ inside a square of area 4 times $\mathbb R^{2n-2}$, and Gromov's nonsqueezing theorem caps any ball at capacity 4. The lower bound is the new idea, and it has two ingredients.

One is a planar object. Gross's conformal map $F(w) = \tfrac{8}{\pi^2}\sum_\ell (-1)^\ell w^{2\ell+1}/(2\ell+1)^2$ sends the unit disk onto a convex lens whose boundary height is uniformly distributed when the boundary angle is. The paper proves a uniform estimate on horizontal slices of the lens for $g = F^{-1}$. Its Proposition 3.2 says the slice integral of the area density of $g^k$ is at most $(\pi/4)|g|^{2k} + \epsilon_k$, with $\epsilon_k \to 0$. The other ingredient is pluripotential theory. If a holomorphic map $f$ has an isolated zero of order $k$ at the origin, then $\lbrace |f|^2 < 1\rbrace$ with the form $dd^c(|z|^2 + |f|^2)$ contains symplectic balls of every capacity below $\pi k$. Assemble $f_j(z) = g(\langle b_j, z\rangle)^k$ over the facet normals $b_j$ of a polytope approximating $K$ from outside. An explicit map then pulls $\omega_0$ back to exactly that form, landing the ball in $\operatorname{int}K' \times (\pi k/4)\operatorname{int}K'^\circ$. Divide the momentum by $\pi k/4$, and a ball of capacity $\pi k$ becomes one of capacity about $\pi k / (\pi k/4) = 4$. The 4 is forced by the arithmetic.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig8.png"
  alt="Left, the unit disk with equally spaced points on its boundary; an arrow labelled F maps it to a lens-shaped region on the right whose boundary points sit at equally spaced heights between minus one and one."
  caption="Gross's conformal lens, the one planar object both symmetric proofs are built on: equal angular steps on the circle become equal vertical steps on the lens boundary, so a uniform angle gives a uniform height (The symmetric Mahler conjecture and its equality cases, Figure 1)."
/>

The second symmetric proof, ["The symmetric Mahler conjecture and its equality cases"](https://github.com/openai/math/blob/main/preprints/The-symmetric-Mahler-conjecture-and-its-equality-cases-September-22-2026/paper.pdf) (26 pages), uses the same lens and the same holomorphic zero. It turns them into a bound on a sum of "feasibility probabilities" through a generalized Lelong number, with no symplectic geometry at all. It also does something the symplectic paper explicitly does not: it classifies the equality cases. It lifts $K$ to $K \oplus_1 [-1,1]$, extracts metric medians for every triple of points in the norm whose unit ball is $K^\circ$, and then invokes Hansen and Lima's 1981 classification of finite-dimensional spaces with the 3.2 intersection property. The result is that equality holds exactly for linear images of Hanner bodies.

The general case is a third, unrelated proof, and the longest at 72 pages: ["The Mahler Conjecture for General Convex Bodies"](https://github.com/openai/math/blob/main/preprints/The-Mahler-Conjecture-for-General-Convex-Bodies-September-22-2026/paper.pdf). It uses Klartag's reduction to a cone over $K$ and its dual cone. Then it builds biased Gaussian projections onto both cones, fixing the bias and covariance together with a Brouwer fixed point. An entropy estimate turns those into a lower bound for the volume product. The error terms reduce to averages over pairs of eigenvalues of a Gaussian random matrix, and a strict one-variable "segment inequality" absorbs them. That inequality is where I would expect trouble in a human paper. Its constants are fixed decimals (0.602, 0.365, 0.22 and so on), and its proof uses rational certificates and the Bernstein-coefficient positivity of polynomials of degree 30 and 32.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig9.png"
  alt="Dependency diagram: cone formulation, then Gaussian projections and entropy estimate, then quadratic identities, then linear-field estimate and equality; a separate scalar branch of profile bounds and a segment variance estimate feeds into the middle steps."
  caption="The general (nonsymmetric) proof's dependency chain; the right-hand scalar branch holds the numerically certified inequalities, which the Lean development also covers (The Mahler Conjecture for General Convex Bodies, Figure 1)."
/>

Here is what moves this from "extraordinary claim" to "extraordinary claim with a receipt". Four Comparator challenges cover the symmetric inequality, its equality classification, the general inequality with its simplex equality case, and the Gromov width. Each pairs a short, readable statement with a solution tree. Together the Mahler and polar-product trees run to 591 Lean files and about 81,000 lines, and they allow only the three standard axioms. The challenge file ends in `sorry` by design: it is the statement the checker compares the solution against. Here is the symmetric one in full:

```lean
-- lean/ComparatorChallenges/MahlerConjecture.lean:9-20
def coordinatePolar (K : Set (I → ℝ)) : Set (I → ℝ) :=
  {p | ∀ v ∈ K, (∑ i, p i*v i) ≤ 1}

theorem symmetric_mahler {n : ℕ} (hn : 1 ≤ n)
    {K : Set (Fin n → ℝ)} (hK : IsCompact K) (hconv : Convex ℝ K)
    (hsym : ∀ x ∈ K, -x ∈ K) (hint : (interior K).Nonempty) :
    (4:ℝ)^n/(Nat.factorial n:ℝ) ≤
      (volume K).toReal*(volume (coordinatePolar K)).toReal := by
  sorry
```

I read it for traps and found none. `toReal` of an infinite volume is 0, so an unbounded polar would make the statement false, not vacuously true. The compactness, convexity, symmetry and nonempty-interior hypotheses are exactly Mahler's. The general version (`GeneralMahler.lean:24`) takes the infimum of the volume product over interior centres, which is equivalent to using the Santaló point. It also states the equality case as "$K$ is the convex hull of $n+1$ affinely independent points". The Gromov-width challenge (`SymmetricPolar.lean:40`) defines symplectic embeddings from scratch. The solution proves the cylinder nonsqueezing theorem it needs in Lean (`OAI/Geometry/PolarProducts/Nonsqueezing.lean:111`) rather than assuming it. That is a substantial formalization in its own right.

What is not covered: the functional Mahler inequalities and the entropy-transport corollaries in the summary, and an embedding at capacity exactly 4. There is one bookkeeping gap. The general-Mahler challenge has a JSON config and appears in `lean/docs/087.md` but is missing from `lean/formalization.yaml`. Family 101, later in this section, would independently imply the general inequality through Klartag's 2018 reduction. So the release proves general Mahler twice as well, though only once in Lean.

One more thing deserves saying, because it bears on how to read the whole release. This is one of the ten families with a published reasoning summary (`reasoning_traces/symmetric-and-general-mahler-conjectures.pdf`, 45 pages). It is not a clean march to the answer. The prompt for part one asked for dimension four and up. The model found the triangle-wave series, the one whose boundary real part is uniformly distributed, by way of Lundin's extremal function. Part two, the equality cases, was posed with the sharp lower bound handed over as a premise. Part three, the general case, goes through roughly fifty sections of abandoned routes: tensor amplification, Gale transforms, Brascamp–Lieb deficits, Riccati flows. Only then does it land on the Gaussian-projection argument, and it closes by saying its conclusion "remains dependent on the uniform scalar inequalities". Those are exactly the inequalities the Lean development then checks. One of the reviewers I worked with also evaluated the key slice estimate numerically. The excess over $(\pi/4)|g|^{2k}$ stays positive but shrinks, from about 0.13 at $k = 5$ to about 0.004 at $k = 1280$. That fits a slowly vanishing $\epsilon_k$ and is a useful sanity check, not a proof.

My read: if the Lean checks pass, this is the most consequential convex-geometry result in decades. Mahler is now settled in every dimension, with a symplectic corollary (every normalized capacity equals 4 on symmetric polar products) that is interesting in its own right.

### $\theta(p_c) = 0$, and who got there first

Bernoulli bond percolation keeps each edge of a graph independently with probability $p$. Below a critical value $p_c$ every cluster is finite; above it an infinite cluster appears. The question that defines the field is what happens exactly at $p_c$. In two dimensions Harris and Kesten showed there is no infinite cluster. In high dimensions the lace expansion gives the same answer (Hara and Slade, then Fitzner and van der Hofstad down to $d \ge 11$). Dimension 3 was the famous hole, and Benjamini and Schramm's 1996 conjecture asks for the same conclusion on every quasi-transitive graph with $p_c < 1$. Before this autumn the general case was known for nonamenable unimodular graphs (Benjamini, Lyons, Peres and Schramm) and for graphs of exponential growth (Hutchcroft).

Family 213 claims both: bond and site percolation on $\mathbb Z^3$ have no infinite cluster at their critical points, and bond percolation on every infinite, connected, locally finite quasi-transitive graph with $p_c < 1$ has none either.

Here is the context I did not expect to need. On 28 August 2026 Justin Leder of Anthropic posted a Claude-generated, Lean-verified proof of Kozma and Nitzan's "near-one gluing" conjecture, which Kozma and Nitzan had shown implies $\theta(p_c) = 0$ on $\mathbb Z^d$ for every $d \ge 2$ ([arXiv:2401.12397](https://arxiv.org/abs/2401.12397)). Ahmed Bou-Rabee followed within days with claimed proofs of the remaining Kozma–Nitzan conjectures and of the site case, using ChatGPT and Claude, and Gil Kalai [wrote it up on 3 September](https://gilkalai.wordpress.com/2026/09/03/amazing-there-is-no-percolation-at-the-critical-probability-in-all-dimensions-solved-by-ai-via-a-conjecture-of-gady-kozma-and-shahaf-nitzan/). The OpenAI papers are dated 24 September and cite all of this. The cubic-lattice paper calls itself "a self-contained cubic argument" and says that Leder's first-contact comparison "already implies the finite joint connection inequality for bonds." I respect that. It also means the $\mathbb Z^3$ result in this family is a second proof, and the new mathematics is the general quasi-transitive theorem.

The finite ingredient both papers share is a gluing inequality. For independent, possibly hyperedge, percolation on a finite graph, a starting node $o$, a relay set $A$ and a target $T$,

$$
\mathbb P(o \leftrightarrow A,\ o \leftrightarrow T) \;\ge\; \mathbb P(o \leftrightarrow A)\,\min_{a \in A} \mathbb P(a \leftrightarrow T).
$$

Read it as a relay race. If every point of $A$ reaches $T$ with probability at least $m$, then reaching $A$ at all gives you at least an $m$ chance of reaching $T$, even though the configuration that brought you to $A$ is correlated with what lies beyond. Taking $T$ to be a single vertex gives Kozma and Nitzan's multiplicative gluing conjecture. The widget below checks it by simulation on a small piece of the square lattice. In this geometry every path from $o$ to $T$ crosses the relay column, so the left side is just $\mathbb P(o \leftrightarrow T)$; the inequality held at every setting I tried, with visible slack.

<GluingRelay />

The rest of the $\mathbb Z^3$ argument is a contradiction in the style of Grimmett and Marstrand. Suppose an infinite cluster exists at $p_c$. Then large seed boxes connect, with high probability, to each of the 24 quarter-faces of a surrounding cube. Only finitely many radii are involved, so those estimates survive at some $q < p_c$. The gluing inequality lets a connection pass through fresh territory, thirteen elementary geometric moves carry it from one coarse cube to the next, and an exploration indexed by $\mathbb Z^2$ keeps the conditional failure probability of each cube uniformly small. A planar counting argument then produces an infinite cluster at $q < p_c$, which is impossible.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig1.png"
  alt="A cube drawn in perspective with each visible face split into four quarter-faces; one quarter-face on the right face is shaded and labelled F with indices 1,+ and +,+, with coordinate axes x1, x2, x3 at lower left."
  caption="The 24 quarter-faces of a box. If an infinite cluster existed at p_c, a large seed box would connect to each of them with high probability, and that finite fact survives at a slightly smaller p (Critical bond and site percolation on the cubic lattice, Figure 1)."
/>

The quasi-transitive paper has to handle graphs that look nothing like a lattice. It uses Hutchcroft's theorem for exponential growth as a black box and splits the subexponential remainder into two cases. If balls grow faster than every polynomial, the paper bisects a path between two large clusters until it finds a single edge where they almost touch, chooses edges by how much a search's conditional success probability drops, and controls the total cost with relative entropy; the result contradicts a bound of order $s^{-1/2}$ on two distinct clusters of size $s$ meeting across an edge. If growth is polynomial along a subsequence, the Tessera–Tointon structure theorem supplies a nilpotent quotient, an induction on its infinite cyclic factors produces two integer coordinates, and the cubic-lattice corridor construction is rerun in those coordinates.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig8.png"
  alt="Two panels. Left: a disk of ellipse parameters with short blue line segments around its boundary, labelled s=1 at the top and s=0, A=WI at the centre, with the caption 'Boundary center lines follow the minor axes'. Right: axes t and z with two blue arrows from the origin to points (1,1) and (1,-1), captioned 'A linear map normalizes two separated directions'."
  caption="The two-direction argument for subexponential-growth graphs: a winding-number count on the parameter disk forces two well-separated favoured directions, which a linear map turns into the moves (1,1) and (1,-1) used to build corridors (No percolation at criticality on quasi-transitive graphs, Figure 2)."
/>

Both theorems have Comparator statements, and the cubic one is listed among the release's main formal results. I read the statement line by line. It defines $\mathbb Z^3$, nearest-neighbour bond and site configurations, the product Bernoulli laws and the critical parameters from scratch, then asks for almost-sure finiteness of every cluster:

```lean
-- lean/ComparatorChallenges/CriticalZ3.lean:50-59
noncomputable def bondCritical : ℝ :=
  sInf {p : ℝ | p ∈ Set.Icc (0 : ℝ) 1 ∧ 0 < bondLaw p (bondInfiniteAt 0)}

noncomputable def siteCritical : ℝ :=
  sInf {p : ℝ | p ∈ Set.Icc (0 : ℝ) 1 ∧ 0 < siteLaw p (siteInfiniteAt 0)}

theorem critical_no_infinite :
    (∀ᵐ ω ∂bondLaw bondCritical, ∀ x : Vertex, (bondCluster ω x).Finite) ∧
    (∀ᵐ ω ∂siteLaw siteCritical, ∀ x : Vertex, (siteCluster ω x).Finite) := by
  sorry
```

The solution modules are about 10,000 lines for $\mathbb Z^3$ and 68,000 for the quasi-transitive theorem. If they compile, this family is as solid as anything in the release, and the general Benjamini–Schramm criticality conjecture (bond form) is settled. Priority for the lattice case belongs elsewhere.

### Donaldson's tamed-to-compatible question (formalized)

This is the one I would put first, because it combines a famous problem, a short paper and a machine-checked proof.

On a 4-manifold with an almost complex structure $J$, a symplectic form $\omega$ *tames* $J$ if $\omega(v, Jv) > 0$ for every nonzero $v$. It is *compatible* if it is also $J$-invariant, $\omega(Ju, Jv) = \omega(u, v)$. Compatible implies tamed. [Donaldson asked in 2006](https://arxiv.org/abs/math/0607083) whether, on a closed 4-manifold, the existence of a taming symplectic form forces the existence of a compatible one. The naive fix, averaging $\omega$ with $\omega(J\cdot, J\cdot)$, makes the form invariant and destroys closedness. Before this release the answer was known for integrable $J$ (Li–Zhang, via the theory of complex surfaces), for a residual set of tamed $J$ when $b^+ = 1$ ([Taubes](https://arxiv.org/abs/0910.5440)), and for a few rational surfaces.

[The paper](https://github.com/openai/math/blob/main/preprints/Taming-implies-compatibility-on-four-manifolds-September-23-2026/paper.pdf) proves the general statement in 18 pages. The idea is a duality argument with a hard analytic core. If no compatible form exists, Hahn–Banach produces a nonzero positive current $P$ that kills every closed invariant form. An elliptic correction gives $T = P + Q$ that kills every closed form, with $Q$ anti-invariant. The work is to show that the part of $P$ sitting on points of positive 2-dimensional density is itself closed. That comes down to one boundary term in a cutoff argument, of size $r^{-3} \cdot O(r^2) \cdot o(r) = o(1)$: an angular error and a transverse error whose squared integrals are $O(r^4)$ and $o(r^2)$, multiplied together. Once that piece is closed, pairing the residual forces $Q = 0$, and a closed positive current on complex lines cannot coexist with the taming form.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig6.png" alt="Left: a plane L_ell with a displacement z/r and its normal part z-perp/r in the tangent space. Right: two unit bivectors ell and the transported tau m forming a chord delta, with 1 minus their inner product equal to delta squared over two." caption="The decisive estimate in the tamed-to-compatible proof multiplies an angular error and a transverse error; their squared integrals are O(r^4) and o(r^2), which is just enough to kill an r^-3 cutoff term (Taming implies compatibility on four-manifolds, Figure 1)." />

What made me take it seriously is the Lean side. The solution lives in `OAI/Geometry/TamingCompatibility` (582 files, about 83,600 lines, no `sorry`), and the Comparator statement reads like the textbook definition:

```lean
-- lean/ComparatorChallenges/TamingCompatibility.lean:58
theorem taming_implies_compatibility
    [T2Space X] [SecondCountableTopology X] [CompactSpace X] [ConnectedSpace X]
    (J : AlmostComplexStructure X)
    (h : ∃ α : TwoForm X, IsSymplectic α ∧ Tames α J) :
    ∃ η : TwoForm X, IsSymplectic η ∧ Compatible η J := by
  sorry
```

`IsSymplectic` there means smooth, closed and nondegenerate, with smoothness and closedness tested by pulling back along smooth maps from open subsets of $\mathbb R^4$. `Tames` and `Compatible` are exactly the two conditions above. I found nothing in the file that trivializes the statement. If the library compiles, this is a formally verified answer to a question Donaldson posed twenty years ago. The compatible form's cohomology class may differ from the taming one; with Li–Zhang's cone comparison the paper gets the same class when $b^+ = 1$.

### Kadison's similarity problem

Kadison asked in 1955 whether a bounded representation of a C\*-algebra that respects multiplication but not necessarily adjoints can always be made into an honest \*-representation by a change of inner product, $\rho(a) = S\pi(a)S^{-1}$. It was the first problem on most lists of open questions in the field for decades.

The known reductions were sharp. Haagerup ([Ann. Math. 1983](https://doi.org/10.2307/2007028)) solved the cyclic case and showed similarity is equivalent to complete boundedness. Nuclear algebras (Bunce, Christensen), algebras without tracial states (Haagerup) and II₁ factors with property Γ (Christensen) were done. Kirchberg proved in 1996 that the similarity property of a C\*-algebra is equivalent to every derivation into its representations being inner, and Pisier built the theory of similarity degree around it. What remained, roughly, were II₁ factors like $L(\mathbb F_2)$ that have neither property Γ nor any amenability to lean on.

The paper (31 pages) attacks exactly that. Its key estimate is a universal constant $C$ such that for every von Neumann algebra $P \subset B(K)$, every $Y \in B(K)$ and every matrix size $h$,

$$
\big\lVert [Y^{(h)}, X] \big\rVert \;\le\; C\, g_P(Y)\, \lVert X \rVert \qquad (X \in M_h(P)),
$$

where $g_P(Y)$ is the norm of the commutator map $a \mapsto [Y, a]$ on the unit ball of $P$. In words: an inner derivation's matrix amplifications cost no more than the derivation itself, up to one absolute constant. The hard case is a finite factor, and the proof builds a tracial free product $D * A_0$ with $D$ a II₁ factor generated by matrix algebras, averages over the unitary groups of those matrix algebras (Schur–Weyl duality kills every contraction except permutations), approximates the operator by a right multiplier on small supports, and reads off a Hankel matrix whose norm stays bounded below as the corner shrinks. Popa's relative independence theorem then plants this free product inside the ultrapower of an arbitrary separable II₁ factor. Kirchberg's equivalence and Paulsen's theorem turn the estimate into similarity.

A bonus corollary: every von Neumann algebra is hyperreflexive with one universal constant, which Eleftherakis and Paulsen had shown is equivalent to a positive answer.

The Lean statement is clean and general, over arbitrary Hilbert spaces (`lean/ComparatorChallenges/KadisonSimilarity.lean`, lines 15–21 and 79–82):

```lean
abbrev BoundedUnitalHom :=
  {π : A →ₐ[ℂ] (H →L[ℂ] H) // Continuous π}

def SimilarToStar (π : BoundedUnitalHom A H) : Prop :=
  ∃ S : (H →L[ℂ] H)ˣ, ∀ a : A,
    (S : H →L[ℂ] H) * π.1 (star a) * (↑S⁻¹ : H →L[ℂ] H) =
      star ((S : H →L[ℂ] H) * π.1 a * (↑S⁻¹ : H →L[ℂ] H))

def SimilarityTheorem : Prop :=
  ∀ (A : Type u) [CStarAlgebra A] (K : Type v) [NormedAddCommGroup K]
    [InnerProductSpace ℂ K] [CompleteSpace K] (π : BoundedUnitalHom A K),
    SimilarToStar A K π
```

The same file states the uniform commutator estimate and universal hyperreflexivity, with the constant spelled out as `3 * factorConstant + 2`. Since the proof leans on Popa's independence theorem and Kirchberg's equivalence, the formalization had to carry those too, which is a lot of operator algebra to have in Lean.

### The crossing number of $K_n$ and $K_{m,n}$

Draw the complete graph $K_n$ in the plane with as few edge crossings as you can. Guy (1960) and Harary and Hill (1963) found drawings with

$$
Z(n) = \tfrac14 \left\lfloor \tfrac n2 \right\rfloor \left\lfloor \tfrac{n-1}2 \right\rfloor \left\lfloor \tfrac{n-2}2 \right\rfloor \left\lfloor \tfrac{n-3}2 \right\rfloor
$$

crossings and conjectured nobody could do better. The bipartite version is older still: Turán's brick-factory problem from 1944, with Zarankiewicz's 1954 formula $\lfloor m/2 \rfloor \lfloor (m-1)/2 \rfloor \lfloor n/2 \rfloor \lfloor (n-1)/2 \rfloor$. Before this release the exact value of $\mathrm{cr}(K_n)$ was known only up to $n = 14$, the last cases by computer (Pan and Richter for $n \le 12$, Aichholzer for 13 and 14). The best general lower bound was about 98.56% of $Z(n)$ (Balogh, Lidický and Salazar, using flag algebras). Kleitman settled Zarankiewicz for $\min(m,n) \le 6$ in 1970, and the general case had a 0.9118 asymptotic ratio for $K_{n,n}$.

Family 165 claims both formulas, exactly, for every $n$. The upper bounds are the classical drawings. The widget below is the two-page drawing the paper uses: vertices on a line, edge $ij$ above it when $(i+j) \bmod n$ is less than $\lfloor n/2 \rfloor$, below otherwise. I recomputed its crossing count for $n = 5$ to $17$ and it lands on $Z(n)$ every time, including the known values $\mathrm{cr}(K_{11}) = 100$ and $\mathrm{cr}(K_{13}) = 225$.

<TwoPageCrossingDrawing />

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig2.png"
  alt="Two panels showing the same seven labelled vertices on a horizontal line. The left panel draws the edges whose endpoint sum mod 7 is 0, 1 or 2 as blue semicircles above the line, with 2 crossings; the right panel draws the remaining edges as orange semicircles, with 7 crossings."
  caption="The upper-bound half: the endpoint-sum two-page drawing of K7, with 2 + 7 = 9 = Z(7) crossings. The lower bound is the new part (The crossing number of complete graphs, Figure 3)."
/>

The lower bound is the interesting half, and it is algebraic rather than a search. Take a drawing that has already been normalised so adjacent edges never cross. Tutte-style signed crossing counts plus the cyclic order of edges at each vertex define a bilinear form that vanishes on every pair of cycles. For $n = 2s+1$, the paper weights cycles by polynomials in four variables, symmetric in the endpoints of each edge and of degree at most $s-2$ in each variable. That space has dimension exactly $Z(2s+1)$. Evaluating a polynomial at every pair of edges that actually cross gives a linear map, and two facts finish it. The map is injective, which is the hard part (a Lagrange-interpolation "detector" forces anything in the kernel to vanish when endpoints coincide, and dividing out those factors lowers the degree forever). And its image is totally isotropic for a nondegenerate form, so its dimension is at most half the number of coordinates. Each coordinate is a crossing pair, so there are at least $Z(n)$ of them. Deleting a vertex gives the even case. The $K_{m,n}$ paper runs the same cycle pairing into a subspace-intersection inequality.

I expected a proof of a sixty-year-old conjecture to be long. It is 13 pages for $K_n$ and 16 for $K_{m,n}$. That brevity is a warning sign for a human paper, and here it is offset by the formal statement, which I read line by line. It quantifies over arbitrary continuous drawings, counts crossing points, requires proper crossings via a local chart, and forbids triple points:

```lean
-- lean/ComparatorChallenges/CompleteCrossing.lean:75-82
def ordinaryCrossingNumber (graph : Graph Vertex Edge) : ℕ :=
  sInf (Set.range (@crossingCount Vertex Edge graph))

def hill (order : ℕ) : ℕ :=
  (order / 2 * ((order - 1) / 2) * ((order - 2) / 2) * ((order - 3) / 2)) / 4

theorem complete_graph_crossing_number (order : ℕ) (at_least_three : 3 ≤ order) :
    ordinaryCrossingNumber (completeGraph order) = hill order := by
  sorry
```

The solution modules are about 18k lines for $K_n$ and 21k for $K_{m,n}$, and both families are listed among the release's main formal results. If those build, the ordinary (point-count) crossing numbers of complete and complete bipartite graphs are settled. Pair and odd crossing numbers are not addressed. Hebbar and Mangam claimed both formulas in 2018, and their proof was not accepted; the difference this time is a kernel-checkable statement. Of everything in the group, this is the result I would most like to see independently compiled.

### Nagata's conjecture (family 039)

Nagata's conjecture is the cleanest open problem I know in classical algebraic geometry. Pick $r\ge 10$ very general points in the plane. Then any plane curve of degree $d$ with multiplicity at least $m_i$ at the $i$-th point satisfies $\sum_i m_i < d\sqrt r$. Nagata proved it in 1959 when $r$ is a perfect square, as part of his counterexample to Hilbert's fourteenth problem. For every non-square $r\ge 10$ it has been open since. The case $r=10$ is the famous one. The best general bounds, such as Roé's $d > m(\sqrt r - 1 - \pi/8)$, fall short by a constant.

The paper claims the full statement for every $r\ge 10$, with unequal multiplicities, and the Lean statement matches it quantifier for quantifier. It asks for a countable family of proper Zariski-closed exceptional sets, a configuration avoiding all of them, and the strict inequality for every effective plane curve:

```lean
-- lean/ComparatorChallenges/Nagata.lean (statement; proof at OAI/AlgebraicGeometry/PlaneCurves/Nagata.lean:148)
theorem nagata_conjecture :
    ∀ count : ℕ, 10 ≤ count →
      ∃ exceptional : ℕ → Set (OrderedDistinctPoints count),
        (∀ index, IsConfigurationZariskiClosed (exceptional index) ∧
          exceptional index ≠ Set.univ) ∧
        (∃ points : OrderedDistinctPoints count, ∀ index, points ∉ exceptional index) ∧
        ∀ points : OrderedDistinctPoints count, (∀ index, points ∉ exceptional index) →
          ∀ curve : Workers.W10.EffectivePlaneCurve, ∀ multiplicities : Fin count → ℕ,
            (∀ index, multiplicities index ≤ W15.curveOrdinaryMultiplicity curve (points.val index)) →
            (∑ index, (multiplicities index : ℝ)) <
              (Workers.W10.curveDegree curve : ℝ) * Real.sqrt (count : ℝ)
```

The proof is an interpolation argument made explicit. It specializes the points to a structured configuration and bounds how many conditions a degree-$d$ form can satisfy by counting lattice rows under sheared polygons. The figure shows one of those polygons and its slice widths.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig3.png"
  alt="Two panels. Left: a shaded quadrilateral D1 in the X-Y plane with vertices at (0,0), (0,a), (1,0) and (1+a/9, -a(1+a/9)), and a red horizontal slice of length s at height Y-minus. Right: a tent-shaped graph of slice width w(Y) peaking at 1 over Y=0, with the interval from Y-minus to Y-plus where the width is at least s marked in red."
  caption="The lattice-counting step in the Nagata proof: slice widths of the polygon D1 and the heights where they exceed s, which bound the integer rows of conditions (Nagata's Conjecture for Plane Curves, Figure 1)."
/>

The paper is 24 pages, which would make me nervous for a 65-year-old problem if it weren't for the Lean statement. A companion proves the qualitative Nagata–Biran conjecture: on every polarized surface $(S,L)$, very general $r$-tuples have the maximal Seshadri constant $\sqrt{L^2/r}$ once $r$ passes a threshold depending on $(S,L)$. That one is formalized too (`OAI/AlgebraicGeometry/Seshadri/Main.lean:25`), but its threshold is not effective. Two more papers extend it to higher dimension and positive characteristic, without Lean.

### Hilbert's tenth problem over Q (family 004)

Matiyasevich showed in 1970 that no algorithm decides whether an integer polynomial has an integer root. Over the rationals the question stayed open, because the obvious transfer needs a polynomial definition of $\mathbb Z$ inside $\mathbb Q$, and Mazur's conjecture on the topology of rational points predicts there is no such definition. Julia Robinson (1949) defined $\mathbb Z$ in $\mathbb Q$ with quantifiers, Poonen cut it to two universal quantifiers, Koenigsmann to one, and Koymans–Pagano and Alpöge–Bhargava–Ho–Shnidman settled every ring of integers of a number field in 2024–25. $\mathbb Q$ itself was untouched.

The claim is that no algorithm decides rational solvability, with the number of variables part of the input. What I like about the argument is that it does not fight Mazur's conjecture. Instead of defining $\mathbb Z$, it builds for each polynomial $f$ an effectively generated sequence of finite tests, each answerable by rational-solvability queries, with the property that $f$ has an integer zero exactly when every test passes. Multiples $nP$ of a point on a rank-one elliptic curve stand in for the integers $n$. If all tests pass but there is no integer zero, compactness gives a nonstandard model in which some "integer" index is nonstandard, and a height bound (from an abelian-surface family, modularity, and Faltings heights) produces the contradiction.

The 62-page paper cites three other release manuscripts as theorems it uses: a pointwise 2-converse for curves with rational two-torsion (in this family, 91 pages), Fontaine–Mazur modularity at 2 (family 010) and the Goldfeld 2-converse (family 006). None is formalized. So this is a famous problem whose claimed solution sits on roughly 400 pages of unrefereed arithmetic geometry from the same source. I would not call it solved until number theorists have gone through that chain.

### Kaplansky's zero-divisor conjecture fails (family 196)

Kaplansky asked in the 1950s whether the group algebra $K[G]$ of a torsion-free group over a field can have zero divisors. It can't for orderable groups, unique-product groups, elementary amenable groups (Kropholler, Linnell and Moody) or many 3-manifold groups. Giles Gardam disproved the neighbouring unit conjecture in 2021 with an explicit unit over $\mathbb F_2$, but the zero-divisor conjecture stayed open.

Family 196 claims a finitely presented torsion-free group $G$ with a finite 2-dimensional $K(G,1)$ and nonzero $\alpha,\beta\in\mathbb F_2[G]$ with $\alpha\beta=0$. The construction immerses two finite graphs in a rose, a graph with one vertex, builds a graphical small-cancellation group from them, and arranges that every coefficient of $\alpha\beta$ comes from an even number of products. Torsion-freeness and asphericity are proved separately by cone surgeries. The Comparator statement includes the "moreover": torsion-free in the plain sense, finitely presented, a finite Hausdorff 2-dimensional CW complex with contractible universal cover, and $\alpha\beta=0$ in `MonoidAlgebra (ZMod 2) G` (`lean/ComparatorChallenges/TorsionFreeZeroDivisors.lean`, proof at `OAI/Algebra/GroupRing/Main.lean:13`).

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig7.png"
  alt="Two horizontal number lines labelled source and target. A shaded block from ks to (k+1)s on the source line is translated by x maps to x plus nu s plus r, landing as an orange block that overlaps the target block from (k+nu)s to (k+nu+1)s in a stretch of length s minus r."
  caption="One step of the cancellation bookkeeping: a translated block overlaps its target in all but o(s) letters, which pins down most of the next word (A Torsion-Free Group Algebra with Zero Divisors, Figure 1)."
/>

This is over $\mathbb F_2$ only. Characteristic zero, where the conjecture connects to the Atiyah conjecture, is untouched.

### Kaplansky's direct finiteness, surjunctivity and the Determinant Conjecture (family 197)

The direct-finiteness conjecture says $ab=1$ implies $ba=1$ in every group algebra. Kaplansky proved it in characteristic 0 for all groups. In characteristic $p$ it is known for sofic groups (Elek and Szabó, 2004), and no non-sofic group is known. So a counterexample does two things at once: it disproves Kaplansky, and it proves that non-sofic groups exist, one of the best-known open questions in group theory. It also refutes Gottschalk's surjunctivity conjecture, because the same elements define a cellular automaton that is injective but not surjective, and the Lück-style Determinant Conjecture, through an integral matrix with Fuglede–Kadison determinant strictly between 0 and 1.

The family has four papers, and the formal coverage is uneven in a way the catalogue blurb hides. The headline, a torsion-free group over $\mathbb F_2$ (manuscript of October 4), is not in Lean. What Lean has is the September 23 version: a finite field of characteristic 2 and a finitely presented group with an element of odd prime order (`OAI/RingTheory/DirectFiniteness/FinitelyPresented.lean:35`). It also has an odd-characteristic version over a field of order $p^4$, where $p$ is the smallest prime factor of $\bigl(\binom{1200}{600}!\bigr)^2+1$, which is a fine way to say "some prime, don't ask" (`OAI/Algebra/OddKaplansky/Main.lean:81`). And it has the Determinant counterexample, derived in Lean from the characteristic-2 one. Non-soficity is not stated in Lean; it follows on paper from Elek–Szabó. The release includes an abridged reasoning trace for the characteristic-2 case.

### Kakeya in three and four dimensions

The four Fourier-analysis families that follow (Kakeya, restriction, Bochner–Riesz, local smoothing) are best read together. They form a chain that harmonic analysts have studied for fifty years. Local smoothing for the wave equation implies Bochner–Riesz, Bochner–Riesz implies restriction, and restriction implies Kakeya. The release claims every link in $\mathbb R^3$. None of the four is formalized. And they lean on each other.

The Kakeya set conjecture asks whether a set containing a unit segment in every direction must have full dimension. In $\mathbb R^3$ that is human work. Hong Wang and Joshua Zahl proved it in February 2025 ([arXiv:2502.17655](https://arxiv.org/abs/2502.17655)), and it is one of the results of the decade. The new three-dimensional content in family 074 is the *maximal function* version, the quantitative statement harmonic analysis actually consumes:

$$
\|K_\delta f\|_{L^3(S^2)} \le C_\epsilon\, \delta^{-\epsilon} \|f\|_{L^3(\mathbb R^3)} .
$$

It measures how much $\delta$-tubes pointing in different directions can pile up. Wang and Zahl's method gives a density dependence of $\lambda^{K(\epsilon)}$ where the maximal conjecture needs $\lambda^3$, and they flag the upgrade as open. The OpenAI paper imports their set estimate and closes that gap through an extremal problem. Any counterexample produces a stationary arrangement of planes and plates, which entropy increments and planar projection estimates then rule out. The second paper claims the set conjecture in $\mathbb R^4$, in the Hausdorff-dimension form and with no measurability assumption. Before this, the best bound was Katz and Zahl's 3.059. The method fits quadratic polynomials to tube families across scales.

Caveats: 272 pages between the two, no Lean, and the four-dimensional paper uses a lemma (a weighted plank bound) from the three-dimensional one.

### Hadwiger's conjecture, disproved on paper and linearised in Lean

Hadwiger's conjecture (1943) says that a graph with chromatic number $t$ contains $K_t$ as a minor. It generalises the four-colour theorem, it is known for $t \le 6$, and it is arguably the central open problem of graph minor theory. Family 157 bundles three manuscripts, and they point in opposite directions.

The first two (104 and 131 pages) claim counterexamples. For arbitrarily large $m$ there are $m$-vertex graphs with independence number at most 2 whose largest "connected matching" has fewer than $m/100$ edges. With $\alpha(G) \le 2$ you need at least $m/2$ colours even fractionally, while a known inequality turns the small connected matching into a clique-minor bound $h(G) \lt 26m/75 + 2/3$, which is below $m/2$. So $\chi_f(G) > h(G)$, and Hadwiger fails even in its fractional form. The second paper does the same for Colin de Verdière's conjecture $\chi \le \mu + 1$. The construction samples points carrying algebraic labels and declares some pairs to be "holes" (non-edges), arranged so the holes form no triangle, which is what gives $\alpha \le 2$. A distribution theorem then forces any matching to contain two edges with all four cross pairs missing, so they cannot both sit in a connected matching.

The $\alpha \le 2$ case is the best-known stress test of the conjecture; Seymour's 2016 survey singles it out. In December 2025, Kühn, Sauermann, Steiner and Wigderson disproved the *odd* Hadwiger conjecture with exactly such graphs (arXiv 2512.20392), so the ground was softened. Still, a counterexample to Hadwiger itself would be the biggest graph-theory news in decades, and here it has no Lean statement, no explicit graph, and 235 pages of new probabilistic and algebraic machinery that nobody outside OpenAI has checked.

The third manuscript is formalized, and it is a big deal in its own right: an absolute $C$ with $\chi_{\text{list}}(G) \le C \cdot h(G)$ for every graph. That is the Linear (List) Hadwiger conjecture. The best general bounds before this were $O(t \log \log t)$ (Delcourt and Postle) and, in September 2026, $O(t \log \log \log t)$ (Liu and Luo, arXiv 2609.06867). The formal statement is `OAI.LinearListHadwiger.main_theorem` (`ComparatorChallenges/ListHadwiger.lean:30`), and the solution proves the Reed–Seymour input rather than assuming it. It is not listed in `formalization.yaml`'s main results, though, which looks like a bookkeeping gap rather than a signal.

So the honest summary is that the family's lead claim is its least verified one. The positive linear bound is probably the most solid big result in the family. I also noticed that the list-colouring paper is dated September 23 but cites a Gu–Xu preprint "posted on October 1, 2026", so at least that manuscript was revised after its stated date.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig6.png"
  alt="Four vertices: A1 joined to A2 and B1 joined to B2 by solid edges, with all four cross pairs A1–B1, A1–B2, A2–B1, A2–B2 drawn as dashed red lines labelled holes."
  caption="The conflict the Hadwiger construction forces: two matching edges whose four cross pairs are all non-edges cannot belong to one connected matching (A counterexample to Hadwiger's conjecture, Figure 1)."
/>

### Erdős's \$5000 progression problem

Erdős conjectured that any set $A$ of positive integers with $\sum_{a \in A} 1/a = \infty$ contains arithmetic progressions of every length (erdosproblems.com #3, with a \$5000 prize attached). The primes are the motivating case, and Green–Tao handles them by other means. What the conjecture needs is a Szemerédi bound $r_k(N)$ that decays faster than $N/(\log N)^{1+\epsilon}$. For $k = 3$ that arrived with Kelley and Meka in 2023. For $k \ge 5$ the best was Leng, Sah and Sawhney's $N \exp(-(\log \log N)^{c})$ from 2024, which is not summable over dyadic scales.

Family 159 is a single 198-page manuscript claiming

$$
r_k(N) \le C_k \, N \exp\!\big(-c_k (\log N)^{\epsilon_k}\big) \quad \text{for every } k \ge 3,
$$

a Kelley–Meka-strength bound for every progression length, with Erdős's conjecture as a one-line corollary. It is one of the ten families whose reasoning summary OpenAI published, and that 41-page trace is worth reading for texture: the first attempt aimed at $N/(\log N)^3$, and many steps are flagged "unverified" along the way. The proof runs a density increment over "triangular" polynomial cells and shows that the precision lost at each weight depends only on higher-weight data, so the increments cost polynomially in $\log(1/\alpha)$.

The Lean situation needs a careful reading. The Comparator statement (`ErdosReciprocal.lean:31`) is the qualitative Erdős corollary: a set with non-summable reciprocals has a $k$-term progression for every $k$. The solution is enormous (about 1.12 million lines in 4,774 files), and it reaches the corollary through a weaker internal bound, $N \exp(-c(\log \log N)^{1+\eta})$, which is just enough for summability. If that builds, Erdős #3 is settled formally, which is remarkable. The paper's headline quantitative Theorem 1.1 is not checked at all, and the family does not appear in `formalization.yaml`.

## The rest, by discipline

What follows covers every remaining family, grouped the way the catalogue groups them. Within each group the order is mine, most important first. Majors and notables get prose; technical results get a line. The [searchable catalogue](#the-searchable-catalogue) at the end lists every family in every group, with its grade, kind, Lean status and main caveat.

The grades are mine, not OpenAI's. "Landmark" means a famous named problem that the paper claims to settle fully. It is a statement about the claim, not about whether the claim is right. That is why there are so many of them.

## Theoretical computer science

Complexity theory is where famous conjectures go to stay conjectures, and the field has a long history of confident "proofs" of things like P versus NP that fall apart on page three. That is the backdrop for one of the densest parts of the release: 40 families, 73 manuscripts, about 2,800 pages, and a list of headline claims that reads like the open-problems slide at the end of a STOC keynote. The Unique Games Conjecture. L = BPL. Matrix multiplication with $\omega \le 9/4$. Integer multiplication below $n\log n$. Approximate counting of perfect matchings in general graphs. Garey and Johnson's three-machine scheduling problem. Subset Sum below $2^{n/2}$. Deterministic polynomial factoring over $\mathbb F_p$ without GRH.

Any one of those would be the paper of the decade. I can't check a 108-page derandomization proof, so I sorted the claims by how much independent evidence stands behind them, and that sorting matters more than the headlines: 24 of the 40 families have their main theorem stated in Lean against a Comparator challenge, and 8 have nothing formal at all, and those two groups are not the same kind of claim.

"Main theorem formalized" here means what it means everywhere on this page: OpenAI says a kernel accepts it against a Comparator statement you can read yourself. Most preprints offer far less; a refereed proof offers more. The map below shows where the approximation results land if they hold.

<ApproximationThresholdMap />

### 2-to-1 games and colouring a 3-colourable graph

Two families sit right next to Unique Games and are easiest to understand together.

[Perfect completeness for 2-to-1 games](https://github.com/openai/math/blob/main/preprints/Perfect-completeness-for-2-to-1-games-September-23-2026/paper.pdf) (family 105, 53 pages) proves Khot's 2-to-1 Games Conjecture in its perfect-completeness form: for every fixed $\delta$, it is NP-hard to tell a 2-to-1 game that is *fully* satisfiable from one of value at most $\delta$, with every right label having exactly two preimages. Unique Games needs completeness near one; colouring needs completeness exactly one, because a graph is either 3-colourable or it isn't. Austrin, O'Donnell, Tan and Wright had perfect completeness with soundness stuck at $23/24 + \varepsilon$. The Lean statement (`PerfectCompleteness.lean`, solution about 137,000 lines) is formalized.

Then the human timeline gets interesting. On 14 September 2026, nine days before this manuscript's date, Yumou Fei, Dor Minzer and Shuo Wang posted [ECCC TR26-179](https://eccc.weizmann.ac.il/report/2026/179/), proving the 4-to-1 Games Conjecture with perfect completeness and deriving that finding a $k$-colouring of a 3-colourable graph is NP-hard for every constant $k$. The OpenAI paper cites them accurately and says a step of its own has "a close counterpart" in their Lemmas 5.15–5.16. Since 2-to-1 is the strongest of Khot's $d$-to-1 conjectures, this is still an extension of their result. It is a smaller step than "proves Khot's 2-to-1 conjecture" sounds when you don't know about the paper from the week before.

[Hardness of finding large independent sets in three-colorable graphs](https://github.com/openai/math/blob/main/preprints/Hardness-of-finding-large-independent-sets-in-three-colorable-graphs-September-24-2026/Hardness-of-finding-large-independent-sets-in-three-colorable-graphs-September-24-2026.pdf) (family 106, 21 pages) is where I think the release's own summary oversells. The catalogue entry leads with "It is NP-hard to color a three-colorable graph using any fixed number $c \ge 3$ of colors." True, and first proved by Fei, Minzer and Wang, as the paper itself says in its history section. What is new is the stronger form: for every fixed $\delta \lt 1/3$ there is a reduction from 3SAT to graphs that are 3-colourable when the formula is satisfiable and have *no independent set of $\delta n$ vertices* otherwise. Fei–Minzer–Wang get the same independent-set soundness only with an 8-colourable promise. This paper keeps 3-colourability of the whole graph and goes straight from ordinary Label Cover, without 105.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/theoretical-cs-fig2.png" alt="Left: a commutative diagram where a phase vector x is pulled back along an answer projection pi by a zero-length link, and evaluating either side at the satisfying answers gives the same value. Right: a circle divided into three colored arcs (blue, green, red) of length one third, with a point phi(x) and its antipode phi(x)+1/2; a black arc near the antipode shows where the phases of neighbours can lie, at circle distance at least 3/8." caption="The completeness half of the colouring reduction: satisfying labels keep phases consistent across projection links (left), edges force phases nearly opposite, and three equal arcs of the circle give a proper 3-colouring (Hardness of finding large independent sets in three-colorable graphs, Figure 1)." />

The construction is pleasingly geometric. Every vertex carries a phase on the circle $\mathbb R/\mathbb Z$, computed from a "phase assignment" to product answers along chains of Label Cover questions. Edges join a point to the half-shift of another, so in the honest case adjacent vertices sit at circle distance at least $3/8$, more than the $1/3$ length of an arc, and colouring by which third of the circle you land in is proper. Soundness turns an independent set into Lipschitz functions on tori and uses Tim Austin's dimension-free junta theorem (a continuous cousin of Friedgut's) to show they depend on few coordinates, then a Ramsey-style alignment lemma to decode a Label Cover labeling. The main theorem is formalized (`IndependentSets.lean`, `OAI.LargeIndependentSets.mainTheoremReal`, about 94,000 lines). For the colouring algorithms people actually run (DSatur, greedy with good orderings, SAT encodings) nothing changes; the best polynomial-time guarantee is still a small power of $n$: about $n^{0.197}$ colours from Kawarabayashi, Thorup and Yoneda ([arXiv 2406.00357](https://arxiv.org/abs/2406.00357)), $n^{0.195}$ from Bansal, Huang and Lee ([arXiv 2602.05904](https://arxiv.org/abs/2602.05904)). The gap between "any constant is hard" and "about $n^{0.2}$ is achievable" is as wide as ever.

### A cubic lower bound for the permanent

The permanent of an $m \times m$ matrix is the determinant with every sign set to plus, and the difference is the whole of algebraic complexity theory. The determinant is computable by Gaussian elimination; the permanent is #P-hard (Valiant, 1979), and Valiant's VP versus VNP question, the algebraic version of P versus NP, asks whether the permanent can be written as the determinant of a polynomial-size matrix whose entries are affine-linear forms in the variables. The smallest such size is the determinantal complexity $\mathrm{dc}(\mathrm{perm}_m)$; showing it is superpolynomial would separate VP from VNP.

For twenty years the best lower bound was quadratic: $m^2/2$, by Mignon and Ressayre (IMRN 2004) using Hessian ranks, extended by Landsberg, Manivel and Ressayre (2013) to the *border* version, where you only need a sequence of determinants converging coefficientwise to the permanent. The best upper bound is Grenet's $2^m - 1$. [A cubic lower bound for border determinantal complexity of the permanent](https://github.com/openai/math/blob/main/preprints/A-cubic-lower-bound-for-border-determinantal-complexity-of-the-permanent-September-24-2026/A-cubic-lower-bound-for-border-determinantal-complexity-of-the-permanent-September-24-2026.pdf) (family 108, 35 pages) raises the exponent to three: $\overline{\mathrm{dc}}(\mathrm{perm}_m) \ge m^3/(5529600e)$ for every $m \ge 1408$, which also gives cubic lower bounds for affine algebraic branching programs.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/theoretical-cs-fig3.png" alt="Proof map in four boxes. Top: one integral coefficient polynomial G_D(X), the coefficient of s^q t^q in per_k(X + sI + tD). Left branch over the complex numbers: the permanent projects onto a0 G_D(X) with m at least 11k - 2q. Right branch over the algebraic closure of F_2: gradient zeros of G force a plane-curve delta invariant at least T, giving codimension at least T - k. They meet in a smooth affine restriction G = per_m composed with an affine map, of degree r with r - 1 at least k/4 in d variables with d - 1 at least k^2/200, and the polar count gives border dc of per_m at least (r-1)(d-1)/(4e) at least k^3/(3200e)." caption="The two uses of one coefficient polynomial: a permanent projection over C, and plane-curve geometry in characteristic two that certifies smoothness, combined by a polar-degree count (A cubic lower bound for border determinantal complexity of the permanent, Figure 1)." />

The mechanism is a clean two-step. First, a general lower bound: any form of degree $r$ in $d$ variables whose projective zero set is smooth has (border) determinantal complexity at least $(r-1)(d-1)/(4e)$, proved by counting simple polar intersections that must persist in any nearby determinant. Second, a construction: take the coefficient $G_D(X) = [s^q t^q]\,\mathrm{per}_k(X + sI + tD)$, which has degree about $k/3$ in $k^2$ variables, show it is a projection of a permanent of size about $11k$, and show (by a detour through plane curves over $\overline{\mathbb F}_2$, where permanent and determinant coincide) that a generic linear section of it is smooth. Degree linear in $k$ times dimension quadratic in $k$ gives cubic. Figure 1 above is the whole proof in one picture.

The main statement is formalized, and I find it the most satisfying Lean snippet in the group because there is nothing to hide in it: the permanent as a sum over permutations, an affine matrix pencil, coefficientwise limits, and the inequality.

```lean
-- lean/OAI/Algebra/Determinantal/Main.lean:384-389
theorem permanent_cubic_lower_bounds :
    (∀ (m n : ℕ), 1408 ≤ m → HasBorderDeterminant n (permanentPolynomial m) →
      (m : ℝ)^3 / (5529600 * Real.exp 1) ≤ (n : ℝ)) ∧
    (∀ (m n : ℕ), 1408 ≤ m → HasExactDeterminant n (permanentPolynomial m) →
      (m : ℝ)^3 / (5529600 * Real.exp 1) ≤ (n : ℝ)) :=
  ⟨permanent_border_cubic_lower_bound, permanent_exact_cubic_lower_bound⟩
```

One thing I could not resolve: the paper leans on a deep external theorem (the arbitrary-characteristic Severi dimension bound of Christ, He and Tyomkin), but the 57-file, 21,000-line Lean development never mentions Severi, so either it found a more elementary route or the dependency is hidden under other names. With no extra axioms allowed, it has to be one of those. Also keep the constant in view: $m^3/(5529600e)$ beats $m^2/2$ only once $m$ is in the millions. As mathematics this is a real move on a famous problem; as a step toward VP versus VNP it is a polynomial improvement on a road that needs a superpolynomial one.

### Counting perfect matchings in any graph

Of the theoretical-CS landmarks, this is the one I expect to survive review intact, because the Lean statement is a complete specification of a randomized algorithm and its guarantee. [A fully polynomial randomized approximation scheme for perfect matchings in general graphs](https://github.com/openai/math/blob/main/preprints/A-Fully-Polynomial-Randomized-Approximation-Scheme-for-Perfect-Matchings-in-General-Graphs-September-23-2026/main.pdf) (family 113, 45 pages) gives an algorithm that, for any simple graph and any $\varepsilon, \delta$, returns the number of perfect matchings to within a factor $1 \pm \varepsilon$ with probability $1 - \delta$, in worst-case time polynomial in the input, $1/\varepsilon$ and $\log(1/\delta)$, and returns exactly zero when there are none.

Exact counting is #P-complete even for bipartite graphs (Valiant). Jerrum and Sinclair (1989) gave a Markov-chain FPRAS when near-perfect matchings don't outnumber perfect ones by more than a polynomial, and explicitly asked about the general case. Jerrum, Sinclair and Vigoda (JACM 2004) solved the bipartite case, the permanent of a nonnegative matrix, and that algorithm is one of the jewels of the field. General graphs, with their odd cycles, resisted for 35 years, and Štefankovič, Vigoda and Wilmes showed why: on some graphs every JSV-style chain either mixes exponentially slowly or puts exponentially little mass on perfect matchings. The new algorithm changes the state space instead of the weights, sampling from a product of perfect-matching spaces on an enlarged, coloured graph and moving by exchanging alternating cycles between coordinates, with marked edges exposing the partition-function ratios and vertex scaling keeping them balanced. A companion paper proves the Anari–Oveis Gharan–Vinzant perfect-matching entropy conjecture, $F(x) - (2 - 2/m)B(x) \le H(x) \le F(x)$ at every point of the matching polytope.

The Comparator statement (`MatchingFPRAS.lean`) quantifies over all graphs and rational $\varepsilon, \delta$, bounds every run of a fixed random machine by one polynomial, requires zero on graphs with no perfect matching, and requires at least a $1 - \delta$ fraction of random tapes to land in the $(1 \pm \varepsilon)$ window: the whole definition of an FPRAS, written out. Approximate counting is not something most engineers do, but sampling is: dimer models in statistical physics, random perfect matchings as null models, hafnian-type quantities in boson sampling. The polynomial is unspecified and surely large, so the first practical consequence is that someone will now try to make it small.

Two neighbouring families extend the same line, and I'll keep them short. [Family 114](https://github.com/openai/math/blob/main/preprints/Approximate-counting-of-common-bases-of-two-matroids-September-23-2026/main.pdf) gives an FPRAS for the number of common bases of two matroids given by independence oracles, which contains bipartite perfect matchings and the single-matroid counting breakthrough of Anari, Liu, Oveis Gharan and Vinzant (STOC 2019) as special cases; the two-matroid version is in Lean, the polymatroid extension is not. [Family 115](https://github.com/openai/math/blob/main/preprints/Exact-Uniform-Sampling-of-Contingency-Tables-with-Arbitrary-Margins-September-24-2026/main.pdf) samples contingency tables with arbitrary binary-encoded row and column sums exactly uniformly in expected polynomial time, and counts them with cell bounds; statisticians use exactly this distribution as the null model behind Fisher-type exact tests, so far with MCMC heuristics of unknown mixing time. Both are formalized.

### Integer multiplication below $n \log n$

In 1971 Schönhage and Strassen multiplied $n$-bit integers in $O(n\log n\log\log n)$ steps on a multitape Turing machine and conjectured that $n\log n$ is optimal. Fürer got within $2^{O(\log^\ast n)}$ of it in 2007, and Harvey and van der Hoeven reached $O(n\log n)$ exactly (Annals of Mathematics, 2021), which most people took to be the end of the story. [Integer multiplication below n log n](https://github.com/openai/math/blob/main/preprints/Integer-multiplication-below-n-log-n-September-23-2026/paper.pdf) (family 109, 73 pages) claims one fixed multitape Turing machine that multiplies exactly in $O(n(\lg n)^{1-\kappa})$ time with $\kappa = 2^{-182}$, refuting the conjecture in the model it was stated in.

The idea is network coding. On a tape machine, just rearranging $\Theta(n)$ bits at each of $\Theta(\log n)$ FFT levels already costs $n\log n$, so beating it requires saving on both data movement and arithmetic. The paper builds two procedures from fixed linear networks: one swaps two address fields of an array in $O(V u^\tau)$ time for $\tau \lt 1$ by passing XOR combinations of the data around, and one applies a layer of butterflies $(u+v)/2, (u-v)/2$ to selected coordinate bits more cheaply than a full pass. Both recurse on themselves with a strict saving ($s \lt Wm$ smaller calls per level), and that strictness is where the power of $\log n$ comes from. The rest is the Harvey–van der Hoeven pipeline (Gaussian resampling, Bluestein chirps, Nussbaumer's synthetic transforms) with every representation change charged. This is the same engine the DFT family (130) reuses for its "exact Fourier transform below $n\log n$" claim, which has its own section.

Afshani, Freksen, Kamma and Larsen (2019) showed an $\Omega(n\log n)$ circuit lower bound for multiplication would follow from the network-coding conjecture. That does not clash with this result, because a time-$T$ Turing machine only gives circuits of size $O(T\log T)$, but it is a good sign that the paper's new ingredient is precisely linear combinations of data in transit. There is no Lean statement. And for anyone with a bignum library: $(\lg n)^{2^{-182}}$ is indistinguishable from 1 for any number that fits in the universe. GMP is safe. The interest is purely that $n \log n$ was a natural barrier and, if this holds, it isn't one.

### Three machines, unit jobs, precedence constraints: in P

Some open problems are famous because they are deep; this one is famous because it is on a list. Garey and Johnson's 1979 book closes with a handful of problems whose complexity was unknown, and $P3 \mid \mathrm{prec}, p_j = 1 \mid C_{\max}$ (schedule unit-length jobs with precedence constraints on three identical machines to minimize makespan) is "OPEN8". Two machines are easy (Coffman–Graham, 1972); an unbounded number is NP-complete (Ullman, 1975). Three stayed open for 47 years. The approximation side got close: Levey and Rothvoss (STOC 2016) gave a $(1+\varepsilon)$-approximation via LP hierarchies in slightly superquasipolynomial time, later improved to quasipolynomial, and Nederlof, Swennenhuis and Węgrzycki gave an exact $2^{O(\sqrt{n\log n})}$ algorithm.

[A polynomial-time algorithm for three-machine unit-job scheduling](https://github.com/openai/math/blob/main/preprints/A-polynomial-time-algorithm-for-three-machine-unit-job-scheduling-September-24-2026/paper.pdf) (family 124, 22 pages) puts it in P. Feasible schedules are reorganized into intervals whose job sets have bounded-size descriptions, and a dynamic program searches a family of polynomially many such descriptions, with global boundary conditions and a simplification of inherited information keeping the descriptions bounded through the recursion. The Lean statement fixes the exponent in the theorem itself, which made me laugh and then made me trust it more:

```lean
-- lean/ComparatorChallenges/ThreeMachine.lean:86-96
def ReleaseTheorem : Prop :=
  ∃ (k q g : ℕ) (M : Machine k q g) (C : ℕ), 0 < C ∧
    ∀ n (G : Instance n), 1 ≤ n → G.acyclic →
      ∀ deadline : Option ℕ,
        (∀ T ∈ deadline, 1 ≤ T ∧ T ≤ n) →
        ∃ out, CorrectOutput G deadline out ∧
          M.Produces (encodeInput G deadline) out
            (C * ((encodeInput G deadline).length + 2) ^ 150020)

theorem main_theorem : ReleaseTheorem := by
  sorry
```

Time $O((L+2)^{150020})$ on a multitape machine. Nobody will run it. It is still exactly the right kind of answer to a classification question, and a 22-page paper with a Lean proof (about 18,600 lines in `OAI/Computability/Scheduling`) is about the most checkable package a claim like this could come in. One oddity: this challenge, like the matching FPRAS and k-median ones, is not listed under `status.main_results` in `formalization.yaml`, although its Comparator JSON and solution module exist. I'd read that as bookkeeping, not as a retraction.

### Subset Sum below $2^{n/2}$

Horowitz and Sahni's meet-in-the-middle algorithm solves Subset Sum on $n$ integers in $2^{n/2}$ time (1974), Schroeppel and Shamir got the space down to $2^{n/4}$ (1981), and in the worst case nothing better than polynomial factors has been known since. Random instances fall much faster to the representation technique of Howgrave-Graham and Joux and its descendants (about $2^{0.291n}$), but those arguments need randomness in the input. The most recent worst-case gain was Chen, Jin, Randolph and Servedio's $2^{n/2}/n^{0.5023}$, cited by the paper.

[Subset Sum in time $O(2^{0.49n})$](https://github.com/openai/math/blob/main/preprints/Subset-Sum-in-Time-2-power-0-49n-October-4-2026/subset-sum.pdf) (family 138, 33 pages) claims a randomized algorithm that is correct with probability $2/3$ and runs in $O(2^{0.49n})$ word-RAM time on every execution, for integers of polynomially many bits. The idea mixes two old tools in a new way. An isolation lemma (Mulmuley–Vazirani–Vazirani) makes some transformed target have a unique solution. Each subset then gets a random phase $e_P(\sum_{i \in I}\phi_i)$ modulo a random prime $P$ larger than the running time, and the sum of phases over subsets hitting the target exactly equals the sum over subsets hitting it modulo $P$, minus the sum over "aliases" that only hit it modulo $P$. The modular sum is estimated from a sample of Fourier frequencies, cleverly enough to avoid the obvious $2^n$ second-moment cost, and the aliases are listed explicitly with representation-style enumeration and subtracted. A nonzero remainder means a solution exists. A companion paper keeps $2^{n/2}$ time but needs only $2^{n/5}$ words of space, below Schroeppel–Shamir.

No Lean. The saving is $2^{0.01n}$, so at $n = 100$ it is a factor of two, and the word size is $\Theta(n)$ bits, generous but standard here. Cryptographic knapsacks are not affected (they fall to lattice attacks anyway). The reason this matters is that $2^{n/2}$ has been the worst-case barrier for half a century, and a lot of fine-grained complexity quietly assumes it. It is also not the only fine-grained barrier claimed this month: Alman and Vassilevska Williams's subquadratic 3SUM and subcubic APSP ([arXiv 2610.06783](https://arxiv.org/abs/2610.06783)) is a separate paper, not part of this release, and has [its own article](/articles/subquadratic-3sum-apsp).

### Deterministic factoring over $\mathbb F_p$, conditional on a companion

Factoring a polynomial over a prime field is routine with randomness: Berlekamp (1970) and Cantor–Zassenhaus (1981) do it in expected polynomial time, and every computer algebra system ships it. Doing it deterministically, in time polynomial in both the degree and $\log p$, has been open the whole time. Even the degree-2 case contains the problem of deterministically finding a quadratic non-residue mod $p$, which is only known under the Generalized Riemann Hypothesis. Under GRH the best general bound was Evdokimov's quasipolynomial $(n^{\log n}\log p)^{O(1)}$.

[Deterministic polynomial factorization over prime fields](https://github.com/openai/math/blob/main/preprints/Deterministic-Polynomial-Factorization-over-Prime-Fields-October-4-2026/Deterministic-Polynomial-Factorization-over-Prime-Fields.pdf) (family 142, 48 pages) claims it unconditionally, in $O(((n+1)\lceil\log_2 p\rceil)^{10^{12}})$ bit operations. The structure is a clean split. Theorem 1.2 is purely algebraic: given, for each prime $q \le n$, an auxiliary prime $\ell \equiv 1 \pmod{12q}$ modulo which $p$ is not a $q$-th power, the algorithm factors $f$ deterministically, using geometry over curves to separate factors. Theorem 1.1 then needs such $\ell$ to be polynomially small, and that comes from a uniform zero-free region for Hecke L-functions proved in a different family entirely: 029, "Primitive roots for every admissible integer base", which claims the infinitude part of Artin's primitive root conjecture. Neither has a Lean statement.

So I classify this one as conditional. If the companion falls, the paper notes that its reduction plus the zero-free strip that follows from GRH still gives a deterministic polynomial algorithm under GRH, which would itself improve on Evdokimov. If both hold, this is a landmark in computational algebra stacked on a landmark in analytic number theory, with no machine check for either. An exponent of $10^{12}$ means Cantor–Zassenhaus stays in every library you use.

### Sakoda–Sipser: two-way automata need exponentially many states

In 1978 Sakoda and Sipser asked how many states a two-way deterministic finite automaton (one whose head can move back and forth) needs to simulate a nondeterministic one. Both recognize exactly the regular languages, so the question is about succinctness, and it became the flagship open problem of descriptional complexity. Lower bounds were known only for restricted deterministic machines (Sipser's sweeping automata, 1980; Kapoutsis for few reversals), and for unrestricted ones only quadratic bounds.

Family 129 claims exponential separations for unrestricted machines. [The one-way liveness paper](https://github.com/openai/math/blob/main/preprints/An-exponential-two-way-deterministic-state-lower-bound-for-one-way-liveness-September-25-2026/main.pdf) takes Sakoda and Sipser's own complete family: letters are binary relations on $h$ points, and a word is accepted when the product of its relations is nonempty. A one-way NFA with $h + 3$ states recognizes it; any two-way DFA with $s$ states needs $4(s+2)^2 \ge 2^{\lfloor (h-2)/31\rfloor}$. [The complementation paper](https://github.com/openai/math/blob/main/preprints/An-exponential-state-lower-bound-for-two-way-nondeterministic-complementation-September-25-2026/paper.pdf) shows two-way NFAs cannot be complemented with polynomially many states. The key representation is the Brauer diagram monoid: a deterministic machine's repeated visits to a cell are encoded as perfect matchings on ports, multiplication glues and contracts paths, and the algebra shows matchings cannot track arbitrary relation products without exponentially many ports. All three statements are in Lean.

The caveat is the standard one for this problem, and the paper is careful about it. Alphabets grow (exponentially) with $h$, which is the usual setting, and the hard inputs are long. Kapoutsis showed that an exponential lower bound on *polynomially long* inputs would separate L/poly from NL, and nothing here claims that, so this is the language-theoretic question answered, not a complexity-class separation in disguise.

### The Courtade–Kumar conjecture

Take $n$ fair bits $X$, send them through independent binary symmetric channels with crossover $\varepsilon$ to get $Y$, and choose any Boolean function $f$ of $X$. How much can $f(X)$ tell you about $Y$? Courtade and Kumar conjectured in 2013 (IEEE Transactions on Information Theory, 2014) that the answer is at most $1 - h_2(\varepsilon)$ bits, achieved by just sending one coordinate, the "dictator". It looks like a homework problem and resisted a decade of serious effort: Samorodnitsky (2016) proved it in a high-noise range, Yu pushed the balanced case to correlation $0.914$ with computer assistance, and Anantharam, Bogdanov, Chakrabarti, Jayram and Nair (2017) proposed a stronger "Hellinger conjecture" that implies it.

[Sharp binary-information contraction on the discrete cube](https://github.com/openai/math/blob/main/preprints/Sharp-binary-information-contraction-on-the-discrete-cube-September-24-2026/main.pdf) (family 119) proves a sharper statement for all soft binary channels, $I(T_\rho u) \le \psi(|\rho|\,\psi^{-1}(I(u)))$, with the conjecture as the Boolean case, and the companion proves the Hellinger conjecture for every output bias, using the differential approach of Chen, Gohari and Nair (2025) plus exact-arithmetic certificates. The first paper's statements, including the Courtade–Kumar inequality with attainment, are in Lean; the Hellinger paper is not. For anyone designing one-bit quantizers or feature hashes under noise, the folk wisdom "a single coordinate is the most informative bit" now has a theorem behind it.

### The rest of the major results

These families would each be headline news in a normal year. I give each a few sentences: the claim, what it improves, how much is checked.

#### Mean-payoff games in quasipolynomial time (104)

Mean-payoff games, where two players move a token around a weighted graph and fight over the long-run average, underlie controller synthesis and are in NP and coNP, with only pseudopolynomial (Zwick–Paterson, 1996) or subexponential randomized algorithms for binary weights. [The deterministic paper](https://github.com/openai/math/blob/main/preprints/Deterministic-quasipolynomial-time-mean-payoff-games-September-25-2026/paper.pdf) claims $2^{O((\log L)^2)}$ bit operations for exact values and optimal positional strategies, putting them where parity games landed in 2017 (Calude et al.), with extensions to turn-based stochastic games (threshold zero) and mean-payoff parity games. Only the randomized algorithm, and a small counterexample to a step of Truffet's proposed procedure, are in Lean. Polynomial time remains open.

#### Randomized k-server at $O(\log^2 k)$ on every metric (110)

Bubeck, Coester and Rabani killed the $O(\log k)$ randomized k-server conjecture in 2023 ([arXiv 2211.05753](https://arxiv.org/abs/2211.05753)) with an $\Omega(\log^2 k)$ lower bound on some metrics. [This family](https://github.com/openai/math/blob/main/preprints/Squared-logarithmic-randomized-k-server-on-arbitrary-metrics-September-24-2026/Squared-logarithmic-randomized-k-server-on-arbitrary-metrics-September-24-2026.pdf) claims a matching $O(\log^2(k+1))$ competitive ratio on every metric space, finite or not, with no dependence on the number of points, by allocating fractional server mass over partitions driven by the posterior of a hidden offline optimum. A companion makes it uniform on finite rational metrics with polynomial per-request cost but an additive constant that "may be enormous". Both are in Lean. Caching theory, not caching practice.

#### Depth-3 circuits beyond $2^{c\sqrt n}$ (112)

Explicit depth-3 (OR–AND–OR) lower bounds have sat at $2^{c\sqrt n}$ since Håstad, with only the constant $c$ improving. [This 20-page paper](https://github.com/openai/math/blob/main/preprints/Beyond-the-Square-Root-Exponent-for-Depth-Three-Boolean-Circuits-September-23-2026/main.pdf) gives one polynomial-time language that needs $2^{\omega(\sqrt n)}$ gates, through a restriction lemma that rewrites a restricted wide CNF as a signed combination of narrow CNFs with bounded total weight, independent of the clause count. It does not reach $2^{n^{1/2+\varepsilon}}$, the bound people actually want. Formalized.

#### Uniform Sparsest Cut: no constant factor (117)

Arora, Rao and Vazirani's $O(\sqrt{\log n})$ algorithm is the famous one; on the hardness side there was nothing unconditional at constant factors. [The hardness paper](https://github.com/openai/math/blob/main/preprints/Constant-factor-hardness-of-uniform-sparsest-cut-September-24-2026/Constant-factor-hardness-of-uniform-sparsest-cut-September-24-2026.pdf) claims every constant factor is NP-hard for the uniform-demand version, by a direct reduction independent of the Unique Games paper, and [the companion](https://github.com/openai/math/blob/main/preprints/Near-square-root-logarithmic-integrality-gaps-for-uniform-sparsest-cut-September-24-2026/Near-square-root-logarithmic-integrality-gaps-for-uniform-sparsest-cut-September-24-2026.pdf) shows the Goemans–Linial SDP loses $c\sqrt{\log n}/(\log\log n)^3$, nearly matching ARV (previous gaps: $\Omega(\log\log n)$ by Devanur, Khot, Saket and Vishnoi, improved by Kane and Meka). Only the integrality gap is in Lean, so the bigger claim is the unchecked one.

#### Bin packing: no algorithm is within a constant number of bins (118)

The configuration LP is the workhorse of cutting-stock software, and the Modified Integer Round-Up Conjecture said it is never off by more than one bin. [This paper](https://github.com/openai/math/blob/main/preprints/Additive-hardness-and-unbounded-configuration-gaps-in-bin-packing-September-24-2026/Additive-hardness-and-unbounded-configuration-gaps-in-bin-packing-September-24-2026.pdf) builds instances, all items larger than $1/6$, where the LP says $B$ and the optimum exceeds $B + c$ for any $c$, and shows it is NP-hard to tell $B$ bins from more than $B + c$. Since Hoberg and Rothvoss (SODA 2017) achieve $\mathrm{OPT} + O(\log \mathrm{OPT})$, bin packing's additive approximability is now pinned between $\omega(1)$ and $O(\log \mathrm{OPT})$, and Williamson and Shmoys's textbook open problem has a negative answer. Fully formalized, about 159,000 lines.

#### Maximum matching in almost-linear time (120)

Bipartite matching became almost-linear with the 2022 max-flow breakthrough; general graphs, with blossoms, stayed at Micali–Vazirani's $O(m\sqrt n)$ from 1980. [This 86-page paper](https://github.com/openai/math/blob/main/preprints/Almost-Linear-Time-Maximum-Cardinality-Matching-in-Sparse-General-Graphs-September-24-2026/main.pdf) claims a Monte Carlo $(n+m)^{1+o(1)}$ algorithm that outputs a maximum matching with probability $2/3$, plus the same for $f$-factors. No Lean. As with almost-linear max-flow, the $o(1)$ hides constants that will keep push-relabel and blossom codes in production.

#### $(1+\varepsilon)$ edit distance in almost-linear time (121)

Exact edit distance cannot be truly subquadratic unless SETH fails (Backurs–Indyk, STOC 2015), and the best near-linear approximation was a constant factor (Andoni–Nosatzki, FOCS 2020). [This paper](https://github.com/openai/math/blob/main/preprints/An-Almost-Linear-Approximation-Scheme-for-Edit-Distance-September-24-2026/paper.pdf) claims any fixed accuracy $1+\varepsilon$ in $N^{1+o(1)}$ expected time, formalized in Lean. It falls back to exact dynamic programming below an input-length threshold the paper itself calls very large, so aligners and diff tools are unaffected; the theoretical question is closed.

#### Trace reconstruction: superpolynomial below, quasipolynomial above (122)

How many randomly-deleted copies of an $n$-bit string do you need to recover it? Bounds stood at $\tilde\Omega(n^{3/2})$ below (Chase, [arXiv 1905.03031](https://arxiv.org/abs/1905.03031)) and $\exp(\tilde O(n^{1/5}))$ above (Chase, 2021). [The lower-bound paper](https://github.com/openai/math/blob/main/preprints/quantitative-lower-bounds-for-trace-reconstruction-September-24-2026/paper.pdf) claims $n^{\Omega(\log\log n)}$ at every fixed deletion rate, and two upper-bound papers claim quasipolynomial samples, one with an efficient decoder. Both jumps are enormous; only the lower bound is in Lean, so the upper bounds deserve the closer look. DNA-storage people should note the worst case is not polynomial.

#### k-median at exactly $1 + 2/e$ (125)

Jain, Mahdian and Saberi showed in 2002 that k-median can't be approximated better than $1 + 2/e \approx 1.736$; the best algorithm reached $2 + \varepsilon$ ([arXiv 2503.10972](https://arxiv.org/abs/2503.10972)). [This paper](https://github.com/openai/math/blob/main/preprints/The-Approximation-Threshold-for-Metric-k-Median-September-24-2026/main.pdf) claims a deterministic $(1 + 2/e + \varepsilon)$-approximation with specified candidate facilities, closing the gap. All four related statements are in Lean. For clustering practitioners, local search and k-means++-style seeding will still win on real data, but the worst-case answer is now known.

#### No small semidefinite program for perfect matching (126)

Rothvoss proved in 2014 that no polynomial-size linear program describes the perfect matching polytope, and asked about semidefinite ones. [This paper](https://github.com/openai/math/blob/main/preprints/Exponential-PSD-rank-of-positively-shifted-matching-matrices-October-5-2026/shifted-matching-psd.pdf) claims every exact SDP lift needs size $2^{\Omega(n)}$. The Lean statement is weaker, superpolynomial ($\gt n^C$ for every $C$), which already answers Rothvoss's question but not at the stated strength.

#### The Gotsman–Linial bound (127)

A degree-$d$ polynomial threshold function on $n$ bits has average sensitivity at most $8d\sqrt n$ ([13 pages](https://github.com/openai/math/blob/main/preprints/Average-Sensitivity-of-Polynomial-Threshold-Functions-September-25-2026/main.pdf), formalized). Kane's 2013 bound had $\sqrt n$ times factors exponential in $d$; the exact-extremizer form of the conjecture is false, and the paper doesn't claim it. Consequences flow into agnostic learning and noise sensitivity of PTFs.

#### Shortest common superstring within factor 2 (128)

A new deterministic 2-approximation ([19 pages](https://github.com/openai/math/blob/main/preprints/A-Polynomial-Time-2-Approximation-for-Shortest-Common-Superstring-September-24-2026/paper.pdf), formalized), using Euler tours in a hierarchical substring graph with forced-occurrence lower bounds. The previous best was about 2.466 (Englert, Matsakis and Veselý, STOC 2022). It does not prove the Greedy conjecture; the paper even cites a 2026 preprint claiming Greedy fails it, which I did not check.

#### The switch chain mixes for every degree sequence (131)

Network scientists sample graphs with a given degree sequence by repeatedly swapping edge endpoints. Kannan, Tetali and Vempala conjectured in 1999 that this mixes in polynomial time for every sequence; it was known for regular and other special classes. [This paper](https://github.com/openai/math/blob/main/preprints/Polynomial-Mixing-of-the-Switch-Chain-for-Every-Graphical-Degree-Sequence-September-25-2026/main.pdf) claims mixing within $2n^8$ steps for every graphical sequence, formalized, and an exact sampler (not formalized). The bound justifies the method; it is far too loose to set your step count.

#### Block sensitivity beats sensitivity squared (132)

After Huang proved the Sensitivity Conjecture in 2019 ([arXiv 1907.00847](https://arxiv.org/abs/1907.00847)), giving $\mathrm{bs}(f) \le s(f)^4$, the natural guess was that the truth is quadratic, as in every known example (Ambainis–Sun's $\tfrac23 s^2$). [This 11-page paper](https://github.com/openai/math/blob/main/preprints/A-superquadratic-separation-between-sensitivity-and-block-sensitivity-September-25-2026/paper.pdf) builds functions with $\mathrm{bs}(f) \ge s(f)^\alpha$ for a fixed $\alpha \gt 2$, by nesting tournament-indexed predicates. The Lean development is about 4,400 lines, the smallest in the group, which makes it the easiest claim here for an outsider to audit.

#### Generalized star height at most three (134)

With complement allowed in regular expressions, nobody has found a language needing two nested stars, and nobody could prove any fixed bound. [This paper](https://github.com/openai/math/blob/main/preprints/Generalized-Star-Height-at-Most-Three-September-25-2026/article.pdf) proves three always suffice, by encoding finite-monoid computations with a prefix code of word pieces (two earlier papers in the family get 13 and 4). Formalized. The famous version, whether height one always suffices, is still open, and the expressions can be astronomically long.

#### PCP for PPAD (136)

Babichenko, Papadimitriou and Rubinstein conjectured that End-of-Line, the canonical PPAD problem, reduces with only quasilinear blow-up to generalized circuits whose approximate solutions survive an adversary corrupting a constant fraction of gates. [This 147-page paper](https://github.com/openai/math/blob/main/preprints/The-PCP-for-PPAD-conjecture-a-quasilinear-reduction-September-25-2026/paper.pdf), the longest in the group and unformalized, claims exactly that. It is the PPAD analogue of the PCP theorem and sharpens conditional lower bounds for approximate Nash equilibria.

#### Gradient queries for log-concave sampling: $d^\varepsilon$ (139)

For a $d$-dimensional density $e^{-V}$ with $I \preceq \nabla^2 V \preceq 2I$, [this paper](https://github.com/openai/math/blob/main/preprints/Subpolynomial-query-complexity-for-well-conditioned-log-concave-sampling-September-26-2026/article.pdf) claims $C_\varepsilon d^\varepsilon$ value-and-gradient queries suffice for every $\varepsilon \gt 0$, and $\Omega(\log d)$ are necessary, so the optimal dimension exponent is zero; previous upper bounds were small powers of $d$. Formalized. Computation between queries is unbounded, so this is information-theoretic: no MCMC sampler you use gets faster.

#### The existential theory of the reals is in the counting hierarchy (141)

Problems like graph drawing, three-player Nash and neural-network training feasibility are $\exists\mathbb R$-complete, known to lie in PSPACE since Canny (1988). [This paper](https://github.com/openai/math/blob/main/preprints/Existential-universal-real-sentences-in-the-counting-hierarchy-October-4-2026/etr-counting-hierarchy.pdf) puts $\exists\mathbb R$ in $\mathsf C_{26}\mathsf P$, a fixed level of the counting hierarchy, and $\exists\forall$ sentences in some fixed level, in the spirit of Allender et al.'s PosSLP result. No Lean. Structural, not algorithmic.

### Notable results, briefly

[One-sample matroid prophet inequalities](https://github.com/openai/math/blob/main/preprints/One-Sample-Suffices-for-Matroid-Prophet-Inequalities-against-an-Almighty-Adversary-September-23-2026/final.pdf) (111) shows one sample per element buys a constant fraction of the prophet's value on any matroid even when the arrival order is chosen by an adversary who sees every sample, value and random bit. The constant is $2^{-310}$, so it's a qualitative statement, and it is formalized.

[Uniform noncommutative identity testing](https://github.com/openai/math/blob/main/preprints/One-Rational-Matrix-Hitting-Point-for-Noncommutative-Formulas-September-24-2026/One-Rational-Matrix-Hitting-Point-for-Noncommutative-Formulas-September-24-2026.pdf) (116) gives one explicit rational matrix tuple of dimension at most $2ns^2$ on which every nonzero size-$s$ noncommutative formula is nonzero, improving Forbes and Shpilka's quasipolynomial hitting sets, plus versions for each positive characteristic and for formulas with inverses. Lean covers the hitting property but not the size and time bounds.

[Weisfeiler–Leman refinement](https://github.com/openai/math/blob/main/preprints/Unconditional-time-lower-bounds-for-Weisfeiler-Leman-equivalence-September-25-2026/paper.pdf) (133) proves that deciding $k$-WL equivalence needs $n^{\Omega(k)}$ time, unconditionally, and is EXPTIME-complete when $k$ is part of the input. The unconditional bound is of the hierarchy-theorem kind, which is why it can exist at all. For graph learning it says higher-order GNNs, whose power is $k$-WL, cannot escape the $n^k$ cost. All four papers are formalized.

[Homogeneous depth-five circuits for iterated matrix multiplication](https://github.com/openai/math/blob/main/preprints/Homogeneous-depth-five-lower-bounds-for-iterated-matrix-multiplication-September-25-2026/Homogeneous-depth-five-lower-bounds-for-iterated-matrix-multiplication-September-25-2026.pdf) (135) pins the complexity at $n^{\Theta(\sqrt n)}$ gates over characteristic zero, matching upper and lower bounds, formalized. It sharpens the picture after Limaye, Srinivasan and Tavenas's 2021 constant-depth lower bounds without moving the general question.

[One-tape time in two-fifths-power space](https://github.com/openai/math/blob/main/preprints/Simulating-One-Tape-Time-in-Two-Fifths-Power-Space-September-25-2026/article.pdf) (137) answers a question from Ryan Williams's 2025 "time in square-root space" paper for one-tape machines: $T^{2/5}$ polylog space suffices, beating the classic $\sqrt T$. Seventeen pages, unformalized, model-specific.

[Memory–sample lower bounds for noiseless Gaussian regression](https://github.com/openai/math/blob/main/preprints/Memory-and-precision-in-noiseless-Gaussian-regression-September-27-2026/paper.pdf) (140) shows a streaming learner with $o(d^2)$ bits needs $\Omega(d\log(1/\varepsilon))$ exact Gaussian measurements to recover a unit vector to angle $\varepsilon$, sharpening Sharan, Sidford and Valiant's $\Omega(d\log\log(1/\varepsilon))$-type bound. Six overlapping papers and 301 pages for one tight result, all formalized, which tells you something about how the model works when a proof goes well: it writes it several more times.

### What I make of the theoretical-CS shelf

Two patterns stood out once everything was in one table. First, the formalized and unformalized halves are different populations. The results with Lean statements tend to be reductions, constructions and algorithms with precise specifications (Unique Games, the colouring reduction, the matching FPRAS, three-machine scheduling, the permanent bound), where a Comparator statement can say exactly what is claimed. The unformalized ones include several of the biggest claims (L = BPL, integer multiplication, almost-linear matching, Subset Sum, PCP for PPAD, deterministic factoring), where the statement involves resource accounting on machine models that Mathlib does not have yet, or analytic number theory that would take a library of its own. My prior for the second group is much lower than for the first, and it should be.

Second, the release's one-line summaries are written to sound as large as possible, and in at least two places that costs accuracy. Family 106 leads with a corollary that Fei, Minzer and Wang published nine days earlier, and the 2-to-1 result (105) is presented without the context that the 4-to-1 version had just appeared. Both papers credit the human work correctly in their history sections, which is where a careful reader would look and a skimmer would not. For the families with formal statements, I'd want to see Comparator run independently and someone compare each challenge statement to the paper line by line before calling any of them settled. For the rest, the usual process applies: experts read them, and we find out.

## Number theory, logic and group theory: 51 families, and where the Lean stops

This slice has 51 families and 88 manuscripts, about 4,700 pages, and the first thing I noticed was not any single theorem. It was the shape of the formalization.

The group-theory and logic results are almost all in Lean: Cannon's conjecture, Thompson's group $F$, the Artin $K(\pi,1)$ conjecture, rigidity of the Turing degrees, Kervaire, Boone–Higman, the Partition Principle. So are the analytic number theory headliners: the $7/8$ zero-free half-plane, the irrationality exponent of $\pi$, Catalan's constant. The arithmetic geometry is not. The full BSD formula in rank at most one, Hilbert's tenth problem over $\mathbb Q$, modularity over imaginary quadratic fields, the $p$-adic section conjecture, Goldfeld's conjecture and Fontaine–Mazur at 2 have no formal proof at all, and they lean on each other: the Hilbert's-tenth paper cites three other unrefereed manuscripts from the same release as inputs.

So within this group there are two very different kinds of claim. One kind comes with a machine-checkable statement that I could read and judge for faithfulness; for those, the remaining question is whether the Lean actually compiles against the comparator, which nobody outside OpenAI has reported yet (the catalogue says `review: unchecked`, and I did not build it). The other kind is several hundred pages of modern arithmetic geometry that only specialists can referee, stacked three deep.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig1.png" alt="Page 2 of the release's overview PDF: the Number theory heading followed by one-paragraph summaries of families 001 to 010, from Milne's rationality conjecture to Fontaine-Mazur modularity at the prime 2, each with links to its manuscripts." caption="The release's own catalogue entry for the first ten number-theory families; every paragraph is one claimed resolution (openai/math overview.pdf, page 2)." />

### Cannon's conjecture (family 246)

A hyperbolic group has a boundary at infinity. For the fundamental group of a closed hyperbolic 3-manifold that boundary is the 2-sphere. Cannon conjectured in the early 1990s that the converse holds: a hyperbolic group whose boundary is $S^2$ acts geometrically on hyperbolic 3-space. By Perelman's geometrization this is the statement that such groups are, up to finite kernels, Kleinian, and it is the central open problem of the 3-dimensional part of geometric group theory.

The standard route goes through analysis on the boundary. Bonk and Kleiner proved the conjecture when the boundary's Ahlfors-regular conformal dimension is attained; Cannon, Floyd and Parry reduced it to a bound on combinatorial moduli of annuli; Markovic and Haïssinsky had algebraic criteria; Groves, Haïssinsky, Manning, Osajda, Sisto and Walsh reduced the residually finite case to a relative version. The 31-page paper takes the analytic route head on. By Bourdon–Kleiner, it suffices to show that the combinatorial 2-modulus of curve families of a fixed small diameter is bounded at every scale. Assuming it is not, it splits into an exponential-records case and a subexponential case, builds limit functions with separating level continua (the figure below is that step), and derives a contradiction from the action of the group. Sullivan–Tukia straightening then turns the conformal structure into an action on $\mathbb H^3$.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig5.png" alt="A rectangle with a top band where H_N = 0 and a bottom band where H_N = 1, a jagged level continuum C_{N,t} across the middle, and a path eta straddling it, extended by dashed curves into a top-to-bottom crossing." caption="A step in the modulus argument for Cannon's conjecture: any path that straddles a level set must meet it, which is how crossing estimates control the boundary's conformal structure (A Modulus Proof of Cannon's Conjecture, Figure 1)." />

I read the Lean statement carefully because it has to define everything from scratch. Hyperbolicity is uniformly thin geodesic triangles in a Cayley graph for a finite symmetric generating set. The boundary is based geodesic rays modulo bounded synchronous distance, with the quotient of the pointwise topology, which is a standard model of the Gromov boundary. $\mathbb H^3$ is the upper half-space with the explicit $\operatorname{arcosh}$ distance, and the isometry group includes orientation-reversing maps. Proper, cocompact and finite kernel are spelled out. That is a faithful statement:

```lean
-- lean/ComparatorChallenges/CannonGeometricAction.lean:89
theorem cannon (G : Type u) [Group G] [TopologicalSpace G] [DiscreteTopology G]
    (D : CayleyData G) (hthin : D.ThinTriangles)
    (hsphere : Nonempty (Boundary D ≃ₜ SphereTwo)) :
    ∃ ρ : G →* H3Isom, ProperAction ρ ∧ CocompactAction ρ ∧
      (ρ.ker : Set G).Finite := by
  sorry
```

If the formal proof compiles, Cannon's conjecture is settled, in 31 pages.

### Thompson's group F is not amenable (family 248)

Thompson's group $F$ is the group of piecewise-linear homeomorphisms of $[0,1]$ with dyadic breakpoints and slopes that are powers of 2. Whether it is amenable has been open since Geoghegan asked in 1979. It is the natural test case: $F$ has no free subgroups (Brin–Squier), so the usual obstruction to amenability is missing, yet it is not elementary amenable either. Claimed proofs have gone both ways and been withdrawn, and the paper lists them honestly: Shavgulidze's amenability claim, Moore's withdrawn amenability claim, Akhmedov's nonamenability claim.

This one is 13 pages and almost elementary. Fix a Lipschitz map $f$ on the unit ball of a Hilbert space that moves every point by at least $\delta$; Benyamini and Sternfeld showed such maps exist in infinite dimensions, and an appendix builds one on $L^2$ with $\delta = 1/2$. Colour every dyadic partition recursively by applying $f$ to the mean of its parents' colours. $F$ can carry any separated pair of dyadic intervals onto any other, so if Følner sets existed, all averaged correlations between colours of separated intervals would be nearly equal. Expanding $\|z_i - m\|^2$, the common value cancels and only the nested pairs remain, a $1/D$ fraction of the terms, as the paper's one figure shows. That makes the mean colour an approximate fixed point of $f$, contradicting $\delta$.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig4.png" alt="A D-by-D grid of parent intervals I_1 to I_D against children I_i·I_j. One column, I_i, is shaded orange and labelled nested pairs, D terms each at least -1; the rest are separated pairs, D(D-1) terms each with mean in [alpha - eta, alpha + eta]." caption="The single counting step the nonamenability proof turns on: only one column in D pairs a child with its own parent, so those terms are a 1/D fraction (Thompson's group F is nonamenable, Figure 1)." />

A proof this short for a problem this famous makes me want to find what it uses about $F$ that an amenable group of interval homeomorphisms lacks; the transitivity on separated pairs of dyadic intervals is the obvious candidate. But the formal statement is faithful (it defines $F$ as dyadic PL homeomorphisms with slopes $2^k$, insists the group law is composition, and rules out a positive normalized left-invariant mean on $\ell^\infty(F)$), and the solution is about 9,500 lines. Corollaries in the paper: non-unitarizable representations of $F$ (via family 251) and a percolation result from another family.

```lean
-- lean/ComparatorChallenges/ThompsonNonamenability.lean:78
theorem thompson_F_nonamenable_composition :
    ∃ group : Group F, letI := group;
      (∀ (h g : F) (x : UnitInterval),
        (h * g).val.toHomeomorph x = h.val.toHomeomorph (g.val.toHomeomorph x)) ∧
      ¬ Nonempty (InvariantMean F) := by
  sorry
```

### Artin groups: K(π,1), parabolic intersections, and a group that is not CAT(0) (family 254)

Artin groups generalize braid groups: one generator per node of a Coxeter diagram and a braid relation of the prescribed length between each pair. The $K(\pi,1)$ conjecture, going back to Arnold, Brieskorn, Pham and Thom, says a natural finite complex built from the diagram, the Salvetti complex, is already aspherical. It would give every Artin group a finite classifying space, torsion-freeness and computable cohomology. Deligne proved it for spherical type in 1972, Charney and Davis for FC type and dimension two, and Paolini and Salvetti for affine type in 2021. The general case was open.

The 62-page paper claims all finite-rank Artin groups, any labels including $\infty$. The proof filters a poset model of the universal cover by "harmonic heights" and shows each layer can be removed without creating homotopy, using a two-point estimate (a sum of inequalities equals minus a positive-definite Cartan form) and an isolated-layer obstruction. The paper's dependency map is a good summary:

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig6.png" alt="Flowchart of the Artin K(pi,1) proof: framed category and twists, layer calculus and caps feed a two-point estimate and an isolated-layer obstruction, which give spherical-residue sublevels, then contractibility of the poset realization, which with the Salvetti comparison lemma proves X(W,S) is a K(A,1)." caption="The paper's own dependency map for the K(π,1) proof; the left column is the classical comparison with the Salvetti complex, the right column is new (Harmonic heights and the Artin K(π,1) conjecture, Figure 1)." />

Two companions claim the Parabolic Intersection Conjecture (any intersection of parabolic subgroups is parabolic) and an explicit Artin group on 116 generators, with labels in $\lbrace 2, 3, \infty \rbrace$, that admits no geometric action on any proper CAT(0) space. All three are in Lean. One thing to know about the $K(\pi,1)$ target: `SalvettiCover` is defined as the realization of the nerve of a poset of "lifted cells" (an Artin group element paired with a spherical subset, ordered by a face relation built from reduced words). That poset is the standard combinatorial description of the universal cover's cells, so the statement is the conjecture, but the identification is part of what you trust when you read the statement, not something the Lean proves.

```lean
-- lean/ComparatorChallenges/HarmonicArtin.lean:59
theorem salvetti_cover_contractible [Finite S] (M : CoxeterMatrix S) :
    ContractibleSpace (SalvettiCover M) := by
  sorry
```

### Rigidity of the Turing degrees (family 241)

Say $A \le_T B$ when an oracle for $B$ can compute $A$. The question is whether this ordering of all degrees has any symmetry besides the identity. Slaman and Woodin showed that every automorphism fixes everything above $0''$, that there are only countably many automorphisms, and that each is induced by an arithmetic function on reals; they also showed rigidity is equivalent to their biinterpretability conjecture. The rest stayed open.

The paper is ten pages. It takes the Slaman–Woodin representation theorem as an input, cited from their unpublished 2005 manuscript, and finishes with a Baire-category trick: from four values of such a representing function at rational affine combinations of a real $t$ and two generic reals, you can compute $t$, which forces every degree to be fixed. That leans heavily on an input almost nobody has seen in print, which would worry me if it were not for the Lean: the statement uses Mathlib's `TuringReducible` and asks that every order isomorphism of the degrees be the identity, and the solution directory (about 84,000 lines) has subdirectories called `Representation`, `CohenForcing`, `Constructibility` and `SetModels`. That suggests the representation theorem itself was formalized rather than assumed. I did not audit it.

### Kervaire's conjecture in thirteen pages (family 256)

Can you kill a nontrivial group by adding one generator and one relation? Kervaire conjectured in the 1960s that you cannot. Gerstenhaber and Rothaus proved it for finite groups by topology of unitary groups; Klyachko for torsion-free groups in 1993. Kawauchi has published a proposed resolution through knot theory, which the paper mentions and does not use.

The claim proves the stronger unimodular statement: for any group $A$ and any word $w$ in $A * \langle t \rangle$ with exponent sum $\pm 1$ in $t$, $A$ injects into the quotient. Kervaire follows in two lines. The argument uses spectral phase in unitary groups and a theorem that a certain planar surface cannot have exactly one nontrivial boundary. The Lean target is the unimodular injectivity, written with Mathlib's coproduct and normal closure. The family's headline is broader, Howie's conjecture on nonsingular systems of equations over arbitrary groups, and that 16-page paper is not formalized.

### Two hyperbolic groups nobody expected (families 252 and 257)

Gromov asked whether every hyperbolic group is residually finite (every element survives in some finite quotient) and whether every hyperbolic group acts geometrically on a CAT(0) space. Most hyperbolic groups people use are cubulated, and cubulated hyperbolic groups are residually finite by Agol and Wise, so both questions were believed to have positive answers by many people and remained open.

Family 252 claims a torsion-free hyperbolic group that is not residually finite, the fundamental group of a finite Euclidean triangle complex, in 21 pages. The proof is existential: one member of a fixed finite family of elements lies in every finite-index normal subgroup, without saying which. By Malcev's theorem such an element dies in every finite-dimensional linear representation over any field, so the group is not linear either. A consequence the summary does not mention: by Kapovich–Wise, a non-residually-finite hyperbolic group implies some hyperbolic group has no proper finite-index subgroups at all.

Family 257 claims a hyperbolic group with a finite 2-dimensional classifying space and no geometric action on any proper CAT(0) space, in any dimension. The two results point the same way (a non-CAT(0), non-cubulated world inside hyperbolic groups), which is at least consistent. Both main statements are in Lean; for 257 the two-dimensionality is not.

### Eilenberg–Ganea fails (family 249)

Cohomological dimension (the length of a projective resolution) and geometric dimension (the smallest aspherical complex) agree for every group except possibly in dimension 2, by Eilenberg–Ganea in 1957 and Stallings–Swan. Bestvina and Brady showed in 1997 that, for certain kernels of right-angled Artin groups built from a spine of the Poincaré homology sphere, either Eilenberg–Ganea or Whitehead's asphericity conjecture must fail, without knowing which.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/number-theory-logic-groups-fig7.png" alt="Two diagrams: a commuting square and a commuting three-cube with vertices labelled by heights, and in blue the diagonal edge and triangle where the height equals zero, forming the level set X." caption="How the Eilenberg-Ganea counterexample is built: the group is the height-zero level set of a right-angled Artin group's cube complex, sliced cube by cube (A finitely generated counterexample to the Eilenberg-Ganea conjecture, Figure 1)." />

The 25-page paper takes exactly such a kernel, for the presentation $\langle x, y \mid x^2 = y^5,\ x^2 = (xy^{-1})^3 \rangle$, and shows it has cohomological dimension 2 and no 2-dimensional classifying space at all, so geometric dimension 3. It is finitely generated and residually finite but not finitely presented, so the finitely presented version is untouched, and Whitehead's conjecture is not decided. Formalized.

### A finitely presented infinite periodic group (family 247)

Burnside asked in 1902 whether a finitely generated group in which every element has finite order must be finite. Golod said no in 1964, but every known counterexample needs infinitely many relations. The finitely presented version was open.

The claim builds a unital $\mathbb F_2$-algebra $R$ whose Steinberg group $\mathrm{St}_{12}(R)$ is infinite, finitely presented and periodic, and also has property (T). A companion shows a finite-index subgroup is residually finite with every element of 2-power order, and the construction yields a finitely presented nil algebra that is not nilpotent, answering the finitely presented Kurosh question. The Lean covers the periodic group. The family title leads with the residually finite 2-group, which is the unformalized companion.

### Boone–Higman (family 250)

A finitely generated group has solvable word problem if and only if it embeds in a finitely presented simple group. Boone and Higman proved the "if" direction in 1974 and conjectured the converse; Belk, Bleak, Matucci and Zaremsky proved it for hyperbolic groups in 2023. The claim proves it in general, strengthens the target to a simple group of type $F_\infty$, and builds one $F_\infty$ group containing every finitely presented group. All three are formalized.

### Dixmier's unitarizability problem (family 251)

Amenable groups have every uniformly bounded Hilbert-space representation similar to a unitary one. Dixmier asked in 1950 whether that characterizes amenability. Groups containing a free subgroup fail it; nonamenable groups without free subgroups were the hard case (Monod and Ozawa handled free Burnside groups). The 19-page paper claims the equivalence for every discrete group, with non-unitarizable witnesses of norm at most $1+\varepsilon$, and it is in Lean. Combined with family 248 it gives non-unitarizable representations of Thompson's $F$. A companion, not formalized, proves that strong Ulam stability characterizes amenability for countable groups, answering Burger, Ozawa and Thom.

### The Partition Principle does not imply Choice (family 244)

If every surjection $X \to Y$ has a reverse injection $Y \to X$, must the axiom of choice hold? The question goes back to Beppo Levi in 1902 and is often called the oldest open problem in set theory; the paper notes that several recent attempts were withdrawn. The claim builds a symmetric forcing model of ZF plus the Partition Principle plus choice for well-orderable families in which choice fails. The relative-consistency statement $\mathrm{Con}(\mathrm{ZF}) \Rightarrow \mathrm{Con}(\mathrm{ZF} + \mathrm{PP} + \mathrm{AC}_{\mathrm{WO}} + \neg\mathrm{AC})$ is in Lean, which means Lean had to formalize ZF syntax and symmetric extensions; I would want a set theorist to read those definitions. The second formal target assumes a ground model with an internal inaccessible cardinal, which the paper's transitive-model theorem does not.

### Shelah's eventual categoricity, in ZFC (family 240)

Morley proved that a countable first-order theory categorical in one uncountable cardinal is categorical in all of them. Shelah conjectured the analogue for abstract elementary classes, where compactness fails: categoricity in one large enough cardinal implies categoricity in all larger ones. Known results needed amalgamation, tameness or large cardinals (Shelah–Vasey, Vasey, Boney). The 121-page paper claims it in ZFC with a uniform threshold depending only on the Löwenheim–Skolem number.

Here the formalization is misleading if you only look at the family's Lean badge. What is formalized is the companion: under CH, a specific proposed threshold, the Hanf number $\beth_{\omega_2}$, does not suffice. The ZFC theorem itself is not in Lean. AEC theory has a history of long correction cycles, so this one needs human referees.

### Catalan's constant is irrational (family 005)

$G = 1 - 1/9 + 1/25 - 1/49 + \cdots$ is $L(2, \chi_{-4})$, and nobody knew it was irrational. Rivoal and Zudilin showed one of $\beta(2), \beta(4), \ldots, \beta(14)$ is irrational, later cut to $\beta(2)$ through $\beta(10)$; Calegari, Dimitrov and Tang did the conductor-3 analogue. Approximations to $G$ existed, but their denominators grew too fast.

The 44-page proof builds determinants of size $48N$ whose entries are rational combinations of $1$, $G$ and $\zeta(2)$, with polynomial rows chosen to have high Taylor contact so the $\zeta(2)$ term cancels in every entry. If $G$ were rational the determinants would be rational; prime-by-prime denominator control gives a lower bound and an integral formula an upper bound, and they cross. The Lean statement is the bare irrationality of the explicit series. Of every landmark here, this is the one whose statement is easiest to check and whose formal proof (the `OAI/NumberTheory/Catalan` directory is large) leaves least room for interpretation.

### Modularity over imaginary quadratic fields (family 030)

Every elliptic curve over every imaginary quadratic field is modular: it matches an automorphic representation of $\mathrm{GL}_2$ over that field, at every place. Over $\mathbb Q$ this is Wiles, Taylor–Wiles and Breuil–Conrad–Diamond–Taylor. Over imaginary quadratic fields the automorphic side lives on hyperbolic 3-space and has torsion in several degrees, so the method only recently worked: the ten-author potential automorphy paper (2023), Allen–Khare–Thorne for a positive proportion, and Caraiani–Newton for all curves over $\mathbb Q(i)$ and a few other small fields.

The claim is 30 pages. It is a prime-switching argument in the spirit of Wiles's 3–5 switch: Jacobians of cyclic covers whose cohomology links $E$ through congruences mod $p$, mod 3, mod $p$ to a curve already known to be modular, with parameters chosen so each step satisfies the existing lifting theorems. Brevity is consistent with a clever reduction. It is also exactly the kind of argument where one unmet hypothesis (residual image, local conditions at 2 or 3) breaks the chain. No Lean.

### Artin's primitive root conjecture, infinitude for every base (family 029)

Artin conjectured in 1927 that any integer other than $-1$ and the squares is a primitive root modulo infinitely many primes. Hooley proved it on GRH; Heath-Brown showed at most two primes fail, so one of 2, 3, 5 works, but not which. The claim proves infinitude for every admissible base with at least $c_a x/(\log x)^2$ such primes in $(x, 2x)$, short of Artin's predicted $x/\log x$ density. The analytic input is itself huge: Hecke $L$-functions of every cyclotomic field containing the 12th roots of unity are zero-free in $\Re s > 1 - 10^{-6}$, uniformly, from the same cubic-theta machinery as family 003. A companion on simultaneous primitive roots is explicitly conditional on four inputs. No Lean.

### The local p-adic section conjecture (family 019)

Grothendieck conjectured in 1983 that for hyperbolic curves over arithmetic fields, rational points are exactly the splittings of the arithmetic fundamental group. Over $p$-adic fields, Koenigsmann proved the birational version and Pop–Stix showed every section is localized at some valuation. The claim proves the full local statement for every curve of genus at least two over every finite extension of $\mathbb Q_p$, in 24 pages plus a 7-page cover construction: build a finite étale cover with a prescribed exterior sheet and degree divisible by $p$ over small disks, lift the section to a birational one, and apply Koenigsmann. The global corollary is honest about its scope: it covers curves whose finite-descent locus equals their rational points, which includes $X_0(N)$ and $X_1(N)$, not the global conjecture in general. Short for its fame, unformalized.

### Two-point Chowla with plain averages (family 007)

Does the parity of the number of prime factors of $n$ predict that of $n+h$? Chowla conjectured not. Tao proved the two-point case with logarithmic averaging in 2016, which smooths over scales; with ordinary averages it was open. The claim gets $O(X/(\log X)^c)$ for every fixed pair of nonproportional affine forms, plus the corrected binary Elliott conjecture for bounded multiplicative functions that are uniformly nonpretentious. Formalized. The model's reasoning summary for this family is also released.

### Deligne–Drinfeld (family 008)

The Grothendieck–Teichmüller Lie algebra $\mathfrak{grt}_1$ encodes the symmetries of braided associativity. Deligne and Drinfeld conjectured it is free on one generator in each odd weight $3, 5, 7, \ldots$. Brown's 2012 work on mixed Tate motives gives the generators and Willwacher proved freeness; computations confirmed equality through weight 29. The 41-page claim supplies the missing all-weight dimension bound, so there are no extra solutions. In Lean, which for a statement defined by three polynomial identities in free Lie algebras is a meaningful check.

### Zilber–Pink for abelian varieties and curves in A₂ (family 016)

"Unlikely intersections": a subvariety should meet special subvarieties of too-small codimension only finitely often. The claims are the abelian case over $\overline{\mathbb Q}$ in all dimensions (via a non-density theorem and the Barroero–Dill reduction) and the full curve case in the Siegel threefold $\mathcal A_2$ for Hodge-generic curves, in three components. Prior work had curves in abelian varieties (Habegger–Pila) and $\mathcal A_2$ curves under boundary hypotheses (Daw–Orr). 170 pages, non-effective, no Lean.

### Restricted geometric Langlands in characteristic p (family 014)

Eight papers, 348 pages, the largest family in this group. The headline is the $\overline{\mathbb Q}_\ell$-linear restricted geometric Langlands equivalence for curves in characteristic $p$, under explicit conditions on $p$, which proves a full-support conjecture of Gaitsgory and Raskin in those regimes. Consequences include Ramanujan–Arthur decompositions of cusp forms over function fields (conjectures of Gaitsgory–Lafforgue–Raskin) and temperedness at every place for globally generic cusp forms of split exceptional groups. One paper (global Arthur enhancements) is explicitly conditional on a decomposition statement. No Lean, and I am not able to judge it beyond its statements.

### Margulis–Platonov over global fields (family 018)

Normal subgroups of the rational points of a simply connected absolutely almost simple group should come only from the finitely many places where the group is compact. Most types were known; the claim finishes the remaining anisotropic outer type A, triality $D_4$ and $E_6$ cases over number fields and does all global function fields, including characteristic 2. 159 pages, uses finite simple group theory, no Lean.

### Ostmann's inverse Goldbach problem (family 013)

Could the primes, up to finitely many changes, be a sumset $A + B$ with both sets having at least two elements? Ostmann conjectured not in 1956. Sieve methods forced any such $A$ and $B$ to have about $\sqrt x$ elements below $x$, which looked consistent; Green and Harper showed an inverse large sieve conjecture would suffice. The 80-page claim supplies that kind of inverse structure through quadratic characters with moving centres and a finite-field tree comparison. The full statement is in Lean.

### The Gaussian moat (family 028)

Can you walk to infinity on Gaussian primes with steps of bounded length? Gordon asked at the 1962 ICM. The claim says no, more strongly: for each step bound $D$ every connected component has at most $B_D$ primes. The interesting part is how little arithmetic it uses. It shows that just deleting the multiples of finitely many split Gaussian primes already traps every bounded-step walk, by an entropy argument on residues along a long self-avoiding walk. That explains the 29 pages. Formalized, ineffective.

### Squarefree values of quartics (family 020)

Is $n^4 + 2$ squarefree for a positive proportion of $n$? Erdős raised it in 1953. Hooley handled exponent $d - 1$ and Browning's determinant method reached $k = d - 2$ only for degree at least 9. The claim proves $k = d - 2$ for degrees 4 to 8 with the predicted Euler-product density, so with Browning every degree. The Lean statement covers all $d \ge 4$.

### The prime factors of p − 1 (family 011)

For random integers the sizes of prime factors follow the Poisson–Dirichlet law. The claim is that the same holds for $p - 1$ as $p$ ranges over primes (Ford–Konyagin–Luca), together with Erdős's conjecture that some $n$ have more than $n^{1-\varepsilon}$ totient preimages and $x^{1-o(1)}$ primes with $x^\delta$-smooth predecessors for every $\delta$. Prior smooth-shifted-prime results needed thresholds around $x^{0.28}$. This would be a large step in multiplicative number theory, unformalized across 187 pages.

### Duffin–Schaeffer with a shift (family 022)

Koukoulopoulos and Maynard proved the Duffin–Schaeffer conjecture in 2019. The claim proves the weak inhomogeneous version: for any fixed shift $\gamma$, divergence of $\sum \varphi(q)\psi(q)/q$ gives $\|qx - \gamma\| < \psi(q)$ infinitely often for almost every $x$, with numerators unrestricted. Ramírez had shown the unweighted version fails. 87 pages, no Lean.

### Choiceless polynomial time does not capture P (family 243)

Is there a logic that expresses exactly the polynomial-time properties of unordered structures? Choiceless polynomial time with counting was the leading candidate, and Blass, Gurevich and Shelah conjectured it falls short. The claim gives an explicit polynomial-time query, solvability of a linear system over $\mathbb F_3$ on grid-like structures, that CPT cannot define, even with sets of unbounded rank, plus a separation showing one witnessed symmetric choice adds power. Both in Lean, with the caveat that CPT has to be encoded inside Lean.

### Single-fold Diophantine representations (family 242)

Matiyasevich asked in 1974 whether every computably enumerable set has a polynomial representation with exactly one witness per member. The 27-page claim says yes, using multiples of a point on a rank-one elliptic curve (whose height grows quadratically in the index with bounded error) to replace Pell-equation exponentiation. So Hilbert's tenth problem stays undecidable even under a promise of at most one solution. Formalized.

### Weak normalization implies strong normalization (family 245)

For every pure type system, if every legal term has some terminating reduction, then every reduction terminates: the Barendregt–Geuvers–Klop conjecture, with arbitrary sorts, non-functional rules and open contexts. This matters to anyone building proof assistants. 72 pages, formalized.

### A finitely presented simple amenable group (family 253)

Juschenko and Monod gave infinite finitely generated simple amenable groups in 2013, but those topological full groups can never be finitely presented. The claim finds an infinite finitely presented simple amenable group inside polygon exchange transformations. Existence only, formalized.

### Quasi-isometric rigidity of polycyclic groups (family 255)

Gromov showed groups that look like nilpotent groups at large scale are virtually nilpotent. Eskin, Fisher and Whyte conjectured the solvable analogue and proved it for Sol and lamplighters. The claim proves it for all virtually polycyclic groups by controlling the height coordinate of self quasi-isometries of solvable Lie models. 73 pages, with one of the largest Lean developments in the group (about 262,000 lines).

### Gersten's conjecture for one-relator groups (family 258)

A one-relator group with no Baumslag–Solitar subgroup is hyperbolic; a companion shows every hyperbolic one-relator group is virtually compact special, so with Kielak–Linton they are virtually free-by-cyclic. This builds on Louder–Wilton and Linton's hierarchy work. 198 pages, the two papers cite each other, no Lean.

### A group without fixed price (family 259)

Gaboriau asked whether every essentially free measure-preserving action of a group has the same cost. The claim exhibits a finitely generated amalgam whose Bernoulli action costs at least $1 + \eta$ while a sequence of height extensions costs as little as $1 + 99/M$. The constants are explicit and $\eta$ is tiny. 27 pages, no Lean.

### The rest of the number theory

The notable results are each a real problem with a name attached, just smaller ones.

Milne's rationality conjecture (001) says specialized Hodge classes on abelian varieties with good reduction pair rationally with divisors on the reduction, in every realization. The pairing theorem stands alone; the stronger "represented by one algebraic cycle" corollary uses the release's own Hodge-conjecture-for-CM-abelian-varieties claim, which has no Lean.

Function-field reconstruction (009) recovers a function field of dimension at least two over an algebraically closed field from its mod-$\ell$ Milnor K-groups in degrees one and two, and proves Bogomolov–Pop reconstruction from the abelian-by-central Galois quotient. Only the injectivity half is in Lean.

The joint Dickman law (012) says the largest prime factors of $n$ and $n+1$ are independent in natural density, so $P^+(n) < P^+(n+1)$ half the time; Teräväinen had it in logarithmic density, and the best lower density for the ordering was 0.280. Formalized.

Torus-packet equidistribution (015) is the higher-degree Duke theorem for totally real fields of prime degree at least five, primitive quartics and primitive sextics; only the prime-degree case is in Lean.

Jacobsthal's function (021) gets the quadratic bound $h(k) \ll k^2/(\log\log k)^2$, answering Jacobsthal's question; Iwaniec had $k^2 \log^2 k$, and the true order is still unknown.

Patterson's bias for cubic Gauss sums (023), the $X^{5/6}/\log X$ main term, was proved by Dunn and Radziwiłł on GRH; this does it unconditionally with the same cubic-theta toolkit as family 003. Formalized.

The number of totients (024) gets an asymptotic up to a bounded oscillating factor, building on Ford's 1998 order of magnitude, which answers Erdős and Hall's question that $V(cx)/V(x) \to c$. Formalized.

Short Egyptian fractions (025): every $a/b$ is a sum of $O(\log\log b)$ distinct unit fractions, Erdős Problem 304; Vose's $\sqrt{\log b}$ had stood since 1985. Formalized, with explicit constants.

Large prime gaps (026): for every $C$, a positive proportion of gaps exceed $C \log p_n$, which answers Erdős and Prachar's question about $p_n/n$. Only the corollary is in Lean, and the proof is 19 pages.

Integral points on character varieties (027) answers Litt's question for $\mathrm{SL}_r$ character varieties of curves: integral points are Zariski dense after one field extension.

Uchida's conjecture (031): every continuous open homomorphism between Galois groups of solvably closed extensions of number fields comes from a unique field embedding, via Hoshi's cyclotomic criterion.

### What I would check first

If I were refereeing this group with limited time, I would spend it in this order. First, build the Lean for 003, 017 and 248 against their comparator files: the statements are faithful and use Mathlib's own definitions, so a clean build is close to decisive. Second, have a set theorist read the Lean encodings behind 244 and 243, where the statement depends on how ZF and CPT were written down. Third, get number theorists onto the 006 → 002 → 004 chain, because the most famous unformalized claim in the group rests on it. The 240 eventual-categoricity paper needs the same treatment for the opposite reason: its badge says Lean, but the ZFC theorem is not what was formalized.

How I checked: I extracted the text of all 88 PDFs with `pdftotext`, read each family's summary and abstracts and the introduction and main theorem of every principal paper, and grepped bibliographies for citations of other release manuscripts. For Lean status I read `lean/docs/<id>.md`, the `permitted_axioms` in each `ComparatorChallenges/*.json` (only `propext`, `Quot.sound` and `Classical.choice` for every family here), and the target `.lean` statements for the families discussed above; `grep` found no `sorry` in `lean/OAI`. Line counts are of the solution directory named in each comparator config and can miss shared imports. I compiled nothing. Prior-work citations are from the papers' own literature reviews and my own knowledge of the fields; where I could not confirm a citation I left it out.

## Algebraic geometry and algebra

Algebraic geometry is where a confident wrong proof hides best, so I read this group with the most suspicion. The catalogue files 36 families under algebraic and complex geometry and 18 under algebra: 54 families, 118 manuscripts and about 5,700 pages.

The headline names are the kind you see on a graduate syllabus under "open": the Hodge conjecture for CM abelian varieties, abundance, Iitaka's conjecture, Nagata, Fujita, Bloch, Serre's positivity, Lech, Kaplansky's zero-divisor and direct-finiteness conjectures, the Nakayama and finitistic-dimension conjectures, Alperin's weight conjecture. Some are claimed proved and some disproved. If even a third of them survive refereeing, this one slice of the release is a decade of a field.

So the useful question is how much of each claim anything other than the model has checked. The answer splits the group cleanly. Ten families have a Lean statement that matches the paper's headline: Nagata, Zariski cancellation over $\mathbb C$, Abhyankar–Sathaye, Griffiths' positivity, the Kähler splitting theorem, Kaplansky's zero divisors, the finitistic dimension, Auslander–Reiten with Tachikawa, Saxl and the Bass trace conjecture. Six more have Lean for a piece that is not the headline. The other 38, including Hodge, abundance, Serre and Alperin, are paper-only. The Lean side is mostly the counterexamples, and the paper-only side is mostly the deep positive theorems. That makes sense: an explicit polynomial or group is something a proof assistant can hold, and a 781-page induction through the minimal model program is not.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig1.png"
  alt="A page of the OpenAI research catalogue headed 'Algebraic and complex geometry', with catalogue paragraphs for family 032 (rational Hodge conjecture for CM abelian varieties and products of K3 surfaces), 033 (Campana's orbifold Iitaka conjecture) and 034 (log abundance and effective Iitaka fibrations), each followed by links to its manuscripts."
  caption="How the catalogue opens the algebraic geometry section: Hodge, Iitaka and abundance, back to back (openai/math overview.pdf, page 5)."
/>

I graded what each family claims, not whether it is right. That gives 33 landmark, 17 major and 4 notable; nothing in the group read as merely technical. By kind there are 35 proofs, 16 counterexamples and 3 partial results. A landmark here means a famous named conjecture claimed in full, which is why the count looks absurd. The Hodge, Nagata and Kaplansky results are covered above with the headliners; the order below starts with the claims a machine has checked, then the rest by weight.

### The Abhyankar–Sathaye conjecture fails in four variables (family 049)

This is the result I would show a skeptic first. Abhyankar, Moh and Suzuki proved in the 1970s that any embedding of the line into the plane is straight: if $\mathbb C[x,y]/(f)\cong\mathbb C[t]$, then $f$ is a coordinate, meaning some polynomial automorphism sends $f$ to $x$. The Abhyankar–Sathaye conjecture asks the same in every dimension. It has been open for $n\ge 3$ for fifty years, with affirmative results only for special shapes of $f$.

Family 049 gives an explicit counterexample in $\mathbb C[h,u,v,w]$. Start from the cusp $(u^3,-u^2)$ and perturb it:

$$
x = u^3 + hv,\qquad y = -u^2 + hw,\qquad x^2 + y^3 = h\,s,
$$

where $s = 2u^3v + 3u^4w + h(v^2-3u^2w^2) + h^2w^3$. Put $p = -2s^2x + 3sy^2 - 3s^3y$ and $F = h - p - 1$. The paper proves $\mathbb C[h,u,v,w]/(F)\cong\mathbb C^{[3]}$, so the zero set of $F$ is a copy of affine 3-space. That is the hard half, done by a lifting lemma over an arbitrary base ring. The other half takes one line. A coordinate is the first entry of an automorphism, whose Jacobian determinant is a nonzero constant, so a coordinate's gradient never vanishes. But $\nabla F = 0$ at $(2,0,-\tfrac12,\tfrac12)$, a point on the fibre $F=-1$.

I expanded $F$ in exact rational arithmetic to check that half myself. It has 69 terms and degree 17, the cusp identity holds, and at $(2,0,-\tfrac12,\tfrac12)$ I get $F=-1$ with all four partial derivatives exactly $0$. The widget does the same computation in your browser with BigInt rationals.

<AbhyankarSathayeCriticalPoint />

The hard half is in Lean. The Comparator statement is short enough to read in one go, and it says precisely what the conjecture denies:

```lean
-- lean/OAI/AlgebraicGeometry/AbhyankarSathaye/Counterexample.lean:35-41
theorem exists_noncoordinate_polynomial (n : ℕ) (hn : 4 ≤ n) :
    ∃ F : MvPolynomial (Fin n) ℂ,
      Nonempty ((MvPolynomial (Fin n) ℂ ⧸ Ideal.span {F}) ≃ₐ[ℂ]
        MvPolynomial (Fin (n - 1)) ℂ) ∧
      ¬ ∃ (equiv : MvPolynomial (Fin n) ℂ ≃ₐ[ℂ] MvPolynomial (Fin n) ℂ)
        (index : Fin n), equiv (MvPolynomial.X index) = F :=
  ⟨extendedF hn, counterexample.2.2.2.2 n hn⟩
```

The paper is six pages long. Its own remark is that Shpilrain and Yu had asked whether a hypersurface isomorphic to a coordinate hyperplane could have a critical point on another fibre, and this answers them too. A second manuscript in the family gives a different polynomial, $f = x_1 - 2Q\bigl(Q(x_2+x_4)+x_1x_4\bigr)$ with $Q = x_2^2-x_4^2+x_1x_3$. That one becomes a coordinate after adjoining one variable but is not one, the first counterexample to the stable coordinate conjecture. That polynomial is not in Lean. The original question in three variables, a plane in 3-space, is untouched.

### Zariski cancellation fails over $\mathbb C$ (family 047)

Cancellation asks whether $X\times\mathbb A^1\cong\mathbb A^{n+1}$ forces $X\cong\mathbb A^n$. It is true for curves and, by Fujita, Miyanishi and Sugie around 1980, for complex surfaces. Neena Gupta disproved it in positive characteristic in 2014, using Asanuma's threefold, and then in every dimension at least 3. Over $\mathbb C$ in dimension 3 and up it stayed open; the paper cites a July 2026 preprint that still lists it as open.

Family 047 writes down one polynomial in five variables. With $x = s^2+u^3+p^2F$,

$$
H = x^2F - (1+2sx)J - p^2J^2 - pu,\qquad A = \mathbb C[p,s,u,F,J]/(H),
$$

and claims that $A$ is a 4-dimensional domain with $A[w]\cong\mathbb C^{[5]}$ but $A\not\cong\mathbb C^{[4]}$. The Comparator statement says exactly that, with Krull dimension, `IsDomain` and both isomorphism claims spelled out (`lean/ComparatorChallenges/ComplexCancellation.lean:20`). The solution proves non-polynomiality with locally nilpotent derivations and a graded invariant, the Makar-Limanov style of argument, at `OAI/Algebra/AffineCancellation/Main.lean:10-29`. The paper also gets a stable-coordinate counterexample in five variables from the same $H$, which is not formalized. Dimension 3 over $\mathbb C$ is still open.

### The finitistic dimension conjecture fails (family 198)

For a finite-dimensional algebra, the little finitistic dimension is the supremum of the projective dimensions that are finite, taken over finitely generated modules. Bass conjectured in 1960 that it is finite. It is known for monomial algebras, radical-cube-zero algebras and algebras of representation dimension at most 3 (Igusa and Todorov). It implies the Nakayama conjectures, which is why it is the flagship of the homological conjectures for Artin algebras.

Family 198 claims a finite-dimensional complex algebra $A$ and, for every $m\ge1$, a finite-dimensional module $N_m$ with $2m-2\le\operatorname{pd}N_m<\infty$. The Lean statement defines the little finitistic dimension with Mathlib's `projectiveDimension` and asks for exactly that (`OAI/Algebra/Finitistic/Main.lean:44`). A second formalized statement gives an algebra whose little and big finitistic dimensions are infinite on the left and 0 on the right, with injectives failing to generate the derived category on one side only (`OAI/Algebra/FinitisticAsymmetry/Main.lean:48`). This lines up with family 199: a counterexample to the Nakayama conjectures would itself force infinite finitistic dimension somewhere.

### Auslander–Reiten, Tachikawa and the Nakayama conjectures fail (family 199)

These conjectures all say a rigid module is projective, in different forms. Auslander–Reiten (1975): a module $M$ with $\operatorname{Ext}^i(M,M\oplus A)=0$ for all $i>0$ is projective. Tachikawa (1973): over a self-injective algebra, vanishing positive self-extensions force projectivity. Nakayama (1958): if every term of the minimal injective resolution of $A$ is projective, then $A$ is self-injective.

Family 199 works over $k=\mathbb F_2(q,H_1,H_2)$. One paper builds a finite-dimensional algebra and a non-projective Gorenstein-projective module with all the required Ext groups zero. The other builds a finite-dimensional symmetric algebra with a non-projective module whose positive self-Ext all vanish. The endomorphism algebra of $A\oplus M$ then breaks the classical, generalized and strong Nakayama conjectures, the Auslander–Gorenstein conjecture and the Wakamatsu tilting conjecture, and everything survives any extension of $k$, including the algebraic closure. Both base counterexamples are in Lean (`OAI/Algebra/AuslanderReiten/All.lean:738`, `OAI/RingTheory/Tachikawa/Counterexample.lean:293`). The Lean statement for Tachikawa defines "symmetric" by an explicit associative nondegenerate form and uses Mathlib's `Abelian.Ext`. The step to Nakayama and the rest is paper-only, and so is everything outside characteristic 2.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig5.png"
  alt="A flow diagram of four boxes. E = T tensor T with Ext algebra k[tau1, tau2] leads, via derived induction and trace obstruction, to Y with W^a = 0 outside {-3, 0}; two twists and a finite fibre lead to F and v: X to F tensor X; a triangular cone leads to the pair (Lambda, Z) with vanishing Hom complex; trivial extension and induction give the symmetric counterexample A = Lambda semidirect D Lambda, M = A tensor Z."
  caption="How the Tachikawa counterexample is assembled, from a tensor-square seed to a symmetric algebra (A counterexample to Tachikawa's second conjecture, Figure 1)."
/>

The Nakayama conjecture is old, and I'd want an expert to confirm the passage from the symmetric counterexample to $\Gamma_K$. Still, with the base counterexamples checked by a kernel, the remaining step is one corollary in a 39-page paper rather than a whole program.

### Log abundance in characteristic zero (family 034)

Now the claims with no machine check, starting with the biggest one in birational geometry. The minimal model program wants every variety to be birational either to one whose canonical class is nef or to a Mori fibre space. Abundance is the conjecture that a nef log canonical class is semiample, meaning some multiple is generated by sections and defines a morphism. It is known for surfaces and threefolds (Miyaoka, Kawamata, and Keel–Matsuki–McKernan for log threefolds) and in the log general type case by Birkar–Cascini–Hacon–McKernan. In dimension 4 and up, even the first step, nonvanishing, was open.

The family has 14 manuscripts and 781 pages. The core is "Log abundance in characteristic zero", whose Theorem 1.1 is the full rational-boundary conjecture: for a normal projective lc pair $(X,B)$ over an algebraically closed field of characteristic 0, with $B$ effective rational and $K_X+B$ $\mathbb Q$-Cartier, nef implies semiample. The induction proves more along the way: canonical nonvanishing for smooth varieties in every dimension, good minimal models for lc pairs and finite generation. The new tool is a signed boundary criterion. If $L=K_X+D$ is nef, $L|_D$ is semiample, and $L$ is a signed combination of boundary components, then $\kappa(L)=\nu(L)$. It is proved by deforming boundary fibres into the interior with filtered Hodge modules.

The release is inconsistent about inputs. This paper takes log Iitaka subadditivity from family 033 as a finished theorem, and notes that the one case it needs also follows from Hacon–Popa–Schnell. The Kähler version, dated ten days later, still states log subadditivity as "Assumption 1.1", and the family title says "under logarithmic Iitaka subadditivity". A third paper on Kähler fourfolds is conditional on three assumptions. The others claim the effective Iitaka fibration conjecture (one pluricanonical degree $m(d)$ works for all smooth $d$-folds), uniform indices for slc log Calabi–Yau pairs, fourfold nonvanishing, and a form of Birkar's Stein-degree conjecture.

I would hold this whole cluster at arm's length until someone who has worked on abundance reads the induction. Abundance has resisted 35 years of very strong people. One release resolving it, Iitaka's conjecture, existence of minimal models and the effective Iitaka fibration all at once is either the biggest week in birational geometry since BCHM or a chain with a weak link. Nothing here is in Lean.

### Iitaka's subadditivity conjecture (family 033)

Iitaka's conjecture $C_{n,m}$ says Kodaira dimension is superadditive in fibrations: $\kappa(X)\ge\kappa(F)+\kappa(Z)$ for $f\colon X\to Z$ with general fibre $F$. It is a pillar of classification theory. Known cases took decades: curve bases and fibres with good minimal models (Kawamata), bases of general type (Viehweg), general-type fibres (Kollár), total dimension at most 6 (Birkar), abelian-variety bases (Cao and Păun), bases of maximal Albanese dimension (Hacon, Popa and Schnell).

Family 033 claims the conjecture in several forms. There is Campana's orbifold version on compact manifolds in Fujiki class $\mathcal C$ with rational SNC boundary, which implies the ordinary and logarithmic ones. There is a separate projective proof over any algebraically closed field of characteristic 0, through positivity of the entire highest Hodge line of a variation of Hodge structure. There is Popa's logarithmic variation inequality $\bar\kappa(U)\ge\kappa(F)+\max\{\bar\kappa(V),\operatorname{Var}(f)\}$, and log additivity in a stratum-smooth setting. The papers say plainly that earlier preprint claims by Tsuji and Maehara are not used, which I appreciated, because this problem has a history of announced proofs that did not hold. The Lean side covers only the negative-fibre branch of the additivity paper (`LogKodairaFiberNegative.lean`), which is a small edge case.

### Minimal models and generalised abundance (family 036)

Existence of minimal models asks that every pair with pseudo-effective adjoint have a minimal model, and otherwise a Mori fibre space. BCHM proved it for log general type. Family 036 claims it for every projective generalized log canonical $\mathbb Q$-pair in characteristic 0, with the nef part fixed, and therefore for ordinary lc pairs. It also claims the Lazić–Peternell generalised abundance conjecture: if $(X,B)$ is klt, $K_X+B$ is pseudo-effective and $M$ is nef with $K_X+B+M$ nef, then $K_X+B+M$ is numerically equivalent to a semiample divisor. A Kähler analogue lives in Bott–Chern cohomology, and a short paper does numerical dimension one. The generalised abundance paper uses the release's own abundance and termination theorems as black boxes, so it stands or falls with families 034 and 056.

### Serre's intersection multiplicity positivity (family 193)

Serre defined the intersection multiplicity of two modules over a regular local ring as $\chi(M,N)=\sum_i(-1)^i\,\ell\bigl(\operatorname{Tor}_i(M,N)\bigr)$ and conjectured in 1965 that it is positive when the supports meet properly, with $\dim M+\dim N=\dim R$. He proved it when $R$ contains a field or is unramified. Vanishing for $\dim M+\dim N<\dim R$ came from Roberts and from Gillet–Soulé in the 1980s, and Gabber proved nonnegativity in the 1990s with de Jong's alterations. Strict positivity in ramified mixed characteristic was the last open piece.

The paper claims it with no liftability or smoothness hypothesis. The strategy uses current tools. It reduces to two complete local domains $D=R/P$ and $E=R/Q$, then passes to perfectoid-type algebras where parameter Koszul homology has normalized length zero (Faltings, then Cai, Lee, Ma, Schwede and Tucker). Rational K-theory with support, via Land–Tamme descent, shows that ordinary and normalized Euler characteristics agree after summing over Galois twists, and each twist is shown positive. The paper is honest in a footnote I did not expect: a public post in December 2025 credited to "Editan", Copilot and Gemini already claimed the prime-quotient form, which the paper records as a claim, not a theorem. So AI-assisted claims on this exact problem predate the release. 31 pages, no Lean.

### Lech's conjecture (family 194)

Lech conjectured in 1960 that Hilbert–Samuel multiplicity cannot drop along a flat local map: $e(R)\le e(S)$. He proved it in dimension at most 2. Linquan Ma proved dimension 3 in equal characteristic and the graded case over perfect fields, through lim Ulrich sequences. Family 194 claims every dimension and characteristic, with no restriction on residue fields. Lean covers only a supporting characteristic-$p$ lemma comparing Dutta and Hilbert–Samuel multiplicities over a complete domain (`DuttaDomain.lean`). The theorem itself is paper-only.

### Bloch's conjecture for surfaces with $p_g=0$ (family 040)

Mumford showed in 1969 that a surface with a holomorphic 2-form has an enormous group of 0-cycles. Bloch conjectured the converse: if $p_g=0$, the Albanese map on degree-zero 0-cycles is an isomorphism, so when $q=0$ as well, all points are rationally equivalent. Bloch, Kas and Lieberman proved it in 1976 for Kodaira dimension below 2. For surfaces of general type there were only families: Godeaux, Catanese and Barlow surfaces (Voisin), Inoue surfaces, and a list of quotient constructions.

The paper claims every surface with $p_g=q=0$. It runs a virtual count on determinant-fixed Quot schemes with two marked surface factors to get a decomposition $0=c[\Delta_X]+\Gamma$ in $\mathrm{CH}^2(X\times X)_{\mathbb Q}$ with $\Gamma$ a sum of external products. Then it proves $c>0$, using a lattice argument when $K^2\le8$ and a comparison with Cartwright and Steger's surface when $K^2=9$. That split, shown below, is where I would look first, since $K^2=9$ covers the fake projective planes.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig4.png"
  alt="A flow chart. Two marked surface factors in a determinant-fixed Quot calculation give 0 = c[Delta_X] + Gamma in CH^2(X x X). Two branches follow: K^2 at most 8, using a lattice class and at most three blowups to force c > 0; and K^2 = 9, using a Cartwright–Steger comparison with c_mid = 12 and c = 12(s + tau) > 0. Both lead to CH_0(X_0) tensor Q = 0, then by Roitman's theorem and q = 0 to deg: CH_0(S) isomorphic to Z."
  caption="The general-type proof of Bloch's conjecture: one diagonal identity, two numerical branches (Bloch's Conjecture for Surfaces with pg = q = 0, Figure 1)."
/>

### Fujita's freeness conjecture (family 038)

Fujita conjectured in 1987 that for a smooth projective $n$-fold and any ample $L$, the bundle $K_X+mL$ is globally generated for $m\ge n+1$. That is sharp on $\mathbb P^n$. It was known up to dimension 5 (Reider, Ein–Lazarsfeld, Kawamata, Ye–Zhu). In general the best bound was linear, around $1.78\,n$ in a recent preprint, after Angehrn–Siu's quadratic bound and Ghidelli–Lacini's $n(\log\log n+2.34)$. The paper claims the sharp $n+1$ in 25 pages. It controls a minimizing log canonical center by comparing the growth of restriction maps with a tangential first variation, uniformly over singular subvarieties of unbounded degree, and then lifts a section in the standard way. It covers only freeness, not Fujita's very-ampleness conjecture. The paper itself mentions an announced proof by Chan that was withdrawn, which is a fair reminder of how this problem behaves. The Kobayashi paper (family 051) uses this result.

### Kobayashi's canonical ampleness conjecture (family 051)

Kobayashi conjectured around 1970 that a compact Kähler manifold with no non-constant entire curves $\mathbb C\to X$ has ample canonical bundle. All prior results needed a curvature hypothesis: negative holomorphic sectional curvature (Wu–Yau, Tosatti–Yang), quasi-negative curvature (Diverio–Trapani), or Kähler hyperbolicity. The paper claims it with no curvature hypothesis at all, starting from Brody's derivative bounds. Its inputs include two other release results, Fujita freeness and the semialgebraic universal cover theorem, and the latter rests on log abundance and Campana's abelianity conjecture. That puts it downstream of the least-checked parts of the release.

### Alperin's weight conjecture (family 202)

Alperin's 1986 conjecture counts the simple modular representations in a $p$-block of a finite group locally: $l(B)$ equals the number of conjugacy classes of $B$-weights, which are defect-zero characters of $p$-local normalizer quotients. It is known for $p$-solvable groups, symmetric groups, $\mathrm{GL}_n$ and many groups of Lie type. Navarro–Tiep and Späth reduced it to "inductive conditions" on finite simple groups, and those have been checked family by family.

The paper claims every block of every finite group at every prime, and here I read more carefully than anywhere else, because the proof does not use the classification of finite simple groups. It turns both sides into weighted counts of tuples $(x,y,u_1,\dots,u_n)$ with $[x,y]u_1\cdots u_n=1$, then needs a uniform $p$-power divisibility of those counts (its Theorem 3.1). It proves that by building marked genus-one curves of large index in characteristic $p$ and controlling inseparable covers of them. A classification-free proof would be a spectacular new idea. It is also exactly where an error would hide, and Theorem 3.1 is the step I would ask an expert to check. The dependency diagram is the paper's own.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/algebraic-geometry-algebra-fig6.png"
  alt="A vertical dependency chart. At the top, a marked genus-one curve with large divisor index and independent absolute field differentials (Section 4) splits into two branches: a sharp inseparable genus bound (Section 5) and a valued lift with marked-cover descent (Section 6). They merge into a finite-presentation and connectedness argument (Section 7), then uniform tuple divisibility (Theorem 3.1), then chain inversion and central trace limits giving l(B) = |W_p(B)| (Section 3)."
  caption="The Alperin proof runs through arithmetic geometry of curves in characteristic p, not through simple groups (The Blockwise Alperin Weight Conjecture, Figure 1)."
/>

### Saxl's conjecture (family 205)

Saxl conjectured in 2012 that the tensor square of the staircase representation of $S_n$, for $n=m(m+1)/2$ and shape $(m,m-1,\dots,1)$, contains every irreducible representation; equivalently, every Kronecker coefficient $g(\rho_m,\rho_m,\mu)$ is positive. Pak, Panova and Vallejo proved hooks and two-row shapes for large $m$, and Ikenmeyer proved shapes comparable with $\rho_m$ in dominance order. The paper claims all of it by exhibiting, for each target, a vector in one cyclic polytabloid submodule of the tensor square. A companion claims the broader tensor square conjecture: for every $n$ outside $\{2,4,9\}$, some irreducible of $S_n$ has a universal tensor square. Both are in Lean (`OAI/RepresentationTheory/Saxl/Main.lean:88` and `UniversalTensorSquares.lean`). The stronger single-cyclic-vector statement is not. My grep found one `axiom` declaration somewhere under `OAI/RepresentationTheory`, which the Comparator would reject if Saxl depended on it. I did not trace the imports.

### The Bass trace conjecture and Kaplansky's idempotents (family 207)

Bass conjectured in 1976 that the Hattori–Stallings trace of a projective module over a group ring sees only elements of finite order. For torsion-free groups this says idempotents are 0 or 1, which is Kaplansky's idempotent conjecture. Over $\mathbb C$ both were known for groups satisfying Baum–Connes-type hypotheses (Berrick, Chatterji and Mislin for $\ell^1$-Bass). The family headline is the $\ell^1$-Bass conjecture for every discrete group, which is analytic and not formalized. The algebraic companion is formalized: the complex group-ring Bass trace conjecture for every group, and for torsion-free $G$ that the only idempotents of $RG$ are 0 and 1, for every commutative domain $R$ of characteristic 0 (`OAI/RingTheory/BassTrace/Main.lean:23` and `BassTorsionFree.lean`). This is about the group ring, not the reduced C*-algebra, so it does not touch Kadison–Kaplansky. Even so, a kernel-checked idempotent conjecture over $\mathbb C$ for every torsion-free group would close a 70-year-old question. I'd want to see the Lean definition of the Hattori–Stallings trace reviewed by someone who knows the subject.

### Kuznetsov's rationality conjecture fails (family 054)

Kuznetsov conjectured that a cubic fourfold is rational exactly when its Kuznetsov component is the derived category of a K3 surface. All known rational cubics fit, and Addington–Thomas showed the categorical and Hodge-theoretic K3 conditions agree generically. Until 2025 no cubic fourfold was even known to be irrational. Katzarkov, Kontsevich, Pantev and Yu then claimed the very general one is, using Hodge atoms (arXiv:2508.05105), and the paper cites them. The paper claims more: for every admissible discriminant $d$ above an ineffective threshold, a very general cubic in the Hassett divisor $\mathcal C_d$ is irrational, although it has both a geometric K3 category and an associated K3. Rational cubics do exist on $\mathcal C_{26}$, $\mathcal C_{38}$ and $\mathcal C_{42}$ (Russo–Staglianò), so the threshold is not decoration. The tool is a new additive invariant of birational maps that counts divisors whose MRC quotient is birational to a K3 in a fixed lattice class, evaluated two ways through weak factorization.

### Shafarevich's holomorphic convexity conjecture fails (family 046)

Shafarevich conjectured that the universal cover of a projective variety is holomorphically convex, a higher-dimensional uniformization statement. It is known when the fundamental group is linear (Eyssidieux, Katzarkov, Pantev and Ramachandran, 2012), and the possible counterexamples were always expected to need strange fundamental groups. The family claims two counterexamples. One is a smooth projective surface whose universal cover is not holomorphically convex, built from a family of curves with paired punctures that are pinched, with an infinite pro-2 quotient certifying that a rational fibre does not contract. The other is a projective fourfold with large fundamental group whose universal cover has no compact curves and still is not Stein, because a lifted abelian surface covers $\mathbb R\times(S^1)^3$, whose third homology a Stein surface cannot carry. The second argument is pleasingly short. The first one, 26 pages for a disproof of a famous conjecture, deserves a careful reading of its infinitude argument. Campana's abelianity paper (family 057) cites it.

### Lipman–Zariski fails for surfaces (family 048)

Lipman showed in 1965 that a free tangent sheaf forces normality, and the Lipman–Zariski conjecture asks that it force smoothness. It is known for hypersurfaces, graded rings, isolated singularities in dimension 3 and up (van Straten–Steenbrink), local complete intersections, and klt and log canonical spaces (Greb–Kebekus–Kovács–Peternell, Druel, Graf–Kovács). In characteristic $p$ it fails already for $xy=z^p$. The paper claims a normal affine complex surface with $\operatorname{Der}_{\mathbb C}(A)\cong A^2$ and a non-regular point. The point is an isolated Gorenstein singularity built analytically and then algebraized through its completed local ring. It is necessarily not log canonical, which is consistent with what was known.

### Griffiths' positivity conjecture fails (family 050)

Griffiths asked in 1969 whether every ample vector bundle carries a Hermitian metric with Griffiths-positive curvature. It holds for line bundles and on curves (Umemura, Campana–Flenner). Positive results after twisting, such as Berndtsson's Nakano positivity of $E\otimes\det E$, made most people believe it. The paper takes an explicit rank-2 bundle $G$ on $\mathbb P^1\times\mathbb P^1$ and pulls it back by the $m$-th power map. $E_m=f_m^*G\otimes\mathcal O(1,1)$ stays ample for every $m$ but has no Griffiths-positive metric once $m\ge m_0$, with $m_0$ coming from a compactness argument and not explicit. It is formalized (`OAI/Geometry/QuadricBundles/Main.lean:10`). A Lean statement about Hermitian metrics and ampleness needs custom definitions, and those deserve an audit before anyone treats the statement as the textbook one.

### Zariski's multiplicity question fails (family 059)

Zariski asked in 1971 whether the embedded topology of a hypersurface singularity determines its multiplicity. It is true for plane curves and for multiplicity 1 (A'Campo, Lê), and true in so many special classes that most people expected a yes. The family claims two pairs of germs that are ambiently homeomorphic with different multiplicities: 2 and 3 in an ambient dimension divisible by 8, and 4 and 5 in $\mathbb C^4$. The method keeps the analytic order separate from the topology. In high dimensions the embedded link of an isolated singularity is classified by the integral Seifert form of its Milnor fibre, so building two equations with different lowest degrees and congruent Seifert forms is enough. The surgery-theoretic classification is a black box, which is fine, but it confines the method to dimensions where that classification applies. Surfaces in $\mathbb C^3$ are not covered.

### The global spherical shell conjecture (family 060)

Class VII surfaces, compact complex surfaces with $b_1=1$ and Kodaira dimension $-\infty$, are the last unclassified piece of the Kodaira classification. Kato showed that a surface containing a global spherical shell, a holomorphic copy of a neighbourhood of $S^3$ with connected complement, is a degenerate blown-up Hopf surface. Nakamura conjectured every minimal class VII surface with $b_2>0$ has one. Teleman proved $b_2=1$ and $b_2=2$ with gauge theory. The paper claims all $b_2>0$ in 39 pages, which would finish the classification: every such surface deforms to a Hopf surface blown up $b_2$ times and is diffeomorphic to $(S^1\times S^3)\#\,b_2\,\overline{\mathbb{CP}}^2$. A uniform argument where Teleman needed case-specific work is a strong claim and an exciting one if it holds.

### The LeBrun–Salamon conjecture (family 062)

LeBrun and Salamon conjectured in 1994 that every positive quaternion-Kähler manifold is a symmetric Wolf space. The twistor space of such a manifold is a contact Fano manifold, so the complex form is that every contact Fano manifold is the adjoint variety of a simple Lie algebra. It was known in real dimensions 8 and 12 and pushed somewhat further by torus-action methods. The paper claims the complex statement in every dimension, hence the Riemannian one, plus a classification of projective contact manifolds. 40 pages, no Lean.

### Hyperkähler SYZ and Lagrangian bases (family 041)

On a compact hyperkähler manifold, a nef line bundle with Beauville–Bogomolov square zero should be semiample, giving a Lagrangian torus fibration. This is the hyperkähler SYZ conjecture. It was known for the four known deformation types (Bayer–Macrì and Markman for $K3^{[n]}$, Yoshioka for Kummers, Mongardi with coauthors for OG6 and OG10). The family claims it for all of them, and separately that the base of every Lagrangian fibration is $\mathbb P^n$, extending Hwang's smooth-base theorem. With the boundedness theorem of Engel, Filipazzi, Greer, Mauri and Svaldi, it gets finitely many deformation types in each dimension when $b_2\ge5$. The two papers cite each other.

### Campana's abelianity conjecture (family 057)

Campana defined special manifolds as those with no fibration onto an orbifold of general type, and conjectured their fundamental groups are virtually abelian. Results so far applied only to linear representations of $\pi_1$. The paper claims the whole group for every special compact Kähler manifold, through $L^2$ Dolbeault spectral analysis on the universal cover. In particular, every compact Kähler manifold of Kodaira dimension 0 would have virtually abelian $\pi_1$. Companions give an independent proof that linear representations of special quasi-projective varieties are virtually 2-step nilpotent, announced earlier by Cao, Deng, Hacon and Păun, and a root-orbifold case. The Kollár–Pardon and Kobayashi papers use this.

### No small Cohen–Macaulay module (family 195)

Hochster asked whether every complete local domain has a finitely generated module of full depth. Such a module would settle several homological conjectures at once. It is true in dimension at most 2, and big Cohen–Macaulay algebras exist in all characteristics, after André. The paper claims a 3-dimensional complete normal local domain over $\mathbb C$ with no small Cohen–Macaulay module, obstructed by a Chern-character inequality on a resolution of an explicit branched cover of a blown-up plane. Characteristic $p$ is not addressed. It is 17 pages.

### Eisenbud–Green–Harris and lex-plus-powers (family 200)

The EGH conjecture says that among homogeneous ideals containing a regular sequence of degrees $a_1\le\dots\le a_n$, the lex-plus-powers ideal realizes every Hilbert function. The lex-plus-powers conjecture adds that its graded Betti numbers are the largest. Known cases were Clements–Lindström (pure powers), Caviglia–Maclagan (degrees growing fast) and Mermin–Peeva–Stillman (ideals containing squares). The family claims both in characteristic 0 for all degree sequences with $a_i\ge2$, with two independent proofs. Positive characteristic is not claimed.

### Kurosh's problem for division rings fails (family 201)

Kurosh asked in 1941 whether an algebra that is algebraic over its centre is locally finite. Golod and Shafarevich answered no for algebras in 1964, but the division-ring case stayed open. The paper claims a countable division ring of characteristic 0 that is algebraic over its centre and generated by two elements, yet infinite-dimensional. It is built by solving free-series equations step by step inside central division algebras. A Lean directory named `Kurosh` exists, but there is no Comparator entry or scope note for this family, so I treat it as unformalized.

### Donovan's conjecture (family 203)

Donovan conjectured that blocks with defect groups of bounded order come in finitely many Morita equivalence classes. It was known for cyclic and Klein-four defect groups, abelian 2-groups (Eaton–Livesey), $p$-solvable groups and symmetric groups. The family claims it over every algebraically closed field of characteristic $p$, including $p=2$, and in an integral form over Witt vectors and over ramified complete DVRs. Unlike Alperin, this proof uses the classification. It bounds Cartan invariants for quasisimple groups through tilting objects on flag varieties, uniformly in field size and rank, and handles extensions with equivariant Frobenius comparisons.

### Finite lattice representation fails (family 206)

Grätzer and Schmidt proved in 1963 that every algebraic lattice is the congruence lattice of some algebra. Whether every finite lattice is the congruence lattice of a finite algebra is one of the oldest open problems in universal algebra. Pálfy and Pudlák showed in 1980 that it is equivalent to every finite lattice being an interval in the subgroup lattice of a finite group. The family claims a no, an explicit coloured-graph characterization of finite congruence lattices, and undecidability both of that representation property and of recognizing subgroup intervals. Lean has the graph characterization only (`FiniteCongruenceGraph.lean`), not the negative answer. The two papers come to 267 pages.

### The other major results

The remaining majors are real results in their fields, but either a named conjecture's remaining case or a strong theorem short of a famous conjecture. I give each a short entry.

#### The ordinary double point gap (family 037)
Every singular klt germ of dimension $n$ has normalized volume at most $2(n-1)^n$, with equality only at the ordinary double point. This is the Spotti–Sun conjecture, known in dimension 3 through Liu–Xu. The proof inducts from a separately proved fourfold case, which is its only internal input. It bounds the volume of singular K-semistable Fano varieties.

#### Every K3 surface is Oka (family 042)
Every complex K3 surface, projective or not, has the convex approximation property, which answers Forstnerič's question. Through every point and tangent vector there is a dense immersed entire curve. The family title also mentions class VII surfaces, but the paper's theorem is about K3.

#### P = W for $\mathrm{SL}_n$ (family 043)
The perverse filtration of the $\mathrm{SL}_n$ Hitchin fibration equals the weight filtration of the twisted character variety in every composite rank, including the variant part. With Maulik–Shen's prime-rank case, this completes fixed-determinant P = W.

#### Tangent bundle splittings (family 052)
If the tangent bundle of a compact Kähler manifold splits into two integrable subbundles, the universal cover is a product. On rationally connected manifolds every splitting is integrable, as Höring conjectured, so the manifold is a product. This is the two-summand case of Beauville's conjecture. It is formalized (`OAI/Geometry/KahlerSplitting/Main.lean:14`), on custom complex-manifold foundations.

#### Pixton's completeness fails (family 053)
A tautological class outside the span of Pixton's original relations vanishes in Chow and in cohomology. The example has genus $10^{60}$, far beyond any computer check. Extended relation systems may survive.

#### Gepner stability and Calabi–Yau threefolds (family 055)
Toda's Gepner stability condition on every quintic threefold, with phase shift exactly $2/5$. Separately, large-volume Bridgeland stability conditions on every Calabi–Yau threefold, using a corrected Bogomolov–Gieseker-type inequality on all threefolds. Existence on general CY3s was one of the central open problems in the area.

#### Termination of flips on fourfolds (family 056)
Every log canonical MMP on a projective fourfold with rational boundary terminates in characteristic 0, with Kähler analogues for existing flip sequences. Two manuscripts share the title "Finite ordinary minimal model programs on compact Kähler fourfolds" and state different theorems, which is a versioning oddity.

#### Kollár–Pardon (family 058)
A universal cover with a semialgebraic presentation is a bounded symmetric domain times $\mathbb C^m$ times a simply connected projective variety. The bounded-domain half, that semialgebraic bounded domains with compact quotients are symmetric, is formalized on self-contained definitions (`OAI/Analysis/SymmetricDomains/Main.lean:55`). The classification uses the release's abundance and abelianity papers.

#### The generalized Mukai conjecture (family 063)
$\rho(\iota-1)\le n$ for Fano manifolds, with equality only for $(\mathbb P^{\iota-1})^\rho$. It was known up to dimension 5. The proof is a 19-page Gromov–Witten argument through eigenvalues of quantum multiplication by divisors.

#### The $\mu$-constant problem for surfaces (family 064)
Families of isolated surface singularities in $\mathbb C^3$ with constant Milnor number are topologically trivial. This is the one case Lê–Ramanujam's 1976 h-cobordism argument could not reach.

#### Virasoro constraints (family 065)
The full descendant Virasoro conjecture for every smooth complete intersection, including odd and primitive classes, without semisimplicity. A companion shows the constraints pass to projective bundles.

#### Bounded klt complements (family 066)
Klt complements of index bounded by dimension and $\epsilon$ for $\epsilon$-lc Fano contractions, the finite-coefficient form of Shokurov's conjecture, plus the Birkar–Shokurov Cartier-divisor conjecture.

#### Anticanonical nonvanishing (family 068)
If $-K_X$ has a smooth semipositive metric, some power of it has a section. Seven papers form one chain, and several are stated reductions.

#### Quantum geometric Langlands at irrational level (family 069)
$D_c(\mathrm{Bun}_G)\simeq D_{-1/(rc)}(\mathrm{Bun}_{G^\vee})$ for every simple $G$, every curve and every $c\notin\mathbb Q$, in 119 pages. Rational and critical levels are excluded.

#### Saturation for $\mathrm{Spin}(2n)$ (family 204)
Saturation factor 1 in type D, extending Knutson–Tao's type A theorem and the $\mathrm{Spin}(8)$ to $\mathrm{Spin}(12)$ cases. Types E remain.

#### Finite symmetric tensor categories (family 208)
Every finite symmetric tensor category in characteristic $p$ has a fibre functor to some higher Verlinde category $\mathrm{Ver}_{p^n}$, which is the finite case of the Benson–Etingof–Ostrik conjecture.

#### Gersten's conjecture fails integrally (family 209)
Explicit 2-dimensional ramified regular local rings of mixed characteristic $(0,5)$ have $K_3$ and $K_5$ classes that die over the fraction field. The unramified case is untouched.

### Four narrower results

Family 035 proves one case of abundance in positive characteristic: nef log canonical threefold adjoints of numerical dimension one are semiample when $p>3$. The characteristic-$p$ threefold MMP is fully established only for $p>5$, so I'd check how the paper handles $p=5$. Family 044 proves the equivariant Hikita conjecture for every finite quiver, under a regularity hypothesis on the stability character. Family 067 proves the Campana–Peternell conjecture in dimension 6 only, ending with an exact arithmetic certificate. Family 210 proves the sixth case of Foulkes' conjecture and, more interestingly, that the canonical Foulkes–Howe map is surjective once $b\ge a(a-1)$, independent of $\dim V$. That second statement is in Lean (`OAI/RepresentationTheory/FoulkesHowe/Stabilization.lean:15`), but the sixth-power case is formalized only for $b\ge30$.

Nothing in the group is a purely technical lemma. The closest is family 044's regularity-restricted Hikita, and even that is a named conjecture.

### What I would believe today

Ranked by how much I trust them right now, not by importance:

1. The explicit counterexamples with matching Lean statements: Abhyankar–Sathaye, cancellation over $\mathbb C$, zero divisors, finitistic dimension, Auslander–Reiten and Tachikawa, and Nagata, which is the one positive theorem in that tier. If they build under the Comparator, they are true, and building them is mechanical.
2. Short, self-contained papers that rest on established machinery: Fujita, Lipman–Zariski, Griffiths (also in Lean), the Shafarevich fourfold, Lech, Serre, Saxl (in Lean).
3. Large claims that lean on other unrefereed release papers: the abundance, Iitaka and minimal model cluster, Kobayashi, Kollár–Pardon, and generalised abundance. Each is only as good as its weakest companion.
4. Large claims with unusual proofs: Hodge for CM abelian varieties, the K3 companions through Fukaya categories, Alperin without the classification, and Bloch.

## Analysis, PDE and convex geometry

This is the heaviest part of the release by weight. I went through 58 families and 104 manuscripts, about 4,200 pages in all. They cover real and complex analysis, convex and metric geometry, functional analysis and partial differential equations. I started with the titles, and they read like the open-problems chapter of a graduate textbook: Mahler, Falconer, Kakeya, Bochner–Riesz, Koebe, Brennan, Kirk, Tingley, De Giorgi, Lane–Emden, hot spots, Vlasov–Maxwell. Several of these are problems I was told, as a student, that nobody expected to see settled.

Three patterns came out of reading them, and they shape everything below.

The first is that the Lean coverage and the length of the papers pull in opposite directions. 34 of the 58 families have a Comparator challenge matching the headline theorem. Ten are partly formalized and 14 have nothing. The formalized ones are often startlingly short. Tingley's problem takes 12 pages, Lane–Emden 20, log-Brunn–Minkowski 20, Kirk's fixed-point problem 19 and Falconer 64. The unformalized ones are often enormous. The Kakeya pair runs 272 pages, the Schrödinger endpoint 179, local smoothing 165 and Ball–Evans 192. Almost all of classical Fourier analysis (Kakeya, restriction, Bochner–Riesz, local smoothing, the trilinear Hilbert transform, the $L\log L$ problem) sits in the unformalized pile. Those are also the results that cite each other.

The second pattern is priority. Several headline problems had already moved in the weeks before the release, and some of that was AI-assisted human work. Wang and Zahl proved the Kakeya set conjecture in $\mathbb R^3$ in February 2025. The scalar Crouzeix conjecture fell in July and August 2026. A proof of the Mumford–Shah conjecture went up on arXiv two days before OpenAI's. Its author says the strategy and draft came from an OpenAI model. Humans reached De Giorgi in dimension four only days before the release claims dimension eight. In each of these cases the OpenAI paper is either an extension or a second proof, and I say which.

The third pattern is a few places where the official summary says something different from the paper. The details are in the sections.

My grades: 20 families landmark, 28 major, 9 notable and 1 technical. "Landmark" means a famous named problem that the paper claims to settle fully. By kind, 42 are proofs, 10 are counterexamples or disproofs, 4 settle part of a conjecture and 2 are sharp bounds. The order below is my own, most important first. It weighs what is claimed against how much of it a machine has checked.

### Relativistic Vlasov–Maxwell: no blowup for large data in three dimensions

The relativistic Vlasov–Maxwell system models a collisionless plasma. A density $f(t, x, v)$ of charged particles moves under the electric and magnetic fields it creates, and those fields obey Maxwell's equations sourced by $f$. In 1986 Glassey and Strauss proved the fact that organises the whole field. A smooth solution in three dimensions can only break down if some particles reach unbounded momentum in finite time. For forty years the large-data question stayed open: given smooth data with no smallness, symmetry or neutrality, can that happen? People had dimension reductions (Glassey and Schaeffer, the 1990s), weak solutions (DiPerna and Lions, 1989), sharper continuation criteria (Pallard; Luk and Strain) and small-data results. The closest large-data result is Xuecheng Wang's cylindrically symmetric case ([arXiv:2203.01199](https://arxiv.org/abs/2203.01199)), whose follow-up appeared this July. Family 362 claims the full answer in [49 pages](https://github.com/openai/math/blob/main/preprints/Global-classical-solutions-of-the-three-dimensional-relativistic-Vlasov-Maxwell-system-September-23-2026/paper.pdf): every smooth admissible datum has a unique global classical solution, smooth on every finite time interval, with compact momentum support. Nonzero total charge is allowed. The system is the one-species relativistic one, with constants normalized.

The mechanism is a signed impulse estimate. Follow a "receiver" particle whose energy stays within a factor of 8 of some dyadic level $w$ over a time interval of length $I$. If $P$ bounds every momentum present, the paper's Proposition 2.1 says the receiver's momentum changes by at most

$$
M P \sqrt{I} \;+\; A\, I\, \frac{P^2 \log(2+P)}{w}.
$$

Apply that at the first moment the maximal energy reaches $2^n$. Climbing from $2^{n-1}$ to $2^n$ then takes time at least about $c/n$. Those times have a divergent sum, so no finite time is long enough to reach infinite momentum. The hard part is proving the estimate. The force comes from the Glassey–Strauss retarded-field formulas, with sources traced back along the receiver's backward light cone and binned by momentum, distance and angle. Energy flux plus occupation bounds handle most bins. They lose a factor $\sqrt w$ exactly in the thin sector where a source moves closer to the light ray than the receiver does. There the paper integrates the force along source trajectories *before* taking absolute values. An exact identity makes the singular parts cancel, leaving a total derivative, milder bulk terms and a term with the receiver's own acceleration that is fed back through the Lorentz force law. The bootstrap assumption is then used twice, once to count how often particles can change direction and once to win the $\sqrt w$ back.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig5.png"
  alt="Flow chart: an assumed increment bound feeds baseline occupation estimates, a bound on direction changes and improved occupation; a signed-cancellation box is used twice, once for direction changes and once for a selected-bin force estimate, which gives a strict improvement and closes the continuity argument."
  caption="The bootstrap at the heart of the Vlasov–Maxwell paper: signed cancellation along source trajectories is used twice, and the second use buys the strict improvement that closes the loop (Global classical solutions of the three-dimensional relativistic Vlasov–Maxwell system, Figure 1)."
/>

Forty-nine pages for a problem of this stature would make me nervous in any human submission. What changes the picture is the Lean side. The Comparator statement, `OAI.RVM.global_classical_solution` (`lean/ComparatorChallenges/VlasovMaxwell.lean:103`), spells out the equations, the admissible data, classical solutions, smoothness on finite horizons and uniqueness in 112 lines. It matches the paper's Theorem 1.1. The solution is about 90,000 lines in 207 files, with no `sorry`, and it has to contain the Glassey–Strauss continuation theory and local existence, not assume them.

The model's own reasoning summary for this family is candid. It ends by saying the result "rests on the occupation bound, direction-count estimate, and cutoff cancellations needed to close the signed momentum bootstrap". It also records an exponent contradiction it hit, $H < 0.41z$ against $H > 0.496z$, and fixed by shortening its stability windows. If the formal proof checks, the trace's worry is answered. If it does not, the trace tells you exactly where to look.

### Falconer's distance conjecture in every dimension

Take a compact set $E \subset \mathbb R^d$ and look at its set of distances $\Delta(E) = \lbrace |x - y| : x, y \in E\rbrace$. Falconer asked in 1985 how large the Hausdorff dimension of $E$ must be before $\Delta(E)$ has positive length, and conjectured that anything above $d/2$ suffices. He proved $(d+1)/2$. Then came forty years of steady work: Wolff's $4/3$ in the plane, Erdoğan's $d/2 + 1/3$, and Guth, Iosevich, Ou and Wang's $5/4$ in the plane ([Inventiones 2020](https://arxiv.org/abs/1808.09346)). After them came the Du–Zhang line of results with thresholds like $d/2 + 1/4 + 1/(8d-4)$. The conjecture was open in every dimension.

Family 073 claims all of it in [64 pages](https://github.com/openai/math/blob/main/preprints/The-Falconer-distance-conjecture-in-all-dimensions-September-23-2026/paper.pdf): for every $d \ge 2$ and every compact $E$ with $\dim_H E > d/2$, $\Delta(E)$ has positive Lebesgue measure. The strict threshold is all it claims. Nothing is said at exactly $d/2$, and nothing about the pinned version. The paper builds directional "cap bounds" for pairs of points from auxiliary projections and a planar Furstenberg incidence estimate. It runs a multiscale recursion on the mass profiles of two Frostman measures, closed by a "two-depth potential" that pays for each step. A shell-summation lemma then shows the limiting distance measure is absolutely continuous. What surprised me is what it avoids. There is no decoupling and no use of the prior distance theorems, which is the toolkit every recent human result was built from. Odd dimensions need an extra thin-tube upgrade.

A 64-page proof of one of the central problems of geometric measure theory would be hard to believe on its own. The Lean statement is one of the cleanest in the release:

```lean
-- lean/ComparatorChallenges/FalconerAllDimensions.lean:11-15
theorem falconer_distance_conjecture :
    ∀ (d : ℕ), 2 ≤ d → ∀ E : Set (EuclideanSpace ℝ (Fin d)), IsCompact E →
      (d : ℝ≥0∞) / 2 < dimH E →
        0 < volume {r : ℝ | ∃ x ∈ E, ∃ y ∈ E, dist x y = r} := by
  sorry
```

The solution first states the planar Furstenberg input, Orponen and Shmerkin's theorem in the tube form of [Orponen–Shmerkin–Wang](https://arxiv.org/abs/2209.00348), as an explicit hypothesis. Then it proves that input in Lean as well (`publishedPlanarFurstenberg_proved`), so the top-level statement is unconditional. I would put this second only to Mahler in this group for how much the Lean check is worth.

### Koebe's circle-domain conjecture

In 1908 Koebe asked whether every domain in the Riemann sphere is conformally equivalent to a circle domain, one whose complementary pieces are all round disks or points. He proved it for finitely many holes in 1920. He and Schramm proved it for countably many in the Annals in 1993. The uncountable case, with complements as wild as a Cantor set, stayed open. Ntalampekos's survey from March this year ([arXiv:2603.15098](https://arxiv.org/abs/2603.15098)) still lists it.

[Family 071](https://github.com/openai/math/blob/main/preprints/Koebes-Circle-Domain-Conjecture-September-23-2026/paper.pdf) claims the existence statement in full, for every domain with no hypothesis on the complement. A companion proves one direction of He and Schramm's rigidity conjecture: a circle domain with conformally removable boundary is rigid, so any conformal map onto another circle domain is a Möbius map. Uniqueness in general is not claimed, and Rajala showed the converse direction is false. The proof collapses each hole to a point and carries finite-energy test functions through finite circle-domain approximations. Small-energy "barriers" force the unmarked holes to shrink to points, and a period argument stops the marked ones from bulging past their limiting disks. The two papers total 89 pages and are coupled: the existence paper uses a transfer theorem from the rigidity paper.

The Lean development is the largest in this group, about 211,000 lines in roughly 2,600 files. It states both theorems in `KoebeCircleDomains.lean`. Moore's decomposition theorem appears inside it as an explicit hypothesis. Since the Comparator statement itself is unconditional, it must be discharged internally for the check to pass. Anyone auditing should look at how conformality and conformal removability are defined.

### De Giorgi's conjecture in dimension eight, and the result I trust least

De Giorgi conjectured in 1978 that a solution of the Allen–Cahn equation $\Delta u = u^3 - u$ that is monotone in one direction must be one-dimensional, a flat transition layer, in dimensions up to 8. It is the diffuse version of Bernstein's problem for minimal graphs. Ghoussoub and Gui proved dimension 2 (1998), Ambrosio and Cabré dimension 3 (2000), and Savin dimensions up to 8 (Annals 2009), but only under an extra assumption that $u \to \pm 1$ as the monotone coordinate goes to $\pm\infty$. Del Pino, Kowalczyk and Wei showed it fails from dimension 9 on (Annals 2011). Removing Savin's assumption was the open problem.

The human frontier moved in September 2026. Two groups proved stable rigidity in $\mathbb R^3$, and with it the monotone case in $\mathbb R^4$: Liu, Luo, Wang, Wei, Wei and Wu ([arXiv:2609.21680](https://arxiv.org/abs/2609.21680)), and Chan, Fernández-Real, Figalli, Florit-Simon and Serra ([arXiv:2609.30194](https://arxiv.org/abs/2609.30194)). Hong, Li and Wang had just proved the stable Bernstein theorem for minimal hypersurfaces in $\mathbb R^7$ ([arXiv:2609.15720](https://arxiv.org/abs/2609.15720)).

[Family 375](https://github.com/openai/math/blob/main/preprints/De-Giorgis-conjecture-in-dimension-eight-September-26-2026/article.pdf) claims the endpoint. Every monotone solution in $\mathbb R^8$ is a planar $\tanh$ profile, with no condition on the limits and no energy bound. The key step is that every stable solution in $\mathbb R^7$ is flat, again with no energy-growth assumption. Since monotone solutions in lower dimensions extend trivially, this would settle De Giorgi in every dimension where it is true. The reduction from $\mathbb R^8$ to stable solutions in $\mathbb R^7$ uses classical tools: Alberti–Ambrosio–Cabré, Jerison–Monneau, Federer's dimension reduction and Wang's near-unit-density theorem. That is the most reliable part. The new content is ruling out unbounded energy density for stable solutions in $\mathbb R^7$. It works through curvature-moment estimates on the zero set, a "penalized reach" that measures both how sheets bend and how close neighbouring sheets come, and a radial stability test that leaves a strict margin in interface dimension 6. The paper itself flags that margin as tight. It also checks consistency with the known counterexamples in an appendix, with no contradiction.

This is 97 pages with no Lean formalization. It jumps from the human frontier of $\mathbb R^3$ to $\mathbb R^7$ in a single paper. Put plainly: it claims the Allen–Cahn analogue of a theorem that humans proved for the much simpler minimal-surface case only three weeks ago. If it is right, it is one of the best results in the release. Until an expert has read the curvature-moment section, I would not cite it as settled.

### Fourier restriction for positively curved surfaces in three dimensions

Stein's restriction conjecture asks for which $p$ the Fourier extension of a function on a curved surface lands in $L^p$. For surfaces in $\mathbb R^3$ the conjectured range is $p > 3$. The human record was the Wang–Wu bound $p > 22/7$ ([arXiv:2411.08871](https://arxiv.org/abs/2411.08871)), after Guth's polynomial partitioning reached $13/4$. Family 077 claims the full range in two forms. One is the bounded-data estimate for the sphere, upgraded to $L^q \to L^q$ through Bourgain's factorization. The other is the diagonal $L^p \to L^p$ estimate for every compact surface with definite second fundamental form. Corollaries include sharp Schrödinger local smoothing in $\mathbb R^2$. The new tool is an "elliptic capacity" that tracks how wave-packet mass concentrates when tested against eccentric ellipses, and it is shown to propagate along lines. Caveats: 128 pages, no Lean, and the general-surface paper plugs in the 3D Kakeya maximal theorem from family 074 and a packet theorem from its own companion. If 074 has a hole, so does this.

### The Bochner–Riesz conjecture in three dimensions

Fefferman's 1971 ball-multiplier theorem says that cutting off the Fourier transform sharply at a sphere is unbounded on $L^p$ for every $p \ne 2$. The Bochner–Riesz conjecture says that smoothing the edge to $(1 - |\xi|^2)_+^\delta$ makes the operator bounded for every $\delta$ above a critical line. Carleson and Sjölin settled the plane in 1972. In $\mathbb R^3$ the best range was $\max(p, p') \ge 22/7$ (Gao, Wu and Xi, [arXiv:2509.01116](https://arxiv.org/abs/2509.01116)). Family 078 claims boundedness on $L^3(\mathbb R^3)$ for every $\delta > 0$, and hence the full strict range by interpolation. Its argument is an information-theoretic "positive transport" scheme. It bounds how fast information about masses moving along lines can grow across scales, then transfers that to oscillatory packets. It is 114 pages with no Lean. Zipeng Wang has an unrefereed all-dimension claim ([arXiv:2501.12742](https://arxiv.org/abs/2501.12742)), which the paper cites. As far as I can tell that claim is not accepted.

### Local smoothing in three space dimensions

Sogge's local smoothing conjecture says that averaging a wave over a short time window recovers almost all the derivatives that a fixed-time $L^p$ estimate loses. Guth, Wang and Zhang settled two space dimensions in 2020. Do not confuse that with this. In three space dimensions the best human result, Gan, He, Li and Wu ([arXiv:2502.05973](https://arxiv.org/abs/2502.05973)), was sharp only for $p \ge 10/3$ and still lost a fixed $1/12$ of a derivative at the critical exponent $p = 3$. Family 079 claims the $\epsilon$-loss estimate at $p = 3$:

$$
\|e^{it\sqrt{-\Delta}} f\|_{L^3(\mathbb R^3 \times [1,2])} \le C_\epsilon \|J^\epsilon f\|_{L^3}.
$$

Because local smoothing at $p = 3$ implies the other three, section 11 of this 165-page paper re-derives Bochner–Riesz, restriction for the sphere and the Kakeya maximal estimate by a route independent of 074 and 078. I read that two ways. If all four papers are right, the release has produced the whole chain twice, which is a striking internal consistency check. But nothing outside the release has checked any of it, and a shared blind spot would propagate through all four. Of everything in this group, this cluster most needs human referees, and it will be the slowest to get them.

### Brennan's conjecture, and a 1996 formula that does not hold

A conformal map $\varphi$ from a simply connected domain onto the disk can distort area near the boundary as badly as the boundary is wild. Brennan conjectured in 1978 that $\int |\varphi'|^s\,dA$ is still finite for every $4/3 < s < 4$. The lower end and $s < 3$ were classical. Bertilsson's 1999 thesis pushed the upper end to about 3.42, and it stuck there for a quarter-century. In the language of integral-means spectra the conjecture is $B_S(-2) = 1$. Family 072 claims it, in the uniform form $M_{-2}[f'](r) \le C_\epsilon (1-r)^{-1-\epsilon}$ over the whole schlicht class. The proof moves to a compact family of normalized maps and encodes boundary growth as the growth rate of a positive transfer operator. It then derives a contradiction from a hypothetical too-fast rate through a Morse-index argument pairing saddle points with maxima.

The companion paper is more surprising. Kraetzer conjectured in 1996, from numerics on Julia sets, that the bounded-class spectrum is exactly $t^2/4$ for $|t| \le 2$. The paper proves $B_b(-1) < 1/4$, so Kraetzer's formula fails at $t = -1$. It does not say by how much: the gap $\epsilon$ is never computed. That overturns a long-held, numerically supported belief, and I found no rigorous lower bound that contradicts it. Both headline statements are in Lean (`OAI/Analysis/IntegralMeans/Main.lean:10`, `OAI/Analysis/StrictMeans/Main.lean:7`). The summary's extra claim, $B_S(t) = |t| - 1$ for every $t \le -2$, is proved in the paper but is not in the formal statement. 46 pages in total.

### David–Semmes in higher codimension

David and Semmes asked in the early 1990s whether a single analytic fact forces geometry. Suppose the $n$-dimensional Riesz transform, a vector-valued singular integral, is bounded on $L^2(\mu)$ for an Ahlfors-regular measure $\mu$. Must $\mu$ be uniformly rectifiable, built out of Lipschitz pieces? For $n = 1$ this was Mattila, Melnikov and Verdera (1996). For codimension one it was Nazarov, Tolsa and Volberg (Acta 2014, [arXiv:1212.5229](https://arxiv.org/abs/1212.5229)), whose proof rests on a maximum principle with no known higher-codimension analogue. Tolsa's ICM survey this year still lists $1 < n < d - 1$ as open. Family 081 claims it for $d \ge 4$ and $2 \le n \le d - 2$, with constants depending only on dimension, the regularity constant and the Riesz bound. It replaces the maximum principle with a flatness-improvement scheme. Normal heights are rescaled by the excess, and in the limit, small Riesz pairings force the rescaled heights to satisfy an equation whose only solutions are affine. It is 56 pages, and the quantitative theorem is in Lean (`Uniformity/Rectifiability.lean:89`, about 824 files). The variation and principal-value corollaries are not.

### Ultraflat Littlewood polynomials exist

Sign a polynomial's coefficients $\pm 1$, $P(z) = \sum_{k < N} \epsilon_k z^k$. On the unit circle its average size is $\sqrt N$. Littlewood and Erdős asked whether the maximum can be pushed down to $(1 + o(1))\sqrt N$, with the minimum up to $(1 - o(1))\sqrt N$. Erdős conjectured not: he expected a fixed gap. Coding theorists conjectured the related "merit factor" is bounded. Kahane built ultraflat polynomials in 1980, but with complex unimodular coefficients, not $\pm 1$. Balister, Bollobás, Morris, Sahasrabudhe and Tiba showed in 2020 that *flat* $\pm 1$ polynomials exist, within constant factors. Family 076 claims ultraflat ones: for every $\epsilon$ and large $N$, a $\pm 1$ polynomial with $(1-\epsilon)\sqrt N \le |P| \le (1+\epsilon)\sqrt N$ on the whole circle. That would refute Erdős's conjecture and make the largest binary merit factor tend to infinity. The construction builds a smooth unimodular function from quadratic-phase waves on short arcs, arranges its first $N$ Fourier coefficients to be nearly $\pm 1/\sqrt N$, and rounds them to exact signs with discrepancy theory (Spencer, Lovett–Meka).

There is a live dispute. el Abdalaoui posted a preprint in 2025 ([arXiv:2504.21499](https://arxiv.org/abs/2504.21499)) asserting that ultraflat $\pm 1$ sequences cannot exist. An appendix of the OpenAI paper argues his reasoning is wrong. Lean covers the upper bound $\max |P| \le (1+\eta)\sqrt N$ and a finite-$p$ flatness statement (`OAI/Analysis/Littlewood/Main.lean:187`), and the flatness statement alone already contradicts el Abdalaoui's claim. The headline two-sided theorem, from a paper dated 5 October, is not formalized.

### Planar and bounded-treewidth graphs embed into $L_1$

This is a computer-science problem wearing geometry's clothes. Embedding a graph's shortest-path metric into $L_1$ with distortion $C$ is equivalent to the multicommodity flow–cut gap being at most $C$. Gupta, Newman, Rabinovich and Sinclair conjectured in 1999 that every minor-closed family has bounded distortion. Planar graphs and bounded-treewidth graphs were the flagship cases. Rao's 1999 bound for planar graphs was $O(\sqrt{\log n})$, and treewidth 2 was the most anyone had for treewidth. Family 089 claims both: a universal constant for every planar graph with arbitrary positive edge lengths, and a constant $C(k)$ for treewidth below $k$. $L_1$ metrics are positive combinations of cuts, so the task is a random cut distribution that separates pairs in proportion to distance without cutting edges often. The planar proof evolves monotone cuts from coarse to fine scales with exactly-gluing interpolation so losses never accumulate. Both theorems are in Lean (`PlanarL1/Embedding.lean:96`, `TreewidthL1/Main.lean:135`), with planarity defined through topological drawings. Constants are not explicit, and the full GNRS conjecture for all minor-closed families remains open. 75 pages.

### The triangular lattice is universally optimal

Why do vortices in superconductors, Coulomb gases and packed disks all settle into a hexagonal pattern? Cohn and Kumar conjectured in 2007 that the triangular lattice is "universally optimal" in the plane. It would minimize energy for every completely monotone interaction of squared distance, as $E_8$ and the Leech lattice do in dimensions 8 and 24 (Cohn, Kumar, Miller, Radchenko and Viazovska, Annals 2022). Family 090 claims the planar case for every locally finite configuration of density one. It follows the Viazovska-style method: for each Gaussian, build a sharp auxiliary function below it whose Fourier transform is nonnegative and which touches it at lattice points, then integrate over Gaussians. The same philosophy gives a sharp Cohn–Elkies certificate, so the planar linear-programming bound equals $\pi/(2\sqrt3)$. A separate Voronoi-cell argument claims Sandier and Serfaty's Abrikosov-lattice conjecture for the renormalized energy $W$, and through Bétermin and Sandier the Brauchart–Hardin–Saff conjecture on the sphere's logarithmic energy.

Universal optimality and the Cohn–Elkies certificate are in Lean, which covers their interval-arithmetic certificates. The Sandier–Serfaty paper (49 pages, also interval arithmetic) and the Riesz and log jellium results are not. Only the minimum value is proved, not uniqueness of minimizers. 199 pages over four papers.

### The log-Brunn–Minkowski inequality and the B-conjecture

Brunn–Minkowski says volume is concave under Minkowski averaging. Böröczky, Lutwak, Yang and Zhang conjectured in 2012 a stronger version for origin-symmetric bodies. Average the support functions *geometrically*, $h_K^{1-\lambda} h_L^\lambda$, take the Wulff body, and its volume should still be at least $|K|^{1-\lambda}|L|^\lambda$. They proved it in the plane. Saroglou did unconditional bodies, Kolesnikov and Milman local versions, and there were partial symmetry results as recently as August. Family 091 claims every dimension in 20 pages, and the main inequality is in Lean, in a single file of about 24,000 lines (`OAI/Geometry/LogVolume/BrunnMinkowski.lean`). The core is a variance inequality for even test functions in "moment coordinates", essentially the Kolesnikov–Milman local spectral formulation made global. Corollaries follow through Saroglou's transfer theorem: the symmetric $L_p$ Brunn–Minkowski inequality for $0 < p < 1$, and the (B)-conjecture of Banaszczyk and Latała for even log-concave measures. Those corollaries are not formalized, and equality cases are not claimed. The paper notes that Stancu had earlier announced an all-dimensional log-Minkowski inequality with details deferred.

### Hyperbolicity cones that are not slices of the PSD cone

This one matters to anyone who uses semidefinite programming. Hyperbolicity cones generalize the cone of positive semidefinite matrices, and interior-point methods run on them. The generalized Lax conjecture says every hyperbolicity cone is spectrahedral, a linear slice of a PSD cone. If true, hyperbolic programming would be no more expressive than SDP. Helton and Vinnikov settled three variables in 2007. Kummer and Netzer proved the conjecture for strictly hyperbolic polynomials only last month ([arXiv:2609.24542](https://arxiv.org/abs/2609.24542)), with a proof they say an AI system found. In January, González Nevado posted a claimed proof of the full conjecture ([arXiv:2601.12267](https://arxiv.org/abs/2601.12267)).

Family 095 disproves it. The explicit counterexample is $p(X, Z, y) = \det\big((\det X) Z - \Phi_y(\operatorname{adj} X)\big)$ on $S^4 \times S^4 \times \mathbb R^3$, built from the Choi–Lam biquadratic form. That form is nonnegative but dominates no nonzero bilinear square, and a PSD pencil pushed to rank-one limits would produce such a square. It is in Lean (`OAI/Analysis/HyperbolicCones/Main.lean:25`). An appendix points to the step in González Nevado's preprint that fails. A second paper claims much more: a huge existential cone with *no* semidefinite lift at all, which disproves Netzer and Sanyal's projected version and would make hyperbolic programming strictly more expressive than SDP even with extra variables. That one is not formalized. A third paper shows the explicit cone *does* have an exact semidefinite lift, so the no-lift claim rests only on the big existential example. A small inconsistency: the second paper's abstract says degree 16, while its theorem and the Lean doc say 20. Both are right, since $p$ has degree 20 and $p/\det X$ has degree 16 with the same cone.

### The sharp simplex conjecture for isotropic constants

Bourgain's slicing problem asked whether the isotropic constant $L_K$ of a convex body is bounded by an absolute constant. Klartag and Lehec settled that in December 2024 ([arXiv:2412.15044](https://arxiv.org/abs/2412.15044)); Bizeul gave a second proof ([arXiv:2501.06854](https://arxiv.org/abs/2501.06854)). The sharp form asks which body is worst, and conjectures the simplex. Family 101 claims exactly that in every dimension:

$$
L_K \le \frac{(n!)^{1/n}}{(n+1)^{(n+1)/(2n)}\sqrt{n+2}}, \quad \text{with equality iff } K \text{ is a simplex}.
$$

By Fradelizi and Marín Sola it is equivalent to a sharp entropy inequality: among log-concave laws with given covariance, products of one-sided exponentials have least entropy. The proof transports a Gaussian onto the density and controls the entropy deficit through a Gaussian chaos decomposition. Through Klartag's 2018 reduction the theorem implies the general Mahler inequality. The paper says so on its second page, and the summary does not mention it. So the release proves general Mahler by two unrelated routes. This one is 42 pages of hard analysis with no Lean at all. I would treat it as the riskiest landmark in convex geometry here, even though its corollary is formally checked through family 087.

### Tingley's problem

Mazur–Ulam says a distance-preserving map between normed spaces is affine. Tingley asked in 1987 whether knowing only the unit *sphere's* metric already determines the space linearly. If $f$ maps the unit sphere of $X$ isometrically onto that of $Y$, does it extend to a linear isometry? The answer was known for specific families: $\ell_p$, $L_p$, $C^*$-algebras, von Neumann algebras, and every two-dimensional space (Banakh, 2022). It was open even for general three-dimensional spaces. Family 322 claims yes for every pair of real Banach spaces, with no separability, reflexivity or smoothness assumption. The extension is the radial one, $T(x) = \|x\| f(x/\|x\|)$. The proof measures the worst failure to preserve distances between different radii and realizes a hypothetical maximal defect in an ultrapower-type enlargement. It then propagates extremal chords to an invariant convex set and reaches a contradiction with Darbo's fixed-point theorem. Twelve pages, and the full statement is in Lean (`OAI/Analysis/SphereIsometry/Extension.lean:101`). For complex spaces the extension is real-linear, which is the most anyone expected.

### The separable quotient problem is independent of ZFC

Every infinite-dimensional Banach space has a separable infinite-dimensional subspace. Must it also have a separable infinite-dimensional *quotient*? The question goes back at least to Rosenthal's 1969 paper and is usually credited to Banach. Johnson and Rosenthal answered it for separable spaces in 1972, and Argyros, Dodos and Kanellopoulos for duals in 2008. Family 323 says the answer depends on set theory. If the continuum is real-valued measurable, every such space has a separable quotient. Under the continuum hypothesis there is a counterexample of density $\aleph_1$. So, assuming a measurable cardinal is consistent, the problem is independent of ZFC. Only the CH half is in Lean (`OAI/Analysis/SeparableQuotients/Main.lean:9`). It is the striking half, because no counterexample was known in any model of set theory. The positive half from a measurable cardinal is not formalized. 39 pages.

### Lipschitz-equivalent separable Banach spaces need not be isomorphic

If two Banach spaces are the same metric space up to bounded distortion, must they be the same linear space? Aharoni and Lindenstrauss said no for huge non-separable spaces in 1978. The separable case is the central question of nonlinear Banach space theory; Kalton listed it as Problem 3 in his 2008 survey. Godefroy, Kalton and Lancien had shown $c_0$ is determined by its Lipschitz structure. Family 324 claims separable real spaces $X$ and $Y$ with a bijection distorting distances by between $4/21$ and $76/25$. $X$ contains an isometric copy of $c_0(\ell_2)$ and $Y$ contains no linear copy of it. The construction builds a nearly isometric nonlinear change of variables on Hilbert space, linearizes it on the Lipschitz-free space, and assembles graph spaces around it. It is in Lean with the explicit constants (`OAI/Analysis/LipschitzEquivalence/Main.lean:8`), as is a companion result on absorbing $c_0$. 48 pages.

### Kirk's problem: nonexpansive maps on reflexive spaces have fixed points

Banach's fixed-point theorem needs strict contractions. A map that merely does not increase distances need not have a fixed point, and whether it must depends on the space. Browder and Göhde proved it for uniformly convex spaces in 1965, and Kirk under normal structure the same year. The central open question since then: is reflexivity alone enough? Hanebaly published claimed proofs in 2019 and 2021 that were not accepted, and Nourouzi still listed the problem as open this year. Family 328 claims yes. In every real reflexive Banach space, every nonexpansive self-map of a nonempty closed bounded convex set has a fixed point. The Lean statement is exactly the problem:

```lean
-- lean/ComparatorChallenges/ReflexiveFixedPoints.lean:21-29
theorem exists_fixedPoint_of_canonicallyReflexive
    {X : Type u} [NormedAddCommGroup X] [NormedSpace ℝ X] [CompleteSpace X]
    (hX : CanonicallyReflexive X) (C : Set X)
    (hne : C.Nonempty) (hclosed : IsClosed C)
    (hbounded : Bornology.IsBounded C) (hconvex : Convex ℝ C)
    (F : C → C)
    (hF : ∀ a b : C, ‖(F a : X) - (F b : X)‖ ≤ ‖(a : X) - (b : X)‖) :
    ∃ a : C, F a = a := by
  sorry
```

The 19-page proof starts from a minimal fixed-point-free invariant set, where approximate fixed points are almost diametral (Goebel and Karlovitz). It builds a tree of convex combinations whose telescoping differences produce a vector detected along an infinite branch, a vector that weak compactness says must also tend weakly to zero. The machinery is new rather than a classical reduction. Given the history of failed claims, the Lean check is what this one stands on.

### Calderón's problem: one boundary patch, and a buried headline

Calderón asked in 1980 whether voltage-to-current measurements at a body's boundary determine its conductivity, which is the mathematics of electrical impedance tomography. Sylvester and Uhlmann settled smooth isotropic conductivities in dimension 3 and up in 1987. Three harder versions stayed open. Anisotropic conductivities are the Lee–Uhlmann conjecture, a smooth Riemannian metric determined up to diffeomorphism, known in the real-analytic case. Partial data means measuring inputs and outputs on one patch only. And rough conductivities: $L^\infty$ is open in dimension 3 and up, while Astala and Päivärinta did the plane.

Family 365's title is about recovering a metric and a unitary connection together. What it actually claims is bigger. A companion paper claims the smooth anisotropic Calderón problem for $n \ge 3$ from a single boundary patch with inputs and outputs on the same patch. That is the Lee–Uhlmann conjecture together with the same-patch partial-data problem, and as a corollary, smooth scalar uniqueness from one patch. Each would be a landmark. The proof matches harmonic coordinates and Green kernels outward from the patch and extends the isometry across hypersurfaces with an estimate on shrinking spheres. That extension step is exactly where the field has always been stuck for non-analytic metrics. None of it is formalized.

What *is* formalized (`OAI/Analysis/Conductivity/Main.lean:73`, about 75,000 lines) is the opposite direction at low regularity. Two distinct bounded measurable conductivities on a ball in $\mathbb R^3$, both equal to 1 near the boundary, have identical boundary measurements. The trick is a nested fractal of cells on whose surfaces every harmonic function has equal averages. That lets you rescale the conductivity by $\rho^2$ with $\rho = 1 + \delta w$ without changing any boundary response. This contrasts sharply with uniqueness in the plane and for Lipschitz conductivities, and it is a major result on its own. The two directions are consistent: smooth means unique, merely measurable does not.

### The planar Mumford–Shah conjecture, proved for the second time in a week

The Mumford–Shah functional is the classic variational model for image segmentation. Mumford and Shah conjectured in 1989 that a minimizer's edge set is tidy: smooth arcs, crack tips, and triple junctions at 120 degrees. Decades of work (Bonnet, David, Ambrosio–Fusco–Pallara, De Lellis–Focardi) reduced it to excluding one scenario, a bounded island of edge floating in a whole-plane blow-up limit. Family 366 excludes it in 20 pages. It splits the gradient into a field carrying the jumps across the island and a harmonic part, and shows the energy interaction is harmonic in the island's translation. By minimality that interaction is then constant. So the field vanishes, and deleting the island saves length. It also gets local weak-$L^4$ gradient bounds.

Two days before this paper, Francesco Deangelis posted "Solution of the Mumford–Shah conjecture" ([arXiv:2609.26732](https://arxiv.org/abs/2609.26732)), which rules out bounded components with second-variation inequalities. His acknowledgement says the initial strategy and draft were generated with ChatGPT Astra, an OpenAI model. The OpenAI paper cites him. This is a second, independent proof, not the first. The core compact-piece exclusion step is formalized (`OAI/Analysis/MumfordShah/CompactExclusion.lean:354`), but the family has no Comparator config and is not listed in the Lean docs at all. I found that tree by grepping. The reduction and the regularity inputs are not formalized. Interior statement only, absolute minimizers, $C^{1,\alpha}$ arcs.

### The Lane–Emden conjecture

The scalar equation $-\Delta u = u^p$ has positive entire solutions exactly at or above the Sobolev critical exponent (Gidas and Spruck). The Lane–Emden conjecture says the coupled system $-\Delta u = v^p$, $-\Delta v = u^q$ obeys the same threshold, now a hyperbola: no positive solution when $\frac{1}{p+1} + \frac{1}{q+1} > \frac{n-2}{n}$. Liouville theorems like this feed blow-up analysis and a priori estimates across nonlinear PDE. It was proved up to $n = 3$ by Serrin–Zou and Poláčik–Quittner–Souplet, and $n = 4$ by Souplet in 2009. Above that there were only partial regions, as recently as October 2025. Family 370 claims every dimension, plus the weighted Hénon–Lane–Emden version with $|x|^A$ and $|x|^B$ weights, with no symmetry, decay or stability assumption. The method is unconventional. A virial identity is localized with a compactly supported kernel and controlled by an elementary "interval-pair" inequality, giving a uniform localized energy bound that dilation then contradicts. It is 20 pages for a problem open in dimension 5 and up for thirty years, which I would not have believed without the formalization. The full nonexistence statement is in Lean (`OAI/Analysis/HenonEmden/Main.lean:9`). The existence half of Phan's classification is imported from Bidaut-Véron and Giacomini.

### The hot spots conjecture for smooth simply connected domains

Heat an insulated plate and wait. Rauch conjectured in 1974 that the hottest and coldest points drift to the edge. In mathematical terms, the first non-constant Neumann eigenfunction takes its extremes on the boundary. Burdzy and Werner showed this fails for domains with holes (1999). De Dios Pont showed it can fail for convex domains in high dimension ([arXiv:2412.06344](https://arxiv.org/abs/2412.06344)). In the plane, Judge and Mondal did triangles (Annals 2020), and convex or simply connected shapes in general were open. Family 369 claims every bounded simply connected planar domain with smooth boundary. An eigenfunction's gradient never vanishes inside, so even interior saddle points are excluded.

The proof is the most elegant I read in this group. Map the domain conformally to the disk. Rohleder's principle says the gradients of first eigenfunctions are exactly the minimizers of a div–curl energy over tangent vector fields. Multiplying a gradient's boundary trace by a scalar function keeps it tangent, and the energy of the product is a double integral against a positive kernel. The new analytic core is a kernel theorem. For each interior point $p$, the kernel $K(p,s)K(p,t)/N(s,t)$ is conditionally negative semidefinite, which the paper proves through McCullough and Quiggin's characterization of complete Nevanlinna–Pick kernels. Schoenberg's theorem then embeds the circle in a Hilbert space whose coordinates, used as multipliers, sum to $\tfrac12|\nabla u(p)|^2$. At an interior critical point every multiplied gradient would be another eigenfunction gradient. That is more eigenfunctions than Nadirashvili's multiplicity bound of two allows.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig2.png"
  alt="Unit disk with plus signs at left and right boundary points joined by a solid crosscut, and minus signs at top and bottom joined by a dashed path that must cross it."
  caption="The topological step in the hot spots proof, pulled back to the disk: a positive crosscut separates the two negative boundary points, so a negative path joining them would have to cross it (Strict hot spots and absence of interior critical points on smooth simply connected planar domains, Figure 1)."
/>

It is 21 pages and in Lean (`OAI/Analysis/HotSpots/Main.lean:78`, about 29,000 lines). The formal statement assumes the eigenfunction is smooth up to the boundary, which elliptic regularity supplies but the proof does not derive. I grade it major rather than landmark only because corners are excluded. Convex polygons other than triangles are not covered.

### $L \log L$ functions have almost-everywhere convergent Fourier series

Carleson proved in 1966 that Fourier series of $L^2$ functions converge almost everywhere, and Hunt extended it to $L^p$ for $p > 1$. Kolmogorov had shown in the 1920s that it fails for $L^1$. The question since is how close to $L^1$ you can get. Sjölin reached $L\log L \log\log L$, Antonov $L \log L \log\log\log L$, and $L\log L$ itself is the natural conjecture, recorded by Lie. Family 075 claims it, together with a weak-type bound for the Carleson maximal operator on $L\log L$. The method is what caught my eye. Instead of Fefferman or Lacey–Thiele time-frequency analysis, it uses an information budget: entropy compression, Griffiths–Ginibre correlation inequalities and "thermal labels" for where input mass sits. That is either a new idea or a long way round a hole. It is 76 pages with no Lean, and I can't tell which.

### The triangular Hilbert transform at the symmetric point

The triangular Hilbert transform pairs functions $F(x+t, y)$ and $G(x, y+t)$ against $dt/t$. It is the simplest "entangled" singular integral, and it defeats standard time-frequency analysis; no Lebesgue bound of any kind was known. Thiele's problem list asks for $L^3 \times L^3 \times L^3$ (Problem 13). Family 082 claims that point, with the maximal operator over both truncation endpoints bounded from $L^3 \times L^3$ to $L^{3/2}$, plus $r$-variation for $r > 2$ and a dyadic model with constant 40. The method is short and unusual: a matrix-trace energy that decreases under Gaussian heat flow, controlled by a dimension-free trace inequality. The maximal bound is in Lean (`OAI/Analysis/TriangularHilbert/Main.lean:38`). The variation result and the dyadic model are not. Only the symmetric exponent triple is claimed, not the rest of the conjectured range.

### Stein's conjecture for Hilbert transforms along Lipschitz directions

Averaging along short segments whose direction varies from point to point is easy when the direction field is analytic and can fail when it is merely Hölder. Stein asked whether Lipschitz is enough for the singular, Hilbert-transform version, at lengths up to the reciprocal of the Lipschitz constant. Lacey and Li had it conditionally on a Lipschitz–Kakeya maximal bound, and Bateman and Thiele for fields depending on one variable. Family 083 claims a uniform $L^2$ bound for every 1-Lipschitz unit field at a fixed short scale, uniform in the inner truncation, hence a bounded principal value. It is in Lean (`OAI/Analysis/LipschitzHilbert/Main.lean:17`). The short-scale restriction is part of Stein's conjecture as Lacey and Li state it. Do not read this as boundedness at all scales, or as the Zygmund differentiation conjecture. 87 pages of tile-and-tree analysis.

### The Erdős similarity conjecture for geometric sequences

Erdős asked in 1974 whether every infinite set $A$ can be "avoided": is there a set of positive measure containing no scaled and shifted copy of $A$? Sequences converging slowly were handled by Falconer and Eigen in the 1980s. Even $A = \lbrace 1/2, 1/4, 1/8, \dots\rbrace$ was open for fifty years, because its points crowd together exponentially fast and defeat random constructions. Family 084 claims every geometric sequence. For each ratio $q$ and each $\eta$ there is a compact set of measure above $1 - \eta$ avoiding every copy of $\lbrace q^n\rbrace$, with dilations of either sign. The trick is a finite tree of random "routing" tables on nested grids, with sequence indices assigned to tree edges so that each test touches only boundedly many scales. A union bound then becomes affordable. Only the dyadic case $q = 1/2$ is in Lean (`OAI/MeasureTheory/DyadicAvoidance/Main.lean:103`). The general-$q$ paper is dated 5 October, the day before release. The set depends on $q$, and the full conjecture for all infinite sets is untouched.

### A first $L^p$ bound for the trilinear Hilbert transform

Lacey and Thiele tamed the bilinear Hilbert transform in the 1990s. Adding a third function, $\int f_1(x-t) f_2(x-2t) f_3(x-3t)\, dt/t$, brings in quadratic-phase structure that time-frequency methods cannot see. Since then no $L^p$ bound has been known at any exponent. Tao got only sublogarithmic gains on truncations, and Hu and Lie handled curved variants. Family 086 claims the bound $L^3 \times L^3 \times L^3 \to L^1$, using Leng, Sah and Sawhney's quantitative quadratic inverse theorem from additive combinatorics, continuous "quadratic charts" and a scale-summation scheme that avoids a per-scale loss. If it is right, it is historic, and it is graded major only because it covers a single exponent tuple. It is also, in my view, the highest-risk claim in this group's harmonic analysis: 93 pages, no Lean, dated the day before the release.

### Petty's projection conjecture in dimension four and up

The projection body $\Pi K$ records the areas of all of $K$'s shadows. Petty conjectured in 1971 that, at fixed volume, ellipsoids make it smallest. That is an affine isoperimetric inequality stronger than the classical one, with Lutwak–Petty and affine Sobolev inequalities downstream. Family 088 proves it for $n \ge 4$ with equality only for ellipsoids (`OAI/Geometry/ProjectionBodies/Main.lean:11`). It uses spherical-harmonic estimates that kill harmonics of degree 4 and up, and an affine position chosen by a fixed point that kills degree 2. The "all $n \ge 3$" in the summary depends on a separate human result for $n = 3$ by Chen, Feng, Li, Xi and Xu. The companion counterexample, two 10-simplices beating one 20-simplex by the exact ratio 22,355,476/22,020,096, is formalized. It is not the first disproof of Brannen's simplex-maximum conjecture, though: Feng, Hu, Liu and Xu had counterexamples in every dimension from 9 up. The summary's "in contrast" phrasing could leave the wrong impression.

### Covering space with convex bodies costs $\Theta(n \log n)$

Rogers showed in 1957 that translates of any convex body can cover $\mathbb R^n$ with density about $n \log n$. The best lattice coverings were much worse until Li and Liu reached $n \log n$ times a power of $\log\log n$ this July ([arXiv:2607.28429](https://arxiv.org/abs/2607.28429)), with a matching lattice lower bound for symmetric bodies. Family 092 removes the $\log\log$ factor, so a single lattice achieves $C n \log n$ for every body. It also proves the qualitatively new half: some symmetric body needs density above $c n\log n$ even with arbitrary, non-lattice placements. The lower bound uses a random body cut out by many random caps and a Poisson-witness argument. Both halves are in Lean (`OAI/Geometry/CoveringDensity/Main.lean:79`). The lattice half is a sharp but incremental improvement on two-month-old human work.

### Dimension reduction in $L_p$ with $n^{o(1)}$ coordinates

Johnson–Lindenstrauss squashes $n$ points of Euclidean space into $O(\log n)$ dimensions. For other $L_p$ norms it was not known whether you can beat dimension linear in $n$. Brinkman and Charikar showed dimension reduction fails badly at $p = 1$, and Naor and Ren showed this September that $O(\log n)$ is impossible for $p > 2$. Family 094 claims $n^{o(1)}$ dimensions suffice for every $1 < p < \infty$, $p \ne 2$, at any fixed distortion above 1. Exact isometric embeddings need about $n^2$. The construction is a single distribution of random one-dimensional projections, chosen by a minimax argument and sampled with Maurey's empirical method. Its key ingredient is an electrical-flow localization bound of Gurel-Gurevich, Nachmias and Sachdeva ([arXiv:2605.24130](https://arxiv.org/abs/2605.24130)), whose abstract says its initial proofs came from ChatGPT 5.5 Pro. The OpenAI paper re-proves it, and the Lean development (`OAI/Analysis/LpDimension/Main.lean:13`) is self-contained. 19 pages. That makes it an AI result built on an AI-assisted result, with a machine check at the end.

### The Gaussian propeller

Split Gaussian space into $k$ pieces and add up the squared lengths of each piece's unnormalized centre of mass. Khot and Naor conjectured the maximum is $9/(8\pi)$, attained by three 120-degree sectors in a plane, however many pieces and dimensions you allow. The number fixes the optimal approximation ratio for kernel clustering. Heilman, Jagannath and Naor proved it in $\mathbb R^3$ with a computer-assisted argument (2013). Family 096 claims every dimension and every $k$. It reduces to finitely many Gaussian linear scores, uses the $\mathbb R^3$ result for up to four active cells, and rules out five or more with a deletion estimate built on Ehrhard–Borell concavity. In Lean (`OAI/Probability/GaussianPropeller/Main.lean:78`), the four-cell case appears to be re-proved rather than assumed. The NP-hardness corollary in the summary depends on another manuscript in the same release titled "The Unique Games Theorem". That is a separate extraordinary claim, so read the complexity consequence as conditional.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig6.png"
  alt="The plane split into three 120-degree sectors S1, S2, S3, each with a bold arrow showing the direction of its Gaussian centroid."
  caption="The propeller: three 120-degree sectors whose squared Gaussian centroid lengths sum to 9/(8π), the claimed maximum over every partition in every dimension (The Gaussian Propeller Bound in Every Dimension, Figure 1)."
/>

### The Euclidean Steinitz constant is $\Theta(\sqrt d)$

If unit vectors sum to zero, can you order them so the running sum never strays far? Steinitz's lemma says yes, within $d$ for any norm (Grinberg and Sevastyanov). In Euclidean space the conjectured answer is $C\sqrt d$, matching a simplex lower bound; Banaszczyk had $O(\sqrt d + \sqrt{\log N})$. Family 097 claims the bound with no dependence on $N$: one choice of signs keeps every prefix sum of a sequence within $C\sqrt d$, and Chobanyan's transference turns signs into an ordering. The proof adapts the directional-density and symmetrization method that Guo, Fang and Lu used to settle Komlós-type balancing in September ([arXiv:2609.11189](https://arxiv.org/abs/2609.11189)), a proof they credit to an AI research agent. It is in Lean (`OAI/Analysis/Steinitz/Main.lean:111`). It is existential, with no algorithm and no explicit constant, unlike Dutta, Jha and Jiang's efficient bound for $d \ge \log^7 N$.

### A doubling subset of Hilbert space that fits in no $\mathbb R^k$

Lang and Plaut asked in 2001 whether every doubling subset of Hilbert space bi-Lipschitz embeds into some $\mathbb R^k$. It looks finite-dimensional at every scale, after all. Lafforgue and Naor gave counterexamples inside $L_p$ for $p > 2$, and Schioppa's 2017 attempt at the Hilbert case was withdrawn. Family 098 gives one: a subset of $\ell_2$ with doubling constant at most 76,800 that embeds into no Euclidean space at any distortion. A universal version produces such a compact set inside every infinite-dimensional Banach space. The construction stacks orthogonal displacements over a base plane at sparse scales. Differentiating a would-be embedding on "sheets" forces too many separated points into a bounded ball. Thirteen pages, fully in Lean (`OAI/Geometry/DoublingHilbert/Main.lean:57`).

### Edit distance into $\ell_1$: the exponent is right

A low-distortion embedding of edit distance into $\ell_1$ would make nearest-neighbour search over strings fast. Ostrovsky and Rabani achieved distortion $\exp(O(\sqrt{\log d \log\log d}))$ in 2007. The best impossibility result was Krauthgamer and Rabani's $\Omega(\log d)$, an exponential gap that stood for twenty years. Family 099 closes it up to the constant in the exponent: the distortion is $\exp(\Theta(\sqrt{\log d \log\log d}))$, uniformly over alphabets. The lower bound builds hierarchical words from tagged payloads with components cycling at different prime periods, plus a Fourier inequality over cut metrics. Three papers give four independent lower-bound constructions, all in Lean (`OAI/Combinatorics/EditDistance/Main.lean:13`). I count that redundancy as a point in its favour.

### The complete Crouzeix conjecture

Crouzeix conjectured in 2004 that for any square matrix $A$ and polynomial $p$, $\|p(A)\| \le 2 \max_{W(A)} |p|$, where $W(A)$ is the numerical range. In other words the numerical range controls a non-normal matrix almost as well as the spectrum controls a normal one. Crouzeix and Palencia got $1 + \sqrt 2$ in 2017. The scalar conjecture was then solved by humans this summer: Shanmu Jin's preprint, prepared with GPT-5.6 Sol and checked by Greenbaum, Townsend and Crouzeix, and independently Lorist and Schwenninger ([arXiv:2608.03841](https://arxiv.org/abs/2608.03841)). Family 325 proves the stronger "complete" version, where the polynomial's coefficients are themselves matrices, on any Hilbert space including non-separable ones. That was known only for base matrices of order up to 3 ([arXiv:2608.27346](https://arxiv.org/abs/2608.27346)). It is the version operator theory needs: one similarity with condition number at most 2 turns the conformal image of $A$ into a contraction. Two independent proofs (12 and 22 pages) and Lean (`OAI/Analysis/DirectCrouzeix/CompleteBound.lean:65`) make it one of the best-supported results here. The papers themselves are honest about priority. A headline saying "OpenAI solved Crouzeix" would not be.

### Cotype–cotype under the approximation property

A Banach space is K-convex exactly when it does not contain $\ell_1^n$ uniformly (Pisier). The cotype–cotype question asks whether finite cotype of both $X$ and its dual already forces K-convexity. It is false without some approximation structure, by Pisier's own construction, and was conjectured under the bounded approximation property. Family 326 proves it under the ordinary approximation property, which is weaker, so the result is stronger than the conjecture as posed. It shows Walsh transforms of finite-rank operators decay at a rank-independent rate. It is 21 pages and in Lean (`OAI/Analysis/Cotype/Main.lean:26`). Its entropy-duality corollary is the positive counterpart of family 329's counterexample, and the two are consistent.

### Markov type characterizes superreflexivity

Markov type, introduced by Ball in 1992, asks whether a stationary reversible random walk mapped into a space spreads no faster than diffusively. Naor, Peres, Schramm and Sheffield showed superreflexive spaces have it. Naor asked in his Ribe-program survey (2012, Question 4) whether the converse holds. Family 327 says yes. In a non-superreflexive space, James–Enflo finite representability gives a spreading norm in which you can build reversible chains whose displacement grows linearly in time. Fourteen pages, with the full equivalence in Lean (`OAI/Analysis/MarkovType/Main.lean:36`).

### Metric-entropy duality is false in general

Pietsch's duality conjecture (1972), in the geometric form of Artstein-Avidan, Milman and Szarek, says the number of translates of $L$ needed to cover $K$ is comparable, up to universal constants in the logarithm, to the number of translates of $K^\circ$ needed to cover $L^\circ$. It was known when one body is an ellipsoid. Family 329 disproves the dimension-free form. For every pair of constants there is a symmetric body that breaks the duality against the cube. The construction uses symmetric multilinear forms over a large finite field. It is 15 pages, in Lean (`OAI/Analysis/MetricEntropy/Main.lean:48`), and it does not touch the ellipsoid or K-convex cases.

### Ball's extension problem from Hilbert space into $\ell_1$

Ball's 1992 extension theorem extends Lipschitz maps from a Markov-type-2 source to a Markov-cotype-2 target, losing a constant factor. Whether $\ell_1$ works as a target was open, because $\ell_1$ fails Ball's linear cotype condition. Mendel and Naor asked in 2013 whether $\ell_1$ has the right *metric* Markov cotype. Family 332 says yes, with $N_2(\ell_1) \le 12\sqrt{21}$. So every Lipschitz map from a subset of Hilbert space into $\ell_1$ extends to all of Hilbert space at constant-factor cost. The proof is nine elementary pages: cut decompositions, the cubic smoothing $\varphi(r) = 3r^2 - 2r^3$ and a fourth-power martingale potential. It credits the cut-smoothing idea to a human preprint from September (Cheng, Wang and Xiang). It has no Lean and is dated the day before release, but it is short enough that an expert could check it in an afternoon.

### Boltzmann's equation: two solutions from one initial gas

DiPerna and Lions built global solutions of the Boltzmann equation in 1989. They are so weak that nobody knew whether they are determined by their initial data, and Gismondi, Golding and Novack conjectured this August that they are not. Family 363 confirms it for hard spheres on the torus. One carefully built initial density, cold jets on a hot background concentrated at ever smaller scales, admits two distinct global solutions that keep entropy dissipation. The summary's headline goes further: both solutions satisfy *exact* local conservation of mass, momentum and energy. That comes from a 78-page paper dated 5 October that is not formalized. Lean covers the earlier, weaker renormalized-solution version (`OAI/MathematicalPhysics/Boltzmann/Main.lean:13`, about 160,000 lines). That is still a genuine nonuniqueness theorem, just not the headline one.

### The critical dimension of the one-phase Bernoulli problem is 7

Free boundaries of Alt–Caffarelli minimizers are smooth away from a singular set whose size is governed by $d^*$, the first dimension with a non-flat minimizing cone. Jerison and Savin showed $d^* \ge 5$, and De Silva and Jerison found a non-flat cone in $\mathbb R^7$ in 2009, so $d^* \in \lbrace 5, 6, 7\rbrace$. Family 367 claims $d^* = 7$. It is the free-boundary analogue of Simons' theorem that makes 8 critical for minimal surfaces. It uses a Simons-type stability argument with a hand-built test function and vector field, certified by exact rational tensor inequalities. The catch is in the formalization. Lean covers only the existence of the non-flat cone in $\mathbb R^7$ (`OAI/Analysis/BernoulliCone/Main.lean:10`), which is the 2009 half. The new part, flatness in dimensions 5 and 6, is not formalized.

### The Ball–Evans approximation problem in three dimensions

Nonlinear elasticity models deformations as Sobolev homeomorphisms, and both numerics and variational methods want to approximate them by smooth injective maps without losing derivative accuracy. Convolution destroys injectivity. The plane was settled by Iwaniec, Kovalev and Onninen and by Hencl and Pratelli; dimensions 4 and up have counterexamples. Hencl's 2025 survey ([arXiv:2502.01336](https://arxiv.org/abs/2502.01336)) called dimension 3 "wildly open". Family 368 claims every $1 \le p < \infty$ in $\mathbb R^3$, for bounded domains with no boundary regularity, across two self-contained papers. The constructions count how image curves cross a target triangulation and rebuild the derivative on patches of rank 3, 1 and 2. It is 192 pages with no Lean, which is exactly the kind of intricate geometric construction where gaps hide.

### Stable blowup for a defocusing Schrödinger equation

For the defocusing (repulsive) nonlinear Schrödinger equation, conserved energy stops controlling regularity in high dimension. Merle, Raphaël, Rodnianski and Szeftel showed smooth data can blow up (Inventiones 2022), but only on a thin, non-generic set of data. Family 371 finds blowup that survives every small high-regularity perturbation. There is an open set of data on the 12-dimensional torus that blows up in self-similar fashion, so random data blow up with positive probability. It works in the large-power limit, where the profile becomes a flat core plus a free exterior solved by confluent hypergeometric functions with exactly certified spectra. The Lean development is enormous, about 191,000 lines (`OAI/MathematicalPhysics/DefocusingNLS/MainTheorems.lean:20`). One overstatement: the summary says "a sufficiently large odd power", but the theorem and the Lean give *some* odd power that can be taken arbitrarily large, not every large one.

### Calderón's problem for isotropic elasticity

The elastic version of Calderón's problem asks whether the boundary displacement-to-traction map determines the two Lamé moduli inside. Nakamura and Uhlmann claimed global uniqueness in 1994 and had to restrict it to near-constant moduli in a 2003 erratum. Family 372 claims the general smooth case in $\mathbb R^3$, with no analyticity or closeness assumption. The paper diagnoses where the earlier proofs broke, which is a good sign. Its key step makes a comparison matrix vary holomorphically over a projective null conic, where the absence of global holomorphic sections forces it to be the identity. It is 24 pages, fully in Lean (`OAI/MathematicalPhysics/Elasticity/Uniqueness.lean:94`), and gives uniqueness only, with no reconstruction or stability.

### Infinity-harmonic functions are $C^{1,\alpha}$ in every dimension

The infinity-Laplacian describes the best Lipschitz extension of boundary data and governs tug-of-war games. Savin proved solutions are $C^1$ in the plane (2005), Evans and Savin $C^{1,\alpha}$ (2008), and Evans and Smart everywhere differentiable in every dimension (2011). Whether the gradient is continuous in dimension 3 and up has been open since then. Aronsson's example caps the exponent at $1/3$. Family 377 claims a uniform $C^{1,\alpha_d}$ estimate for every $d \ge 3$, with a non-explicit exponent. Assuming no power-rate affine approximation holds, it rescales thin fitting cylinders anisotropically into an entire non-affine solution of a limit equation. A new Liouville theorem, built on an extremal "horizontal Jensen defect" and a fair-choice drift inequality, says no such solution exists. It is 46 pages, the newest paper in the group (4 October), and has no Lean.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig4.png"
  alt="Left, a thin tilted cylinder of axial radius r and transverse radius root-lambda times r around a point z; an arrow labelled stretch transverse directions leads to a sheared parallelogram box on the right."
  caption="The rescaling at the core of the infinity-harmonic paper: thin fitting cylinders are stretched until a small tilt becomes a finite shear, producing the limit equation the new Liouville theorem rules out (Uniform Interior C^{1,α} Estimates for Infinity-Harmonic Functions, Figure 1)."
/>

### Nine smaller results worth knowing

These are real results, graded notable rather than major because they are a special case, a sharpening of recent human work, or an engineered counterexample.

#### Schrödinger convergence at the exact endpoint (080)

Carleson asked how smooth initial data must be for $e^{it\Delta} f \to f$ almost everywhere. Bourgain showed $s \ge n/(2(n+1))$ is necessary. Du, Guth and Li (planar, 2017) and Du and Zhang (higher dimensions, 2019) proved everything strictly above it. These two papers claim the endpoint itself, through a maximal estimate with $L^1$ output. That is real but narrow, and it took 179 unformalized pages.

#### Centered disk maximal function in $W^{1,1}$ (085)

It answers Hajłasz and Onninen's 2004 question for centered disks in the plane: the gradient of the maximal function is controlled by the gradient of $f$ in $L^1$. A "tangent-disk sweep" carries the argument. In Lean (`OAI/Analysis/DiskMaximal/Main.lean:144`). Higher dimensions and uncentered versions are not claimed.

#### Dimension-free log-Sobolev for subgaussian log-concave measures (093)

It proves Bizeul's 2023 conjecture that subgaussian linear marginals give a dimension-free log-Sobolev inequality, by a 65-page contradiction argument using Gaussian-channel identities. No Lean.

#### Cylinder coverings below the half-area bound (100)

Bang's plank theorem has a three-dimensional cylinder analogue. Cylinders covering a body were conjectured to need total cross-section at least half its smallest shadow, with two cylinders on a regular tetrahedron as the equality case. Tilting into many nearby directions saves normalized area $\tfrac{13}{6000}\epsilon^2$ to leading order. At the largest allowed $\epsilon = 1/2000$ that is about $5 \times 10^{-10}$, bought with $2\lceil 2/\epsilon^2\rceil$ = 16,000,000 cylinders. That refutes the $1/2$ constant and leaves the truth somewhere between Bezdek and Litvak's $1/3$ and just under $1/2$. In Lean.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig7.png"
  alt="Left, a regular tetrahedron with two cylinders parallel to opposite edges covering its lower and upper halves; right, a plot of neighbouring angular sectors whose sides meet on a common supporting line after a small tilt."
  caption="Bang's two-cylinder cover of the regular tetrahedron (left) and the small tilt that lets many nearby directions share boundaries and save a second-order amount of area (Finite angular cylinder covers below the half-area bound, Figure 1)."
/>

#### Lipschitz-free spaces: AP without BAP (330)

It answers Kalton's question negatively. A uniformly discrete metric space whose free space has the approximation property but not the bounded one, and as a corollary an equivalent norm on $\ell_1$ failing the metric approximation property. It sharpens Smith's 2025 discrete example. In Lean.

#### Midpoint convexity and diamonds (331)

Seven papers, 225 pages, one theme. The main result is a separable reflexive space that is asymptotically midpoint uniformly convex but has no equivalent AUC norm, with diamond graphs embedding only with growing distortion. This answers Baudier and Lancien's Problem 39. The official summary gets the consequence backwards. It says midpoint convexity "does not force uniformly bounded diamond distortion", but midpoint convexity already *prevents* uniform diamond embeddings. The new point is that failing AUC-renormability does not force them, even in reflexive spaces. Baudier did the non-reflexive version himself five days before the earliest OpenAI paper. In Lean.

#### Boltzmann–Grad limit over the regular lifespan (364)

Deng, Hani and Ma derived the Boltzmann equation from hard-sphere dynamics for as long as the kinetic solution stays regular ([arXiv:2408.07818](https://arxiv.org/abs/2408.07818)), and listed smooth potentials as open. This extends their result to stable radial potentials with attractive wells, and proves Gaussian fluctuations for hard spheres. Like theirs, it is conditional on a regular Boltzmann solution existing. 123 pages, no Lean.

#### The three-marginal Coulomb Monge problem has no Monge minimizer (373)

In the strong-interaction limit of density functional theory, electrons are assumed to sit at positions that are functions of one electron's position. This builds a smooth density for three electrons in $\mathbb R^3$ where no such map is optimal. Every optimal plan must split the central mass between two equally good families, but each outer region holds only half the mass a deterministic choice would send there. The density is engineered, not physical, and the paper says so. Fourteen pages, in Lean.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/analysis-pde-fig3.png"
  alt="A central box B near 0 with mass one third, and four outer boxes near plus and minus e1 and e2, each with mass one sixth, connected by arrows labelled Y1, Y2, W1, W2."
  caption="The engineered density behind the Coulomb Monge counterexample: any deterministic map from the central piece would push mass one third into a region that only holds one sixth (A counterexample to the Monge ansatz for the three-marginal Coulomb cost, Figure 1)."
/>

#### Brenier maps are $1/3$-Hölder stable, and no better (374)

For a uniform source on a convex body, optimal transport maps satisfy $\|T_\mu - T_\nu\|_{L^2} \le C\, W_2(\mu, \nu)^{1/3}$ uniformly over targets. Tilting the middle piece of a maximum of three affine functions shows $1/3$ is sharp, which disproves Letrouit's conjectured $1/2$. Both halves are in Lean. This one is useful for anyone estimating transport maps from samples.

### One technical result

#### Turing machines in forced Navier–Stokes flows (376)

Nine papers, 258 pages. For any Turing machine and input there is a smooth external force on the 3-torus such that the Navier–Stokes solution from rest is globally smooth, and a tagged fluid particle enters a fixed region exactly when the machine halts. The papers admit the catch on page 2. Any smooth divergence-free velocity becomes an exact Navier–Stokes solution once you define the force as its material acceleration minus the viscous term. So the content is a kinematic construction of incompressible flows that simulate a machine, not anything about unforced fluids or Tao's blowup programme. Parts are in Lean, including a compactly supported version in $\mathbb R^3$.

### What I would tell a colleague

If you read one thing from this part of the release, read the Mahler family. Its main statements are short enough to audit by eye in Lean. It has two unrelated proofs of the symmetric case, and the symplectic one is a genuinely new idea. Next I would read Vlasov–Maxwell, Falconer and the hot spots paper. Each has a mechanism you can follow in an hour and a formal statement behind it.

The results I would hold at arm's length are the long, unformalized ones that cite each other: the Kakeya–restriction–Bochner–Riesz–local-smoothing chain, the trilinear Hilbert transform and De Giorgi in dimension eight. They may well be right. The release has produced so many checkable results in this area that a prior of "probably wrong" no longer fits. But nothing yet stands behind those papers except the papers themselves. The priority cases deserve care too. Kakeya in $\mathbb R^3$ is Wang and Zahl's, scalar Crouzeix is Jin's and Lorist–Schwenninger's, and Mumford–Shah was posted two days earlier by Deangelis with an OpenAI model's help. Any headline that credits these to this release alone is wrong.

## Combinatorics

Combinatorics is where the release is most crowded with names I grew up on. The catalogue lists 37 combinatorics families built from 50 manuscripts, and reading the titles felt like reading a list of problems I had been told would outlive me: Hadwiger, Harary–Hill, Zarankiewicz, Barnette, Seymour, Sidorenko, Ryser, the chromatic number of the plane, Erdős's \$5000 progression problem.

Two things surprised me. The first is how short some of these papers are. Barnette's conjecture takes 11 pages, the circulant Hadamard conjecture 15, Seymour's second-neighbourhood conjecture 15, and the crossing number of $K_n$ 13. The second is how much of the combinatorics comes with a Lean statement. 27 of the 37 families have a Comparator statement that matches the paper's headline theorem. Six more are partly formalized, and four have nothing at all. That is a much better ratio than the release as a whole. It also makes the four unformalized ones stand out, and the partial ones need careful reading: the headline claim is sometimes exactly the part Lean does not cover.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig8.png"
  alt="A page of the OpenAI research catalogue listing combinatorics families 157 to 168, each with a bold title, a one-paragraph summary and links to its manuscripts."
  caption="How the release presents its combinatorics: one catalogue paragraph per family, here 157 (Hadwiger) through 168 (Kazhdan–Lusztig) (openai/math overview.pdf, page 18)."
/>

I graded 13 families landmark, 19 major and 5 notable; none of them looked merely technical to me. By kind, 24 are proofs, 8 are counterexamples and 5 are improved bounds. The order below is my own, with the most important first. It weights what is claimed against how much of it has a machine-checked statement.

### The plane is not five-colourable

How many colours do you need so that no two points at distance exactly 1 get the same colour? Since the 1950s the answer has been known to lie between 4 and 7 (Nelson, Isbell, the Moser spindle), and it stayed there until Aubrey de Grey's 1581-vertex graph pushed the lower bound to 5 in 2018 (arXiv 1804.02385). Polymath16 shrank such graphs but never found a 6-chromatic one.

Family 158 claims the lower bound 6 for arbitrary colourings, measurable or not, so $6 \le \chi(\mathbb R^2) \le 7$. What surprised me is that it gets there without exhibiting a finite graph. A finite 6-chromatic unit-distance graph must exist by de Bruijn–Erdős, and the paper says plainly that it does not produce one. The route has two halves. First, any proper $k$-colouring can be averaged over algebraic rotations and translations, and a rigidity theorem for invariant measures turns it into a measurable colouring that is proper almost everywhere. Going from arbitrary to measurable colourings was always the obstacle: Falconer proved measurable colourings need 5 colours back in 1981, and that never transferred. Second, for measurable colourings, at most two colours meet along any interface. Planar topology gives a cycle of colours around a point, and cycles of length 3, 4 and 5 are ruled out, the last by placing a Moser spindle inside a region that only three colours can reach.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig3.png"
  alt="A unit-distance graph with seven labelled vertices drawn over shaded circular regions: two blue unit rhombi sharing an edge, a dashed rotated copy, and a highlighted orange unit edge between z_T and z_uT."
  caption="The last step: a Moser spindle placed so every vertex lies in a region restricted to three colours, which forces a contradiction on the orange unit edge (The Euclidean plane is not five-colorable, Figure 3)."
/>

The Lean statement is clean: no function `ℂ → Fin 5` separates every pair at distance 1 (`EuclideanFiveColor.lean:10`), with no measurability assumption anywhere. A second statement formalizes the seven-colouring upper bound. Both are in the main-results list. My one reservation is size. About 15k lines seems small for the ergodic-rigidity machinery the paper describes, so how much of it comes from Mathlib is the first thing I would check. Wikipedia's Hadwiger–Nelson page already records the claim as not yet independently verified.

### Barnette's conjecture

Tait tried to prove the four-colour theorem in 1884 by claiming that every cubic 3-connected planar graph has a Hamiltonian cycle. Tutte killed that in 1946. Barnette's conjecture, recorded by Grünbaum in 1969, is the surviving piece: add "bipartite" and the claim should hold. Before this release it was known for graphs whose faces have size at most 8 (Schnieders, arXiv 2508.03531) and checked by computer for small orders.

Family 180 claims all of it in 11 pages. The proof dualises. A cubic bipartite planar graph becomes a sphere triangulation whose faces alternate black and white, and a Hamiltonian cycle corresponds to splitting the triangulation's vertices into two sets that each induce a tree. Fix one black triangle as the outer face. A "state" assigns every other black triangle one of its corners, using each non-root vertex exactly once. Take pairs of states that disagree everywhere. For the second state $s$, collect the edge opposite its chosen corner in each triangle; the goal is a pair where those edges form a forest.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig1.png"
  alt="Two triangles. Left: triangle t with corners r(t), s(t), u(t), a solid blue arrow from r(t) to s(t) and a dashed arrow from r(t) to u(t). Right: a triangle with a centre point t and a shaded sub-region, a dashed edge v to w replaced by a solid two-spoke route v to t to w."
  caption="The local moves in the Barnette proof: a pair of states picks a directed edge inside each black triangle (a), and the spoke detour (b) is the route the positive-circulation weights live on (Paired states and Hamiltonian cycles in cubic bipartite planar graphs, Figure 1)."
/>

The existence argument is the clever part. Build a finite exponential sum over all pairs. Grouped by the second state, every pair whose edge set has a cycle cancels against the pair obtained by reversing that cycle, because a "disk identity" (any directed cycle in a state's chosen edges encloses a fixed signed count of black faces) makes their phases opposite. Grouped instead by the union of the two choices, positive circulations on spokes make the sum nonzero. So some pair survives with a forest, and the forest becomes the two induced trees. It is non-constructive: there is no algorithm here.

The formal statement is the thing to look at, because planarity is easy to get wrong in Lean:

```lean
-- lean/ComparatorChallenges/BarnetteHamiltonian.lean:28-30
/-- Planarity means existence of a crossing-free topological plane embedding. -/
def Planar {V : Type u} (G : SimpleGraph V) : Prop :=
  Nonempty (PlaneEmbedding G)

-- lean/ComparatorChallenges/BarnetteHamiltonian.lean:41-49
/-- Cubic bipartite three-vertex-connected plane graphs have a Hamiltonian cycle. -/
def MainStatement : Prop :=
  ∀ (V : Type u) [Fintype V] [DecidableEq V] (G : SimpleGraph V)
    [DecidableRel G.Adj],
    G.IsRegularOfDegree 3 → G.IsBipartite → Planar G →
      ThreeVertexConnected G → HasHamiltonianCycle G

theorem main : MainStatement.{u} := by
  sorry
```

`PlaneEmbedding` (lines 10-26) asks for injective continuous paths in $\mathbb R^2$ whose interiors are pairwise disjoint, so planarity is topological. A Hamiltonian cycle is a single spanning cycle (the comment in the file is explicit that a 2-factor does not count), and the solution's `Main.lean` builds the combinatorial dual from the topological embedding itself (`exists_exact_dual`), importing only Mathlib. That step, from curves in the plane to faces, is where I expected an assumption to hide. I did not find one. The family is not in the main-results list.

### Sidorenko's conjecture is false

Sidorenko's conjecture (1993) says that among graphs with edge density $p$, the random graph minimises the density of every bipartite pattern $H$: $t(H,G) \ge p^{e(H)}$. It underpins a lot of extremal and quasirandom graph theory and had been proved for trees, even cycles, hypercubes, and many structured families (Hatami; Conlon, Fox and Sudakov; Kim, Lee and Lee; and others). Computer searches in 2025 found nothing.

Family 161 names a specific $H$: the point–triple incidence graph of 22 triples on 13 points in which every covered pair of points lies in exactly two triples, so 35 vertices and 66 edges. It claims a finite host $G$ with $t(H,G) \lt p(G)^{66}$. The host is not explicit. It is sampled from a kernel built from differences of half-rank symmetric matrices over $\mathbb F_q$, in a large fixed dimension, with $q \to \infty$. Each pair lying in exactly two triples is what lets a sign model produce a negative four-way correlation, and the Lagrangian geometry of symmetric-matrix graphs controls the rank-deficient terms. The forcing conjecture falls as a corollary.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig7.png"
  alt="A dependency flowchart: a fixed triple system feeds activation supports, a transverse contribution and a weighted singular bound; these lead through a rank-layer limit to a negative averaged coefficient, then dilution and a crossed product, then sampling a finite host graph G."
  caption="How the Sidorenko counterexample is assembled; Sections 4 to 7 are the delicate singular-configuration bounds (A counterexample to Sidorenko's conjecture, Figure 1)."
/>

You cannot check this one by computing: there is no graph to count in. The Lean statement hard-codes the 22 triples and asserts strict inequality against `edgeDensity ^ 66` (`SidorenkoCounterexample.lean:41`), which makes the claim verifiable in principle. The forcing-conjecture part is not formalized, and the family is missing from the main-results list.

### Seymour's second-neighbourhood conjecture

In 1990 Seymour conjectured that every oriented graph (no loops, no 2-cycles) has a vertex with at least as many vertices at distance exactly two as at distance one. Socially: someone has at least as many friends-of-friends as friends. Fisher proved it for tournaments in 1996. The best general ratio was about 0.657 (Chen, Shen and Yuster, 2003), raised to roughly 0.7155 more recently, and a 2026 computer-assisted preprint handled minimum out-degree 7 (arXiv 2606.30588).

Family 173 claims the full conjecture in 15 pages. Take a minimal counterexample, first by order and then by arc count. Deleting arcs carefully shows every nonempty proper vertex set $S$ satisfies $|F^2(S) \setminus F(S)| \lt |F(S) \setminus S|$, where $F$ is the out-neighbourhood map. Blowing every vertex up into a large transitive tournament turns that into a strict inequality $|U| + |F^2(U)| \lt 2|F(U)|$ for all proper $U$. An extremal argument over two families of vertex pairs then contradicts it, using a new pruning lemma for arbitrary binary relations whose cost depends on the points left uncovered.

The Lean statement (`SeymourSecondNeighborhood.lean:34`) is twelve lines of definitions and one theorem. Distance two is defined as "not equal to $v$, not an out-neighbour, reachable in two steps", which is exactly the conjecture. If I had to pick one result in this group for a human referee to read first, it would be this one: it is short, elementary, and the formal statement leaves nowhere to hide.

### The Erdős–Gallai cycle decomposition

Every graph on $n$ vertices should split into $O(n)$ edge-disjoint cycles and single edges (erdosproblems.com #184). Erdős, Goodman and Pósa recorded the problem with an $O(n \log n)$ bound in 1966. Conlon, Fox and Sudakov reached $O(n \log \log n)$ in 2014 and Bucić and Montgomery $O(n \log^* n)$ in 2022. Family 181 claims the linear bound.

I like the diagnosis in the introduction. The previous argument lost $\log^* n$ because it paid a cost proportional to the full graph at each of many density scales. The fix is to find one scale layer whose expanding pieces are big enough to pay for a whole prefix of scales, identify pairs of low-degree vertices to shrink the graph, solve the quotient by induction, and lift each quotient cycle back by routing through reserved edges. The bookkeeping uses $\phi(n) = \max(1/2, 1 - 1/\sqrt{\log n})$ so that the $O(1/\log D)$ overlap cost is absorbed. The statement `OAI.ErdosGallai.erdos_gallai` (`CycleDecomposition.lean:31`) is formalized and listed; the constant is not explicit.

### Circulant Hadamard matrices and Barker sequences

A circulant Hadamard matrix is a $\pm 1$ sequence whose cyclic shifts are mutually orthogonal. Ryser conjectured in 1963 that only order 4 works (plus the trivial order 1). Turyn showed any other order is $4u^2$ with $u$ odd and not a prime power, and later work left 4,489 candidate orders below $4 \cdot 10^{30}$. The conjecture also decides the even case of Barker sequences, the low-autocorrelation codes radar engineers have used since the 1950s. The problem has a long history of claimed proofs that did not survive, several of them on arXiv.

Family 179 claims it in 15 pages: a group-ring form of Turyn's 2-adic descent, then, for $u > 1$, an alternating product of character values in $\mathbb Z[i][C_{u^2}]$ that must be a root of unity of odd order, compared against residues at a prime above 2. The corollary is that Barker sequences exist exactly at lengths 2, 3, 4, 5, 7, 11 and 13. The circulant statement is formalized and listed (`CirculantHadamard.lean:24`), and so is the even-length Barker statement. The odd-length list rests on the classical Turyn–Storer theorem, which is cited rather than formalized. Given the problem's track record, the Lean statement is what makes me take this one seriously.

### Combinatorial invariance of Kazhdan–Lusztig polynomials

Kazhdan–Lusztig polynomials encode the singularities of Schubert varieties. Lusztig and Dyer conjectured in the 1980s that a polynomial $P_{u,v}$ depends only on the Bruhat interval $[u,v]$ as an abstract poset. This is the conjecture DeepMind's 2021 Nature paper was built around. It was known for lower intervals, short intervals, elementary intervals (Barkley and Gaetz), and all intervals of length at most 6 (Barkley, Gaetz and Lam, arXiv 2601.07793).

Family 168 claims it for arbitrary Coxeter systems, infinite and non-crystallographic ones included, in 31 pages. The written proof transports a reflection order from one system to an edge schedule on the other, bounds how much each removed edge condition can grow kernel modules of a moment-graph sheaf, and notices that the product of these bounds equals Dyer's path formula for the $R$-polynomial. Reciprocity forces every bound to be tight.

There is a wrinkle the reviewers of this release should look at. The paper cites the Elias–Williamson character theorem and Braden–MacPherson sheaves as inputs, but comments inside the Lean development (`Frobenius.lean`, `Decomposition.lean`) say those are not formalized. Yet the formal statement (`KLInvariance.lean:88`) is proved with only the standard axioms. So the formal proof cannot be following the paper's route as written, and someone should work out what it does instead. The family is not in the main-results list.

### Ryser's conjecture fails

Ryser's conjecture generalises König's theorem: an $r$-partite hypergraph has cover number at most $(r-1)$ times its matching number. For intersecting hypergraphs that means $r-1$ vertices always cover all edges. Aharoni proved $r = 3$ in 2001, and the intersecting case was known for $r \le 5$.

Family 162 claims counterexamples: intersecting $(q+1)$-partite hypergraphs with cover number $q+1$ for every large prime $q$, and a second family at ranks $s^n + 1$. The construction perturbs the affine plane over $\mathbb F_q$ by randomly splitting and merging lines, and uses the Szőnyi–Weiner stability theorem to rule out small covers. Gyárfás's tree-cover conjecture falls with it. Both main theorems are formalized (`BalancedRyser.lean:39`, `RyserOddExtensions.lean:61`). The thresholds are not explicit, so nobody knows the smallest rank at which Ryser actually fails. Small ranks are untouched.

### Clique-free graphs: AEKS and Alon–Krivelevich–Sudakov

Triangle-free graphs of average degree $d$ have independent sets of size about $n \log d / d$ (Ajtai, Komlós and Szemerédi, then Shearer). Excluding $K_r$ for $r \ge 4$ should give the same, which is the Ajtai–Erdős–Komlós–Szemerédi conjecture from 1981 (erdosproblems.com #802). For 45 years every argument lost a $\log \log d$ factor. The colouring analogue, $\chi = O(\Delta / \log \Delta)$ for $K_r$-free graphs, is the Alon–Krivelevich–Sudakov conjecture.

Family 184 claims both. The independence paper (21 pages) maximises an entropy-like functional $\sum_v w_v(1 - \log w_v)$ minus edge mass, shows its triangle mass is small by splitting vertices along coupled random walks, and reads off a large independent set. It is formalized with a faithful statement (`CliqueFreeLog.lean:11`) and listed. The colouring paper, which proves the stronger correspondence-colouring version, is dated October 5, two days before release, and has no Lean. The family title leads with the unformalized half.

### Hindman's finite sums and products

Colour the positive integers with finitely many colours. Hindman's 1979 conjecture says you can find $a_1 \lt \dots \lt a_m$ whose subset sums and subset products all share one colour. Even $m = 2$, a monochromatic $\{x, y, x+y, xy\}$, was open for general colourings. Moreira got $\{x, x+y, xy\}$ in 2017, and Bowen and Sabok, then Alweiss, did the full statement over the rationals.

Family 164 claims the integer case in 70 pages with no Lean at all. It models each colour class by piecewise nilsequences and chooses rational scale adjustments from a fixed finite list so that additive shifts do not destroy the predicted colour. It leans on corrected or revised versions of deep inputs: a 2024 erratum to the Green–Tao–Ziegler inverse theorem and Green–Tao's quantitative Leibman theorem with its erratum (Section 6.1 is literally titled "The corrected quantitative equidistribution input"). I treat it as unverified until someone in the area reads it.

### Hypercube Ramsey numbers are linear

The cube $Q_n$ has $2^n$ vertices and degree $n$, so the general theorems that give linear Ramsey numbers for bounded-degree graphs blow up. Burr and Erdős asked in 1975 whether $R(Q_n) = O(2^n)$ anyway. The best bound was Tikhomirov's $2^{(2-c)n}$ from 2022.

Family 171 claims $R(Q_n) \le C \cdot 2^n$ in a single 172-page argument by contradiction: a colouring with no monochromatic cube must be almost perfectly pseudorandom, and pseudorandomness itself lets you embed a cube by tiling it into subcubes and matching through Hall's theorem. Going from exponent $2-c$ straight to linear in one unformalized step is a very large jump, and this is one of the four families with no Lean. I put it in the same bucket as the Hadwiger counterexample: important if true, and nothing but the manuscript to go on.

### Borsuk's conjecture fails in dimension 9

Can every bounded set in $\mathbb R^d$ be cut into $d+1$ pieces of smaller diameter? Kahn and Kalai said no in dimension 1325 in 1993. The record dropped to 65, then 64 (Jenrich and Brouwer, 2014), then 63 in 2026 (Grinsztajn with GPT-5.5 Pro, arXiv 2608.12561). Family 156 claims 9.

The witness is a classical object: unit vectors $u \in \mathbb R^4$ mapped to projectors $uu^\top$, which live in the 9-dimensional space of trace-one symmetric $4 \times 4$ matrices, with diameter $\sqrt 2$ attained exactly at orthogonal lines. The obstruction is topological, not Frankl–Wilson counting. A cover by 10 smaller pieces gives a labelling of $\mathbb{RP}^3$ in which orthogonal lines never share a label, a mod-2 degree argument forces a rigid tetrahedral structure, and a finite case analysis kills it. That analysis is in Lean as 145 generated case files, and the full statement (`BorsukNine.lean:25`) is listed. Short, formal, and built on a set everybody knows: I find this one easy to believe. It does not find the smallest failing dimension, which is now somewhere between 4 and 9.

### Periodic tiling fails in dimension 3

The periodic tiling conjecture said any finite tile that tiles $\mathbb Z^d$ by translations also tiles it periodically. Greenfeld and Tao disproved it in 2022 in some huge, unspecified dimension (Annals 2024). It holds in dimensions 1 and 2. Family 155 claims a counterexample in $\mathbb Z^3$, and in $\mathbb R^3$ via the cube thickening, even with arbitrary real translations, so 3 is the least dimension where it fails.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig4.png"
  alt="A four-box flow from left to right: word arrays with nonconstant columns, a common solution in Z squared times a cyclic group V, one tile F in Z squared times Z mod QZ, and one finite tile T with its cube thickening. A return arrow below reads: a fully periodic tiling would give a forbidden vertical period."
  caption="The construction runs left to right; any periodic tiling, decoded right to left, would produce a forbidden vertical period (A translational tile with no fully periodic tiling in dimension three, Figure 1)."
/>

The method is Greenfeld–Tao's "p-adic Sudoku" with two new moves: make the finite factor cyclic and then unfold it into a third integer coordinate. The Lean statement is stronger than the paper's theorem, since it includes minimality and therefore re-proves the positive results in dimensions 1 and 2. The tile is presumably astronomically large; no size is given.

### Unit and pinned distances in the plane

This is the family where the release talks to OpenAI's own earlier result. In May 2026 an OpenAI model disproved Erdős's unit distance conjecture, and Sawin pushed the construction to $u(n) \ge n^{1.014}$ (arXiv 2605.20579, with a digested account by Alon, Bloom, Gowers and others in arXiv 2605.20695). The upper bound was still Spencer–Szemerédi–Trotter's $n^{4/3}$ from 1984. Family 167 claims $u(n) \le C n^{\beta}$ for some $\beta \lt 4/3$, and separately that all but $o(n)$ points of any planar set see at least $n^{1-\epsilon}$ distinct distances, which settles Erdős #604 in its $n^{1-o(1)}$ form (Katz and Tardos had exponent 0.864).

The shared idea is arithmetic. Specialise the configuration into a number field, write squared distance as $(z - z')(w - w')$ with $z = x + iy$ and $w = x - iy$, and use the product formula across all places to turn many repeated distances into contradictory height constraints. That has to be the shape of any proof: $n^{4/3}$ is tight for unit circles in other norms, so a power saving must use Euclidean arithmetic somewhere. Both theorems are formalized, but neither exponent is explicit. One documentation slip: `lean/docs/167.md` says the unit-distance theorem "is not included" and then links its Comparator file.

### Distinct distances in every dimension

Family 166 claims that $n$ points in $\mathbb R^d$, $d \ge 3$, determine at least $c_d \, n^{2/d}$ distinct distances, sharp up to the constant. That is stronger than the $n^{2/d - o(1)}$ form of the conjecture (erdosproblems.com #1083). It appeared six weeks after Tidor, Yu and Zakharov's 147-page $n^{2/3 - o(1)}$ result for $\mathbb R^3$ (arXiv 2608.14454), uses their kind of flat geometry, and claims more than they did, in every dimension, in 103 pages, with no Lean. The paper also mentions a withdrawn earlier constant-factor claim by someone else. This is the family I trust least among the majors.

### Strong thin trees

Goddyn's thin tree conjecture asks whether high edge connectivity forces a spanning tree that uses only a small fraction of every cut. The strong form asks for a $C/k$ fraction in any $k$-edge-connected graph, which is the right order. Anari and Oveis Gharan (2015) got $\mathrm{poly}(\log \log n)/k$. Family 174 claims $C/k$ and, in a companion paper, a deterministic polynomial-time algorithm. The method keeps a large packing of disjoint spanning trees while halving the graph, combining effective-resistance hierarchies with a Marcus–Spielman–Srivastava selection. Both statements are formalized. The original motivation, constant-factor asymmetric TSP, was already achieved by Svensson, Tarnawski and Végh in 2018, so this is now about the structure rather than the algorithm.

### Talagrand's threshold conjectures

Family 175 is three short papers (10 to 12 pages each). The first proves Talagrand's conjecture that fractional and integral expectation thresholds agree up to a universal factor, here $25 \cdot 512^4$. The second proves his discrete-convexity conjecture with $k = 2^{75}$ unions. The third proves a graph-decomposition conjecture that Ascoli, He, Park and Talagrand posed in August 2026 (arXiv 2608.11183), which Jia-Qi Yang had partially settled a month later (arXiv 2609.20864). The constants are absurd and irrelevant; the conjectures only ask for universal ones. Both Talagrand statements are formalized with those exact constants.

### The second Kahn–Kalai conjecture

Park and Pham proved the abstract Kahn–Kalai conjecture in 2022. The graph version says the threshold for $G(n,p)$ to contain $H$ is within a single log factor of the obvious "some subgraph is expected less than once" bound. Dubroff, Kahn and Park had $\log^3 n$ (arXiv 2508.14269). Family 176 claims $p_c \le 2048 e^{50} p_E (1 + \log_2 e(H))$ in 12 pages, by keeping every extension history in a tree and running Bayesian resampling at every node with the same random set. Formalized and listed.

### Van der Waerden numbers grow superexponentially

Erdős asked whether $W(k)^{1/k} \to \infty$ for two colours (erdosproblems.com #138). Fox and Hunter did three colours in June 2026, and Campos, Fox and Schildkraut showed $W_2(k)/2^k \to \infty$ in August. Family 160 claims $W_r(k) > k^{10^{-5} k \lfloor \log_2 r \rfloor}$ for all large $k$, two colours included, via a Behrend-style colouring of high-dimensional tori and the local lemma. It is 25 pages and the exact statement is formalized and listed. This is one of the best-supported results in the group.

### The log exponent of $r(s,t)$

Bradač settled the polynomial exponent of off-diagonal Ramsey numbers in 2026, $r(s,t) \ge c\, t^{s-1}/(\log t)^{2s-4}$ (arXiv 2605.28793). Family 170 sharpens the log exponent to the optimal $s - 2 + o(1)$ for every $s \ge 5$, using the same random projective-flag graphs and an entropy-compression argument. The conceptual step was Bradač's, and $r(4,t)$ is untouched. Both papers are formalized.

### Square-difference-free sets

How large can a subset of $\{1, \dots, N\}$ be with no two elements differing by a perfect square? Green and Sawhney got $N \exp(-c\sqrt{\log N})$ and asked for a power saving. Family 182 claims $N^{1-c}$, then extends it to intersective polynomials and to polynomials at prime arguments. Only the square case is formalized. The prime-argument paper depends on another claim in the same release, a zero-free half-plane $\mathrm{Re}(s) > 7/8$ for all Dirichlet $L$-functions, so it inherits whatever status that claim ends up with.

### Halving lines

Dey's 1998 bound of $O(n^{4/3})$ halving lines had not been improved by a power in 28 years. Family 183 claims $C n^{4/3 - \epsilon}$ with an ineffective $\epsilon$, via a limiting-measure argument showing the crossing-lemma packing cannot be tight. Theorems 1.1 and 1.2 are formalized. The $k$-set corollary the summary leads with is not.

### Sharp thresholds for graph properties

Friedgut and Kalai proved in 1996 that symmetric monotone graph properties have threshold width $O(1/\log n)$ and conjectured $(\log n)^{-2}$. Bourgain and Kalai got $(\log n)^{-2+\eta}$. Family 186 claims $2^{19} \log(1/2\epsilon)/(\log n)^2$ in a 10-page paper. It restricts to a random block of about $\sqrt n$ vertices, which keeps the permutation symmetry, and converts low-degree Fourier control on the block into control on the whole graph. The graph case is formalized with that explicit constant. The gain over Bourgain–Kalai is removing $\eta$, and the hypergraph companion is unformalized.

### Random triangle removal

Delete the edges of random triangles from $K_n$ until none is left. Bohman, Frieze and Lubetzky showed about $n^{3/2 + o(1)}$ edges survive. Family 188 claims the constant: the leave is $(1/(2\sqrt 2) + o(1))\, n^{3/2}$ in $L^2$, which is what a conjecture of Joos and Kühn (arXiv 2412.15039) predicts for triangles. The whole theorem is formalized, including the Joos–Kühn prefix input rather than assuming it.

### The Heilbronn triangle problem

Place $n$ points in a unit square to make the smallest triangle as large as possible. Komlós, Pintz and Szemerédi's $\log n / n^2$ lower bound has stood since 1982. Family 191 claims $c\, n^{-2+\eta}$, a power improvement, consistent with the known upper bound $n^{-7/6+o(1)}$. The $\eta$ is comically small: Lean uses $\eta = 1/(10^5 K)$ where $K$ is built from $\binom{163}{41}$. Lean proves it along an unbounded sequence of $n$, not every large $n$. The paper is candid that two earlier claimed power bounds on arXiv were flawed or heuristic.

### Shareshian–Wachs positivity

Family 169 proves that chromatic quasisymmetric functions of unit interval graphs are $e$-positive with coefficients in $\mathbb N[q]$. The $q = 1$ case is the Stanley–Stembridge conjecture, which Hikita proved in 2024, so this is the graded refinement plus an explicit permutation formula. The Lean challenge states the full identity as a definition returning a witness structure, which is why it lists no theorem name; it is a complete formalization.

### Euclidean Ramsey sets, classified

Family 172 gives an algebraic criterion, a tensor condition over the coordinate field, for which finite point sets are Euclidean Ramsey. Consequences include that any five concyclic points are Ramsey and that some kite shapes are Ramsey without being subtransitive, which refutes one direction of Leader, Russell and Walters's conjecture. Graham's spherical conjecture had been refuted on September 20, about two weeks before the release went public, by Pálvölgyi, in a paper he says ChatGPT wrote (arXiv 2609.23327), so that headline is not new here. The classification and several corollaries are formalized, in surprisingly few lines.

### Coboundary expanders and Ramanujan graphs

Two results are about building expanders. Family 177 constructs bounded-degree $\mathbb F_2$ coboundary expanders in every dimension $d \ge 3$; Chapman and Lubotzky had dimension 2. The complexes are congruence quotients of a coset complex over a polynomial ring, and the new content is the cohomology vanishing. It is formalized, which surprised me given the building theory involved. Family 178 gives a deterministic polynomial-time construction of nonbipartite $d$-regular Ramanujan graphs on every large even order. Existence was already known from Huang, McKenzie and Yau (about 69% of random regular graphs are Ramanujan, arXiv 2412.20263), so the contribution is derandomisation. It is 85 pages of resolvent estimates with no Lean.

### Five smaller results

Family 185 refutes the unrestricted infinite-matroid intersection and packing/covering conjectures with two self-dual matroids on a countable set, built using an ultrafilter. The examples are neither finitary nor cofinitary, so Nash-Williams' original finitary conjecture is untouched, as the paper itself says. It is formalized using Mathlib's infinite matroids.

Family 187 settles the last open hexomino in Harary's achievement game: Maker builds Snaky on the empty grid within 21 moves. The proof is a certificate of 728 conditional positions plus a soundness lemma, and the strategy is formalized as an explicit policy.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/combinatorics-fig5.png"
  alt="Six grey unit squares forming the Snaky hexomino: four cells in a row from (0,0) to (3,0), then two cells one row up from (3,1) to (4,1)."
  caption="Snaky, the hexomino whose Maker–Breaker status had been open (Snaky in 21 Maker moves, Figure 1)."
/>

Family 189 finishes the Erdős–Faudree–Rousseau–Schelp cycle–clique Ramsey conjecture, $R(C_m, K_n) = (m-1)(n-1)+1$ for $m \ge n \ge 3$ except $(3,3)$. Keevash, Long and Skokan had already reduced it to finitely many cases; the residue is 3,099 computer-checked patterns, and the Lean development is about 214 MB of generated certificates.

Family 190 gives an explicit $66 \times 66$ binary pattern for which ordered matrix removal cannot be polynomial, answering Alon and Ben-Eliezer's question negatively. It is fully formalized.

Family 192 disproves the Gopalan–Servedio conjecture: for every $C$ there is a Boolean function whose linear Fourier coefficients sum to more than $C\sqrt{\deg f}$. It is formalized but says nothing about how fast the violation can grow.

## Probability, statistical mechanics and dynamics

I opened this part of the release expecting the usual mix of solid technical lemmas and a few overreaching titles. What I found reads like the contents page of a fantasy problem book. In 41 families and 124 manuscripts, OpenAI's model claims $\theta(p_c) = 0$ on every quasi-transitive graph, $p_c < p_u$ for every nonamenable one, Polyakov's mass gap for the 2D $O(n)$ models, Rokhlin's 1949 multiple-mixing problem, Sinai's conjecture for the standard map, ergodicity of every irrational triangle billiard, the uniform bound in Hilbert's sixteenth problem, Cardy's formula for bond percolation on $\mathbb Z^2$, the $3/4$ exponent for self-avoiding walk, and the Mézard–Parisi formula for diluted spin glasses. Each of these, alone, would be the paper of someone's career.

Checking what the Lean library actually states changed my view more than reading the papers did. 15 of the 41 families have a Comparator statement that matches the headline theorem, 13 have a formal statement covering only part of the family, and 13 have nothing. The partial cases need care, because the headline is sometimes exactly the piece Lean leaves out: the Hilbert's sixteenth family formalizes its quintic Liénard companion, not the uniform bound.

One piece of outside context matters here, and the release handles it honestly: the famous three-dimensional percolation case was settled by other AI-assisted work a month before these papers were written.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig10.png"
  alt="A page of the OpenAI Research Catalog headed 'Dynamical systems and ergodic theory', listing entries 143 to 147: uniform limit-cycle bounds in Hilbert's sixteenth problem, Banach's simple Lebesgue-spectrum problem, Rokhlin's multiple-mixing problem, Sinai's positive-entropy conjecture for the standard map, and the near-boundary Birkhoff conjecture, each with a one-paragraph summary and paper links."
  caption="How the release itself presents the dynamics results: one-paragraph claims, each linking to its manuscripts (OpenAI Research Catalog, overview.pdf page 16)."
/>

My grades: 13 landmark, 24 major, 4 notable, none merely technical. 39 families prove something and 2 are counterexamples. The manuscripts add up to about 8,000 pages, and the distribution is lopsided: the honeycomb self-avoiding-walk family alone is 1,067 pages over 13 papers, the critical SK dynamics family 720, the random planar maps 762, while Rokhlin's problem takes 40 pages and ergodicity of irrational triangles takes 12. I ordered the sections below by what is claimed weighted by how much of it a kernel could check, so the formally stated landmarks come first.

### $p_c < p_u$ on every nonamenable graph

On a tree-like graph percolation has a second transition. Between $p_c$ and a uniqueness threshold $p_u$, infinitely many infinite clusters coexist, kept apart by the graph's exponential boundary. Benjamini and Schramm conjectured in 1996 that $p_c < p_u$ on every nonamenable quasi-transitive graph. Each special case cost a paper: one-ended planar graphs (Benjamini and Schramm), large girth (Nachmias and Peres), graphs with nonconstant harmonic Dirichlet functions (Gaboriau), Gromov hyperbolic graphs (Hutchcroft, 2019), acylindrically hyperbolic groups (Choi and Seo), and for every nonamenable group some well-chosen generating set (Pak and Smirnova-Nagnibeda; Thom). Hutchcroft also proposed the stronger statement $p_c < p_{2\to2}$, where $p_{2\to2}$ is the largest $p$ at which the two-point function $T_p(x,y) = \mathbb P_p(x \leftrightarrow y)$ is a bounded operator on $\ell^2$.

Family 214 claims all of it in 53 pages: $\sup_{p<p_c}\|T_p\|_{2\to2} = \|T_{p_c}\| < \infty$ and $p_c < p_{2\to2} \le p_u$, on every infinite, connected, locally finite, nonamenable quasi-transitive graph, hence every Cayley graph of every nonamenable finitely generated group with any finite generating set, which is the point: earlier group-theoretic results let you pick the generators. Mean-field critical exponents fall out: a finite triangle diagram, susceptibility of order $(p_c-p)^{-1}$, $\theta(p)$ of order $p-p_c$, cluster-volume tail $n^{-1/2}$. The proof bounds the critical operator through a weighted trace that averages the choice of root, and uses switching constructions to show that the $\ell^2$ threshold sits strictly above $p_c$. A Comparator statement exists (`BenjaminiSchramm.lean`, about 31,000 lines of solution), though it is not in the release's main-results list. 53 pages for a thirty-year-old conjecture is short. I would want a percolation specialist to read the switching argument before believing it, and the Lean is the main reason to take it seriously now.

### Polyakov's mass gap for the 2D $O(n)$ models

Put a unit vector in $\mathbb R^n$ at every site of $\mathbb Z^2$ and reward neighbours for pointing the same way. For $n=2$ (the XY model) there is a Berezinskii–Kosterlitz–Thouless transition into a phase with power-law correlations. For $n \ge 3$ Polyakov argued in 1975 that the non-abelian symmetry makes correlations decay exponentially at every temperature, however low: a mass gap generated from nothing, the lattice cousin of the Yang–Mills mass gap. It has been one of the signature open problems of mathematical physics ever since; rigorous results covered high temperature or large $n$ (Kupiainen's $1/n$ expansion), and Patrascioiu and Seiler argued for decades that a massless phase might exist after all.

Family 215 is 379 pages, and its title leads with the continuum: a canonical interacting massive continuum limit of the 2D $O(3)$ model satisfying the Osterwalder–Schrader axioms, with a unique vacuum, a mass gap, a nonzero connected four-point function and an isolated one-particle pole, plus the exact low-temperature transfer gap of the square-lattice $O(4)$ model,

$$
m_{\mathrm{lat}}(\beta) \sim 32\, e^{\pi/4 - 1/2}\, \sqrt{\beta}\, e^{-\pi\beta},
$$

which is the form Hasenfratz, Maggiore and Niedermayer predicted in 1990 from the Bethe ansatz. Those are 380 unformalized pages. The piece I find most consequential is a 20-page paper underneath: exponential decay of two-point correlations for every $n \ge 3$ at every finite $\beta$, uniformly over finite free-boundary subgraphs and edge strengths in $[0,\beta]$. That is Polyakov's conjecture in its spin-correlation form, and it is in the release's main formal results.

The proof is short enough to describe. Work with three-component spins on a finite graph. Rotating spins by a spatially varying angle gives exact identities for derivatives of the partition function. Rotating about two different axes produces a response along their commutator, and summing a fourth derivative over a family of test functions yields a nonnegative square plus local error terms. Test this on a square annulus whose inner and outer layers are pinned to directions an angle $t$ apart: signed frequency families produce a $\log N$ factor in the commutator response, and the only small quantity left to control is the probability that a sign cluster crosses the annulus, which a Campbell–Chayes style representation handles. The case $n > 3$ reduces to $n=3$ by conditioning on all but three coordinates, as Aru, Garban and Sepúlveda did for their projection.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig7.png"
  alt="A square annulus of side 4N: outer boundary layer labelled 'outer layer: s_x = s_*', an inner square labelled 'norm of x less than N' with 'inner layer: s_x = exp(tc) s_*', and a shaded band between them labelled 'compact support of the test profiles'."
  caption="The twist test behind exponential decay in 2D O(n) models: pin the inner and outer layers of an annulus to directions rotated by an angle t and measure the response with test profiles supported in between (Exponential decay in two-dimensional classical O(n) models, Figure 1)."
/>

The formal statement quantifies the way the paper does, with constants depending only on $n$ and $\beta$:

```lean
-- lean/ComparatorChallenges/ClassicalON.lean:46-54
def ExponentialDecay : Prop :=
  ∀ (n : ℕ), 3 ≤ n → ∀ (β : ℝ), 0 < β →
    ∃ A m : ℝ, 0 < m ∧ ∀ (G : LatticeGraph) (b : G.edges → ℝ),
      (∀ e, 0 ≤ b e ∧ b e ≤ β) → ∀ x y : G.vertices,
        0 ≤ correlation n G b x y ∧
          correlation n G b x y ≤ A * Real.exp (-m * siteDistance x.val y.val)

theorem main : ExponentialDecay := by
  sorry
```

`correlation` is defined a few lines above as the ratio of two integrals against the uniform measure on the sphere, so nothing physical is smuggled in. If this compiles, a conjecture physicists have used as a working assumption for fifty years is a theorem, with a 13,000-line certificate. The continuum $O(3)$ construction and the $O(4)$ asymptotics are a separate and much larger claim, and the $O(4)$ "sharp bounds" paper says its continuum spectral interpretation needs extra convergence hypotheses.

### Hilbert's sixteenth problem: the unformalized half

The second part of Hilbert's sixteenth problem asks how many limit cycles (isolated periodic orbits) a planar polynomial vector field of degree $d$ can have. Dulac claimed finiteness for each individual field in 1923; a gap was found, and finiteness was finally proved by Écalle and by Ilyashenko around 1991–92. Whether there is a bound $B(d)$ depending only on the degree is open for every $d \ge 2$, including quadratic fields. Partial results bound cycles near particular degenerations: Bautin's three for quadratic foci, Ilyashenko–Yakovenko cyclicity of elementary polycycles, Binyamini–Novikov–Yakovenko for the infinitesimal problem.

Family 143 claims the full existence statement: for every $d$ there is a finite $B(d)$ such that every real planar polynomial field of degree at most $d$ has at most $B(d)$ limit cycles in the whole plane. It gives no formula for $B(d)$, not even for $d=2$.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig3.png"
  alt="A flow chart of boxes and arrows. Three top boxes: 'Abstract isolated-zero criterion, Sections 2-5', 'Verified passage packets, Sections 6-8', 'Finite geometric templates; reductions to transfer models, Sections 9-10' feed a wide box 'Matching systems: absolute isolated-zero finiteness for every finite system in the required differential and algebraic closure, Theorem 11.4'. That and 'Projection count for regular fibers, Theorem 11.2' feed 'Uniform hyperbolic-cycle bound in a fixed square', which via 'rotation and rescaling' gives 'A bound depending only on degree, throughout the plane, Theorem 1.1'."
  caption="The proof architecture of the claimed uniform bound in Hilbert's sixteenth problem: everything funnels through an absolute isolated-zero finiteness theorem for the matching systems (Uniform bounds for planar polynomial limit cycles, Figure 1)."
/>

The strategy is to follow nearby trajectories once around a cycle and write the return map as a chain of local passages near singular points. Near a saddle a passage has a flight time $\log(1/r)$ and a contraction exponent $\lambda\log(1/r)$, and as coefficients vary these clocks can diverge at different rates. The paper represents each return by equations matching passage endpoints, lets the cuts move so representations form continuous families, shows different hyperbolic cycles land in different connected components, and reduces everything to "absolute isolated-zero finiteness": each auxiliary system has finitely many isolated solutions even when every variable, coefficients included, is free. It does not use individual finiteness as an input, and it notes a published claim (Yeung) of a gap in a leading-term step of Ilyashenko's 1991 monograph.

That is a coherent plan, and it is exactly the territory where Dulac's proof failed: asymptotic expansions of transition maps at degenerate singular points. It runs to 160 pages and none of it is formalized. The Lean in this family is for the companion paper, which settles the degree-5 case of the Lins Neto–de Melo–Pugh conjecture: a classical Liénard system $\dot x = y - F(x)$, $\dot y = -x$ with $\deg F \le 5$ has at most two limit cycles, and two occur. (The conjecture's general formula is false from degree 6 on, by De Maesschalck and Dumortier; degree 4 was settled by Li and Llibre in 2012, so degree 5 was the open case left.)

```lean
-- lean/ComparatorChallenges/QuinticLienard.lean:39-42
theorem main :
    (∀ F : Polynomial ℝ, F.degree ≤ 5 → {C : Set Plane | IsLimitCycle F C}.encard ≤ 2) ∧
    (∃ F : Polynomial ℝ, F.degree ≤ 5 ∧ {C : Set Plane | IsLimitCycle F C}.encard = 2) := by
  sorry
```

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig4.png"
  alt="A plot with axes y (horizontal) and u (vertical) showing an arch-shaped curve from y< to y> at base height h, peaking at (phi(t), t). Arrows mark the half-width r on each side of phi(h)+M, and an offset M between phi(h) and phi(h)+M."
  caption="An orbit arc in the quintic Liénard analysis, parametrised by base height h and peak height t; two such half-orbits close into a periodic orbit exactly when their endpoints match. This is the half of the family that is machine-checked (Two limit cycles for quintic Liénard systems, Figure 1)."
/>

The catalogue entry and the family title put the uniform bound first. My reading: the quintic Liénard theorem is a real, checkable result; the uniform bound is an unrefereed, unformalized claim on a 126-year-old problem, and I would hold it at arm's length until someone in the area has read Sections 2 to 5.

### Rokhlin's multiple-mixing problem

A measure-preserving transformation $T$ is mixing if $\mu(A \cap T^{-n}B) \to \mu(A)\mu(B)$: events far apart in time become independent. In 1949 Rokhlin asked whether that forces three or more events to become jointly independent when all the gaps grow. Ledrappier's 1978 example shows the answer is no for $\mathbb Z^2$ actions. For a single transformation it has been proved only under extra structure: rank one (Kalikow), singular spectrum (Host), finite rank (Ryzhikov), some smooth flows (Kanigowski and Ravotti).

Family 145 says yes, for every invertible mixing transformation of any probability space, standard or not, in 40 pages. The argument assumes a least failing order $k$, reduces to a zero-entropy finite-alphabet process with a limit joining whose proper marginals are independent but which is not a product, and then builds an ordered family of "array laws" by taking repeated limits along sums of layouts. The work is in controlling every residual interaction of that family, not just one joining.

```lean
-- lean/ComparatorChallenges/Rokhlin.lean:23-33
def MixingOfOrder (μ : Measure Ω) (T : Ω ≃ᵐ Ω) (k : ℕ) : Prop :=
  ∀ A : Fin k → Set Ω, (∀ i, MeasurableSet (A i)) →
    Tendsto
      (fun gaps : Fin (k - 1) → ℕ+ =>
        μ (⋂ i : Fin k, timeMap T (layoutTime gaps i : ℤ) ⁻¹' A i))
      atTop (𝓝 (∏ i : Fin k, μ (A i)))

theorem mixing_all_finite_orders (μ : Measure Ω) [IsProbabilityMeasure μ]
    (T : Ω ≃ᵐ Ω) (h_pres : MeasurePreserving T μ μ) (h_mix : IsMixing μ T) :
    ∀ k : ℕ, 3 ≤ k → MixingOfOrder μ T k := by
  sorry
```

`atTop` on functions `Fin (k-1) → ℕ+` means every gap tends to infinity, which is the right notion, and the measurable space carries no standardness assumption. The solution is about 28,000 lines; the challenge is not in the release's main-results list, though it has a scope page. Two other families lean on this one: the pointwise ergodic averages in family 154 and the mixing corollary of the Banach construction in family 144. If it holds, it is the cleanest landmark in the group: a one-line question, a one-line answer, and a formal statement short enough to read in a minute.

### The Mézard–Parisi formula for diluted spin glasses

In a diluted spin glass each spin interacts with a Poisson number of others through random couplings. Mézard and Parisi's cavity method describes the free energy through a hierarchy of laws of "cavity fields", their sparse-graph version of replica symmetry breaking. On the rigorous side, Franz and Leone adapted Guerra's interpolation to get upper bounds, and Panchenko and Talagrand (2004) proved that for every finite hierarchy depth $r$ the trial functional $\Phi_r$ bounds the limiting pressure from above, conjecturing equality for the infimum. Panchenko's later work gave exact variational formulas in other languages (exchangeable spin distributions, finite-RSB structure), but the matching lower bound in the hierarchical form stayed open.

Family 221 proves it, for Poisson-diluted Ising models of even arity $p \ge 2$ in the Panchenko–Talagrand class, with only first moments of the interaction and field:

$$
\lim_{N\to\infty} \frac1N\, \mathbb E \log Z_N \;=\; \inf_{r \ge 0} \Phi_r .
$$

The class matters. Interactions must factor as $e^{\theta} = a\,(1 + b\prod_\ell f_\ell(s_\ell))$ with $\mathbb E[(-b)^n] \ge 0$ for all $n$. That covers Viana–Bray, symmetric diluted even-$p$ spin glasses and soft even-$K$ SAT at positive temperature. It does not cover odd $K$-SAT, hard constraints, or fixed-degree (Bethe lattice) graphs, which is where much of the physics interest sits. This is one of ten families with a released reasoning summary, and the summary shows that OpenAI's prompt posed exactly this class, quoting the Panchenko–Talagrand hypotheses verbatim. The model was not choosing a convenient special case; it was asked a precisely scoped question and answered it.

The reasoning summary is also the best window I found into how the model works. It records failed attempts (pinning that selects states with the wrong weights; an exact purity claim that "proves too strong" because Poisson–Dirichlet clustering can hide inside an exponent jump) and then the idea that unlocked it: "shift every internal branching depth of the target tree by one level." Small marked Poisson perturbations produce identities for prescribed replica branching patterns. Delaying each branching by one level creates new branching vertices, each carrying a small coefficient, and each can be charged to its own small factor. Averaging over depths turns the identities into concentration of all multi-overlaps, which is what a diluted model needs (pair overlaps alone do not determine replica spin patterns). Then a cavity insertion is compared with independent hierarchical messages built from a sampled site's magnetization.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig6.png"
  alt="A tree diagram: a vertical 'common stem' splits at depth e into two branches ending at 'anchor' and 'old leaf'. A blue dot at depth e+1 on the anchor branch has a dashed blue edge labelled minus delta leading to 'new child'. Dotted horizontal lines mark depths e and e+1."
  caption="The key move in the lower bound: delaying a replica split by one level creates a new branching vertex with a small coefficient minus delta, and the proof charges each such vertex to its own small factor (The Mézard–Parisi formula for diluted spin glasses, Figure 1)."
/>

The Comparator statement is a single convergence, with the admissibility conditions spelled out in about 140 lines of definitions above it:

```lean
-- lean/ComparatorChallenges/DilutedSpin.lean:146-150
/-- main:theorem. No boundedness of interactions, fields, or messages is added.
Existence of the thermodynamic limit is part of this convergence assertion. -/
theorem mezard_parisi {p : ℕ} (M : Model p) (hM : Admissible M) :
    Tendsto (pressure M) atTop (𝓝 (variationalValue M)) := by
  sorry
```

The solution is about 56,000 lines, not in the main-results list. I graded this major rather than landmark only because of the class restriction. Within that class it is the exact conjecture Panchenko and Talagrand stated, proved in 36 pages, with a formal statement.

### The Thorp shuffle mixes in $\Theta(\log n)$ steps

This is the result an engineer is most likely to have touched. The Thorp shuffle cuts a deck of $n = 2^d$ cards into two halves and, for each $i$, flips a fair coin to decide whether the $i$-th card of the left half or of the right half goes on top in the next two positions. It is the round function behind several format-preserving encryption schemes for small domains (Morris, Rogaway and Stegers used it that way), where the number of rounds needed for a near-uniform permutation is a security parameter. After $d$ shuffles every single card is uniform. The whole deck takes longer, and how much longer was open for decades: Morris's first bound was $O(d^{44})$, Montenegro and Tetali got $O(d^{29})$, Morris then $O(d^4)$ and $O(d^3)$. The conjectured answer was $O(d)$.

Family 238 proves it, twelve ways. The headline paper shows the full permutation law after $1600d$ shuffles converges to uniform in total variation as $d \to \infty$, from any starting order, and a counting argument gives a lower bound: one shuffle uses only $n/2$ coins, so after $t$ shuffles at most $2^{tn/2}$ orders are reachable and the distance from uniform is at least $1 - 2^{tn/2}/n!$, which forces about $2d$ shuffles. Mixing time is $\Theta(d) = \Theta(\log n)$.

The proof is a two-stage transfer. First, after $200d$ shuffles, the joint positions of a uniformly random list of $7n/8$ labels are close to uniform on average over the list. Second, a separate paper shows that this kind of partial information, plus a vanishing sign bias, controls the Fourier transform of the full law on every irreducible representation of $S_n$, so the product of two such permutations is near uniform. The figure shows the bookkeeping that makes the first stage tractable: once some cards' paths are revealed, a switch with two free slots averages the conditional mass and a switch with one free slot transports it.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig2.png"
  alt="Two diagrams of a single card-pair switch. Left, 'Two free positions': inputs u and v each connect to both outputs, which receive (u+v)/2 each, labelled 'Unrevealed fair switch'. Right, 'One free position': input u goes along a solid arrow to an output labelled u while a filled dot (an exposed card) goes along a dashed arrow, labelled 'Switch fixed by the observed path'."
  caption="How conditional mass moves through one Thorp switch when some card paths are revealed: two free slots average, one free slot transports. This bookkeeping is how the proof controls partial permutation information (Optimal-order mixing of the Thorp shuffle, Figure 1)."
/>

You cannot see an asymptotic constant on a small deck, but you can see the two regimes. The widget computes the exact law of the Thorp shuffle on 4 and 8 cards by pushing all $8! = 40{,}320$ orderings through every coin pattern. On 8 cards the distance is 1.0000, 0.9996, 0.9937 and 0.8984 after 0 to 3 shuffles, exactly the coin-count floor each time: the shuffle is uniform on the orders it can reach, and there are not enough of them. At 4 shuffles the floor drops to zero, and the deck is within 1/4 of uniform after 5.

<ThorpExactMixing />

The Lean side is unusual for its breadth. Thirteen Comparator statements cover this family, four of them in the release's main-results list. One states the optimal order directly, with `mixingTime d` defined as the least $t$ with total-variation distance at most $1/4$ for a Lean model of the shuffle (positions are $d$-bit strings; one shuffle is a coin-controlled switch on the leading bit followed by a rotation of the bits):

```lean
-- lean/ComparatorChallenges/ThorpRemaining.lean:536-538
def OptimalOrderMain : Prop :=
  Asymptotics.IsTheta atTop (fun d : ℕ => (_root_.OAI.Thorp.mixingTime d : ℝ))
    (fun d : ℕ => Real.log (2 ^ d : ℝ))
```

It is one conjunct of `remaining_main` at line 543. The constants in the formal versions are far worse than the paper's: the scope page says one formalized route gives full-deck mixing after $16{,}040{,}400\,d$ steps. That does not matter for the order, and it is a nice illustration of how formalization trades constants for certainty. The result covers power-of-two decks only; the paper says the optimal order for general even $n$ is not addressed.

### Random walk in random environment: escape implies speed

Give every site of $\mathbb Z^d$ its own random, independent set of jump probabilities and let a walker move according to them. In one dimension such a walk can escape to infinity with zero speed (Solomon, 1975), because traps slow it down. For $d \ge 2$ it was conjectured that this cannot happen: if the walk is transient in a direction $\ell$, it should move ballistically, with a deterministic velocity $v$ and $v \cdot \ell > 0$. The same circle of problems contains Kalikow's question whether the probability of escaping in direction $\ell$ must be 0 or 1, proved in $d = 2$ by Zerner and Merkl (2001) and false for some dependent environments in $d \ge 3$ (Bramson, Zeitouni and Zerner, 2006).

Family 220 claims the ballisticity conjecture for iid uniformly elliptic environments in every $d \ge 2$, in the form Fribergh and Kious state it, and the 0–1 law in $d \ge 3$ for iid strictly elliptic environments (no uniform lower bound on jump probabilities) and for finite-range-dependent uniformly elliptic ones. In $d \ge 3$ the directions of positive escape probability form exactly the open hemisphere around $v$. The machinery is the Sznitman–Zerner regeneration structure: cut the path at times after which it never returns below a new height record, and show these pieces have finite mean duration.

The ballisticity theorem is in the release's main formal results (`DirectionalBallisticity.lean`, about 99,000 lines of solution), and the 0–1 law and hemisphere statement have their own challenges. I would put this next to the percolation and $O(n)$ results as the strongest landmark claims in the group: a well-known conjecture, a faithful statement, and a large certificate.

### Self-similar measures: the dimension formula with overlaps

Take finitely many contractions $x \mapsto r_i x + t_i$ of the line and probabilities $p_i$; the self-similar measure is the law of a random infinite composition. Its dimension should be the entropy of the randomness divided by the average log-contraction, capped at 1. Overlaps spoil this. If two different words compose to the same map, you lose entropy, and the exact-overlaps conjecture says that is the only way to lose it. Hochman's 2014 inverse theorem proved the formula under exponential separation, and since then a series of results by Varjú, Breuillard–Varjú, Rapaport, Rapaport–Varjú and Feng–Feng covered algebraic parameters and special families.

Family 148 claims the general statement, Varjú's "Conjecture 3": for every finite self-similar system on $\mathbb R$, with signed unequal ratios and exact overlaps allowed,

$$
\dim_H \mu = \min\Bigl\{1, \frac{h_{\mathrm{RW}}}{\chi}\Bigr\},
$$

where $h_{\mathrm{RW}}$ is the entropy rate of the random walk on composed maps (counting maps, not addresses) and $\chi$ the Lyapunov exponent. With no exact overlaps $h_{\mathrm{RW}} = H(p)$, so this contains the exact-overlaps conjecture on the line. The proof compares a finite law's exact entropy with its entropy at a grid scale, shows hidden information must be witnessed by fair two-point laws at finer scales (an estimate that does not depend on the smallest gap, which matters because Baker and Bárány–Käenmäki built systems with no exact overlaps and arbitrarily fast superexponential near-coincidences), and conditions on block symbol counts so each block's contraction is fixed. It is 22 pages and its main theorem is in the release's main formal results (`SelfSimilar.json`). That combination of short, famous and formalized is the pattern I kept seeing, and it is the one that makes me most impatient for someone to compile the library.

### Sinai's conjecture for the standard map

The Chirikov standard map $f_k(x,y) = (x + y + k\sin 2\pi x,\ y + k\sin 2\pi x)$ on the torus is the textbook picture of Hamiltonian chaos. For large $k$ pictures show a chaotic sea dotted with small islands, and Sinai conjectured that area has positive metric entropy for a positive-measure set of parameters. No one could prove it, because elliptic islands keep reappearing (Duarte) and hyperbolic sets of full dimension (Gorodetski) can still have zero area. Berger and Turaev got positive entropy only after a $C^\infty$ perturbation.

Family 146 claims positive metric entropy for every $k$ above some non-explicit $k_0$, stronger than Sinai asked. By Pesin's formula the entropy equals the integral of the positive Lyapunov exponent, so the work is derivative growth on a positive-area set, controlling the repeated losses near the lines where $\cos 2\pi x$ vanishes. A corollary gives a positive-area ergodic component whose cyclic pieces are Bernoulli. It is 45 pages, with Comparator statements for entropy, Lyapunov exponent and components (not in the main-results list).

### Every irrational triangle billiard is ergodic

A ball bouncing in a triangle has no curvature to make it chaotic. With rational angles, unfolding the reflections turns the motion into straight lines on a flat surface, and Kerckhoff, Masur and Smillie (1986) used that to show ergodicity for a dense $G_\delta$ set of polygons. For a specific triangle with an irrational angle, ergodicity was known only under extremely fast rational approximation (Vorobets), and numerical studies have repeatedly raised doubts for some irrational triangles.

Family 150 claims ergodicity for every triangle with at least one angle irrational relative to $\pi$, in a 12-page paper, and weak mixing in a 13-page companion. Ergodicity has a Comparator statement that builds the billiard flow itself, discarding vertex-hitting trajectories. I graded it landmark because it is a long-standing problem in polygonal billiards. I also flag it: 25 pages in total is very little for a problem this hard, the main issue (low regularity on surfaces with cone points) is delegated to an analytic framework of Forni and Moll, and the weak-mixing half has no formal statement.

### Banach's simple Lebesgue spectrum

Banach asked, in the Scottish Book, for a transformation whose Koopman operator $f \mapsto f \circ T$ has simple Lebesgue spectrum: a single orbit $\{f \circ T^n\}$ that is an orthonormal basis. Lebesgue spectrum is common (Bernoulli shifts have it), but with infinite multiplicity; the closest earlier construction, by Mathew and Nadkarni, has a Lebesgue component of multiplicity two. Family 144 constructs a $C^\infty$ volume-preserving diffeomorphism of the 3-torus with simple Lebesgue spectrum on the whole mean-zero $L^2$ space, as a limit of explicit conjugated twists in the Anosov–Katok spirit. The paper is careful to say that the historical real-line question Ulam recorded differs from the probability-space form it solves. 30 pages, with a Comparator statement (`ThreeTorus.lean`) outside the main-results list.

### First-passage percolation: no bigeodesics, and a smooth limit shape

Put iid random travel times on the edges of $\mathbb Z^2$. A bigeodesic is a doubly infinite path every finite piece of which is a fastest route. Their nonexistence is a conjecture going back to Furstenberg's question via Kesten, with partial results by Licea–Newman and Damron–Hanson (bigeodesics in fixed directions under curvature assumptions). Family 212 claims none exist, for any nonatomic weights whose minimum of four copies has a finite second moment, and separately that the exponential model's limit shape is strictly convex with a $C^1$ boundary, two other long-standing conjectures from Kesten's era. The no-bigeodesics argument runs on Busemann functions and a "label budget" along horizontal lines. Only differentiability of the time constant for Gamma weights is formalized; the two conjectures the family is named for are not.

### Cardy's formula on the square lattice, and FK interfaces for $q \le 4$

Smirnov's 2001 proof of Cardy's formula and his later Ising work rely on a discrete holomorphic observable that exists only on special lattices. Bond percolation on $\mathbb Z^2$, the model most people picture, has had no proof of conformal invariance. More generally, Rohde and Schramm conjectured that critical random-cluster (FK) interfaces on the square lattice converge to $\mathrm{SLE}_\kappa$ with $\kappa = 4\pi/\arccos(-\sqrt q/2)$ for $0 < q < 4$.

Family 223, 602 pages in six papers, claims that for every $1 \le q \le 4$ (bounded Jordan domains, including Cardy's formula for $\mathbb Z^2$ bond percolation with free boundary) and for $0 < q < 1$ (smooth Jordan domains), plus full nested CLE limits for $1 \le q \le 4$, quenched SLE$_{16/3}$ under weak bond disorder, massive SLE for thermal FK–Ising, and the natural occupation measure of the interface. None of it is formalized. If right, this is one of the largest advances in 2D statistical mechanics since Smirnov, and it is the result in this group I most want independent readers on.

### The 3/4 exponent for self-avoiding walk on the honeycomb lattice

Flory's 1949 argument and Nienhuis's 1982 Coulomb-gas calculation predict that an $n$-step self-avoiding walk in the plane has diameter about $n^{3/4}$. Duminil-Copin and Smirnov computed the honeycomb lattice's connective constant $\sqrt{2+\sqrt2}$ in 2012 using a parafermionic observable, but nothing close to the $3/4$ exponent was proved.

Family 237, the largest in the group at 1,067 pages over 13 papers, claims that a uniformly chosen $n$-step honeycomb self-avoiding walk has diameter $n^{3/4+o(1)}$ with high polynomial probability for every large $n$, with local-mass and covering exponents $4/3$. The route goes through strip-crossing masses of order $N^{-1/4}$, bridge lengths $h^{4/3+o(1)}$, cylinder partition functions with exponent $1/6$, and a renewal argument that changes from all-length critical measures to fixed-length walks. The Lean covers supporting estimates (finiteness of bridge sums, the free energy, and the strip-crossing mass) but not the $3/4$ law. It is honeycomb only, an $o(1)$ exponent and not a scaling limit, and the sheer size makes this the hardest family in the group to assess.

### Spin-glass dynamics across the SK transition

Glauber dynamics is the sampler everyone reaches for, and the Sherrington–Kirkpatrick model is the place where it is supposed to break down. Eldan, Koehler and Zeitouni proved a spectral gap for $\beta < 1/4$, and spectral- and entropic-independence methods reached the same range. Family 227 (720 pages, eight papers) claims the whole picture for zero-field Gaussian SK with rate-one heat-bath updates: a dimension-free spectral gap and worst-start cutoff at $\log n / (2\lambda(\beta))$ throughout $\beta < 1$; mixing time $n^{2/3+o(1)}$ at $\beta = 1$; a stretched-exponential obstruction from typical starts at time $\exp(n^{1/10000})$ for $\beta > 1$; and, at criticality, universal random scaling limits for stationary and quench autocorrelations (the mathematical version of aging). The Lean is generous but partial and sometimes weaker than the papers: the gap for $\beta < 1$, cutoff only for $\beta < 1/2$, critical mixing between $n^{2/3-\varepsilon}$ and $e^{\varepsilon n}$ rather than $n^{2/3+o(1)}$, and the low-temperature obstruction. I would cite the formal statements, not the abstracts, if I used this.

### Low-temperature SK fluctuations

The SK free energy is known (Parisi, Guerra, Talagrand), and at high temperature its fluctuations are Gaussian and of order one (Aizenman, Lebowitz and Ruelle, 1987). At low temperature physicists (Kondor; Crisanti, Paladin, Sommers and Vulpiani; Parisi and Rizzo) predicted fluctuations of $\log Z_n$ of order $n^{1/6}$, and Chatterjee's superconcentration theory only proved they are $o(\sqrt n)$. Family 217 claims, for every $\beta > 1$, that $\mathrm{Var}\log Z_n \sim c_\beta n^{1/3}$ with $c_\beta > 0$ and that the standardized free energy has a nondegenerate limit law, characterized uniquely but not identified with any named distribution. 206 pages, no Lean.

### Perceptrons and the jamming exponents

Family 222 gives Gardner-type variational formulas for the Ising perceptron (bounded Borel log-potentials) and the spherical perceptron (bounded continuous potentials, and bi-orthogonally invariant disorder) at every positive temperature and density. The part that surprised me is the fourth paper: for the spherical perceptron with margin $-1$ and quadratic penalty, a sharp feasibility threshold and microscopic gap and force laws with exponents $0.4126930 < \gamma < 0.4126934$ and $0.4231063 < \theta < 0.4231088$, $\gamma = 1/(2+\theta)$. Those are the jamming exponents Franz, Parisi and collaborators derived from full replica-symmetry breaking, and the paper certifies the intervals with a finite numerical computation and error bounds. Limits are taken in a fixed order (size, then temperature, then density). Lean covers the spherical pressure formula and a finiteness lemma for the Ising one.

### Random planar maps: from surfaces to trees

Critical Fortuin–Kasteleyn planar maps are random surfaces decorated by loops, and physics predicts they look like Liouville quantum gravity for $q \le 4$ and collapse to trees for $q > 4$. Sheffield's hamburger–cheeseburger bijection (2016) gave a "peanosphere" limit, and Gwynne and Miller conjectured graph-distance convergence. Family 211, 762 pages in seven papers, claims convergence of spherical FK maps to CLE-decorated LQG spheres for $0 < q < 4$, the finite spherical cases of the Gwynne–Miller conjecture (and for spanning-tree maps), a constructed critical LQG sphere with CLE$_4$ at $q = 4$, the Brownian continuum random tree for $q > 4$, and for $q = 2$ and spanning-tree maps that random walk converges to Liouville Brownian motion with eigenvalues and heat trace. The spectral paper names "stated Brownian/LQG inputs" and leans on a metric companion revised on 3 October. Only the $q > 4$ tree limit, the least novel regime, is formalized.

### The XY model's fine structure

Fröhlich and Spencer proved the BKT phase exists in 1981. The fine predictions stayed open: at the critical point correlations should decay as $r^{-1/4}(\log r)^{1/8}$, and the correlation length should diverge like $\exp(c/\sqrt{\beta_c - \beta})$. Family 216 claims both for the nearest-neighbour cosine XY model, with the free-box limit taken first, plus Gaussian free field limits for discrete Gaussian heights throughout the rough phase with the universal effective temperature $8\pi$. Two of the six papers (center magnetization and the critical spin field) say outright that they assume inputs from the companions. 461 pages, no Lean.

### GOE statistics for random regular graphs

A random 3-regular graph looks locally like a tree, yet Jakobson, Miller, Rivin and Rudnick predicted in 1999 that its eigenvalue spacings follow the GOE. Universality was proved for growing degree (Bauerschmidt, Huang, Knowles and Yau; Bourgade and Huang for $(\log n)^{24} \ll d$), and Bourgade and Huang stated the fixed-degree case as a conjecture. Family 219 proves it for every fixed $d \ge 3$ at every fixed energy in the open Kesten–McKay bulk, and extends it to sufficiently weak diagonal (Anderson) disorder. 104 pages, no Lean.

### Weakly perturbed and disordered Ising models

Family 218 pushes conformal invariance of the 2D Ising model beyond the exactly solvable nearest-neighbour case: weak finite-range even multispin perturbations keep the critical spin and energy correlations; weak contour interactions keep SLE$_3$ interfaces; and weak iid bond disorder of any bounded nondegenerate mean-zero law gives quenched SLE$_3$ in probability over environments. A fourth paper, conditional on stated deterministic estimates, finds the relative variance of quenched correlations growing like $(\log r)^{1/4+o(1)}$, matching the marginal Harris-criterion picture. Everything is perturbative, and Lean has one comparison lemma.

### Voronoi percolation obeys Cardy

Poisson–Voronoi percolation is the standard test of whether conformal invariance is a lattice accident. Benjamini and Schramm conjectured it in 1998, and Tassion's RSW theory and the quenched noise-sensitivity results made it the natural next case after Smirnov. Family 224 claims the annealed Cardy formula in every bounded Jordan quadrilateral, an $\varepsilon^{-3/4}$ asymptotic for the expected pivotal count, and quenched near-critical universality matching the triangular lattice. Annealed means averaged over the tessellation; two of the three papers take Cardy as an input. No Lean.

### The six-vertex model's Gaussian free field

The height function of the six-vertex model with $a = b = 1$ and $0 < c \le 2$ should converge to a Gaussian free field whose variance the Coulomb gas fixes. This was known only at or near the free-fermion point (Kenyon; Giuliani, Mastropietro and Toninelli), with logarithmic delocalization for $1 \le c \le 2$ (Duminil-Copin, Karrila, Manolescu and Oulamara). Family 225 claims the full regime, endpoint $c = 2$ included, with squared multiplier $1/\arcsin(c/2)$, for the plane state obtained from balanced tori. One 93-page paper, no Lean.

### Double dimers become CLE$_4$

Superimpose two random domino tilings and you get loops expected to converge to CLE$_4$. Kenyon proved conformal invariance of loop observables, and Dubédat and Basok–Chelkak identified which punctures loops surround. Family 226 upgrades that to convergence of the loops as curves, but only for the Temperleyan square lattice in the half-plane. 33 pages, no Lean.

### Lipschitz height functions

Schramm's 2006 problem list asks whether random Lipschitz height functions on the triangular lattice converge to the Gaussian free field, and whether their level lines are SLE$_4$. Family 232 claims the field part of Problem 2.2 (odd integer heights with two-arc boundary data), a GFF limit for weighted integer Lipschitz heights at $x \in [1/\sqrt2, 1]$, and both halves of Problem 2.3 for real-valued heights, with SLE$_4$ at one tuned boundary amplitude. The integer interface (the SLE half of 2.2) is not claimed. Building on Glazman and Manolescu's delocalization; no Lean.

### Orthogonally invariant spin glasses

Replace the Gaussian coupling matrix of SK by a Haar-random rotation of any spectrum with a compact limit and no outliers. Physicists (Marinari, Parisi and Ritort; Parisi and Potters) predicted the free energy; rigorous results stopped at high temperature (Bhattacharya and Sen; Fan and Wu). Family 234 claims a variational formula at every temperature, almost surely and in expectation, with external fields and the ground-state energy as $T \to 0$, and the whole thing has a Comparator statement. I did not check the formula against the physics predictions.

### Random $k$-SAT thresholds, with someone else's priority

Chvátal and Reed conjectured in 1992 that random $k$-SAT has a sharp satisfiability threshold at a fixed density; Friedgut proved a sharp threshold that might drift with $n$, and Ding, Sly and Sun proved the conjecture for large $k$. On 5 October 2026 Gaia Carenini posted [ECCC TR26-229](https://eccc.weizmann.ac.il/report/2026/229/) proving it for every $k \ge 3$, and OpenAI credits her with priority in two abstracts. Family 235 offers another proof, plus $\mathrm{Var}(H_n) = \Theta_k(n)$ for the hitting time of the first unsatisfiable prefix, and a proof that the 3-SAT threshold is a computable real. Threshold existence, variance and computability all have Comparator statements. The variance and computability results are new; the threshold is not this family's.

### Symmetric sign matrices are singular at rate $(1/2)^n$

How likely is a random symmetric $\pm1$ matrix to be singular? Two equal rows already cost about $2^{-n}$, and the conjecture was that nothing else matters at exponential scale. Costello, Tao and Vu proved singularity has probability $o(1)$; Campos, Jenssen, Michelen and Sahasrabudhe proved $e^{-cn}$; Tikhomirov had settled the non-symmetric analogue in 2020. Family 239 claims $\mathbb P(\det A_n = 0) = (1/2 + o(1))^n$, and $(p^2 + (1-p)^2 + o(1))^n$ for bias $p \ne 1/2$. 101 pages, no Lean.

### Potts reconstruction on trees

Family 229 proves that the Kesten–Stigum bound $d\lambda^2 > 1$ is the exact reconstruction threshold for the symmetric three-state channel (both signs of $\lambda$) and the ferromagnetic four-state Potts channel, on regular and Poisson trees, with nonreconstruction at equality. Mézard and Montanari and Sly had conjectured this; it is known to fail for five or more states. The consequence is the exact weak-recovery threshold $(a-b)^2 > 3(a+2b)$ for the three-community sparse block model. The four-state proof uses exact-arithmetic verification of polynomial inequalities, and the Lean covers only the direction above the threshold, which is the classical Kesten–Stigum half.

### The free uniform spanning forest is a factor of IID

Lyons asked whether the free uniform spanning forest can be built by a local, equivariant rule from iid labels. Family 231 says yes on every infinite connected locally finite graph, with one Borel rule and no root, and extends it to every translation-invariant strongly Rayleigh process on any countable group, amenable or not (Lyons and Thom had the amenable determinantal case). 19 pages, and both statements have Comparator challenges.

### Dynamics: four more structural results

Family 151 disproves the general $C^1$ self-map form of Shub's entropy conjecture: a noninvertible $C^1$ map of $S^1 \times (S^2)^{q+1}$ with zero topological entropy and eigenvalue 2 on second homology. One sphere coordinate mostly squares, a circle coordinate is a clock that can stall, and the other spheres store small signals whose high winding is paid for by tiny coefficients. Yomdin's $C^\infty$ theorem and the diffeomorphism case are untouched. The Comparator statement builds the manifold concretely inside Euclidean space.

Family 152 constructs a zero-entropy ergodic system with no smooth positive-volume model on any compact manifold, the first obstruction to Anosov–Katok smooth realization other than infinite entropy (Kushnirenko). The formal statement proves only the weaker finite-entropy version.

Family 153 proves the Bernoulli convolution $\nu_\lambda$ is singular when $1/\lambda$ is any quartic Salem number in $(1,2)$, the first singular parameters beyond Erdős's Pisot numbers, plus an "arithmetic classification" that is an infinite approximation condition, not an algorithm. The catalogue's word "classifies" oversells the second part. No Lean.

Family 154 proves almost-everywhere convergence of consecutive multiple ergodic averages for every mixing transformation, a case of Furstenberg's pointwise problem. It imports Rokhlin's theorem from family 145 and an $L^3$ bound for the trilinear Hilbert transform from another family of the release, which is itself a famous open problem in harmonic analysis. One of its papers mentions in passing that the same oscillation estimate gives triple pointwise convergence without any mixing assumption, a much larger claim than the title, made by reference.

The near-boundary Birkhoff conjecture (family 147) is a real rigidity theorem, not the Birkhoff conjecture: a smooth convex billiard with a full continuous collar of caustics next to its boundary is an ellipse.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/probability-dynamics-fig5.png"
  alt="Left: an ellipse with two foci and several thin nested elliptical curves near its boundary, shaded, labelled 'A full caustic collar' and boundary label partial Omega. Right: a boundary arc with impact point Gamma, outward normal n_psi, incoming and outgoing unit velocities e_{psi-d} and e_{psi+d}, and angles d."
  caption="The hypothesis of the near-boundary Birkhoff theorem: a full collar of convex caustics next to the boundary, as confocal ellipses provide. A Cantor family of caustics, which KAM theory gives generically, does not qualify (Rigidity of smooth billiards with a continuous caustic collar, Figure 1)."
/>

Family 149 proves boundedness, persistence and classwise permanence for weakly reversible mass-action reaction networks with fixed rates, conjectures that predate most of chemical reaction network theory's modern tools. The Lean covers boundedness and persistence with bounds that depend on the starting point, not the uniform permanence in the title.

### Four smaller results

Family 228 constructs radial pair potentials in $\mathbb R^3$, one with a hard core and one bounded with $|\phi(r)| \le C r^{-3-1/32}$, whose canonical free energy has a first-order kink in temperature on an interval of densities. That is a continuum phase transition of the kind Simon's problem list asks for, but for potentials designed for the proof rather than Lennard-Jones; both are formalized. Family 230 answers Schramm's Hausdorff-measure question for SLE$_\kappa$, $\kappa < 8$: the exact gauge is $r^d(\log\log 1/r)^{(2-d)/2}$ with $d = 1 + \kappa/8$, not the exponent-one gauge Schramm suggested (Lean has the positivity half). Family 233 identifies the joint scaling limit of Ashkin–Teller heights and current clusters on the critical line, a conjecture of Alcalde López, Heeney and Lis. Family 236 shows the free Ising state on the $d$-regular tree is a factor of IID exactly when $\tanh\beta \le 1/\sqrt{d-1}$, confirming Nam, Sly and Zhang, with the main theorem in the release's main formal results.

### What I would believe today

If I had to rank by evidence rather than by fame: the percolation pair ($\theta(p_c) = 0$ on quasi-transitive graphs, $p_c < p_u$), Polyakov's exponential decay, Rokhlin's multiple mixing, ballisticity, the self-similar dimension formula, the Thorp shuffle and the Mézard–Parisi formula all have short formal statements that I read and found faithful, sitting on solution files with no `sorry`. Hilbert's sixteenth (uniform half), Cardy on $\mathbb Z^2$, the six-vertex GFF, the honeycomb $3/4$ exponent, the SK fluctuation law and the Salem Bernoulli convolutions are equally famous and have no formal backing at all. And two famous items arrive with honest priority notes: the lattice case of $\theta(p_c) = 0$ (Leder and Bou-Rabee, a month earlier) and the $k$-SAT threshold (Carenini, a day before the release).

For what machine-checking does and does not buy, the [Thomson $N = 7$ article](/articles/thomson-n7-lean) walks through a smaller case in detail. The short version applies here: a kernel-checked proof of a faithful statement is a proof, and the remaining risk sits in whether the definitions say what the paper means.

## Mathematical physics and operator algebras

Operator algebras is a small field with a short list of famous problems, and mathematical physics problems tend to be long, technical and resistant to clever tricks. This group reads like someone went down Kadison's problem list and Barry Simon's list of open problems and ticked every box. The group has 44 families and 90 manuscripts, about 4,700 pages. Twenty-five families are filed under mathematical physics, from quantum information and quantum complexity through spin chains, Bose gases, atoms and general relativity. The other nineteen are operator algebras.

The titles alone include the free group factor isomorphism problem, Kadison's similarity problem, the Kadison–Ringrose and Kadison–Kastler conjectures, the generator problem, Connes' bicentralizer problem, Kaplansky's quasitrace question, the hyperinvariant subspace problem, Baum–Connes, Kadison–Kaplansky, Haldane's gap, the ionization conjecture, Anderson localization in two and three dimensions, Bose–Einstein condensation of an interacting gas, the Penrose inequality and parity in QAC⁰. Any one of these would be the result of the year in its field.

So the useful question is not which ones are important but how much of each claim a machine has actually checked. Two facts organize everything below.

The first is that the operator-algebra half is unusually well formalized. Twelve of its nineteen families have a Comparator challenge matching the headline theorem, and the papers behind them are short: 23 pages for the free group factors, 31 for Kadison similarity, 28 for Kadison–Ringrose, 18 for the generator problem, 16 for Kaplansky's quasitraces, 14 for Kirchberg's embedding problem. In mathematical physics only seven families have the main theorem in Lean, and the long, famous ones (Anderson, Haldane, the Penrose inequality, the 2D area law) have nothing that touches the headline.

The second is the exception that matters most. The counterexamples to the Baum–Connes and Kadison–Kaplansky conjectures, which I would rank as the biggest claim in the group for anyone working in noncommutative geometry, have no formalization at all, and they run to 184 pages of graphical group constructions.

My grades: 20 families landmark, 14 major, 9 notable and 1 technical. By kind, 32 are proofs, 9 are counterexamples or disproofs, 2 settle part of a conjecture and 1 is conditional on strong hypotheses. By Lean, 19 have the main theorem formalized, 12 are partly formalized and 13 have nothing. The order below is mine: it weighs what is claimed against how much of it has been checked.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/physics-operator-algebras-fig4.png"
  alt="Page excerpt from the release overview headed Operator algebras, with entries 285 (counterexamples to Baum-Connes and Kadison-Kaplansky), 286 (arithmetic rigidity of lattice von Neumann algebras) and 287 (all nonabelian free group factors are isomorphic)."
  caption="How the release itself announces the operator-algebra results; entry 287 is the free group factor isomorphism (OpenAI Research Catalog overview, page 31)."
/>

### The Kadison–Ringrose cohomology conjecture

This one is a consequence of the same circle of ideas, and I read it as the similarity paper's sibling. Kadison and Ringrose conjectured in 1971 that every von Neumann algebra has vanishing bounded Hochschild cohomology with coefficients in itself, $H^k(M, M) = 0$ for $k \ge 1$. Degree one is the Kadison–Sakai theorem that every derivation is inner. The higher degrees were known for hyperfinite algebras (Johnson–Kadison–Ringrose, 1972), for types I, II∞ and III and McDuff factors through completely bounded cohomology (Christensen–Effros–Sinclair, [Invent. Math. 1987](https://doi.org/10.1007/BF01388706)), and for II₁ factors with property Γ (Christensen–Pop–Sinclair–Smith, [Ann. Math. 2003](https://doi.org/10.4007/annals.2003.158.635)). Factors without Γ, $L(\mathbb F_2)$ again, were the gap.

The 28-page paper claims every degree $k \ge 2$ for every complex von Neumann algebra, with no separability or type hypothesis, and Lean states exactly that: every bounded cocycle of degree $n+2$ is the coboundary of a bounded cochain one degree lower (`KadisonRingrose.lean`, lines 30–35). The standard route needs every bounded cocycle to be completely bounded, which is precisely the kind of uniform-in-matrix-size control the similarity paper supplies, so the two results stand or fall together.

### Strong Kadison–Kastler stability

If two von Neumann algebras on the same Hilbert space are close (their unit balls are within $\delta$ in operator norm), are they conjugate by a unitary close to the identity? Kadison and Kastler conjectured yes in 1972. It was known for injective algebras (Christensen), for separable nuclear C\*-algebras (Christensen–Sinclair–Smith–White–Winter, Acta Math. 2012) and for some II₁ factors (Cameron, Christensen, Sinclair, Smith, White and Wiggins, [Duke 2014](https://arxiv.org/abs/1209.4116)), who also connected it to the similarity problem.

The family claims the strong form for all von Neumann algebras with one universal tolerance: for every $\varepsilon > 0$ there is $\delta(\varepsilon)$, independent of the algebras, representations and Hilbert space, such that distance below $\delta$ gives a conjugating unitary with $\lVert u - 1\rVert < \varepsilon$. That main theorem is in Lean (`StrongKadisonKastler.lean`). Two companion papers mark the edges: one-sided near inclusions need not be implemented by small unitaries, and for every $\varepsilon$ there are norm-separable C\*-algebras on a separable Hilbert space within $\varepsilon$ of each other, with the same weak closure, that are not unitarily conjugate at all. Those two counterexamples are paper-only.

### The generator problem

Is every von Neumann algebra with separable predual generated by a single operator? Kadison put it on his 1967 problem list. Type I (Pearcy), properly infinite algebras (Wogen), factors with Cartan subalgebras (Popa), with property Γ or tensor decompositions (Ge–Popa, [Duke 1998](https://doi.org/10.1215/S0012-7094-98-09405-4)) and many more were done; by the established direct-integral reduction only II₁ factors were left, and $L(\mathbb F_n)$ was the obvious holdout.

The 18-page paper proves every II₁ factor with separable predual is singly generated, and more: for any irreducible subfactor $P \subset M$, the unitaries $u$ with $M = W^*(P, u)$ form a dense $G_\delta$ in the trace 2-norm. Take $P$ to be an irreducible hyperfinite subfactor (Popa's embedding theorem); $P$ is singly generated, so two operators generate $M$, and two combine into one. Both statements are in Lean (`FactorGeneration.lean`, `RelativeGeneration.lean`). I checked whether it quietly uses family 287; it does not. It only mentions free group factors for a free-entropy corollary. The two results are consistent: under 287, $L(\mathbb F_\infty) \cong L(\mathbb F_2)$ is finitely generated anyway.

### Spontaneous magnetization in the quantum Heisenberg ferromagnet

Iron is magnetic. The textbook quantum model of a ferromagnet is the nearest-neighbour Heisenberg Hamiltonian $H_\Lambda = -\sum_{\lvert x-y\rvert_1 = 1} \mathbf S_x \cdot \mathbf S_y$, and proving that it actually magnetizes at low temperature in three dimensions has been open for as long as people have tried. Lieb posed it as Problem A in his 1999 list. The antiferromagnet was done in 1978 by Dyson, Lieb and Simon with reflection positivity and infrared bounds, but that argument does not give the ferromagnetic infrared bound, as Dyson, Lieb and Simon themselves explained. The best results were about the free energy: Conlon–Solovej upper bounds, Tóth's exchange-cycle representation, and Correggi, Giuliani and Seiringer's proof that spin-wave theory gets the free energy right ([arXiv:1312.7873](https://arxiv.org/abs/1312.7873)). None of those produces an ordered state.

The family claims it for every dimension $d \ge 3$ and every spin $S$: at every sufficiently low positive temperature there is a translation-invariant zero-field KMS state with $\omega(S^z_0) \ge S/4$. The proof is probabilistic. Split each spin $S$ into $\ell = 2S$ spin-½ slots, write the Gibbs state as a random exchange-path model where each cycle has weight two (its two colours), and pin every boundary slot to "up". Then show, by downward induction on the number of pins, that the probability $m$ sparse test points all sit on cycles that miss the pins is at most $p^m$ for a small fixed $p$. The induction step uses stable multiaffine polynomials (Borcea–Brändén–Liggett negative dependence) to reduce to a return estimate, an exterior-power determinant identity to compute it, and heat-flow and Nash-inequality bounds around blocked regions to close it. A single interior slot is then up with probability bounded away from ½, uniformly in the volume.

Lean formalizes exactly the headline, with $\ell = 2S$ so the bound reads $\ell/8$ (`lean/ComparatorChallenges/Heisenberg.lean`, lines 206–211):

```lean
theorem spontaneous_magnetization (d ℓ : ℕ) (hd : 3≤d) (hℓ : 1≤ℓ) :
    DynamicsConverges d ℓ ∧
      ∃ β₀ : ℝ,0<β₀ ∧ ∀ β : ℝ,β₀≤β →
        ∃ ω : State d ℓ,TranslationInvariant ω ∧ IsKMS β ω ∧
          (ω.functional (spinZ d ℓ 0)).im=0 ∧
          (ℓ:ℝ)/8≤(ω.functional (spinZ d ℓ 0)).re := by
```

The file builds its own quasi-local algebra, infinite-volume dynamics (as a limit over boxes, with convergence part of the theorem) and KMS condition from scratch. I read them and they are the standard definitions, but this is the kind of statement where a definition audit matters more than usual. The family's other three papers, unformalized, go further: Bloch's $T^{3/2}$ law with the exact coefficient for any finite-range coupling generating $\mathbb Z^3$, the first lattice correction $3\zeta(5/2)(\beta S)^{-5/2}/(128\pi^{3/2})$ for nearest neighbours, and a spherical magnetization law in periodic cubes. OpenAI released a reasoning summary for this one too.

### The hyperinvariant subspace problem, and the transitive algebra problem with it

The invariant subspace problem asks whether every operator on a separable Hilbert space has a nontrivial closed invariant subspace. Its stronger cousin asks for a *hyperinvariant* one, invariant under every operator commuting with $T$. Both have been open since the 1950s for Hilbert space (Enflo and Read built Banach-space counterexamples to the weaker one). Lomonosov's 1973 theorem gives hyperinvariant subspaces whenever $T$ commutes with a nonzero compact operator.

Family 293 claims a negative answer: on every separable infinite-dimensional Hilbert space there is a nonzero quasinilpotent operator whose commutant is transitive (no nonzero proper closed subspace invariant under all of it), proper, unital and strongly closed. A proper WOT-closed transitive algebra is precisely a negative answer to Arveson's 1967 transitive algebra problem too, a consequence the companion abstract states and the family summary underplays. It does *not* settle the invariant subspace problem: $T$ may well have invariant subspaces that its commutant does not preserve.

Lean has exactly this (`HyperinvariantSubspaces.lean`, lines 27–31):

```lean
def FullClaim : Prop :=
  ∃ T : H →L[ℂ] H, T ≠ 0 ∧
    Tendsto (fun n : ℕ => ‖T ^ n‖ ^ (1 / (n : ℝ))) atTop (𝓝 0) ∧
    TransitiveCommutant T ∧ commutant T ≠ ⊤ ∧
    @IsClosed (H →L[ℂ] H) strongOperatorTopology (commutant T : Set (H →L[ℂ] H))
```

The companion answers a question of Zhu, Fang and Shi inside the hyperfinite II₁ factor: for every irrational angle, a weighted rotation whose continuous weight has a single zero with logarithmic integral $-\infty$ is nonzero, quasinilpotent and has no nontrivial invariant projection. That fits the known landscape: Haagerup and Schultz showed operators in II₁ factors whose Brown measure is not a point mass do have hyperinvariant projections, and the formalized product model here has Brown measure $\delta_0$, the one case their theorem cannot reach. That consistency is the kind of signal I like.

### Kaplansky's quasitraces are not traces

A 2-quasitrace is additive only on commuting elements. Kaplansky asked whether quasitraces on C\*-algebras are automatically traces. Blackadar and Handelman showed stably finite unital algebras carry quasitraces, so the question is equivalent to "does every stably finite C\*-algebra have a tracial state?". Haagerup proved yes for exact algebras ([arXiv:1403.7653](https://arxiv.org/abs/1403.7653)), a result that underpins the classification programme.

The 16-page paper builds a separable unital C\*-algebra with normalized 2-quasitraces and positive contractions $a, b$ such that every such quasitrace misses additivity by at least $1/144$, so there is no trace at all. By Haagerup's theorem the algebra has to be non-exact, and it is. A consequence: two unital simple stably finite C\*-algebras can have a properly infinite minimal tensor product, even when one of them is $C^*_r(\mathbb F_2)$. Both statements are in Lean (`KaplanskyQuasitrace.lean`, `KaplanskyStableFiniteness.lean`), with the $1/144$ in the statement itself.

### Parity is not in QAC⁰

Classically, constant-depth circuits with unbounded fan-in AND and OR cannot compute parity (Furst–Saxe–Sipser, Ajtai, Håstad). Moore asked in 1999 ([quant-ph/9903046](https://arxiv.org/abs/quant-ph/9903046)) whether the quantum analogue holds: can constant-depth circuits of one-qubit gates and many-qubit Toffolis compute parity? Lower bounds crept up for decades and stalled at limited ancillas: the Pauli-spectrum bounds of Nadimpalli, Parham, Vasconcelos and Yuen ([STOC 2024](https://arxiv.org/abs/2311.09631)) and the barely-superlinear-ancilla result of Anshu, Dong, Ou and Yao ([STOC 2025](https://arxiv.org/abs/2410.06499)).

The family claims the full thing with polynomially many ancillas: for every fixed depth and polynomial qubit budget, every large input length has an input on which the measured output qubit agrees with parity with probability below $1/2 + \varepsilon$. Ancillas start at zero, one output is measured, garbage is unrestricted, which is the standard bounded-error reading. The proof localizes the circuit with a product-projection estimate and inducts over depth. Both the advantage statement and the $2/3$ specialization are in Lean (`QACParity.lean`, `RegularParity.lean`), only 48 pages of paper behind them. Through Xu–Li reductions, strict majority follows.

### Three mutually unbiased bases in dimension six

Two orthonormal bases of $\mathbb C^d$ are mutually unbiased if every vector of one has overlap $1/d$ with every vector of the other; at most $d+1$ such bases exist, and $d+1$ exist in prime-power dimensions, where they give optimal state tomography. Dimension six is the smallest unknown case, and Zauner conjectured in 1991 that only three exist. Numerics strongly agreed. Nobody could even rule out a complete set of seven.

The headline paper claims $N(6) = 3$, and is honest that the upper bound is computer-assisted: under stated binary64 arithmetic and compiler conditions, a complete run of the documented pipeline excludes four bases. The search covers candidate vectors by small phase balls, discards ball pairs that cannot be orthogonal or unbiased using outward-rounded inequalities, and finally shows that two extra bases would need a pattern that never survives.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/physics-operator-algebras-fig3.png"
  alt="Two six-vertex cliques labelled A and B, each drawn with blue orthogonality edges, joined by all 36 orange cross edges marking unbiasedness."
  caption="What two more mutually unbiased bases in dimension six would need: two orthogonality 6-cliques with all 36 cross pairs unbiased. The computer search shows no such pattern survives (The maximum number of mutually unbiased bases in dimension six, Figure 2)."
/>

The Lean part is smaller but, to me, more interesting. The companion gives an exact integer-and-rational certificate excluding seven bases and, through Weiner's completion theorem, an upper bound of five, and proves the Matolcsi–Ruzsa–Weiner Fourier-vanishing conjecture for Hadamard matrices of order six outside Tao's class. Those are formalized (`lean/ComparatorChallenges/MUBSix.lean`, lines 50–59):

```lean
def IsMUBFamily {n : ℕ} (B : Fin n → Basis) : Prop :=
  ∀ r s, r ≠ s → ∀ i j,
    Complex.normSq (inner ℂ (B r i) (B s j)) = (1 : ℝ) / 6

def Attainable (n : ℕ) : Prop := ∃ B : Fin n → Basis, IsMUBFamily B

theorem fourier_and_family_bound :
    (∀ H : CMatrix, IsHadamard H → ¬Equivalent H tao →
      ∀ π : Equiv.Perm Coord, g H (permuteCharge π alpha) = 0) ∧
    (∀ n : ℕ, Attainable n → n ≤ 5) := by
```

So the machine-checked statement is "no complete set of MUBs in dimension six", itself a famous open question, and the claimed value 3 rests on a floating-point computation. I would describe it exactly that way.

### PPT² is false, and entanglement without secret key

Christandl's PPT-squared conjecture says that composing a PPT channel with itself always destroys entanglement. It was proved in dimension three and asymptotically, and it was a favourite open problem in quantum information. The family gives a trace-preserving PPT channel on $M_{21}(\mathbb C)$ whose square is not entanglement breaking, plus two PPT maps on $10 \times 10$ matrices whose composition is not entanglement breaking (disproving the two-map version, Conjecture IV.1 of Christandl, Müller-Hermes and Wolf, [AHP 2019](https://doi.org/10.1007/s00023-019-00774-7)). Both counterexamples are in Lean.

The same construction gives an entangled state on $\mathbb C^{10} \otimes \mathbb C^{10}$ from which no secret key can be distilled, answering Problem 24 of Krüger and Werner's open-problem list ([quant-ph/0504166](https://arxiv.org/abs/quant-ph/0504166)). That part is not formalized, and it holds for "the local-instrument protocols specified here" (joint processing of all copies, unlimited authenticated two-way public communication, protocols completing almost surely). Whether that class captures every operational key-distillation protocol is a question for a specialist; I would not repeat the zero-key claim without that qualifier.

### Counterexamples to Baum–Connes and Kadison–Kaplansky

This claim would reshape noncommutative geometry if it holds, and it is the one I trust least, for one reason: nothing about it is formalized.

The Baum–Connes conjecture predicts that the K-theory of a group's reduced C\*-algebra is computed by an index map from geometry. It was proved for huge classes (Higson–Kasparov for a-T-menable groups, Lafforgue and Mineyev–Yu for hyperbolic groups). Counterexamples were known only for the version with coefficients (Higson, Lafforgue and Skandalis, [GAFA 2002](https://doi.org/10.1007/s00039-002-8249-5)), using expanders. A well-known consequence of the coefficient-free conjecture is the Kadison–Kaplansky conjecture: the reduced C\*-algebra of a torsion-free group has no projections besides 0 and 1.

The family claims three things. A finitely generated torsion-free group whose degree-zero assembly map has a kernel class of infinite order. A finitely generated torsion-free group with a projection $e \in C^*_r(G)$ of trace strictly between 0 and ½, killing Kadison–Kaplansky. And a finitely generated group with a projection of irrational trace, which cannot be in the image of assembly because Lück's theorem makes those traces rational. The mechanism in each case is spectral: build a group from an infinite graphical presentation over labelled expander-like graphs so that a finite matrix over the group algebra has a uniform spectral gap in the regular representation; continuous functional calculus then produces a genuine projection in the reduced C\*-algebra whose trace you can compute. The projections are not in the group ring, which a companion shows has no nontrivial idempotents.

This is the first thing I would ask experts to referee. It is 184 pages across three papers, there is no Comparator challenge, and the groups are by necessity far from every class where the conjecture is known.

### Anderson localization in two dimensions and delocalization in three

Simon's "Schrödinger operators in the twenty-first century" ([2000](https://doi.org/10.1142/9781848160224_0014)) opens with the extended-states conjecture: for the Anderson model on $\mathbb Z^d$, $d \ge 3$, weak disorder should leave some purely absolutely continuous spectrum. Nobody had ever proved ac spectrum for the lattice model in any dimension; it was known only on trees. The other half of the folklore picture is that in two dimensions any disorder localizes everything; rigorous 2D results covered only band edges, such as Ding and Smart's Bernoulli result ([arXiv:1809.09041](https://arxiv.org/abs/1809.09041)).

The family claims both, for i.i.d. uniform potentials: in $d \ge 3$ at small fixed disorder, almost surely purely ac spectrum on a fixed open interval with nonzero weight; in 2D at every disorder, almost surely pure point spectrum everywhere. The 3D proof is a multiscale scheme that marks uncertain "pending" sites, retries them with fresh randomness, and controls their effect with unitary scattering matrices and martingale estimates; the 2D proof gets an impedance estimate uniform in exterior data and finishes with the Simon–Wolff rank-one criterion.

The Lean here formalizes only that the almost-sure spectrum on $\mathbb Z^2$ is $[-4-h, 4+h]$, a textbook fact. The spectral type, which is the entire content, is not formalized, and neither is anything in $d \ge 3$. Two of the most famous problems in mathematical physics, 171 pages, no machine check. Hold this one loosely.

### The spin-one Haldane gap

Haldane predicted in 1981 that antiferromagnetic chains of integer spin have a spectral gap while half-integer chains do not. The half-integer side has a rigorous antecedent (Lieb–Schultz–Mattis, Affleck–Lieb 1986). The integer side was only known for the AKLT model with an extra biquadratic term, and for small perturbations of it (Yarotsky). Numerics put the spin-1 Heisenberg gap at about 0.4105 (White–Huse), and neutron scattering sees it.

The paper claims the pure bilinear chain $H_L = \sum_j \mathbf S_j \cdot \mathbf S_{j+1}$ on even periodic rings has a unique ground state for $L \ge 60$ and a gap above $\log(20)/784$ for even $L \ge 2304$. The proof writes one partition function two ways: as a sum over physical eigenvalues, and through a self-adjoint spatial transfer operator as a moment in the length. Both give probability distributions, and their "purity" (sum of squared weights) improves quadratically when you double length and inverse temperature together. Two finite certificates, rigorous enclosures of matrix-exponential traces of small twisted chains at inverse temperatures $21/2$ and $49/4$, start the bootstrap, and interpolation fills in every even length.

<Figure
  src="https://ai.thesatyajit.com/articles/openai-math/physics-operator-algebras-fig2.png"
  alt="Plot of inverse temperature b against chain length L: three solid horizontal segments at b0, 2b0 and 4b0 covering L from n0 to 2n0, 2n0 to 4n0 and 4n0 to 8n0, joined by dashed arrows that double both coordinates."
  caption="The Haldane-gap bootstrap: doubling length and inverse temperature together improves both purity estimates, and interpolation covers every even length in between (The periodic spin-one Haldane gap, Figure 1)."
/>

It is 30 pages, which is either a sign of a genuinely new idea or a warning. The bound is about 100 times smaller than the numerical gap, which is fine for existence. There is no Lean, and the computer-assisted certificates need to be rerun independently. A companion handles odd open chains with endpoint field $h = 3/5$ (gap above $\log(10)/392$ for $L \ge 960$), which via Tasaki's index theorem gives index −1 for the boundary-selected states.

### The ionization conjecture

How many electrons can a nucleus of charge $Z$ bind? Experiment says $Z+1$ or $Z+2$. The best theorems said roughly $2Z$ (Lieb 1984) and then $1.22Z + 3Z^{1/3}$ (Nam, [CMP 2012](https://doi.org/10.1007/s00220-012-1479-y)), with asymptotic neutrality $N/Z \to 1$ from Lieb–Sigal–Simon–Thirring and Fefferman–Seco. Solovej proved the conjecture inside Hartree–Fock theory ([Ann. Math. 2003](https://doi.org/10.4007/annals.2003.158.509)), not for the Schrödinger equation.

The family claims the real thing: for the full nonrelativistic Coulomb Hamiltonian with two spin states, $M$ nuclei of charges at least one and total charge $Z$ bind at most $Z + CM$ electrons, with a universal constant $C$; for neutral atoms the first ionization energy and the radius outside which half an electron remains are pinned between universal constants. Two companions prove the generalized ionization conjecture: removing $m$ electrons costs $a_{\rm TF}\, m^{7/3}$ asymptotically, and the outer radii follow $(81\pi^2/2)^{1/3} m^{-1/3}$.

The Lean covers the two asymptotic companions (`CoulombIonization.lean`, `CoulombRadii.lean`), which are hard theorems in their own right. The $Z + CM$ bound itself, the actual ionization conjecture, is not formalized, and $C$ is not explicit.

### Bose–Einstein condensation at positive temperature

Proving that an interacting Bose gas condenses in the thermodynamic limit is one of the best-known open problems in the field; Lieb, Seiringer, Solovej and Yngvason's monograph treats it as the central question. Proofs existed only in scaling limits: Gross–Pitaevskii ([Lieb–Seiringer 2002](https://arxiv.org/abs/math-ph/0112032)), dilute trapped gases at positive temperature ([Deuchert–Seiringer–Yngvason, CMP 2019](https://doi.org/10.1007/s00220-018-3239-0)) and lattice models. Even at zero temperature, condensation in the true thermodynamic limit was open.

The family claims, for the 3D hard-sphere gas at small fixed density, a fixed temperature $T > 0$ at which the exact canonical Gibbs state has positive condensate fraction in the constant orbital as volume goes to infinity. Companions prove condensation in every ground state with a fraction uniform in the gas parameter, and the Bogoliubov depletion law $\tfrac{8}{3\sqrt\pi}\sqrt{\rho a^3}$ with the thermodynamic limit taken before the dilute limit. The proofs use random routes between cells in a path representation with exponential-moment bounds on encounters.

Lean has the ground-state condensation statement (`HardSphere.lean`): absolute constants $\varepsilon_0, c_0 > 0$ so that $\rho a^3 < \varepsilon_0$ gives condensate fraction at least $c_0$ along every thermodynamic sequence. If that statement is faithful, it alone answers a famous question. The positive-temperature theorem in the family's title is paper-only, and the temperature it produces is small and not meant to locate the transition.

### The spacetime Penrose inequality

Penrose proposed in 1973 that a spacetime containing trapped surfaces must have ADM mass at least that of a Schwarzschild black hole of the same size, as a test of cosmic censorship. The time-symmetric (Riemannian) version was proved in dimension three by Huisken–Ilmanen ([JDG 2001](https://doi.org/10.4310/jdg/1090349447)) and Bray, and for dimensions below eight by Bray–Lee ([arXiv:0705.1128](https://arxiv.org/abs/0705.1128)). The general spacetime version was open; Ben-Dov's and Carrasco–Mars' counterexamples showed it must be stated with a minimal enclosing area rather than the horizon's own area.

This is the largest family in the release I looked at: 13 manuscripts and about 1,355 pages. The core claim is the enclosing-area inequality

$$
m \;\ge\; \frac12 \left(\frac{A_{\min}}{\omega_{n-1}}\right)^{\frac{n-2}{n-1}}
$$

for one-ended asymptotically flat data in every spatial dimension $n \ge 3$, under the dominant energy condition, weak future trapping ($\theta_+ \le 0$ only), stated decay and positive enclosing area, with arbitrary second fundamental form. Around it are rigidity (equality forces a Schwarzschild–Tangherlini slice under extra horizon hypotheses), charged versions, a Kerr–Newman inequality with angular momentum, counterexamples showing the Kerr–Newman version fails if $J$ is the bare ADM angular momentum with slowly decaying fields, anti-de Sitter versions, and a Riemannian Penrose inequality in every dimension.

"Every dimension" includes $n \ge 8$, where even the Riemannian case was open because minimal surfaces can be singular. The Kerr–Newman version assumes zero ADM momentum, Coulomb asymptotics and the physical-area condition $A \ge 4\pi\sqrt{Q^4 + 4J^2}$ as hypotheses. And the Lean covers only a supporting construction: replacing a Cha–Khuri–Sakovich hyperboloidal end by asymptotically flat ends, plus Schwarzschild equality examples; `lean/docs/260.md` says the general inequality remains unformalized. I cannot tell you whether 1,355 pages of geometric analysis are right, and nobody else will be able to for a while.

### An area law for two-dimensional gapped systems

Hastings proved in 2007 that gapped one-dimensional ground states obey an entanglement area law ([arXiv:0705.2024](https://arxiv.org/abs/0705.2024)), which explains why matrix-product methods work. In two dimensions the conjecture was open except for frustration-free systems with local gaps ([Anshu, Arad and Gosset, STOC 2022](https://arxiv.org/abs/2103.02492)). The family claims the general 2D statement from a global gap alone: for a finite-range Hamiltonian with a unique ground state on any finite induced subgraph of $\mathbb Z^2$, each region's entanglement entropy is at most a constant times its boundary edges. A companion gives PEPS approximations with bond dimension polynomial in $L$ and vector error at most $1/L$ on $L \times L$ squares, an existence result, not an algorithm. No Lean; 146 pages.

### Connes' bicentralizer problem

Connes asked in 1980 whether the bicentralizer of every faithful normal state on a type III₁ factor is trivial. Haagerup's proof for the hyperfinite case is what gave uniqueness of the hyperfinite III₁ factor, which shows the stakes. Ando, Haagerup, Houdayer and Marrakchi introduced relative bicentralizers and proved many cases ([arXiv:1804.05706](https://arxiv.org/abs/1804.05706)). The family claims the general relative conjecture (every expected inclusion $N \subset M$ with separable preduals contains an expected amenable $P \subset N$ with $P' \cap c(M) = N' \cap c(M)$), hence Connes' problem for every III₁ factor with separable predual, in 35 pages over two papers. Lean formalizes only a supporting spectral-recovery theorem, and its docs say the bicentralizer conjecture itself is not included.

### Toms–Winter: strict comparison implies Jiang–Su stability

The Elliott classification programme classifies simple nuclear C\*-algebras that absorb the Jiang–Su algebra $\mathcal Z$. Toms and Winter conjectured that $\mathcal Z$-stability, finite nuclear dimension and strict comparison are equivalent. The first two were shown equivalent by Castillejos, Evington, Tikuisis, White and Winter ([Invent. Math. 2021](https://doi.org/10.1007/s00222-020-01013-1)); strict comparison implies $\mathcal Z$-stability was known only under conditions on the trace space or with uniform property Γ.

Family 291 bundles four papers: the unital Toms–Winter implication (strict comparison gives $\mathcal Z$-absorption for simple separable unital nuclear algebras), a nonunital Cuntz-semigroup version, Robert and Tikuisis' Conjecture C1 ($\dim_{\rm nuc}(A \otimes \mathcal Z) \le 1$ for every separable nuclear $A$), and the unital stably finite case of Szabó's equivariant Conjecture A. The Lean doc says the linked comparison statements prove Jiang–Su absorption with strict comparison in the unital case. The family summary leads with the equivariant result; I think the comparison theorem is the bigger news.

### Strong cosmic censorship near Kerr

The C⁰ version of strong cosmic censorship is false near Kerr: Dafermos and Luk showed the Cauchy horizon is $C^0$-stable ([arXiv:1710.01722](https://arxiv.org/abs/1710.01722)). Christodoulou's reformulation asks instead that generic data admit no extension with locally square-integrable connection. Luk and Oh proved generic $C^2$ inextendibility in spherical symmetry ([arXiv:1702.05715](https://arxiv.org/abs/1702.05715)). Family 264 claims, with no symmetry, that near each fixed rotating subextremal Kerr bridge the data admitting a future $C^0 \cap W^{1,2}_{\rm loc}$ extension form a meagre set. The design is clever: it avoids needing full Kerr stability by testing extensions with averaged holonomies of small loops, injecting late gravitational wave packets that make a curvature block measurably large, and closing with a Baire category argument. It is local, generic and unformalized, 278 pages; a major result if right, not "strong cosmic censorship".

### Sharp one-dimensional Lieb–Thirring constants

The 1D Lieb–Thirring conjecture, that the sharp constant for $\tfrac12 < \gamma < \tfrac32$ is the one-bound-state constant, is claimed in the remaining range, with matrix-valued potentials and all equality cases (direct sums of $\operatorname{sech}^2$ solitons). The $\gamma = 1$ case had been settled earlier in September by Read and Schulz ([arXiv:2609.10478](https://arxiv.org/abs/2609.10478)). The scalar theorem is in Lean; the matrix version is not.

### The Laughlin spectral gap

The Laughlin spectral-gap conjecture for the $V_1$ pseudopotential at filling ⅓ on the sphere gets a gap of at least $1/25$ for large systems, via the Fock-space inequality $H^2 \ge \gamma H$, and stability under weak projected disorder. Earlier work (Nachtergaele, Warzel and Young, [CMP 2021](https://doi.org/10.1007/s00220-021-03997-0)) handled truncated thin-cylinder models. The unperturbed gap is formalized, the disorder stability is not.

### The BFSS threshold bound state, and a surprise

For the BFSS matrix model, the family claims exactly one normalizable zero-energy state for SU(N) at every $N$, which is the D0-brane threshold bound state that matrix theory needs and which was known rigorously only for SU(2) ([Sethi–Stern](https://arxiv.org/abs/hep-th/9705046)). The companion's claim that relative SU(2) BFSS has infinitely many normalizable positive-energy eigenstates contradicts a statement in the original BFSS paper ([hep-th/9610043](https://arxiv.org/abs/hep-th/9610043)) and depends on the precise closure of the supercharge form. I would want physicists to look at that one. No Lean.

### The entropy photon-number inequality

The entropy photon-number inequality of Guha, Erkmen and Shapiro ([arXiv:0710.5666](https://arxiv.org/abs/0710.5666)), the photon-number analogue of the entropy power inequality for beam splitters, is proved for finite-energy multimode inputs in 32 pages and formalized. It is one of the cleanest wins in the group.

### Vertex operator algebras and conformal nets

Every simple unitary strongly rational vertex operator algebra gives a completely rational conformal net, with matching braided tensor categories, answering the strongly rational case of the strong-locality conjecture of Carpi, Kawahigashi, Longo and Weiner ([arXiv:1503.01260](https://arxiv.org/abs/1503.01260)). The net construction is in Lean, the tensor equivalence is not.

### Unitary synthesis from a Boolean oracle

Aaronson and Kuperberg's unitary synthesis problem gets a positive answer in its constant-error form: a uniform polynomial-size oracle circuit reaches every $n$-qubit unitary within diamond distance ½ for a suitable target-dependent Boolean oracle, whose efficient construction is not claimed. Short and unformalized.

### The randomized-versus-quantum query exponent is 4

The randomized-versus-quantum query exponent for total functions is claimed to be exactly 4: a nearly quartic separation disproves the cubic conjecture of Aaronson, Ben-David, Kothari, Rao and Tal ([arXiv:2010.12629](https://arxiv.org/abs/2010.12629)), whose $O(Q^4)$ upper bound is thereby optimal. At 19 pages this is the fastest family in the group for experts to check.

### Kirchberg's embedding problem

Kirchberg and Phillips showed every separable exact C\*-algebra embeds in the Cuntz algebra $\mathcal O_2$, and Kirchberg asked whether every separable C\*-algebra at least embeds in its norm ultrapower $\mathcal O_2^\omega$ (see Goldbring and Sinclair, [arXiv:1404.1861](https://arxiv.org/abs/1404.1861)). The answer is claimed to be no, and the 14-page paper is formalized: the full group C\*-algebra of $\mathbb Z[\tfrac12]^3 \rtimes (\mathrm{SL}_3(\mathbb Z) \times \mathbb Z)$, a group Sauers and Eckhardt had already used to show non-finiteness, embeds in no norm ultrapower of any nonzero unital nuclear algebra. The obstruction is representation-theoretic: the paper tests how a fixed finite-dimensional representation can be approximately matched inside a nuclear ultrapower, shows every match needs a predecessor, and gets a contradiction because finite chains of matches are forced to be periodic.

### Microstates and non-microstates free entropy differ

Voiculescu defined free entropy twice, once by counting matrix approximations ($\chi$) and once through a free Fisher information ($\chi^*$). Biane, Capitaine and Guionnet proved $\chi \le \chi^*$ ([Invent. Math. 2003](https://doi.org/10.1007/s00222-002-0281-4)), and whether they agree when finite was open. The 15-page paper builds a bounded self-adjoint tuple, with a large but fixed number of variables, with $-\infty < \chi < \chi^* - \tfrac12 < \infty$. It is formalized, which means someone put microstates free entropy into Lean; that definition deserves an audit of its normalizations (operator-norm cutoff, limsup over matrix sizes).

### The Kirchberg–Rørdam character criterion

$\mathcal Z$-stability is the regularity property behind classification. Kirchberg and Rørdam asked whether it is detected by the absence of characters on the norm central-sequence algebra ([arXiv:1409.1395](https://arxiv.org/abs/1409.1395)), and Dadarlat and Toms asked whether infinite tensor powers of algebras without characters absorb $\mathcal Z$. Both are claimed for every nonzero unital separable C\*-algebra, with no nuclearity, simplicity or trace hypothesis, and the character criterion is formalized for every free ultrafilter.

### Weak pure infiniteness is strong pure infiniteness

Kirchberg and Rørdam introduced several notions of pure infiniteness for non-simple algebras and asked how they compare. The family proves that if every positive element is properly infinite, the algebra is strongly purely infinite, with no exactness, nuclearity, unitality or simplicity assumption; for exact algebras one fixed amplification suffices, and separable nuclear algebras of this kind absorb $\mathcal O_\infty$. Both formal statements are in Lean.

### Radius of comparison equals half the mean dimension

For a minimal homeomorphism $h$ of a compact metric space $X$, Phillips and Toms conjectured that the radius of comparison of $C(X) \rtimes_h \mathbb Z$ is $\tfrac12\,\mathrm{mdim}(X, h)$. Elliott and Niu had the zero-mean-dimension case ([arXiv:1406.2382](https://arxiv.org/abs/1406.2382)) and Hirshberg and Phillips lower bounds ([arXiv:2009.13045](https://arxiv.org/abs/2009.13045)). The family claims the equality for every minimal $\mathbb Z$-system, infinite values included, and the key input is unusual: a coherent vanishing theorem in complex cobordism that compresses cube-valued maps while preserving boundaries. No Lean, and it needs homotopy theorists as well as operator algebraists to check.

### Naimark's problem in ZFC, second by two days

Naimark's problem in ZFC (a simple non-elementary C\*-algebra with a unique irreducible representation, with no extra set-theoretic axiom) is formalized, but it is not first. Akemann and Weaver had a counterexample assuming ◇ ([2004](https://arxiv.org/abs/math/0312135)), and Ryotaro Tanaka posted a ZFC construction ([arXiv:2609.26930](https://arxiv.org/abs/2609.26930)) two days before OpenAI's manuscript is dated. OpenAI's own summary calls theirs "an alternative" to Tanaka's.

### Notable results

Approximating the electronic ground-state energy of molecules is QMA-hard even with only unit-charge clamped nuclei in the continuum, with no basis, magnetic field or extra potential, strengthening Schuch and Verstraete's 2009 result ([arXiv:0712.0483](https://arxiv.org/abs/0712.0483)); formalized. The classical capacity of every generalized amplitude-damping channel equals its one-shot Holevo value, closing a question Leditzky, Kaur, Datta and Wilde listed as open ([arXiv:1709.01111](https://arxiv.org/abs/1709.01111)); formalized, except additivity with arbitrary partner channels. Exponential threshold parallel repetition for all finite entangled games is formalized, though with a $\delta^{13}$ rate where the paper proves $\delta^5$, and the paper credits the hard all-wins step to Chapter 6 of OpenAI's earlier *Ten Advances in Mathematics and Theoretical Computer Science* and to Zhao Song's August 2026 ECCC report, so its own contribution is the transfer to thresholds. A three-electron molecule whose ground-state density no noninteracting ensemble in an $L^{3/2} + L^\infty$ potential reproduces breaks Kohn–Sham ensemble representability, the assumption under density functional theory, with a nuclear charge given only by a formula. Exact zero-error factoring over a fixed finite gate set in polynomial time is formalized. QAOA reaches the Sherrington–Kirkpatrick ground-state energy when size goes to infinity before depth, proving the conjecture of Basso, Farhi, Marwaha, Villalonga and Zhou ([arXiv:2110.14206](https://arxiv.org/abs/2110.14206)) with no depth bound; the Lean covers supporting Parisi-measure facts only. Scale invariance implies local conformal invariance in 4D, but only inside an axiomatic framework (bounded local net, discrete scaling spectrum, a physical local dilatation current), so it does not settle the physics question. Popa and Vaes' quadratic paving conjecture over arbitrary maximal abelian subalgebras ([arXiv:1412.0631](https://arxiv.org/abs/1412.0631)) is claimed with $5 \times 10^8\,\varepsilon^{-2}$ projections. And the trace cone classifies separable nuclear algebras after tensoring with the Razak–Jacelon algebra and the compacts, answering Robert's question.

### Technical results

Family 286 classifies finite-index bimodules between twisted group factors of lattices commensurable with property (T) lattices over local fields, recovering the group, cocycle and scale. It is 190 pages, unformalized, and builds on OpenAI's own earlier counterexample to Connes' rigidity conjecture.

### What I would ask an expert to check first

If I could send five things to referees tomorrow, I would send the Baum–Connes and Kadison–Kaplansky constructions (the biggest claim with no formal check), the Anderson pair (two famous problems, 171 pages, nothing formal), the Haldane certificates (short proof, computer-assisted inputs that need rerunning), the Lean statement files for families 287 and 288 (to confirm the definitions are the textbook ones and the build passes), and the BFSS positive-energy claim (it contradicts the physics literature). Everything else either has a formal statement I could read or is small enough that an expert will settle it within weeks.

## Geometry and topology

Geometry and topology are fields where proofs are long, pictures carry the argument, and a referee often needs months. This group is also the densest concentration of famous names in the catalogue: 47 families and 75 manuscripts, about 3,400 pages, and among them the Hilbert–Smith conjecture, the purely cosmetic surgery conjecture, Quillen's conjecture, Wall's D(2) problem, the Gromov–Lawson aspherical conjecture, the Cartan–Hadamard conjecture, Yau's uniformization conjecture, Katok's entropy rigidity conjecture, the nearby Lagrangian conjecture, Donaldson's tamed-to-compatible question, the metric Blaschke conjecture, Allard-type regularity for stationary varifolds and Yau's nodal conjecture.

I cannot certify any of these proofs, and nobody outside OpenAI has had time to. So I sorted the claims by what actually backs them. Twelve families have their headline theorem stated and proved in the release's Lean library, and six more have a piece of it. The others stand on the manuscript alone. That split matters more than the fame of the problem.

My grading is stingy, and it is still unusual: 16 landmarks, 22 majors, 5 notables, 4 technical. In most disciplines a "landmark" is a once-a-decade event. Here the model claims sixteen in one batch. That alone is a reason to read the rest of this section as a map of claims, not a list of theorems.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig1.png" alt="Page 33 of the openai/math overview catalogue: the start of the Topology section, listing families 304 (Hilbert–Smith), 305 (disk embedding and Wall's manifold conjecture), 306 (purely cosmetic surgery) and 307 (maximal coarse assembly)." caption="How the release itself presents the topology results: one-paragraph claims, a link per manuscript, no proof status. Families 304 to 307 open the section (OpenAI Research Catalog, overview.pdf page 33)." />

One warning specific to geometry: a Comparator statement is only as good as its definitions, and here those definitions (currents, Riemannian curvature, symplectic forms on charted spaces) are written from scratch in each challenge file, because mathlib has few of them.

### Donaldson's hypersymplectic conjecture

The companion [hypersymplectic paper](https://github.com/openai/math/blob/main/preprints/Deforming-hypersymplectic-four-manifolds-to-hyperkahler-triples-September-23-2026/paper.pdf) (72 pages) reuses those current estimates. A hypersymplectic structure is a triple of closed 2-forms whose span is positive definite at every point, like a hyperkähler triple without parallelism. Donaldson conjectured in the same 2006 paper that such a triple, normalized so $\int \omega_i \wedge \omega_j = \delta_{ij}$, deforms with fixed cohomology classes to a hyperkähler triple. The paper proves it and draws the striking corollary that a closed 4-manifold with a hypersymplectic structure is diffeomorphic to K3 or $T^4$. That corollary is where I would expect experts to push first. None of the deformation argument is in Lean, so it inherits the verified estimates of the tamed-to-compatible proof and nothing more.

### Yau's nodal conjecture, settled both ways (formalized)

Yau conjectured that on a closed smooth $n$-manifold, the zero set of an eigenfunction $-\Delta u = \lambda u$ has $(n-1)$-dimensional measure between $c\sqrt\lambda$ and $C\sqrt\lambda$. Donnelly and Fefferman proved both bounds for real-analytic metrics in 1988. For smooth metrics, Logunov proved the [lower bound](https://arxiv.org/abs/1605.02589) and a [polynomial upper bound](https://arxiv.org/abs/1605.02587) in 2016; on surfaces the best upper exponent was $3/4 - \beta$ (Logunov–Malinnikova).

The release says the upper bound is true on surfaces and false in every higher dimension:

- [On every smooth closed surface](https://github.com/openai/math/blob/main/preprints/Sharp-nodal-length-on-smooth-surfaces-September-23-2026/paper.pdf), the nodal length is at most $C(M,g)\sqrt\lambda$. The proof bounds the average logarithmic growth of $u$ over squares of side about $1/\sqrt\lambda$, then converts growth to length with Roy-Fortin's planar theorem.
- [On $S^3$](https://github.com/openai/math/blob/main/preprints/Smooth-counterexamples-to-Yaus-nodal-upper-bound-in-dimensions-three-and-four-September-23-2026/paper.pdf), with metrics as close to round as you like, and on $S^2 \times T^2$, there are fixed smooth metrics with eigenfunctions whose nodal measure divided by $\sqrt\lambda$ tends to infinity.
- [On $S^4 \times S^1$](https://github.com/openai/math/blob/main/preprints/Power-law-violations-of-Yaus-nodal-upper-bound-September-23-2026/paper.pdf) the excess is a fixed power, $\lambda^{1/2 + \varepsilon_0}$.

The counterexamples glue localized quasimodes, whose zero sets are packed into small regions, into exact eigenfunctions of one fixed smooth metric. All three directions are Comparator challenges with sorry-free solutions (`NodalLength`, `SmoothYau`, `YauCounterexample`). The answer is clean and a little surprising: Yau's square-root law is a theorem in dimension 2 and an artifact of analyticity above it. The counterexamples say nothing about analytic metrics, where Donnelly–Fefferman stands, and nothing about the lower bound.

### Wall's D(2) problem: the one I checked by hand

Most landmark claims here run to a hundred pages. [This one](https://github.com/openai/math/blob/main/preprints/A-Counterexample-to-Walls-D2-Problem-October-6-2026/wall-d2-counterexample.pdf) is nine, and it is elementary enough that I could check its computational heart.

Wall's D(2) problem (1965; Problem D3 in his 1979 list) asks: if a finite 3-dimensional CW complex $X$ satisfies $H_i(\widetilde X) = 0$ for $i > 2$ and $H^3(X; M) = 0$ for every coefficient module $M$, is $X$ homotopy equivalent to a finite 2-complex? Adding enough 2-spheres makes the answer yes (Cohen; Hambleton). For specific finite fundamental groups people have proved it case by case, most recently Hofmann and Nicholson for quaternion groups of order 24, 28 and 32. Bridson and Tweedale had candidate counterexamples conditional on a relation gap.

The paper's argument has two parts. First, a lemma about any finite 2-complex $Y$: if a character $\rho: \pi_1 Y \to \mathbb C^\times$ has vanishing twisted second homology, $H_2(Y; \mathbb C_\rho) = 0$, then $\rho$ is trivial on every element of finite order. The proof is a tower. A map from $\mathbb{RP}^2$ that sees $\rho(g) = -1$ lifts to infinite cyclic covers again and again (because $b_1(\mathbb{RP}^2) = 0$), each lift has a strictly larger image, and every image is a quotient of the same finite triangulation. The vanishing of twisted $H_2$ survives each cover by a Laurent-polynomial rank argument, which is what keeps the tower going.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig4.png" alt="Commutative triangle: F maps by f_i to A_i and by f_(i+1) to A_(i+1), which includes into the infinite cyclic cover of A_i; below, the simplex counts s(A_i) less than s(A_(i+1)) at most s(F)." caption="One step of the tower in the D(2) counterexample: each lift into an infinite cyclic cover has a strictly larger image, but every image is a quotient of the same finite triangulation, so the tower cannot continue forever (A Counterexample to Wall's D(2) Problem, Figure 1)." />

Second, a concrete group. Take

$$
P = \langle x_1, a_1, x_2, a_2, s \mid x_1 a_1 x_1^{-1} = a_1^4,\ x_2 a_2 x_2^{-1} = a_2^3,\ x_1 = a_2^{13},\ s x_2 s^{-1} = a_1^5 \rangle
$$

and kill $t_1 = x_1^2$ and $t_2 = x_2^3$. The normal closure $K$ of $t_1, t_2$ is perfect, because $15$ and $26$ annihilate its abelianization and are coprime. Quillen's plus construction then gives a finite D(2) complex $X$ with $\pi_1 X = G = P/K$. The character $\rho(x_1) = -1$, $\rho(a_1) = \zeta$, $\rho(x_2) = \zeta^2$, $\rho(a_2) = -1$, $\rho(s) = 1$ (with $\zeta = e^{2\pi i/3}$) respects every relation, kills $t_1, t_2$, and makes the twisted boundary matrix injective. But $x_1$ has order exactly 2 in $G$ and $\rho(x_1) = -1$. No 2-complex can have both, so $X$ has no 2-dimensional model.

You can recompute that matrix yourself below. Pick a sixth root of unity for each generator; the widget checks the relations, computes the Fox derivatives in exact arithmetic, and tells you whether the character obstructs a 2-dimensional model.

<WallD2CharacterCheck />

I did the same check independently in Python: all four relators evaluate to 1 under $\rho$, the matrix matches equation (15) of the paper, and the minor on columns $x_1, a_1, x_2, s$ is $-4(1 - \zeta^2) \neq 0$. I also re-derived the perfectness argument for $K$ and read the tower lemma line by line. I did not find a gap. Two qualifications. The version of the D(2) problem most people work on is for finite fundamental groups, and that remains open; this group is infinite (it surjects onto $\mathbb Z$ through $s$). And it is odd that the bibliography cites Hofmann–Nicholson as appearing in a 2027 journal volume. Neither weakens the argument. If I had to bet on one landmark in this group surviving review intact, it would be this one, precisely because a reader can check it in an afternoon.

### The Gromov–Lawson aspherical conjecture, in every dimension

A metric of positive scalar curvature (PSC) makes small balls a little smaller than Euclidean ones. Gromov and Lawson conjectured that closed aspherical manifolds, such as tori and hyperbolic manifolds, never carry one. It was known in dimension 3 (Schoen–Yau, Gromov–Lawson), for spin manifolds whose groups satisfy the Novikov-type conjectures, and in dimensions 4 and 5 ([Chodosh–Li](https://arxiv.org/abs/2008.11888), [Gromov](https://arxiv.org/abs/2009.05332)), where minimal-surface and $\mu$-bubble methods stop.

[The release](https://github.com/openai/math/blob/main/preprints/Positive-scalar-curvature-forces-rational-inessentiality-September-23-2026/paper.pdf) proves the stronger rational form in every dimension and with no spin hypothesis: a closed oriented manifold with PSC is rationally inessential, $c_*[M] = 0$ in $H_n(B\pi_1 M; \mathbb Q)$. Aspherical manifolds are essential, so they have no PSC metric, and any metric of nonnegative scalar curvature on one is flat. A [second paper](https://github.com/openai/math/blob/main/preprints/An-integral-scalar-curvature-bound-for-real-simplicial-volume-October-5-2026/v126-proof.pdf) uses this to prove Gromov's 1986 integral inequality $\int_M (\mathrm{Scal}^-_g)^{n/2} \ge a_n \lVert M \rVert$ for simplicial volume.

The mechanism avoids Dirac operators and minimal surfaces entirely. A detecting cohomology class becomes a map $G: Z \to \mathbb R^q$ of nonzero local degree on a fibred cover, with small fibre derivatives. Then the metric is reshaped one coordinate at a time: a graph deformation plus one warped circle per coordinate, each solved by a nonlinear elliptic equation. Paying a separate scalar-curvature cost per coordinate would cost a factor $q$ and ruin everything. The paper's key estimate is that the decreases of the inverse metric telescope, so a whole pass costs $C(n+1)(2 + \log q)L^2/w^2$. Repeating passes shrinks the map geometrically until a torical obstruction of Cecchini and Schick is violated.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig5.png" alt="Boxes for original fibre directions of dimension n, parameter directions of dimension q minus n, and stabilizing circles of dimension kq; the base Z of dimension q maps by F_k with nonzero degree to the sphere S^q, with scalar curvature controlled on Z times the torus." caption="The bookkeeping behind the all-dimensions Gromov–Lawson proof: the sphere map lives on the base, the scalar-curvature bound on the base times kq warped circles, and only the n original directions pay for shortening (Positive scalar curvature forces rational inessentiality, Figure 1)." />

No Lean, 121 pages across the two papers. The logarithmic cost estimate is the place where a single wrong inequality would sink the whole thing, and it is exactly the kind of analysis that takes specialists weeks to check.

### Gromov's codimension-2 width conjecture

A separate family proves a different Gromov conjecture with the same corollary. For $n \ge 4$, [any complete $n$-manifold with $\mathrm{Scal} \ge 1$](https://github.com/openai/math/blob/main/preprints/Positive-scalar-curvature-and-uniform-codimension-two-width-September-23-2026/paper.pdf) maps continuously to an $(n-2)$-dimensional complex with every fibre of diameter at most $C_n$. Positive scalar curvature makes a manifold thin in two directions. Companion papers prove it under the weaker spectral condition $-4\Delta + \mathrm{Scal} \ge 1$, with an explicit fibre bound of $500/\sqrt\lambda$ in dimension 3. The corollaries are a filling-radius bound and continuous macroscopic dimension at most $n-2$ for universal covers of closed PSC manifolds, which again rules out PSC on aspherical manifolds of dimension at least 4. Two independent routes to the aspherical conjecture in one release is a mild consistency signal. It is not independent verification: the same model wrote both.

### The Cartan–Hadamard conjecture (CAT(0) part formalized)

In a complete simply connected manifold of nonpositive curvature, can a region enclose its volume more efficiently than a Euclidean ball? The conjecture says no. It was known in dimension 2 (Weil, 1926), 3 (Kleiner, 1992) and 4 (Croke, 1984), and [Ghomi and Spruck](https://arxiv.org/abs/1908.09814) reduced the general case to a total-curvature inequality.

The release proves it twice. [One paper](https://github.com/openai/math/blob/main/preprints/Generalized-Cartan-Hadamard-isoperimetry-and-Euclidean-equality-rigidity-September-23-2026/paper.pdf) does the generalized version in every dimension: under $\sec \le \kappa \le 0$, every finite-volume set has at least the perimeter of the equal-volume ball in the model space of curvature $\kappa$, and for $\kappa = 0$ the bounded equality cases are regions isometric to Euclidean balls. [The other](https://github.com/openai/math/blob/main/preprints/Sharp-integral-fillings-in-CAT%280%29-spaces-September-23-2026/paper.pdf) proves something more general for $\kappa = 0$: in any proper CAT(0) space, every compactly supported integral $n$-cycle ($n \ge 2$) bounds an integral current of mass at most $c_n M(T)^{(n+1)/n}$, the Euclidean constant. The Riemannian statement follows because top-dimensional fillings are unique.

The CAT(0) filling theorem is in Lean (`OAI/Geometry/CAT0Fillings`, 330 files, about 50,500 lines), together with optimality of the constant. I read the Comparator file. It defines metric currents, mass, integer-rectifiable currents and CAT(0) from scratch; the CAT(0) inequality there is the correct comparison-triangle formula, and the current definitions look reasonable, though they are bespoke and would need a geometric measure theorist's review. The $\kappa < 0$ theorem and the rigidity statement are paper-only.

### The Hilbert–Smith conjecture

Hilbert's fifth problem asked how much smoothness a continuous transformation group must secretly have. Gleason, Montgomery and Zippin answered the version about the group itself in 1952. The version about actions is Hilbert–Smith: a locally compact group acting faithfully on a connected manifold must be a Lie group. Classical work reduces it to one case, the $p$-adic integers $\mathbb Z_p$ can't act faithfully. It was known in dimensions 1 and 2 and, [by Pardon](https://arxiv.org/abs/1112.2324), in dimension 3, plus under regularity assumptions (Lipschitz, quasiconformal, Hölder).

[The paper](https://github.com/openai/math/blob/main/preprints/The-Hilbert-Smith-conjecture-in-every-finite-dimension-September-23-2026/paper.pdf) proves every dimension in 46 pages with a new invariant. Take a faithful $\mathbb Z_p$-action on an open set $O$ of a chart, stabilize to $W = O \times \mathbb R^{d-n}$ with $d$ odd and larger than $n$, and look at maps from the sphere $S^d$ through the orbit space into polyhedra. The invariant is a Witt group, signature-like classes of self-dual complexes of real sheaves generated by proper images of polyhedra under arbitrary continuous maps. On the test sphere its integral image sits inside a fixed lattice $L_d^{-1}\mathbb Z u_d$. A faithful action would split one class into $p^k$ equal integral parts summing to $4u_d$, and once $p^k > 4L_d$ that is impossible.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig2.png" alt="Diagram of maps S^d to W+ to Y = W+/G to X, with f = tqc drawn as the composite from the compact polyhedral source S^d to the test space X; c collapses the complement of W to the point at infinity." caption="The Hilbert–Smith proof turns a hypothetical p-adic action into a family of test maps from one fixed sphere, then compares sheaf-theoretic signatures of those maps (The Hilbert–Smith conjecture in every finite dimension, Figure 1)." />

The final contradiction is arithmetic a child could check. Everything hard is in making Verdier duality and Witt-group localization behave on categories generated by non-constructible sheaves, and in the integral (not just rational) comparison that fixes the denominator $L_d$. That is where I would look for a gap. There is no Lean.

### Disc embedding, Wall's PD4 question and the Borel conjecture in dimension 4

Freedman's 1982 work on topological 4-manifolds rests on the disc embedding theorem: when the fundamental group is "good", immersed discs with algebraically dual spheres can be replaced by disjoint embedded ones. Whether every group is good, above all the free group $F_2$, has been the central open question of the field for forty years.

[The release](https://github.com/openai/math/blob/main/preprints/A-boundary-only-obstruction-to-four-dimensional-disk-embedding-September-24-2026/paper.pdf) says no. It builds a compact smooth 4-manifold with finitely many immersed discs whose framed algebraic dual spheres satisfy every algebraic hypothesis, $\lambda(f_i, g_j) = \delta_{ij}$, $\lambda(g_i, g_j) = 0$, $\tilde\mu(g_i) = 0$, yet whose boundary circles bound no disjoint locally flat discs at all. Hence $F_2$ is not good, and no group containing it is. The obstruction is new: hypothetical topological discs are smoothed into 3-dimensional "cuts", deformations of the cut data are encoded by polarized Chern–Simons functions with a nilpotent ideal of end holonomies, and a Frobenius–Koszul residue calculation shows a required tensor can't exist.

Two further papers build on it. [One](https://github.com/openai/math/blob/main/preprints/A-PD4-group-without-an-aspherical-manifold-model-September-24-2026/paper.pdf) constructs a finitely presented Poincaré duality group of dimension 4 with a finite classifying space that is not the fundamental group of any closed aspherical 4-manifold, answering Wall's realization question (Davis had non-finitely-presented examples). [The other](https://github.com/openai/math/blob/main/preprints/Nonhomeomorphic-closed-aspherical-four-manifolds-with-the-same-homotopy-type-October-4-2026/paper.pdf), family 320, builds two closed aspherical 4-manifolds with the same hyperbolic fundamental group that are homotopy equivalent but not homeomorphic, so even the weak form of the Borel conjecture fails in dimension 4. Its combinatorial core is pleasant: a chamber whose $d$-fold cyclic covers have marked fillings exactly when $d$ is even, decided by trying to 2-colour an odd cycle.

My worry is structural. All three headline claims run through one new invariant, and the later papers import it as a black box ("the principal inputs are the marked algebraicity theorem [MT, Theorem 5.1]…"). If the marked tensor obstruction has an error, it takes the disc embedding, Wall and Borel claims down together. 193 pages, no Lean.

### The purely cosmetic surgery conjecture

Dehn surgery on a knot $K \subset S^3$ removes a tubular neighbourhood and glues it back with a twist labelled by a slope $r \in \mathbb Q \cup \{\infty\}$. The conjecture: distinct slopes never give orientation-preservingly homeomorphic results. Floer theory had cornered it. Ni and Wu showed the slopes must be opposite; [Hanselman](https://arxiv.org/abs/1906.06773) reduced them to $\pm 2$ or $\pm 1/q$ with genus 2 needed for $\pm 2$; Daemi, Lidman and Miller Eismeier excluded $\pm 1/q$ with filtered instanton homology. What was left: slopes $-2$ and $2$ on a genus-2 knot.

[The paper](https://github.com/openai/math/blob/main/preprints/Purely-Cosmetic-Surgery-on-Knots-in-the-Three-Sphere-September-23-2026/paper.pdf) closes that case with an SO(3)-monopole cobordism argument in the tradition of Pidstrigach–Tyurin, Okonek–Teleman and Feehan–Leness. A hypothetical diffeomorphism $S^3_{-2}(K) \to S^3_2(K)$ gives a closed parameterized cobordism on which an integer instanton count $\Omega$ is nonzero. Coupling to spinors with $\eta$ phase cuts gives a 1-dimensional moduli space whose boundary contributes $2^\eta \Omega$ from instanton links and zero from everything else, so $2^\eta \Omega = 0$.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig3.png" alt="Flowchart: ordinary units and cap pairing lead to integral repetition with Omega nonzero; metric families and line estimates lead to boundary index inequalities; both feed compatible perturbations, compactness and regular end gluing, ending in instanton links 2^eta Omega and lens pairs 0." caption="The cosmetic-surgery proof computes one integer two ways: a nonzero instanton count on the left branch, and a boundary of a one-dimensional monopole moduli space on the right that forces it to vanish (Purely cosmetic surgery on knots in the three-sphere, Figure 1)." />

The use of the earlier reduction is legitimate and makes the new part sharply targeted. But "everything else contributes zero" takes 144 pages of compactness, toric metric sectors and gluing, and the SO(3)-monopole program is famous for being hard to make rigorous. Gauge theorists will need to read Sections 9 and 10 before anyone should call this settled.

### Yau's uniformization conjecture

Yau asked in 1982 whether a complete noncompact Kähler manifold with positive holomorphic bisectional curvature must be biholomorphic to $\mathbb C^n$, the noncompact analogue of the Frankel conjecture (Mori, Siu–Yau). Every earlier result needed maximal volume growth: Chau–Tam with bounded curvature, [Gang Liu](https://arxiv.org/abs/1606.08958) and then Lee–Tam without it. Recent surface results by Datar, Pingali and Seshadri, and by Wu, removed growth assumptions in special cases; the paper notes in a footnote that those authors credit ChatGPT models for parts of their arguments.

[The release's version](https://github.com/openai/math/blob/main/preprints/Uniformization-of-complete-Kahler-manifolds-with-positive-bisectional-curvature-September-23-2026/paper.pdf) assumes only pointwise strict positivity: no curvature bound, no volume growth, no topology. It runs Kähler–Ricci flow from an initial metric that may have unbounded curvature and collapsed regions, which standard existence theory doesn't cover, by exhausting with complete Hermitian approximations. The Harnack inequality comes from minimizing a dual norm over holomorphic discs, which yields a viscosity inequality without global curvature control. A "volume clock" $s = \rho(o, t)$, the logarithmic volume loss at a base point, then measures how fast centred charts contract, and polynomial control of their transition maps makes the charts cover all of $\mathbb C^n$ instead of a proper domain. 77 pages, no Lean. The order of the estimates (scalar comparison before completeness, completeness from Harnack) is where the paper is most delicate.

### Katok's entropy rigidity conjecture

On a closed negatively curved manifold the geodesic flow has two natural invariant measures: Liouville (volume) and the measure of maximal entropy. Katok conjectured in 1982 that they coincide only for locally symmetric metrics, and proved it for surfaces. In higher dimensions only perturbative results were known (Flaminio; Humbert near hyperbolic metrics), and Flaminio showed the surface strategy cannot extend directly.

[The paper](https://github.com/openai/math/blob/main/preprints/Entropy-equality-and-local-symmetry-in-negative-curvature-September-23-2026/paper.pdf) proves all dimensions $n \ge 3$. Entropy equality, through thermodynamic formalism and Livšic theory, gives $J = h + XF$ for the unstable Jacobian, which produces conormal volume forms transported by the flow. The paper then encodes jets of stable and unstable plaques in finite-dimensional spaces where the dynamics contracts, extracts formal Taylor data, and studies the Lie algebra of formal symmetries. Cartan–Guillemin structure theory leaves two alternatives; excluding the infinite-dimensional one shows the stable and unstable distributions are smooth, and Benoist–Foulon–Labourie then makes the flow algebraic. 66 pages, no Lean. The strategy is the classical one ("prove the foliations are smooth"); the novelty is in the formal-symmetry classification, which I could not evaluate.

### The nearby Lagrangian conjecture fails

Arnold's nearby Lagrangian conjecture says every closed exact Lagrangian in a cotangent bundle $T^*Q$ is Hamiltonian isotopic to the zero section. Fifteen years of Floer theory showed such a Lagrangian is homotopy equivalent to $Q$ (Abouzaid, Kragh), then [simple homotopy equivalent](https://arxiv.org/abs/1603.05431) (Abouzaid–Kragh), with further constraints on its Gauss map and normal invariant. It is known for $S^1$, $S^2$ and $T^2$.

[The paper](https://github.com/openai/math/blob/main/preprints/A-counterexample-to-the-nearby-Lagrangian-conjecture-September-23-2026/paper.pdf) claims a counterexample: for some large even $N$, an exact Lagrangian in $T^*(S^9 \times S^{N-1})$ that is diffeomorphic to the base but not Hamiltonian isotopic to the zero section. The invariant is a stable "tube" class from generating-function theory, connected to Waldhausen's algebraic K-theory of spaces (Álvarez-Gavela–Igusa–Sullivan built Legendrians with nontrivial tube torsion). The new step is passing from Legendrians in the jet space to an embedded Lagrangian without creating double points.

This is the claim I trust least among the landmarks, for a simple reason: 28 pages to overturn a central conjecture of symplectic topology, with a non-explicit dimension. There are in-progress Lean files under `OAI/Geometry/NearbyLagrangian` (93 files) but no catalogued result.

### Quillen's conjecture

Quillen conjectured in 1978 that if a finite group $G$ has no nontrivial normal $p$-subgroup, the poset of its nontrivial elementary abelian $p$-subgroups is not contractible. He proved it for solvable groups. Aschbacher and Smith reduced the problem (for $p > 5$) to specific unitary configurations using the classification of finite simple groups; later work extended the reduction to all odd primes, and at $p = 2$ Piterman and Smith narrowed a minimal counterexample to classical groups in characteristic at least 5.

[The paper](https://github.com/openai/math/blob/main/preprints/Rational-homology-and-Quillens-conjecture-September-24-2026/paper.pdf) finishes those cases and proves the stronger rational form: $O_p(G) = 1$ implies nonzero rational homology. It builds explicit cycles from flags of "projective sign groups" and propagates homology through elementary extensions. The structure is sound practice: it takes a long, human-built reduction "in the precise form recalled in Section 2" and closes what remained. That also means correctness depends on matching hypotheses exactly against Piterman's theorems, which a group theorist can check much faster than the cycle constructions. 63 pages, no Lean.

### Curtis's conjecture

The Hurewicz map $\pi_d^S \to H_d(Q_0 S^0; \mathbb F_2)$ asks which homology classes of the infinite loop space come from actual maps of spheres. Curtis conjectured in 1975 that in positive degrees the only ones are the images of the Hopf-invariant-one classes $\eta, \nu, \sigma$ and the Kervaire-invariant-one classes $\theta_j$. A proposed proof had a gap found by Wellington. [The paper](https://github.com/openai/math/blob/main/preprints/The-Stable-Hurewicz-Image-of-the-Sphere-at-Two-September-25-2026/paper.pdf) proves it in 37 pages by following the successive adjoints of one stable class through the Dyer–Lashof and lambda-algebra structure, and with [Hill–Hopkins–Ravenel](https://arxiv.org/abs/0908.3724) concludes the positive Hurewicz image vanishes outside degrees 1, 2, 3, 6, 7, 14, 30, 62 and 126. Eccles's conjecture for spheres follows. No Lean.

### The Kervaire invariant problem at the prime 3

At $p = 2$, Hill, Hopkins and Ravenel used equivariant homotopy theory to show Kervaire-invariant-one classes exist only in a few dimensions. At odd primes the analogous classes $b_j$ in the Adams spectral sequence were understood for $p \ge 5$ (Ravenel), while at $p = 3$ survival beyond the known cases was open. [The paper](https://github.com/openai/math/blob/main/preprints/The-Kervaire-Invariant-Problem-at-the-Prime-Three-September-24-2026/paper.pdf) settles it: $b_j$ survives exactly for $j = 0, 2, 3$, in stems 10, 106 and 322, each detected by an element of order 3. The nonexistence half runs the program Hill–Hopkins–Ravenel proposed for $p = 3$, using $C_9$ acting on height-6 Morava E-theory, and proves along the way a coefficient model conjectured by Belmont and Ray while refuting another clause of their conjecture. The new class in stem 322 is built from the known stem-106 class. The unstable corollary is clean: the Anick space $T^{2n+1}(3)$ has a homotopy-associative multiplication exactly for $n = 3, 27, 81$. 90 pages of spectral-sequence bookkeeping, no Lean.

### The Grothendieck homotopy hypothesis (key step formalized)

Grothendieck's homotopy hypothesis says weak $\infty$-groupoids, with composition defined only up to coherent higher cells, model all homotopy types. Maltsiniotis made it precise using coherators; strict groupoids provably fail. Henry reduced it to one conjecture: adjoining an $n$-cell and an $(n+1)$-cell from an old cell to the new one (an "elementary expansion") is a weak equivalence.

[The paper](https://github.com/openai/math/blob/main/preprints/The-Grothendieck-homotopy-hypothesis-via-elementary-expansions-September-24-2026/paper.pdf) proves Henry's pushout conjecture for every coherator in the Ara–Henry convention, which gives the semi-model structure and the equivalence with spaces. The 26-page length is explained by Henry's reduction doing the bridging. The pushout theorem itself is in Lean (`OAI.Grothendieck.elementary_expansion`, about 12,800 lines); the semi-model structure and the comparison with spaces are not. The result is for this formalization of the hypothesis, not every proposed model of weak $\infty$-groupoids.

### The metric Blaschke conjecture

A Blaschke manifold has injectivity radius equal to its diameter: every geodesic minimizes until it reaches the "antipode". The round sphere and the projective spaces have this property. The conjecture says they are the only ones up to isometry. Green settled surfaces in 1963; Berger, Kazdan, Weinstein and Yang settled spheres and $\mathbb{RP}^n$ by 1980 through a volume comparison and Kazdan's inequality. [The paper](https://github.com/openai/math/blob/main/preprints/The-metric-Blaschke-theorem-September-23-2026/paper.pdf) handles $\mathbb{CP}^n$, $\mathbb{HP}^n$ and the Cayley plane by extending the volume comparison with a Schur-complement analysis of Jacobi determinants, then upgrading equal volume to isometry. 41 pages, no Lean. The topological inputs for the projective planes (Kramer–Stolz, a result attributed to Reznikov) are cited with hypotheses "verified at their uses", which is the step I would audit.

### Almost-everywhere regularity of stationary varifolds

Stationary integral varifolds are the most general soap films: weak minimal surfaces with multiplicity and singularities. Allard showed in 1972 that their regular set is dense, but whether the singular set can have positive area has stayed open, because multiplicity above one breaks his theorem. Simon's notes record the question and Brena, Decio and De Lellis recently stated it as a conjecture.

[The release](https://github.com/openai/math/blob/main/preprints/Almost-everywhere-regularity-of-stationary-integral-varifolds-September-23-2026/paper.pdf) claims $\mathcal H^m(\mathrm{Sing}\,V) = 0$ in every dimension and codimension, and a [second paper](https://github.com/openai/math/blob/main/preprints/A-Codimension-One-Bound-for-the-Singular-Set-of-a-Stationary-Integral-Varifold-October-5-2026/varifold-singular-dimension.pdf) sharpens it to $\dim_H \mathrm{Sing}\,V \le m - 1$, which two crossing planes show is optimal. The route is the one Brena–Decio–De Lellis proposed: approximate by smooth minimal graphs whose squared height vanishes to infinite order at almost every point, then prove a unique-continuation statement that upgrades contact to containment of the support. Measure-valued blow-ups keep height and slope information through weak limits. This is a central open problem of geometric measure theory and 112 pages with no Lean; a one-file stub under `OAI/Geometry/Varifold` is all that exists.

### Infinitely many closed geodesics on spheres

Every metric on $S^2$ has infinitely many geometrically distinct closed geodesics (Bangert with Franks or Hingston, early 1990s). On $S^n$ for $n \ge 3$ this was known only for generic metrics (Rademacher). [The paper](https://github.com/openai/math/blob/main/preprints/Infinitely-many-closed-geodesic-images-on-every-Riemannian-sphere-September-24-2026/paper.pdf) claims every metric on every $S^n$, every manifold finitely covered by a sphere, and every closed 3-manifold. The difficulty with degenerate metrics is telling new geodesics from iterates of old ones; the paper builds "native cuts" and degree-interval arguments in loop-space Morse theory. It also says Klingenberg's old proof is erroneous (following Asselle–Mazzucchelli) and points to two unsupported steps in a recent general claim by Charles. When a machine-written paper disputes a human one, both need adjudication. 50 pages, no Lean.

### Positively curved Einstein 4-manifolds (formalized for sec > 0)

Yang conjectured that the only closed Einstein 4-manifolds with positive sectional curvature are round $S^4$, round $\mathbb{RP}^4$ and Fubini–Study $\mathbb{CP}^2$. Every earlier result needed extra pinching or topology (Gursky–LeBrun, Cao–Tran, Gursky–Malchiodi). [The paper](https://github.com/openai/math/blob/main/preprints/Positively-curved-Einstein-four-manifolds-September-23-2026/paper.pdf) shows one of the two Weyl curvature blocks must vanish: a volume argument forces their second moments to be equal, a weighted Hessian inequality plus exact polynomial sign certificates gives an integrand with nonnegative mean that is pointwise nonpositive, and so both norms are constant.

<Figure src="https://ai.thesatyajit.com/articles/openai-math/geometry-topology-fig7.png" alt="Flowchart with two columns: scalar and upper volume estimates, equal second moments, exact pointwise bounds, conditional lower volume bound and simple connectivity on the left feed weighted Hessian inequality, E I nonnegative, half-conformal flatness, parallel curvature and global model isometry on the right." caption="Assembly of the positively curved Einstein classification: moment balance and exact polynomial sign certificates force one Weyl block to vanish, then the curvature is parallel and the manifold is a model (Positively curved Einstein four-manifolds, Figure 1)." />

The positive-curvature classification is in Lean (`OAI/Geometry/EinsteinFour`, about 209,000 lines, the largest formalization in this group). The [extension to $\sec \ge 0$](https://github.com/openai/math/blob/main/preprints/Zero-Plane-Rigidity-for-Einstein-Four-Manifolds-October-4-2026/einstein-boundary.pdf), where a single zero-curvature plane forces $S^2 \times S^2$, and the [$L^2$ gap theorem](https://github.com/openai/math/blob/main/preprints/An-L2-Einstein-Gap-for-Nonnegatively-Curved-Four-Manifolds-October-5-2026/einstein-gap.pdf), which openly takes that classification as a stated assumption, are paper-only.

### Coarse Novikov fails (formalized)

The coarse Novikov conjecture predicts that the coarse assembly map is rationally injective on spaces of bounded geometry. Yu proved much more under finite asymptotic dimension or coarse embeddability into Hilbert space, and Higson, Lafforgue and Skandalis broke surjectivity with expanders. [The paper](https://github.com/openai/math/blob/main/preprints/A-counterexample-to-the-coarse-Novikov-conjecture-September-23-2026/paper.pdf) builds a disjoint union of finite graphs of bounded degree with an infinite-order class in coarse K-homology whose index vanishes. The graphs have short cycles, which is why Willett and Yu's large-girth injectivity theorem doesn't apply. The ordinary (reduced) version is in Lean (`OAI/Topology/CoarseAssembly`, about 34,500 lines); the maximal version, the family's headline, is paper-only.

### Smooth isometric immersions of surfaces into $\mathbb R^4$ (formalized)

Nash's theorem embeds a surface isometrically in $\mathbb R^{17}$, and Gromov reached $\mathbb R^5$ for immersions of closed surfaces. Whether $\mathbb R^4$ suffices was a question Gromov attributed to Chern. [The paper](https://github.com/openai/math/blob/main/preprints/Smooth-isometric-immersions-of-closed-surfaces-into-Euclidean-four-space-September-23-2026/paper.pdf) says yes for every closed smooth surface, orientable or not: start from a short immersion into a small round $S^3 \subset \mathbb R^4$, add oscillatory rank-one corrections for the metric defect, and keep a usable normal direction through every step. The main theorem is in Lean (`OAI/Geometry/SurfaceImmersion`, 2,050 files, about 175,000 lines).

### A smooth metric with no local immersion into $\mathbb R^3$ (formalized)

Analytic surface metrics embed locally in $\mathbb R^3$ (Janet–Cartan); smooth ones do under nonvanishing or cleanly changing curvature. Pogorelov gave a $C^{2,1}$ metric with no $C^2$ realization. [The paper](https://github.com/openai/math/blob/main/preprints/A-Smooth-Metric-with-No-Local-Isometric-Immersion-into-Three-Space-September-24-2026/paper.pdf) gives a smooth metric on a square, flat to infinite order at the origin, with no smooth isometric immersion of any neighbourhood. Shrinking patches where curvature changes sign accumulate at the origin, and on each patch a boundary-saddle argument for the Darboux equation rules out a height function. The paper itself flags that an old Nadirashvili–Yuan preprint had claimed a smooth counterexample which their published version withdrew; it does not use it. In Lean (`OAI/Geometry/IsometricImmersion`, about 57,500 lines). The metric needs both curvature signs, so the $K \ge 0$ and $K \le 0$ versions remain open.

### Symplectic ball packing in dimension 6 and up (formalized)

In dimension 4, packing balls symplectically into a ball is governed by intricate algebraic-geometry obstructions. Siegel and Yao conjectured that from dimension 6 on, only two obstructions matter: volume, $\sum R_i^n < R^n$, and Gromov's two-ball non-squeezing, $R_i + R_j < R$. [The paper](https://github.com/openai/math/blob/main/preprints/Symplectic-Ball-Packings-in-Higher-Dimensions-September-23-2026/paper.pdf) proves sufficiency by explicit folding-style constructions, and the full equivalence is in Lean (`OAI/Geometry/BallPacking`, about 107,000 lines). The Comparator statement defines a symplectic embedding as a smooth embedding of a neighbourhood whose derivative preserves the standard form, which is the right definition.

### Affine Bernstein in dimensions 3 to 9, and a counterexample in 10 (formalized)

The affine maximal equation is a fourth-order cousin of the minimal surface equation. Chern asked whether entire convex solutions must be quadratic; Trudinger and Wang proved it in dimension 2 and found a singular counterexample in dimension 10. The release proves [rigidity for $3 \le n \le 9$](https://github.com/openai/math/blob/main/preprints/The-affine-Bernstein-theorem-in-dimensions-three-through-nine-September-24-2026/main.pdf) under Euclidean completeness and builds a [smooth nonquadratic example in dimension 10](https://github.com/openai/math/blob/main/preprints/Smooth-Nonquadratic-Affine-Maximal-Graph-in-Dimension-Ten-October-5-2026/affine-maximal-dimension-ten.pdf), so the answer flips exactly at 10, much as the classical Bernstein problem flips at 8. The dimensions 3 to 9 theorem is in Lean (about 41,000 lines); the dimension-10 example is a 30-page ODE argument that is not.

### Negative Kähler curvature without bounded coordinates (formalized)

In one complex dimension, curvature bounded above by a negative constant forces a simply connected complete surface to be the disc. Wu and Yau conjectured that two-sided negative curvature forces a bounded domain in higher dimension too. [The paper](https://github.com/openai/math/blob/main/preprints/A-negatively-pinched-Kahler-threefold-without-bounded-holomorphic-coordinates-September-25-2026/paper.pdf) builds a contractible Stein domain in $\mathbb C^3$ with pinched negative Kähler curvature that admits no bounded holomorphic coordinate system, so it is not biholomorphic to a bounded domain. A companion gives a higher-dimensional example with $\sec \le -1$ and only constant bounded holomorphic functions. The pinched example is in Lean (`OAI/Geometry/Kahler`, about 18,800 lines).

### Thomason model structures in every dimension (formalized)

Thomason showed in 1980 that small categories model all homotopy types. Ara and Maltsiniotis conjectured the same for strict $n$-categories for every $n$, proved $n = 2$, and isolated the missing condition: categorical pushouts along certain maps must behave like pushouts of nerves. [The paper](https://github.com/openai/math/blob/main/preprints/Thomason-Model-Structures-in-Every-Strict-Higher-Dimension-September-25-2026/paper.pdf) proves it for every $n \le \omega$, and the model structure and Quillen equivalence are in Lean (`OAI/CategoryTheory/Thomason`, about 160,000 lines).

### The Singer conjecture in dimension 4

Singer predicted that the $L^2$-homology of the universal cover of a closed aspherical manifold sits in the middle dimension. In dimension 4 that reduces to $b_1^{(2)} = 0$. [The paper](https://github.com/openai/math/blob/main/preprints/The-Singer-conjecture-in-dimension-four-September-25-2026/paper.pdf) proves it for every finite aspherical 4-dimensional Poincaré complex, so for every closed aspherical topological 4-manifold, with $\chi \ge 0$ and $\chi \ge |\sigma|$ as consequences. The argument is combinatorial group theory on chains rather than analysis, and it applies directly to the PD4 complex of family 305. 45 pages, no Lean.

### Chromatic splitting at height 3

Hopkins's chromatic splitting conjecture predicted that the overlap between adjacent chromatic layers of the sphere is a wedge of shifted local spheres. [Beaudry](https://arxiv.org/abs/1502.02190) already broke the strong form at $n = p = 2$. Five papers (372 pages) show it fails at height 3 for $p \ge 5$, even as an equivalence of underlying spectra, because a natural rational map that must vanish on the predicted wedge is nonzero on $\pi_{-3}$ of the real overlap. Weak splitting also fails for the derived $p$-complete sphere at heights $p$ and $p + 1$. The positive side: for $p > n + 1$ the overlap has a $2^n$-stage filtration with the predicted pieces, explicitly eight stages at height 3. The mixed verdict reads as genuine research rather than a headline. No Lean.

### Finite generation for the $K(n)$-local sphere

Hovey and Strickland asked whether every homotopy group $\pi_t L_{K(n)} S$ is finitely generated over $\mathbb Z_p$; it was known at height 1 and at height 2 for $p \ge 5$. [The paper](https://github.com/openai/math/blob/main/preprints/Finite-generation-for-the-Kn-local-sphere-September-24-2026/Finite-generation-for-the-Kn-local-sphere-September-24-2026.pdf) proves it at every prime and height, which together with the known rational computation of Barthel, Schlank, Stapleton and Weinstein pins down the free ranks. The proof bounds torsion degree by degree through a descent spectral sequence with a vanishing line at a finite page. 103 pages, no Lean.

### The Hovey–Strickland and Chai conjectures

Hovey and Strickland conjectured that dualizable $K(n)$-local spectra have exactly $n + 2$ thick tensor ideals, the analogue of the Hopkins–Smith thick subcategory theorem. Barthel, Heard and Naumann proved it at height 2 and showed it follows from Chai's hope about invariant ideals of the Lubin–Tate ring. [The paper](https://github.com/openai/math/blob/main/preprints/Stabilizer-Orbits-and-Thick-Tensor-Ideals-of-Dualizable-Kn-Local-Spectra-September-24-2026/paper.pdf) proves that invariant-ideal statement for every open subgroup of the stabilizer, in 34 pages. No Lean.

### Smith–Toda complexes at every height

The Smith–Toda complex $V(n)$ is a finite spectrum that kills $p, v_1, \ldots, v_n$ to the first power. $V(1)$ to $V(3)$ exist for large enough primes; Nave proved $V((p+1)/2)$ does not exist; whether $V(4)$ exists at any prime was open. [The paper](https://github.com/openai/math/blob/main/preprints/Finite-Smith-Toda-Complexes-at-Varying-Primes-September-23-2026/paper.pdf) constructs $V(n)$ for every $n$ at some prime depending on $n$, by building the tower in an ultraproduct of $p$-local categories and descending finitely many maps to one prime. [A separate paper](https://github.com/openai/math/blob/main/preprints/A-finite-Smith-Toda-complex-V4-at-the-prime-1009-September-23-2026/paper.pdf) builds $V(4)$ explicitly at $p = 1009$. Consistent with Nave, since the prime grows with $n$. No Lean.

### Bounded scalar curvature and Ricci flow

Does a Ricci flow on a closed manifold extend as long as scalar curvature stays bounded? Šešum showed bounded Ricci curvature suffices, and Bamler and Zhang developed the theory under scalar bounds. The release says [yes in dimension 4](https://github.com/openai/math/blob/main/preprints/Bounded-scalar-curvature-and-smooth-extension-of-four-dimensional-Ricci-flow-September-24-2026/paper.pdf), by ruling out Ricci-flat ALE bubble trees with a renormalized Einstein–Hilbert estimate proved in a [companion](https://github.com/openai/math/blob/main/preprints/Path-selection-and-an-elliptic-inequality-on-degenerating-Ricci-flat-trees-September-24-2026/paper.pdf), and [no in high dimensions](https://github.com/openai/math/blob/main/preprints/A-closed-Ricci-flow-with-bounded-scalar-curvature-and-finite-time-curvature-blowup-September-24-2026/paper.pdf), extending Stolarski's doubly warped conical singularities. 118 pages, no Lean.

### A finite-time singularity of Calabi flow

Chen conjectured that the Calabi flow, a fourth-order flow toward constant scalar curvature, exists for all time from any smooth Kähler metric. [The paper](https://github.com/openai/math/blob/main/preprints/A-finite-time-singularity-of-Calabi-flow-on-projective-space-September-24-2026/paper.pdf) builds a $U(10)$-invariant metric on $\mathbb{CP}^{10}$, in the Fubini–Study class, whose flow blows up in finite time with scalar curvature growing like $(T_* - t)^{-1/2}$. Symmetry reduces the problem to one radial variable; products give examples in every complex dimension from 10 up. 54 pages, no Lean.

### The Solomon–Yau least-volume conjecture

The equator is the smallest minimal hypersurface of a round sphere; the question in Yau's problem list is what comes next. The conjecture names the smallest minimal Clifford product $S^k(\sqrt{k/m}) \times S^{m-k}(\sqrt{(m-k)/m})$. Marques and Neves settled $m = 2$ through the Willmore conjecture. [The paper](https://github.com/openai/math/blob/main/preprints/The-Solomon-Yau-least-volume-theorem-September-23-2026/paper.pdf) extends their min-max strategy to every dimension, with a dimension induction to rule out singular limits below the Clifford threshold. Min-max regularity in high dimensions is the delicate part. 32 pages, no Lean.

### Unique tangent flows at the first surface singularity

Zooming in on a singularity of mean curvature flow gives a tangent flow, but different zoom sequences could a priori give different limits. Uniqueness was known for compact, [cylindrical](https://arxiv.org/abs/1312.4046) and conical models separately. Using [Bamler–Kleiner's multiplicity-one theorem](https://arxiv.org/abs/2312.02106), [the paper](https://github.com/openai/math/blob/main/preprints/Tangent-flow-uniqueness-2026-09-24/paper.pdf) proves uniqueness for every model at the first singular time of a closed embedded surface in $\mathbb R^3$, through a localized Łojasiewicz inequality that copes with ends of mixed type. Restricted to the first singular time and a fixed centre. 100 pages, no Lean.

### Notable results

The cubic flat 3-torus. The 20-year-old ball-tube-slab conjecture of Hauswirth, Pérez, Romon and Ros is [proved](https://github.com/openai/math/blob/main/preprints/The-Isoperimetric-Conjecture-for-the-Cubic-Flat-Three-Torus-September-24-2026/article.pdf): among regions of volume $V \le 1/2$ in $\mathbb R^3/\mathbb Z^3$, the minimizers are balls up to $V = 4\pi/81$, round tubes around shortest geodesics up to $V = 1/\pi$, and slabs beyond, with both shapes minimizing at each transition. Higher-genus competitors are excluded with reflection symmetry and curvature-area estimates. The full classification, equality cases included, is in Lean (`OAI/Geometry/CubicTorus`, about 141,000 lines), which makes this the most completely verified result in the group even if it is not the most famous.

Arnold's fixed-point bounds. The homological Arnold conjecture is a theorem of Floer theory; the stronger topological versions were open. The release gives [a Hamiltonian diffeomorphism of the complex quadric threefold](https://github.com/openai/math/blob/main/preprints/A-degenerate-counterexample-to-the-critical-number-Arnold-bound-September-23-2026/paper.pdf) with exactly 3 fixed points, while every smooth function on that manifold has at least 4 critical points, and families of simply connected Kähler 22-manifolds where the nondegenerate fixed-point count falls below the stable Morse number by a fixed fraction (stable Morse number $80 + 1968m$, fixed points $80 + 1952m$). Only the quadric example is in Lean (about 5,900 lines). These refute the strong variants and leave the homological conjecture untouched.

No conjugate points without nonpositive curvature. Ivanov and Kapovitch, and the Burns–Matveev survey, asked whether a closed manifold with a metric without conjugate points must admit one of nonpositive curvature. [The paper](https://github.com/openai/math/blob/main/preprints/A-Three-Manifold-Without-Conjugate-Points-and-Without-a-Nonpositively-Curved-Metric-September-24-2026/paper.pdf) answers no in dimension 3. The manifold is a two-piece graph manifold that Leeb already showed cannot be nonpositively curved; the new work is a metric on it whose geodesics never refocus. In Lean (about 25,000 lines).

Weak MTW and optimal transport. Villani conjectured that the weak Ma–Trudinger–Wang condition forces every tangent injectivity domain to be convex; Figalli, Gallouët and Rifford proved it when cut points come before conjugate points. [The release](https://github.com/openai/math/blob/main/preprints/Global-Support-and-Convex-Injectivity-Domains-under-Weak-MTW-September-25-2026/paper.pdf) removes that assumption and [derives bi-Hölder optimal transport maps](https://github.com/openai/math/blob/main/preprints/Uniform-Bi-Holder-Transport-from-Weak-MTW-September-25-2026/paper.pdf) for densities bounded above and below. Both main theorems are in Lean, except an arbitrary-potential local-injectivity claim that the formal version covers only for finite target families.

Harmonic functions of integer growth. Yau asked whether $\mathrm{Ric} \ge 0$ can only reduce the dimension of harmonic functions of growth at most $k$ compared with flat space, which has $(k+1)^2$ of them in $\mathbb R^3$. [The release](https://github.com/openai/math/blob/main/preprints/A-Three-Dimensional-Counterexample-to-Integer-Degree-Harmonic-Dimension-Comparison-September-26-2026/paper.pdf) builds metrics on $\mathbb R^3$ with $\mathrm{Ric} \ge 0$, Euclidean near the origin, with at least $c(k+1)^2$ such functions for any $c < 9v/4$, where $v$ is the asymptotic volume ratio. The metric depends on $k$. Lean covers only an earlier 16-dimensional example with degree 50000.

### Technical results

- Chromatic fixed-point loss (family 314): for a finite $p$-group, the optimal chromatic Smith loss from $H$ to $G$ equals the shortest cyclic subnormal chain, confirming Kuhn and Lloyd's "cautious hope".
- Hahn–Wilson at height 2 (family 319): for large primes, an fp-type-2 spectrum that is not finitely built from $\mathrm{BP}\langle 2\rangle$, though it satisfies both telescope comparisons.
- Gigli's characterization (family 356): Alexandrov curvature bounded below by $\kappa$ is equivalent to $\mathrm{RCD}((n-1)\kappa, n)$ with Hausdorff measure and Gigli's distributional sectional curvature at least $\kappa$; the lemma that weak Hessian bounds hold along every geodesic is in Lean.
- Bi-Lipschitz charts (family 357): every regular point of a noncollapsed $\mathrm{RCD}(K, n)$ space has a bi-Lipschitz chart with constant depending only on $n$, answering a question going back to Cheeger and Colding. A large in-progress Lean development (about 2,900 files) exists without a catalogued result.

### What I would tell a mathematician friend

Start with the four results where the release did the verification work for you: tamed-to-compatible, Yau's nodal conjecture, the cubic torus and the ball packings, all with faithful-looking Lean statements. Then read the D(2) paper, because you can check it yourself. Treat the long analytic landmarks (Hilbert–Smith, cosmetic surgery, uniformization, Katok, varifold regularity, all-dimensions Gromov–Lawson) as serious proposals by an author whose track record on easier problems is good and on problems this hard is unknown. And be most careful with the cluster that shares one new invariant (disc embedding, Wall's PD4 question, Borel in dimension 4) and with the 28-page nearby Lagrangian counterexample.

## The searchable catalogue

Every one of the 372 families, with the claim in a sentence or two, our grade, the kind of result, its Lean status and the main caveat. Search by family number, by a name ("Kakeya", "Hadwiger", "zeta") or by a word in the claim; open a row for the caveat and links to the section above and to the principal manuscript.

<FamilyCatalogue />

## How I checked

Everything here comes from the repository at commit `adc7f1241`, cloned read-only. I did not build the Lean library, run Comparator or execute any code from the release.

For the release as a whole I parsed `CONTENTS.md` (families, manuscripts and Lean links) and the `\cataloguesection` blocks of `overview.tex` (disciplines), counted pages with `pdfinfo` over all 722 PDFs, and read the dates off the directory names. For the Lean library I counted files and lines with `find` and `wc`, searched `lean/OAI` with ripgrep for `sorry`, `admit`, `axiom` declarations, `native_decide`, `implemented_by` and `debug.skipKernelTC`, tallied the imports by top-level package, and parsed all 405 `ComparatorChallenges/*.json` files for their permitted axioms, theorem names, definition holes and kernel options. I read the challenge statements quoted on this page in full, and Comparator's own README for what a successful run guarantees. The reasoning summaries were extracted with `pdftotext`; section, excerpt and reference counts are from the extracted text, and the "audit" count is my reading of the section titles.

The family-by-family reading was split by discipline across parallel reviewers working from the same clone, myself and AI sub-agents, each of whom read the catalogue summary, every abstract, and the introduction and main theorem statements of every principal manuscript in their group, matched each family against `lean/docs/<id>.md`, `formalization.yaml` and the Comparator statements, and searched for the prior state of the art. I edited their drafts into the sections above. Where a section reports a recomputation (the matrix-multiplication constructions in Python, the Abhyankar–Sathaye critical point in exact rationals, the crossing counts of the two-page drawings, the Fox-matrix minor in the D(2) counterexample, the Thorp shuffle distances over all $8!$ orderings), that was done independently of the release's own code. The figures are pages and crops rendered with `pdftoppm` from the release's PDFs, flattened onto white.

Prior results are cited from arXiv abstract pages, journal DOIs and the papers' own literature sections; where a citation could only be confirmed through a paper's bibliography it is named without a link. Expert reaction is limited to what I could read in primary reporting (Scientific American, October 6). As of October 7 I found no published expert assessment of any individual result and no published Comparator run. The grades in the catalogue are judgments about the size of each claim, not about its correctness; nobody can make the second judgment for 372 results in a day, and this page does not pretend to.
