2026-10-08 · 97 min · math · theory · algorithms · agentic-coding · verifiers
Why read this
Solidtop 85%The whole κ race in one place: where the exponent comes from, all 44 PRs timed, PR #13's certificate re-run and recomputed, and why it changes no runtime.
- Original analysis
- A new technique
- Explained from first principles
Agents & harnessesRuns on a laptop CPUApache-2.0Research paper
How this was scored
- Is it new?
- 2 of 3: A real new idea, method or capability
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 1 of 3: API-only, gated or restrictive licence
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 1 of 3: General advice
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 55 of 100, ranked 327 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
The site owner sent me a link to CrocSwap/integer-mult-bounds with the note "something crazy going on in this repo". When I opened it, the headline said . By the time I had read the pull requests, the best open claim was , a little above . When I started writing this paragraph, PR #45 had just arrived. By the time I finished the article, #48 had nudged the claim to , and Swapnil Jain, racing in a separate repository, had reached with its arithmetic checked in Lean.
Then, at 15:18 UTC, the story changed shape. Colkitt merged Rohan Arun's PR #39 into main, with , after an audit he published alongside it. For the first time a community number is the maintainer's number. It is still conditional on OpenAI's unrefereed manuscript, and the audit was Colkitt working with Codex, not a referee; the section on the merge goes through exactly what it checked. The open PRs had already moved past it: Rohan's own #49, under twenty minutes later, claims .
And it kept going after I thought I had finished. By 18:35 UTC main had moved twice more and stood at , after a new idea broke the plateau I describe below; the open PRs reached an hour later. Aurel Prosz, one of the contributors, built a live dashboard of the whole thing. And Ryan Shea took Jain's network across to a second OpenAI problem, the exact discrete Fourier transform, and claimed a saving 730 million times OpenAI's. The evening and the Fourier transfer have their own sections near the end.
Then, overnight, the plateau broke for good. Between 23:12 UTC on 8 October and 05:06 the next morning, icekylinx's PRs replaced the shape of the network itself, and main jumped ninefold to , about . Jain followed twenty minutes later at , by building on the same cover, and the open PRs were at when I stopped, at 11:00 UTC on 9 October. The ninefold night explains what changed, and why κ is still nowhere near 1.
The afternoon was quieter in ratio and busier in every other way. At 12:57 UTC Colkitt moved main again, to . Ten minutes later Jain posted round eleven at , which led every claim anywhere for nineteen minutes, and said he needed an arXiv endorser for a paper collecting the proofs. By 18:50 UTC the best open claim was , and several contributors had stopped chasing κ to prove how far the current designs can go at all. The afternoon covers all of that.
Some background. Two days ago OpenAI's math release included a 73-page manuscript, Integer multiplication below n log n (result family 109), which claims a multitape Turing machine that multiplies two -bit integers in steps with . I covered it briefly in the openai/math roundup, and it has a tile on the AI breakthroughs in mathematics wall. Doug Colkitt (@0xdoug, the founder of CrocSwap/Ambient) forked the argument into a "research draft" and started tightening the constants. Then other people joined in. Counting as the start, the exponent saving has grown by a factor of about in forty hours.
That sounds like a story about multiplication getting faster. It isn't one, and the reason is the most interesting part. What follows covers what is, where it comes from in the construction, what each step of the race actually changed, what I could re-run myself, and why the very funny "linear time by tomorrow" chart that went round on X is wrong in an instructive way.
- license
- Apache-2.0
- branch
- main
- tests
- 33 files
- source
- 500.5 kB
- commit date
- 2026-10-08
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-08 at 1a74950 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history

What measures
Schönhage and Strassen multiplied -bit integers in steps in 1971 and guessed that was the true answer. Fürer got within a factor of it in 2007. Harvey and van der Hoeven reached exactly (Annals of Mathematics, 2021), and most people took that to be the end.
The OpenAI manuscript claims a bound of
So is the power of you get to divide out. At you have Harvey–van der Hoeven. At you would have linear time, which is the floor, since just reading the input takes steps. Every value in between is a strictly faster growth rate than . For the question "is optimal?", is as good an answer as : any refutes the conjecture.
The machine model matters. The theorem is about a deterministic Turing machine with a fixed finite alphabet and a fixed number of one-dimensional tapes, the model in which Schönhage and Strassen stated their conjecture. On a tape, moving data costs time proportional to the distance it travels, so even rearranging bits at each of FFT levels costs . The manuscript says this outright: "Scanning an array of stored bits at each of transform levels already costs . We therefore need savings in both movement and arithmetic" (upstream/build/sections/00-introduction.tex).
Nobody had ever proved an lower bound in this model. The best result is conditional: Afshani, Freksen, Kamma and Larsen (2019) showed that a circuit lower bound would follow from the network-coding conjecture, and a time- Turing machine only gives circuits of size about , so even that does not transfer. The FFT section of the roundup goes through the same gap for the Fourier transform, where the release's family 130 reuses this manuscript's finite gadget. The point to keep in mind is that was a barrier nobody knew how to cross, not a proven wall.
Where comes from
The pipeline is Harvey–van der Hoeven with every change of representation charged on tape: digits go onto prime cyclic axes by a Chinese-remainder map, Gaussian resampling turns those into power-of-two axes, and a phase twist turns the last axis into polynomials modulo , where multiplying by a root of unity is just a signed shift.
![Flow diagram with four boxes stacked vertically: integer product as radix-digit convolution; cyclic convolution on distinct prime axes (via a Chinese-remainder address map); normalized convolution on power-of-two axes (via Gaussian resampling and chirps); convolution over C[y]/(y^r + 1) with synthetic transforms and packed polynomial products (via a last-coordinate phase twist).](/articles/integer-mult-exponent-race/fig4-multiplication-flow.png)
The new part is two tape procedures built from fixed linear networks. One swaps two address fields of an array; the other applies a layer of butterflies to selected coordinate bits. Each network is a finite circuit on "roles" (wires), it exchanges two banks of values while restoring every scratch value, and each edge of the circuit is carried out by smaller recursive calls of the same procedure. The decisive count is the total number of those smaller calls. If , then multiplying the problem parameter by multiplies the normalized cost by , and the recursion runs at with . The manuscript: "This strict inequality is the source of the power saving."
How much below it lands is tiny. In the original construction the relative deficits are (upstream/build/sections/03-motifs.tex:673-678)
at , so the honest savings are about and . The manuscript then rounds them down to and calls its constants "deliberately conservative; no optimization is claimed". Remember that sentence. A lot of what happened next is people taking it at its word.
Those network savings feed a cost table. With the working precision, the paper divides every operation's cost by the data volume and gets a power of . Seven rows dominate, and each has a margin below 1:
The total time is times to the largest power, so the exponent saving is the smallest margin, minus a sliver of slack to absorb stray factors. The upstream parameters are , and , so , and the paper says "Thus " (08-assembly.tex:866). That is where comes from: three small numbers multiplied together, then halved.
The repository's checker transcribes the same seven margins as exact fractions (scripts/certify.py:136-147):
def margins(p, *, layout_model="adjacent", assembly_model="original"):
require(assembly_model in ("original", "tight-gaussian"), "Unknown assembly model")
e, c, t = p.epsilon, p.c, p.tau
return {
"g1": 1-e*(1+c),
"g2": e*c*(1-t),
"g3": e*(1-p.lamp),
"g4": 1-t-e*(layout_degree(layout_model)-t),
"g5": Q(1, 4)-p.delta-(Q(5, 4) if assembly_model == "tight-gaussian" else Q(3, 2))*e,
"g6": 1-p.delta-e,
"g7": e,
}Once you see as a product of factors, the race is easy to follow. Every improvement attacks one of the factors: it makes the network saving bigger (better circuits), lets grow (better routing, better Gaussian resampling), lets grow (cheaper control movement), or changes the recurrence so the same circuit counts for more (batching). The widget below shows the seven margins on a log scale for the original paper, Colkitt's checkpoint and PR #13. In the original, sits at and nothing else comes close. By PR #13, , and are all balanced at about , which is what a tuned parameter set looks like.
- g1: prefix-slot moves, single butterfly rounds · 1 − ε(1+c)
- g2: chunk exchanges for transform layout · ε·c·(1−τ)
- g3: simultaneous butterfly rounds · ε·(1−λ′)
- g4: CRT and axis layouts · 1−τ−ε(2−τ) → (1−τ)(1−ε)
- g5: Gaussian line maps · 1/4−δ−3ε/2 → 1−δ−2ε
- g6: chirps, twists, scalar products · 1−δ−ε
- g7: packed polynomial products · ε
There is also a ceiling built into this shape. Since and , the smaller of the two is at most whatever you choose, and is at most the bit network's saving . In every witness I checked, sits just under that network saving: against at the checkpoint, against in PR #13, and against in PR #44. To get anywhere near linear time this method would need finite networks whose recursive calls nearly vanish, and nothing here suggests that is possible. The overnight witnesses of 9 October keep the same shape: in PR #144's assembly two margins are balanced against each other so that , where is the weaker network's saving, and that saving is in turn about the network's relative rank deficit divided by how far below full width its average child sits on a log scale. The section on the ninefold night works through the numbers.
Colkitt's day and a half
Colkitt's own checkpoints are on main (the last one is also on release/ternary-30). Each came with an X post from @0xdoug, and each post's ratios check out against the exact witnesses.
| commit (UTC) | what changed | |
|---|---|---|
5cf29ec Oct 7 02:22 | , between and | parameters only: same network, sharper recurrence comparison |
52ce3be Oct 7 13:35 | direct nonadjacent axis swaps: layout from to | |
bcd4ebd Oct 7 19:27 | paired circuits with shared sums, tighter guard and Gaussian width | |
6e56487 Oct 8 00:27 | compact dirty controls instead of spaced windows | |
1c09a58 Oct 8 01:10 | compressed complex network, binary phase frames | |
1a74950 Oct 8 13:10 | ternary five-subset bit circuit over |
The first note is the most honest thing in the repository. With the paper's cost accounting left alone, it finds and also proves a ceiling of for that family, so the witness is within 1% of the best that accounting allows. Colkitt's post said so: "Surpassing this ceiling would require improving the network bounds or cost analysis from the original result."

So the next morning he changed the accounting. The original layout row costs because it reverses axes with adjacent swaps; the manuscript already allowed nonadjacent swaps, and using them directly makes it . That changes to , which no longer forces to be smaller than the network saving. The result-history page explains the consequence: the original layout "forced epsilon < a, yielding a cubic constraint kappa < a^3", and the new schedule makes the saving scale "quadratically in a" (docs/research/result-history.md:253-256). Going from to when is how you gain thirty-odd powers of two in one commit.
The step improved the network itself, which raises : rectangle incidence circuits, then shared intermediate sums, cut the side roles per invocation from 41,122,620 to 2,394,438 and then to 577,576 at , and paired-block circuits at reach 509,194. Fewer roles for the same rank deficit means a larger relative deficit, so a larger .
The compact-control step at attacked . Moving whole spaced windows had a spacing penalty that forced to be of order ; moving only compact control fields removes it, so and becomes roughly times . His post that night put it in two sentences: "The improvement came from removing the spacing penalty behind the quadratic bottleneck. This was done by moving compact control bits instead of entire windows." Its "roughly 48 million fold improvement over the previous result" checks out against (about ), and so does "a 2¹⁴⁸ fold improvement over original OAI result" (). Here the README discloses that the idea did not come from Colkitt: "The compact-control proposal originated with a separate research agent; the supplied note develops its tape, layout, repair and assembly arguments" (README.md:138-140).

The last two checkpoints moved the bottleneck between the two networks. At a compressed complex network (, ) overtook the bit network's , so the bit side became binding. The checkpoint then replaced the bit circuit with one that computes over .
The trick is pretty enough to show. Label the circuit's wires by the five-element subsets of 29 points, of them, and let be the incidence matrix of five-subsets against pairs. Then , and for intersection sizes those residues are . So , where is the adjacency of subsets meeting in exactly two points, and the "side" map is just . A central factor of size plus a correction that only touches intersection-two neighbours gives exactly the identity. I checked the residue identity by brute force over all 126 five-subsets of nine points; it holds.

The full circuit has 19,593,239 active additions and 20,780,789 side roles per invocation, and the exact count gives with a relative deficit of . I recomputed the implied saving from the certificate's , and : , just above the certified . The final minimum margin is , which exceeds by about 0.194% of the claimed saving.
That checkpoint was not first, and its README says so: "Zhihao Chen's earlier PR #7 introduces the same ternary five-subset motif with a different circuit and stronger claimed bound. This release records a separate implementation and conditional checkpoint; it makes no priority or strongest-known-bound claim."
Everyone else shows up
The first outside PR arrived at 20:58 UTC on October 7. Between 02:23 and 14:24 UTC on October 8, 42 more arrived, from about ten people. By 11:00 UTC on October 9 there were 179. The chart has all of them that state a κ: Colkitt's checkpoints in red, OpenAI's starting point in grey, the community PRs coloured by author, a green step line for the best claim open at each moment, and a red step line with a square at each move of main (#39 at 15:18 UTC, #49 at 16:47, a reviewed batch at 18:35, and icekylinx's #144 at 05:06 the next morning). Everything after is squeezed into a sliver at the top, so the zoom button redraws everything from 10:00 UTC on October 8 on its own. Jain's separate track is the purple diamonds, with a ring on the rounds whose arithmetic is checked in Lean; more on that below. Click a point (or use the dropdown) for its title and link.
A few PRs changed the shape of the problem rather than its constants, and those are the ones worth understanding.
Start with eumemic's #3 and #5. #3 compressed the complex network, cutting its side wiring from 3,693,800 roles per invocation to 108,195 and taking its saving from to . That made the bit network binding, and the next limit was the Gaussian resampling row: its cost forced and capped below . #5 rewrote the resampling as chirped correlations evaluated by the existing multiplier and bounded the Neumann series more tightly, which let approach . That one change of analysis was worth a factor of about 2.5 on its own.
Zhihao Chen's #7 is the five-subset construction, at , posted at 06:12 UTC, seven hours before Colkitt's independent version. Chen's PR credits "GPT-6 Astra (OpenAI Codex)" for "research, proof-development and implementation assistance" and asks that future work "explicitly acknowledge Zhihao Chen (jacklightChen), cite this PR and its accompanying note". The maintainer did.
icekylinx's #10 is the structural break you can see on the chart: from to in one PR. Upstream, every edge of the network costs singleton recursive calls, one per unit of rank. #10 shows that a contiguous block of address fields from a projector can be done as one recursive call of width , given a common controlled basis. The recurrence condition stops being a count and becomes a moment: with rank-mass weights on calls of relative width , you need . A singleton call has and costs the full ; in #13's profile, a block covering 195/196 of the fields costs almost nothing extra. The same PR fixes an error-enclosure gap it found in #5's fast Gaussian argument, with no change to the exponent.
eumemic's #13, the one the site owner pointed at, is a neat idea. The network's correctness condition only needs each role's endpoint frames to differ by the identity: . A scratch role does not have to start in the zero frame. It can start in any fixed frame as long as it ends in . Starting every stage-two auxiliary role in the frame of its first gate deletes its rank-756 entrance edge and turns its exit into one projector of rank . That projector compiles to singleton pivots plus one contiguous block of fields. The total rank is unchanged. What changes is where it sits: out of singleton calls costing each and into one block. Under #10's moment, the interchange saving goes from to , and with a third whole-residual class on the complex side, goes from to , a factor of 6.26. Since , the claim is about ; the title's "" undersells it.

After #13 the tempo picks up. #15 brought a smaller producer, #18 partial swaps, and #21 and #23 translated frames and "semantic precision", reaching . #29 introduced a two-stage topology from Aurel Prosz's work and reached . #36 added copied retained centers, which halve one rank term ( becomes ) and jump to . Everything after that is composition and tuning: fixed local bases, reversed corners, carrier matchings. PR #44 claims , which it describes as "approximately 0.02681% above #43". The last three PRs are within 0.03% of each other. Three more arrived while I was writing: chafreaky's #46 adds 0.0000048% to #44, Rohan Arun's #47 adds 0.109% with hill-climbed producers, and chafreaky's #48 reorders the leave-one-out sums to save 847,550 wires and reach . Then Rohan's #49, at 15:36, started from #48's orderings and hill-climbed them the way #47 had, for , 0.127% more. The step line flattens around after 12:30 UTC.
When I first wrote this section I called that flattening the most informative thing on the chart, and said it was what you would expect if the cheap structural moves inside this family of constructions had mostly been made. Three hours later Avi Eisenberg's #53 jumped 9.2% in one PR, and the step line climbed again to . Use the zoom to see both plateaus; the evening section explains the move that broke the first one. I'd still read the flat stretches the same way, only with less confidence about how long any of them lasts.
Re-running PR #13
The brief asked me to check the claims, and these are certificate checks I could actually run: pure Python with fractions, plus a small C++ helper on the PR branches, no network. On a shared 16-core machine at nice 19:
On main at 1a74950, the focused verification from the README. scripts/audit_ternary_side.py builds the full DAG (about 21 million addition nodes) and finished with PASS ternary side construction and conditional 2^-30 assembly. in 176 s at 1.6 GB peak memory. make_ternary_patch.py took 17 s, the 12 focused tests passed in 39 s, the patch applied to the pinned upstream source, and the regenerated certificate and patch were byte-identical to the committed ones.
On PR #13 at 3ef246f, scripts/source_frame_network.py runs in 0.1 s and prints the witness:
PASS kappa=7699/10000000000 > 2^-21
bit saving=77/50000000; complex saving=9/5000000
bit moment gap=4.69651719166272e-10
complex moment gap=3.899666592614166e-09
minimum margin=3849992299500001/5000000000000000000000; gap=492299500001/5000000000000000000000The regenerated certificates/source-frame-network.json compared equal to the committed one. The six tests in tests/test_source_frame_network.py, including an exact model in which the new exit is checked to be an idempotent of rank with exactly corner pivots, passed in 20 s. The full make verify on the PR branch took 432 s: 172 tests passed, all 19 manuscript patch checks applied, and the certificates and patches regenerated without a diff. Those numbers match what the PR description claims.
Then I recomputed the arithmetic myself, without the repository's helpers. From the certificate's parameters (, , , , ) the margins are , and , all about ; and are about , and and about . The minimum is , which beats by . So the claimed does follow from the margins.
The bit-moment check is the step that carries the new idea, so I redid it at 80-digit precision rather than with the script's bound. Here is how the script does it (scripts/source_frame_network.py:54-62 on PR #13):
def moment(weights, ratios, a):
"""Upper bound for sum_i w_i ratio_i^(-a), using exp(x) <= 1/(1-x)."""
logs = []
for r in ratios:
simple = SIMPLE_LOGS[r]
require(log_upper(1/r) < simple, 'Rational logarithm bound failed')
require(0 <= a*simple < 1, 'Exponential upper-bound range failed')
logs.append(simple)
return sum((w/(1-a*ell) for w, ell in zip(weights, logs)), Q(0)), logsThe four classes have weights of about 0.0044 (singletons, relative width ), 0.4934 (width ), 0.0076 () and 0.4946 (). With the moment is , below 1 as required, and bisection puts the largest saving this rank profile supports at about . PR #13's leaves about 0.6% headroom.
None of this checks the hard part. The certificates confirm that the counts, ranks and inequalities are what the notes say they are. They do not confirm that a contiguous projector block really can be executed as one recursive call on a fixed-tape machine, that the controlled basis exists at full size (the review guide for #10 says it is "established by the existence argument, not by materializing every h28 factorization"), or that the upstream theorem is true. Colkitt's own review index was clear, that morning, that this was still open: "These changes appear to cross the singleton-rank recurrence limitation; the headline alone does not establish that step." By the afternoon his audit had worked through those obligations for the one chain he merged, #39's (below). For #13's own construction and every PR outside that chain, the sentence still stands.
A second race, one repository over
While those PRs were landing, Swapnil Jain (@SJ_Swapnil_Jain) was running a race of his own. He never opened a PR on Colkitt's repository; none of the 13 accounts that did is his. He kept Swapnil-jain/integer-mult-kappa, posted six numbered updates on X between 07:20 and 14:03 UTC, and the sixth ended with a sentence nobody on the CrocSwap side could write about their frontier: "Both interchange certificates and the full assembly are checked in Lean's kernel." I cloned it and read all of it, and I wanted to know two things. Is this really a separate result, and what exactly did Lean check?
- license
- Apache-2.0
- branch
- main
- tests
- 6 files
- source
- 253.3 kB
- commit date
- 2026-10-08
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-08 at f2176bc — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
I lined the six posts up against what was open on CrocSwap at the same minute. The witnesses are the exact rationals from the commit each post announced; they are the purple diamonds on the chart above.
| post (UTC) | Jain's | best open CrocSwap claim then | what was new |
|---|---|---|---|
| 07:20 | #7, | the "ε → 1 stack" | |
| 08:15 | #9, | Gaussian correction inverted block by block | |
| 10:05 | #18, | batching on a two-stage bit network | |
| 11:10 | #28, | PR #24's bit network, complex source frames | |
| 12:39 | #37, | his own bit network on one flag basis; first Lean file | |
| 14:03 | #43, | copied centres; both savings and the assembly in Lean |
So he held the lead twice: for about an hour after his first post, until icekylinx's #10 at 08:26, and for nineteen minutes after his third, until Zhihao Chen's #21 at 10:24. His fourth post went out one second after #28 was opened. From round five on, the CrocSwap side was ahead, by 2.5 times at 12:39 and by about 12% at 14:03. His own ratios check out: round six is 2.37 times round five, and to is a factor of . The one soft spot is the first post's comparison with Colkitt's ("a roughly 60 fold improvement"): by then Colkitt's main was already at , so against the newest checkpoint it was about ten.
"Independent" is the wrong word for it, and Jain doesn't use it. His NOTICE pins Colkitt's 6e56487 and cites fourteen open CrocSwap PRs (#3 through #36) plus Aurel Prosz's fork. Round four reads PR #24's published network counts from a JSON file pinned by SHA-256. Round six's copied centres are, in the first sentence of his note, "PR #36's copied retained-centre schedule (icekylinx)" applied to his network. The traffic went the other way as well. Chen's #29 and Rohan Arun's #31 credit "Swapnil Jain for the linked two-stage batching development", icekylinx's #36 links Jain's round-three commit with the line "no source code imported from this repository", and #44 and #47 list him with everyone else. These are two forks of the same argument, reading each other's work all morning.

What is his is the part of the stack that turns a network saving into . Through PR #13, sat at about half the bit network's saving: against . His README explains why. A few costs tie the saving either to the number of axes or to the axis width , with , so the saving gets split between them, and the Gaussian resampling's precision budget, about , has to fit inside , which keeps below . His first round made five changes at once. The coefficients become bits long instead of about , so precision no longer competes with address length. The Gaussian correction is inverted block by block (every block is "the same Toeplitz matrix up to a geometric rescaling, so one stored inverse serves the whole line", as the second post puts it). Each axis becomes one chunk, only the low bits of an axis are moved for resampling, and the axis reversal reuses Colkitt's selected-bit additions. With all of that, can go almost to 1, and that is the "ε → 1" in every post.

I recomputed the round-six witness from the numbers in his certificate, without his code. The bit saving is , the complex saving , so and . He sets . Of his five cost margins, , minus a -sized term, and all come to and a bit; is the smallest; and itself is about 1. The claimed sits under the binding margin. The tidy way to say it: is chosen so that , which makes to eleven digits. Every bit of the network's saving reaches .
The CrocSwap side got to the same place by another road. RaD/hipotures' #20 and Chen's #23, at 10:20 and 10:43 UTC, used "completed-child semantic precision" and bulk resampling, and neither cites Jain. Round four is a neat accidental experiment. Jain and Dominik Scholz's #27 took the same PR #24 bit network, with , put it through their two different assemblies, and got and . Both stacks now pass essentially the whole network saving through, so the 12% gap between 3.67 and 4.12 is entirely the bit network: here against #48's .
The ε → 1 stack also has a price that the headline hides. With near 1, the coefficient-depth guard needs . Round six uses the crude guard (from ), so and each coefficient is bits long. In round five the same rule gave . That is legal, because is fixed and the volume stays , and his certificate says as much ("The exponent is huge but fixed, so the O() statement holds", scripts/certificate_round3.py:15). It is one more reason the runtime section below holds for this track too. The stack notes also list "deliberate departures from upstream hypotheses": the prime number theorem replaces the manuscript's explicit interval lemma, the resampling hypothesis is weakened, and so on. So this track rests on more than the CrocSwap PRs do: the manuscript, the fourteen PRs it cites, and its own replaced hypotheses.
What the Lean files prove
There are two, lean/Round5.lean and lean/Round6.lean, and neither is written by hand. lean/gen.py emits them from JSON files that dump.py and dump6.py write out of the Python certificates. They are core Lean 4: no import, no Mathlib, no lakefile and no lean-toolchain, so no version is pinned; make lean runs lean lean/Round6.lean with whatever is on the PATH. Rationals are pairs of naturals with a hand-written gcd normalisation, is a 30-term atanh series with a geometric tail bound, and each file ends in five theorems. Round six's, from lean/Round6.lean:74-81 with the histogram lists cut:
theorem bit_rank_sum : rankSum bitHist = 78410006675 := by decide
theorem bit_moment : momentOK bitHist (36667, 1000000000) 529 148225616 = true := by decide
theorem cx_rank_sum : rankSum cxHist = 119453132304 := by decide
theorem cx_moment : momentOK cxHist (36926111, 500000000000) 576 207387136 = true := by decide
theorem kappa_assembly : kappaOK (36667, 1000000000) (36926111, 500000000000) (1, 1000)
(2499908335861049, 2500000000000000) (14695, 1) (14694, 1)
(3666565558019, 100000000000000000) 576 119453132304 = true := by decide +kernelThere is no sorry, no axiom and no native_decide in either file. Both decide and decide +kernel end in a proof term that the kernel checks; +kernel only skips the elaborator's own evaluation first. native_decide would have compiled the decision procedure and trusted the compiled code, and it isn't used. The kernel does use GMP for natural-number literals, which is part of Lean's normal trusted base. So yes, it really is the kernel.
The more important question is what the theorems say. They say that three particular Boolean functions return true on particular lists of numbers. They do not say those lists are the child-width histograms of the networks; gen.py pastes them in from Python. They do not say lnUp bounds the logarithm or that the moment sum controls the recurrence. The generator's docstring is explicit: "The analytic facts behind (2) and (3) -- the atanh tail bound and (m/w)^a <= 1/(1 - a ln(m/w)) -- are premises, not checked here: Lean core has no real logarithm." And kappaOK encodes his own five-margin cost table, transcribed from certificate_round3.py; whether those five margins are the right costs is a question for the written notes. The upstream theorem isn't in it at all. His README is honest about this; its evidence table lists the full upstream multiplication theorem as "Assumed" and independent review or formalisation as "Not supplied". So "checked in Lean's kernel" means that the arithmetic the Python fraction certificate already did has been redone by a much smaller and better-trusted checker. The check is real, and narrow.
I couldn't run it: there is no Lean toolchain on this machine, and I wasn't going to install one for this. So I wrote the five definitions out again in Python on integer pairs, including Lean's qsub, which is natural-number subtraction and quietly floors at zero, and evaluated the theorems on the committed inputs. All five come out true, and no subtraction ever hits the floor. The bit moment sum is of its bound and the complex one . At 80-digit precision, with the real exponential, the largest bit saving this histogram supports is about , so the claimed leaves 0.009% headroom; round five's file checks out the same way.
Lean wasn't new to the race, either. princezuda's #26 on CrocSwap, at 10:54 UTC, brought Lean 4 v4.21.0 with Mathlib ("no sorry and no native_decide") and 310 theorems over the certificate arithmetic of PRs #3 to #12, and it found one misstated ceiling. alejandrozu's #45 adds a Lean-checked Gaussian parity audit. But no CrocSwap claim after #12 has its arithmetic in Lean, and Jain's last two rounds do.
His Python checkers all ran. make verify took 194 s at 750 MB peak and passed everything, including 23 unit tests and a deliberately broken inverse that has to fail and does. independent/complex-twostage/run.py 24 took 39 s: 140,064 edges, 0 bad labels, and , slightly more than the the assembly uses. scripts/certificate_round6.py gave at . The histograms these scripts regenerate are identical to the ones pasted into Round6.lean, so the Lean file and the Python checks are about the same objects.
Where that leaves the race, at 15:46 UTC: the highest claim is Rohan Arun's #49 on CrocSwap, , 12.5% above Jain's . The number on CrocSwap's main, #39's , is 6% above Jain's. So there are now three kinds of checking on the board: Jain's arithmetic in Lean, main's new steps read through by its maintainer, and the open PRs' own Python certificates. And main's README now credits Jain, with Aurel Prosz, for "attributed two-stage development and the paid copied-stream endpoint construction", so the two races have formally met. I don't think the difference in arithmetic checking matters much, because the arithmetic was never the likely point of failure. Lean makes "the numbers add up" airtight; it says nothing about whether a contiguous projector block really is one recursive call on a tape, or whether the manuscript underneath is right. Jain said it himself, replying to someone who asked for the unconditional value: "Nothing past OpenAI's result is unconditional yet."
PR #39 reaches main
Rohan Arun opened #39 at 12:58 UTC, "Fixed middle basis and copied reversed corners", at . It was never the best open claim for long; his own #40 beat it by 0.82% seventeen minutes later. But at 13:23 Colkitt had pinned it, head 70ae241, in a merge commit (fd8c563) on an integration/community branch, and he spent the next two hours on it. At 15:12 UTC he committed "Complete conditional audit of the PR39 community witness". At 15:18 the release commit 0605a24, "Publish audited community bound with contributor attribution", moved main, the tag community-kappa-15-2026-10-08 went on it, and GitHub marked #39 merged. Since #39 was stacked on earlier work, the same push also marks icekylinx's #10, #18, #24, #32 and #36 and Rohan's #37 as merged; their commits are in its history. Twelve minutes later Colkitt posted:
Validated and merged Rohan's PR. Big gain on a really difficult regime (that frankly I was stuck at). Incredible work.
The maintainer admitting he was stuck, and that someone else's agent-assisted PR got him unstuck, is my favourite sentence of the whole episode. His headline, "κ = 2⁻¹⁵", is the floor the witness clears rather than its value; the exact number is , which the README puts at "27.36% above 2^-15". His "500 thousand fold improvement over the previous result" checks out against , the last exact witness he had posted (about 468,000-fold). Against the checkpoint that main carried until 15:18 it is about 42,000-fold. The "2 ^ 167 fold improvement over the original OpenAI result" is right: .
What did "validated" mean? The audit, docs/research/community-final-audit.md, is plain about who did it. Its header reads "Reviewed by the project's Codex assistant for Douglas Colkitt", and it calls itself "a maintainer mathematical assessment, not independent human peer review or formal verification." Within that, it is a real proof review, and it goes after exactly the obligations my PR #13 check said the certificates never touch. It works through the partial-swap compiler's block algebra; the claim that one fixed basis serves every gate (each required minor is a nonzero rational function on one irreducible family, so a single rational basis avoids every exceptional set); the copied-centre schedule restoring both operands for arbitrary dirty values; the complex side's phase compiler; the semantic-precision recurrence; the arbitrary-coordinate router; the Gaussian inverse; and the bulk tape schedule. In one place it swaps in a published theorem, Baker, Harman and Pintz on primes in short intervals, and labels it "an actual new input". Its verdict: "no unresolved additional construction or inequality was identified that blocks this particular witness."
The executable side is separate. The audit added its own checker, scripts/audit_community_candidate.py, which imports none of the contributor's code, encloses both recursive moments with 80-term rational logarithm bounds, rebuilds all seven margins, gets the same final gap of about , and shows that the next bit saving on the grid fails. It reports 218 tests, 20 upstream patch checks and certificate regeneration in Docker (GCC 13.3, Python 3.12), a GitHub Actions run on Python 3.11, 3.13 and 3.14, a one-header GCC build fix, and 20 files missing from a vendored manifest restored from PR #34 at their recorded hashes. I re-ran make verify on 0605a24 and it passed, 218 tests and all, in about sixteen minutes (How I checked has the details).
Two limits are written into the audit itself. "The original OpenAI #109 framework ... remains assumed." And "PR40 and subsequent submissions are outside this pinned review." So main is deliberately behind the frontier, by 6.1% when #49 arrived.
Main moves again
I thought the merge was the end of the story. When I fetched everything again at 19:46 UTC, main had moved twice more, 28 more PRs had arrived, a contributor had built a live dashboard for the race, and the race itself had jumped to a second problem.
The first move was routine. At 16:47 UTC GitHub marked Rohan's #49 merged, together with #43, #45, #46, #48 and princezuda's Lean PR #26, and the README went to . It was behind again at once: Rohan Gupta's #50 and RaD's #51 had been open for half an hour with higher claims.
The second move is the one that broke the plateau. Avi Eisenberg's #53 (he is ikeboy on GitHub), opened at 16:31, changes nothing except the scalar producer. Every vertex of the network needs its leave-one-out strip sums: given values, all sums that leave one of them out. The usual way builds prefix sums and suffix sums and takes . Each prefix then has two consumers whose envelopes don't nest, so its second use always costs a fresh wire. #53 shifts the bracket by one, , and now the two consumers of a prefix are nested, so the carrier matching from #44 can keep the existing wire going instead of allocating a new one. It costs more additions (40,329 against #48's 37,098 at ), but the carrier links double, from 6,002 to 12,719, and the roles drop from 36,432 to 32,946. It bought 9.2% in one step, after three hours in which every PR had gained a few hundredths of a percent. I like it because it is counter-intuitive: more arithmetic, fewer wires, and fewer wires is what actually pays for.
After that the ideas stacked. eumemic's #57 compiles every group of operations that share a frame as one invertible binary map and pays to reclaim retired wires, which takes the physical wire count from 160,799,739 to 153,481,944. Eisenberg's #62 replaces the strips with every cyclic interval, each built from an interval one item shorter. It needs additions instead of about , but almost every one of them continues a carrier. On its own it claims , and with #57's compiler on top, . Alejandro Zarzuelo Urdiales's #61 then refined the parameters on that same graph, by , with the arithmetic checked in eleven Lean proofs.
At 18:35 UTC, Colkitt merged #50 to #62 (all except Rohan Garg's #59, which is kept as a reviewed alternative) in one push. main now reads
23.7% above the #49 release and 31% above #39's. The network underneath has changed too: the bit side now runs at on dimensions 23 and 25, with 137,151,806 physical roles.
The review behind it, docs/research/community-round2-review.md, again calls itself "a conditional maintainer review with Codex assistance, not independent human peer review or a formal proof of multiplication." It reads more like a replay ledger than the #39 audit did. Each PR gets a disposition ("Full pinned replay passed", "Complete word and profile replay; used by selected witness"), and the places it holds back are stated. #51's separate witness had its "arithmetic … freshly replayed, but its entire standalone source-graph and negative-basis data argument was not", and #62's claims that its layout is optimal "were not independently reproduced". The general residual compiler, the all-size recursion, the tape implementation and the analytic transfer "remain written inherited obligations, not consequences of a successful finite test alone." So this move rests on finite replays plus reading, on top of interfaces that were audited at #39.
The same afternoon, main imported Jain's repository whole: 87 files, pinned by SHA-256, under research/swapnil-parallel/. It did the thing I couldn't. It compiled both Lean files under Lean 4.31.0, checked that the generator reproduces them byte for byte, and inspected all ten theorems, none of which depends on an axiom. Its reading of what they prove matches mine: they "do not formalize logarithmic analytic bounds, construct a Turing machine, or prove integer multiplication end to end", and the fixed-tape side of the ε → 1 stack has "not received a complete independent audit in this pass." It also notices something that matters for the next section. Jain's complex network saves , more than the 7.17e-5 of the complex network main uses, but main's bit side binds at about , so the better complex network would not move the headline. Jain himself has not posted a seventh round. His last commit is still round six, at 14:01 UTC.
The open PRs didn't wait. Fifteen more arrived between 18:05 and 19:46. Dominik Scholz's #63 and rfu08's #64 published the same construction within fifteen minutes of each other, and #64 says so ("Both decompressed words, complete profiles and finite bridge are identical"). The best open claim when I stopped was Scholz's #76, opened at 19:44, which combines three other open PRs (#70, #71 and #74) to reach , 1.4% above main.
So at 19:46 UTC the board reads: main at , the best open claim at , and Jain at , now 28% behind main. There have been 77 PRs from 20 accounts; GitHub marks 25 merged, 47 are open and 5 were closed without merging. The last fifteen open claims sit between 5.10 and 5.17 times . It's a plateau again, and I'm no longer going to predict how long it lasts. (About three hours, it turned out.)
A dashboard, built by one of the racers
At 15:13 UTC Aurel Prosz (@aurel_pr, Paureel on GitHub), whose two-stage topology and endpoint correction run through half the PRs above, posted a live tracker: Beyond n log n. Colkitt shared it fifty minutes later as an "awesome dashboard from the great @aurel_pr". The site owner's reaction was "damn, things are getting serious", which is fair. The page names no author in its text, but its footer links to Paureel, and the source is open at Paureel/beyond-n-log-n.

It is a React app with a Netlify function that re-reads GitHub every 15 minutes. It covers every PR in every state, all 26 public forks and the 6 PRs inside them, each checkpoint in main's history, a "field map" of which areas of mathematics each PR mentions, and an idea-lineage graph built from explicit reuse language in the PR descriptions. The values are parsed from PR titles and bodies, so they are claims, and the methodology box says as much: "No full-theorem verification is inferred from tests or certificates."
I pulled its JSON (/api/research, fetched at 19:33 UTC) and compared it with my chart. All 46 PR claims from #1 to #49 agree to within 0.2%, which is the rounding in my labels. There are two real differences. The tracker puts OpenAI's at the openai/math commit time, 21:58:50 UTC on October 6, where I had used the 3:19 PM Pacific (22:19 UTC) printed on Julian Schiavo's chart. I checked the commit and moved my point to match. And its "repo checkpoint" line steps at the commit times of everything in main's history, including contributors' commits that came in through merges, so it has main at 4.1239e-5 from 16:16 UTC (ed8201c), where I use 16:47, when GitHub marks the merge. Both choices are defensible. Its 19:33 snapshot also predates #75 to #77, so it showed #71 as the leader.
The race reaches the Fourier transform
At 16:52 UTC Ryan Shea (@ryaneshea) posted:
We're publishing a research draft on OpenAI Problem #130: computing the exact discrete Fourier transform below n log n. Our draft proposes an all-length bound of T(n) = O(n(log n)^(1−δ)), with δ = 7.3×10⁻⁵. That's a 730-million-fold increase in the exponent saving over OpenAI's published δ = 10⁻¹³.
The repository was created at 15:26 UTC with Shea's own earlier construction, a reduced centre and shared sums, worth about , or roughly 2,600 times the saving OpenAI's construction actually achieves. An hour later, at 16:27, he replaced it with Jain's round-six complex network and got . Asked whether this was AI-assisted, he replied: "The heavy lifting was AI, and I shepherded it towards the improvements."
- branch
- master
- tests
- none found
- source
- 77.2 kB
- commit date
- 2026-10-08
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-08 at bf38c00 — branch, commit, commitDate, fileCount, hasTests, languages, shallow
shallow clone: counts describe the pinned tree, not the history
Why should a network built for multiplying integers say anything about the FFT? The roundup's FFT section explains it. OpenAI's explicit Fourier paper (family 130) takes its finite gadget from the multiplication manuscript. The saving lives in one place: an algorithm that applies , tensor powers of the fixed matrix , in with . That algorithm is the same frame-labelled network on arbitrary roles that the race has been tuning: swap two banks, restore every scratch wire whatever it held, and count the dimensions of the frame changes. The rest of the Fourier proof (Good–Thomas reindexing, exact-width local words, synchronisation into sectors, a Bluestein chirp) uses the tensor routine only through that cost bound. The draft's whole argument rests on this. By its reading, Proposition 4.2 of the Fourier paper uses the tensor routine only through the bound on arbitrary complex arrays. So a better network gives a better , and is anything strictly below .

What crosses over is narrow, and the draft is careful about it. It takes only Jain's complex network at : 2,024 triples, , roles and a total child width of , with icekylinx's whole-residual batching (#10), the copied centres from #36 and Prosz's two-stage endpoint correction, all credited. It does not take the bit network, the ε → 1 stack, the Gaussian resampling or any of the tape and precision analysis. Its audit says it "does not use Jain's integer-specific epsilon stack, bit-interchange bound, Gaussian recovery or tape precision analysis".
So the Fourier number comes out large next to the multiplication race, and for a structural reason. In #109, is the smallest of seven margins and the bit network binds; Jain's complex saving was slack. The Fourier model is exact complex arithmetic with word-sized addresses. There are no tapes, no bit network and no to share the saving with, so can sit just under the complex network's own saving, . In the draft's words: "There is no integer-multiplication bit-network constraint or extra factor-of-two loss in this transfer." The half of Jain's work that never mattered for his own is the half that matters here. It is also the best complex network on offer: as noted above, main's own is 7.17e-5.
I checked the numbers. From the draft's role counts, with and gives 207,387,136, and gives the stated . The deficit is , and the 25-row histogram in Section 6 sums to the same , the rank sum that Round6.lean checks for the complex side. With exact fractions, , which is the gap that swallows the extra factor. At 80 digits with the real exponential, the moment at is (Colkitt's replay of Jain's file reports the same gap), and bisection puts the critical saving at , the "critical root" the draft mentions. The draft's own rational bound, which uses , is looser at , and it still clears 1. The headline ratios are right: is exactly 730 million, and against the earlier it is 132,727 times. Against OpenAI's actual critical saving of (the paper rounds it down to for its headline) it is about 350 million.
Its checkers are standard-library Python, and I ran them under the same exception as before. verification/run_checks.py took 80 seconds at 245 MB. It reproduced OpenAI's own from the paper's and , passed both earlier suites and matched their archived certificates, and then ran the round-six verifier: hashes of the vendored Jain files, all 4,096,576 local coefficient entries, both physical schedules with the copied centres charged, the full histogram, exact endpoint identities, negative controls, and "PASS: a=36926111/500000000000, delta=73/1000000, rational moment gap=7.919…E-14". The regenerated certificate matched the committed reference.
The draft is plain about the limits, and so is the best reply it got. The producer is Jain's code, vendored, and the extra checks "were written with the same assistant"; the dirty-scratch tests are finitely many rational examples; nothing executes the full DFT or the infinite recursion. Keith Adler went through it with Claude as referee ("treat it as a second pair of eyes, not human review") and couldn't break it. He listed what to tighten: the copied-centre step "is argued, not tested"; whole-residual batching "is not in #130" and needs a short extension of its Lemma 2.2; and the margin is thin, "a = 7.3852e-5 sits just under the critical root 7.3861e-5".
So what is it conditional on? First, on OpenAI's explicit Fourier manuscript, Proposition 4.2 and Section 5.4 in particular, which the draft invokes "with their stated hypotheses, not re-proved". As the roundup notes, the Lean statement for family 130 is a weaker one: a subsequential, non-constructive circuit bound from the second paper, with "no all-length, bounded-coefficient, conditioning, or bit-complexity claim". The explicit algorithm is manuscript-only, and that is the part Shea improves. Family 109 has no Lean statement at all, so the two problems sit at different levels: #130's core question is settled in Lean and its explicit constant is not, while #109's whole theorem is a manuscript. Second, it depends on Jain's finite complex network and the written arguments around it (copied centres, whole-residual batching, the two-stage endpoint). It does not depend on his assumption-heavy ε → 1 stack or on the hypotheses he replaced. Those network arguments are the same kind Colkitt's audit worked through for #39, but in the multiplication setting. Their transfer into the Fourier paper's array model, including a layout lemma extended from one selected slot to and a final address translation for the endpoint, is new and unreviewed. In one way this rests on less than Jain's own does. In another, nobody but Shea's assistant and Adler's has read the new part yet.
It changes nothing about real FFTs. The README says "No practical FFT speedup … is claimed", and the roundup's numbers still apply in spirit. Ignoring constants and the factor, a 1% saving at needs , so . It's a big improvement on the the roundup found at , and it is still not a number anyone will ever transform. What it shows is that the race is not about one problem. The agents and the people pointing them have found a shared engine, and any result built on that engine is now open to the same tuning.
9 October: the ninefold night
The site owner's note the next morning was short: the race to κ = 1 is still on, check the repos again. When I left it, at 19:46 UTC on October 8, main was at , the best open PR at and Jain at , and I had called it a plateau. At 11:00 UTC on October 9 main was at , the best open PR at and Jain at . A hundred and two PRs had arrived in between. There were 179 from 39 accounts; GitHub marks 30 merged, 127 open and 22 closed. So much for plateaus.
What moved main
The ninefold step is one contributor's chain of four PRs, each built on the one before. icekylinx opened #104 at 23:12 UTC (a "stopped product-ring interchange", ), #115 at 00:26 (partial source gauges, ), #130 at 02:29 (three-stage Cayley covers, ) and #144 at 04:19 (paired cubes and shared completed cores, ). Other people were tuning around them all night (eumemic's #114 was the first claim past , at 00:23), but #130 is the step that matters: it came in 2.4 times above the best open claim of the moment. Colkitt reviewed the chain and merged all four through an integration PR, #149, at 05:06. Ten minutes later he posted:
Validated and merged a breakthrough from IceKylin. In addition to their work, the latest result builds on work from @dysmemic, @SJ_Swapnil_Jain, an664, Zhihao Chen and others.
His "~9X from this morning" is right: over yesterday evening's is 9.03. So is "roughly 2¹⁷¹": it is times OpenAI's . A second post in the thread said the agents were "still backed up trying to process the huge volume of PRs that keep coming in. (Keep them coming!)".
Three shears instead of one swap
Every construction in this race is a way of doing one thing: an interchange, which swaps two banks of values and while putting every scratch wire back the way it found it. Until last night each interchange was a single network on about coordinates (576 for ; main's bit side ran at 575), with an enormous number of roles per invocation: 137,151,806 on main's bit side yesterday evening.
#130 does the swap the way you were taught to swap two registers without a temporary:
Three shears, , then , then , give the swap up to a sign, and the sign is a fixed correction at the end. The trick is in where each shear runs. The address space is split into three orthogonal blocks of dimensions , and , so (70 in #130; #144's complex side uses 72). Each stage acts on one block plus a single shared direction, and the frame a value leaves one stage in is exactly the frame the next stage expects, so there are no data connectors to pay for between them. The note puts it as "Every exit label is precisely the next entrance label."

To make every vertex of the network look the same, the invocations are indexed by every element of the orthogonal group of that binary space, which is what "regular Cayley cover" means here. For that group has about elements. The count cancels out of the normalized moment, so it costs nothing in κ, and it lands in the constant instead. eumemic's later README puts the network at "about 2^2570 roles". #144 then added two things on top: a "paired-cube" local word that writes the identity as , three pieces built from different parts of the circuit, and an664's completed-core sharing from #128, in which one auxiliary bank is reused in turn across the three orthogonal blocks.
Why smaller is worth nine times
The moment condition from #10 explains the gain, and once you expand it to first order it says something you can hold in your head. A network with roles on coordinates spends its rank on children of widths ; the total falls short of by the deficit . Then
where is the rank-weighted average of how far below full width a child sits, on a log scale. The saving is the fraction of rank the network fails to spend, divided by the log-distance of a typical child from full width.
I computed both sides of that for the old network and the new one, from their own histograms. Jain's round-six complex network (the one Shea took to the Fourier problem, and better than yesterday's main) has and : 96% of its rank sits in children wider than half of , so wide children are nearly free. That gives , against a certified . #144's complex side has , and , so , 58 times larger. Its children are narrow (only 13.5% of the rank sits in children wider than 36), so , 8.8 times larger. The ratio comes to , against a certified . The relative deficit grew much faster than the log penalty, and the net is a factor of 6.6 on this side. The deficit itself telescopes neatly, per vertex: .
The margin that binds
The assembly under #144 is CrocSwap's semantic assembly: 47 strict constraints, seven cost margins, and two small parameters and . The network saving fed to it is . The bit side binds: against . Then , and , and the binding margin is the compact phase layer, . The original-prefix margin sits exactly above it, which is the point: is chosen to balance those two, so
I redid that with exact fractions and got , with the binding margin above it (the grid swallows the rest), the same gap the maintainer's review reports.
One wrinkle on the bit side: its saving is "stopped". #104 pays the cross-bank address adapters as loops over fixed-width atoms, which costs linear time, and stops the recursion at atom width with . The usable saving is then , a thousandth of it coming from an older, weaker network (). Unstopped, #144's bit network works out to , and the stop costs it about 0.1%.
The review behind the merge, docs/research/community-round6-review.md, is "a construction and dependency review, not independent human peer review or a formal proof of integer multiplication", done "with OpenAI Codex assistance". It rebuilt the complex producer, ran its own controls (exact factorizations over , and , exhaustive frame identities in dimensions two to four, and dirty-value replays of the 26,417-role mixer), read the written proofs, and prepared two small fixes, a verifier guard and one sign in a sentence. It is honest about one gap: a 1.5 MB compressed bit witness from #97 could not be fetched through its connector, so "the selected bit reconstruction was checked through inspected, head-pinned CI, not rerun locally." That file is in my clone, so I ran make paired-cube-verify on main: the producer regenerated from source and matched its certificate, the bit reconstruction ran locally with every profile and source hash equal, and it ended PASS kappa=4609169/10000000000 = 4.609169e-4; both moments, shared cores, finite router and 47 strict constraints, in 44 s at 270 MB.
Jain's rounds seven to ten
Jain posted four more rounds. Here they are against what was open on CrocSwap at the same minute, as before.
| post (UTC) | Jain's | best open CrocSwap claim then | what was new |
|---|---|---|---|
| Oct 8, 20:56 | #86, | deferred garbage readout | |
| Oct 9, 04:17 | #143, | opposite bank orders (#104), signed reclaim (#112), shared cores (#128) | |
| Oct 9, 05:25 | #147, | his bit word inside #144's three-stage cover | |
| Oct 9, 09:50 | #168, | paired-cube words, birth reuse, nested-prefix modules |
Round seven led everything for two hours, by 22% when it was posted. Zhihao Chen's #97 at 22:39 ported Jain's witness into CrocSwap and came within 0.007% of it, and #99 passed it at 22:56. Round eight landed in the middle of the cover jump and was 3.4 times behind. Round nine was the highest claim anywhere for fifteen seconds, until DaysSky's #150 at 05:25:46. Round ten was about 6% behind eumemic's #168.
What, then, is round nine's 9.14-fold jump over yesterday's main made of? Mostly icekylinx's cover. Jain's README says it plainly: scripts/certificate_round9.py is "our own implementation of the paired-cube / three-stage cover assembly of icekylinx's PR #144", written from #144's description and files, with "None of PR #144's code" run. It first reproduces #144's published exactly, all 47 constraints and 7 margins. Then it swaps one thing: the bit side's word becomes his own round-seven witness (deferred readouts at , ), with 1,504 of its 10,922 deferred gauges omitted. Choosing which ones to omit is an exact – minimum cut, a formulation he credits to Th0rgal's #146. That raises the bit network's saving from #144's to ; the complex side is #144's, unchanged. So round nine is 1.18% above main, and that 1.18% is Jain's. The ninefold before it is the cover's. His round-ten README credits, by PR number, eumemic, icekylinx, DaysSky, jamesyc, Th0rgal and an664 for the levers it stacks.
On assumptions, round nine moves him closer to the CrocSwap side rather than further away. Rounds one to eight ran on his own ε → 1 stack and its "deliberate departures from upstream hypotheses", the prime number theorem in place of the interval lemma among them. Round nine's certificate doesn't import that code at all. It runs #144's assembly, which carries the hypotheses main carries. What it adds is a short list of interfaces "stated here but not machine-checked": the cover lifting, frame theorem, exterior rule and stopped recurrence from #130 and #144; a precision grid of the form , because the cubes divide by 3 ("This looks benign …, but only dyadic grids were audited before"); block-factored frame representatives; and two inherited row-reserve constants. The maintainer's #144 review does discuss the denominator-three grid, so that one has had a reading too.
I recomputed round nine with exact fractions, outside his scripts. From the certified , the stopped mix gives , which binds against the complex . Then , the binding margin is , and the floor on the grid is , as claimed, under the margin. of that is , so "2⁻¹¹·⁰⁶⁶" is right, and the ratios in his post (3.70 over round eight, "2¹⁷⁰ fold") check out too, the second one rounded down from . At 80 digits with the real exponential, the critical saving of his bit histogram (rare-class fallback included) is , so the certified grid point leaves almost nothing on the table. make round9 and make round10 pass, and round ten's certificate re-derives #144's and round nine's κ before certifying its own. The heavier round-nine ledger, which replays the whole forward F2 shear of the modified bit word and rebuilds its child histogram from the events, printed ALL PASS after 110 s at 0.96 GB. Rounds seven to ten each ship a Lean file again (lean/Round7.lean to Round10.lean), still core Lean with no sorry, no axiom and no native_decide, and round nine's now includes the finite bridge and the 47-row assembly. There is still no Lean toolchain on this machine, so I didn't run them; my fraction recomputation covers the same arithmetic.
Two things in his thread are worth keeping. Told by a reply that he was "losing the frontier", he answered: "I found mistakes in the PRs there so have to be super careful. Public results I publish are 100% verified correct." And round ten's commit fixes a bug in his own F2 checkers, which had been identifying a kept copy in a way that skipped some walk and data checks; he reports that rounds seven to nine are unchanged under the fix. I'd trust the second more than the first. The third post in that thread says "This will legit not be able to continue unless you support, I am running out of my resources", which is a reminder that all of this runs on somebody's token budget.
Who @dysmemic is
The other post the site owner flagged, "κ = 4.135962 × 10⁻⁴. The race to κ = 1 continues. We must have linear integer multiplication. We will have it!", is from @dysmemic, and the screenshot attached is the tracker. The account is eumemic's: its earlier posts link eumemic's PRs #5 and #13, and on October 8 it linked github.com/eumemic/exoplanets. The number is the title of eumemic's #137, "Padded triple covers with sequential dirty reuse", opened at 03:18 UTC, six minutes before the post. So it lives on CrocSwap as an ordinary PR, not on a fork. It led the open claims for six minutes. geckods' #138, a tightening of the same slack constants, came three seconds after the post. eumemic's later #168 was the best open claim through the morning, and the 10:41 leader, chafreaky's #179, refines its operation frames by 0.0047%.
The tracker now
Prosz's tracker changed its headline overnight. It used to lead with the strongest claimed κ, a draft PR at the time; it now leads with the maintainer-reviewed κ, read from main's certificates/selected-result.json, and shows the open claims separately. He also added a zoom, after people complained that everything since was a flat line at the top of the chart. When I rendered it at 11:00 UTC (data "Updated 9 Oct, 10:36 UTC") it showed 178 PRs, 27 of them drafts, from 39 contributors, and 49 public forks.

Sort its PR list by highest κ, though, and the top entry is #159 at , which is not a claim anyone should read. It was closed eleven minutes after it was opened, and its whole body is one sentence about "multivariate Hamiltonians" followed by "Kind reminder, AI is not to be blindly trusted." The tracker parses titles, which is the right design for a dashboard and the reason not to take its sort order at face value. I left #159 off my chart.
And the Fourier problem again
Shea's draft hasn't moved: his only commits since are a README rewrite at 23:52 UTC that adds "This is not a faster FFT" and a who-did-what list. Its δ is still from Jain's round-six network.
eumemic, though, did the obvious next thing. At 08:54 UTC @dysmemic posted that the community's networks "also apply to OpenAI's new exact DFT algorithm, raising its saving from 2.1·10⁻¹³ to 4.856·10⁻⁴ in the log exponent. About a 2 billion-fold improvement!", linking eumemic/exact-dft-bounds. Forty minutes later the same account added, "Looks like @ryaneshea already had this idea", and the README now says Shea's repository "has priority for the idea of the transfer". The new repository takes #144's complex network as merged on main (, , deficit 1,936) and adds a batched recursion for the Fourier setting (its Theorem 3.1), so that a frame change of residual dimension becomes one recursive call to . The result is , so . Against OpenAI's actual critical saving of that is 2.3 billion times, and against Shea's it is 6.65 times.
The README is careful about scope: it is "a paper proof with an exact finite certificate, and it is not formally verified", the network's "exact complex phases, the Cayley cover and completed-core sharing are written proofs", and "The constants are astronomically large: the network has about 2^2570 roles, and the recursion's base range is of order 10⁹ bits." Its make verify is one exact-rational moment check plus source pins. I ran it (0.34 s; "margin 3.158222e-03 … PASS" out of ), and my 80-digit critical root for the same histogram is , so leaves about 0.012% headroom. Nobody has reviewed the Fourier transfer itself yet, the batched recursion least of all.
Where the ceiling stood at 11:00
At 11:00 UTC on October 9: main is at , reviewed by its maintainer with Codex; the best open claim is chafreaky's #179 at , 42% above main; and Jain is at , with its arithmetic in Lean. Everything is still conditional on the OpenAI manuscript.
None of this has touched the ceiling. #144's assembly still has a margin equal to and another equal to , so κ cannot pass ½ whatever the networks do. What bounds the current witnesses is much lower than that, and it is the chain above: κ is , about , so 0.09% under the weaker network's saving; and that saving is about divided by . To get to you need networks that waste a couple of percent of their rank at a similar log spread. One draft PR (mikey1201's #172, a "zeta family" on the Boolean lattice) projects κ up to at through the in-tree assembly, but it is labelled a research record with named open obligations, not a claim, and I haven't checked it.
In runtime terms, takes 0.26% off at and 2.0% when , and halving the time needs . The calculator further down has it as a preset.
9 October, afternoon: reusing dead registers
The site owner's second note of the day was "there have been more updates to Jain, check his tweets". There were two. At 12:50 UTC Jain asked for an arXiv endorser in cs.CC (Computational Complexity) for a paper that "collects proofs from our work on OpenAI problem #109", and at 13:07 he posted round eleven, . He ended that post differently from the other ten: "We seem to be approaching a wall. Gains now come in fractions of a percent. The next real jump needs a big breakthrough which we are unable to find yet." By the evening the CrocSwap side was saying the same thing, with proofs.
Rounds ten and eleven do one kind of thing, most visibly on the complex side: they stop paying for registers the network has finished with.
Round ten's version is birth reuse, jamesyc's idea from #124 with eumemic's late compensation reads from #143, which Jain re-implemented as an exact maximum-weight matching. A gauged slot is a role whose register has to be read at a known frame, and in the plain ledger it pays a tail child to get its register into place. Born on the register of a slot that has just died, it doesn't: the donor's last climb and the newborn's first are charged as one ascending chain. On round ten's complex word all 3,180 gauges found a donor. With the rest of that stack, at (so and deficit per vertex), the word came to 14,592 roles and a certified .
Round eleven's new piece is the one his thread describes as "two retired copies of the same value are subtracted in place, and a fresh slot is born on the zeroed register with no read". His word has 428 pairs of registers that each finish holding the same value and are never touched again. For such a pair, one gate , run at a frame that contains both copies' last frames, leaves with no signal at all, only dirty scratch. A fresh role is then born on with no read. Its starting content is a linear function of the initial dirty values, and the frame-zero compensation of every ungauged role is recomputed over the new transcript to cancel it. The kernel is from icekylinx's #184 ("zero-fresh recycling through a containing frame"), a PR that was closed five minutes after it was opened. Jain's own part is applying it to this word, choosing 386 hosts by another exact matching, and writing the ledger and the checker.
The ledger shows why it helps and why by so little. Each host removes one role, so falls from 13,308 to 12,922, and the deficit stays at 1,320, because the children removed and the children added differ in total width by exactly per stage. In the commonest shape (copies last at dimensions 2 and 3, mixed at dimension 4, ) the removed children have widths 20, 4 and 19 and the added ones 2, 1 and 18. In the first-order picture from the ninefold night, rises 3.0%, but the new children are narrow, which costs about 1.5% in , and the net is the 1.43% his gate reports: goes from to . From the round-eleven histogram I get and , so against the certified .
Most of round eleven's 9% over round ten comes from elsewhere. The complex side switched to eumemic's #168 v4 modules (used as data), a re-chosen cube-local circuit Jain calls C1, which he puts at about 1.3% better than #168's, and 42 deleted terminal outputs. The bit side got a new cyclic all-but-one module, 1,762 birth-reuse pairs and 24 terminal sinks. Its went from about 21,491 to 19,017 with the deficit fixed at 1,936, so rose 13% and 6.7%, and the coarse saving went from to .
Which side binds, and which margin
The two rounds bind on opposite sides. In round ten the complex network was the weaker one, against a stopped bit saving of . Round eleven's complex side jumped past the bit side, so now the bit side binds: , stopped at to . Inside #144's assembly nothing moved. The binding margin is still the compact phase layer , with the original-prefix margin exactly above it. I redid the assembly with exact fractions: , , and the floor on the grid is , under the margin. Its is , so "2⁻¹⁰·⁵⁴⁷" is right, and "1.09 fold" is 1.0865. The same arithmetic reproduces round ten's , bound by the complex side.
make round11 passed in under a second. It reproduces #144's κ, round nine's and round ten's before certifying round eleven, then runs eight unit tests. I also ran the two new heavy gates by hand. The complex recycling checker on the word ended ALL PASS after 44 s at 0.96 GB, with its controls rejected, among them a skipped mix and a flipped mix sign. The JSON-only bit ledger replayed the word forward and reflected, matched the release inventory, and ended ALL PASS after 2 min 18 s at 0.98 GB. lean/Round11.lean is there. His gen.py regenerates it byte for byte from the histogram file, it has no sorry, axiom or native_decide, and it covers both moments, the stopped mix, the finite bridge and the 47-row assembly. There is still no Lean toolchain on this machine, so I didn't run it. The README is plain about what is new and "stated here but not machine-checked": the recycling rests on #184's argument and "There is no literal frame replay of the mix gate at any size"; the terminal deletions and sinks rest on jamesyc's lemma from #166; and the C1 circuit, the per-operation descent and the recycling are checked at full size only by an F2 recount and a scalar replay.
Nineteen minutes in front
When round eleven went up, at 13:07:49 UTC, it was the highest claim anywhere. main had moved ten minutes earlier (below) to , and the best open PR was ikeboy's #194 at , so Jain was 1.0% above main and 0.17% above the best open claim. It held until 13:27:16, when evmckinney9's #197 claimed by packing the bit residuals of rohanarun's #187 into shared banks. That reads #197 by its title as it stands tonight, and PRs here get edited in place, though its body's own comparison with #194 agrees with the title. By 18:50 Jain was 2.4% behind the open frontier.
The arXiv paper
The request itself is two lines. arXiv asks people submitting to a category for the first time to be endorsed by someone who already has, so this is the normal gate for a new author. I went looking for the paper. It isn't in his repository at the newest commit (1a580dc, 13:01 UTC), which holds the ten TeX notes and their PDFs, the certificates, checkers and Lean files, but nothing that reads as one paper, and none of his other public repositories has it. So I can't tell you what it proves.
What I can say is what a paper collecting these proofs has to stand on, because the README sets it out round by round. It assumes OpenAI's manuscript, which has no Lean statement and no published review. Since round nine it runs inside icekylinx's cover assembly from #144 and inherits that assembly's interfaces (the three-stage cover lifting, the exterior rule, the stopped recurrence), and on top of those sits his own list of stated-but-unchecked interfaces, which grew again in round eleven. What Lean checks is the arithmetic: the moments, the stopped mix, the finite bridge and the assembly. A single written paper would at least give the proofs one place to be refereed, and refereeing is what this race has had least of.
main jumps again, by import
At 12:57:37 UTC Colkitt pushed a morning's worth of reviewed commits to main in one go, taking it from to , about and 43.6% higher. The witness is Dugongue's #186. It takes chafreaky's shared-edge complex supplier from #181 and eumemic's #168 v4 bit word and adds a compensated rematch, eleven coordinated frame cuts, and "completed entrance banks", which on the bit side pack 2,200 rank-20 entrance gauges into shared scratch banks. Here the complex branch binds, at against an effective bit saving of about . The push also carried an overnight candidate that never reached main on its own, the recycled-bit composition of #147, #150 and #151 at . GitHub marks those three, and two verification PRs, #64 and #101, merged at 12:57:38. #186 itself came in as a pinned package and is still open on GitHub.
The review, community-round8-review.md, is again Colkitt "with OpenAI Codex assistance", and this time it found a gap and filled it. #186's identity "does not alone specify an efficient conflict-free schedule for the physical invocations", so the maintainer wrote a bank-scheduling supplement, checked exhaustively in a small model over (3,888 invocations, 15,552 endpoint basis columns), and says plainly that it "is not a new Lean-checked theorem". It also explains why Rohan's #185 can't be stacked on top: #186 is limited by its complex side, so better bit leaves change nothing. I ran make entrance-bank-verify on main at 3b6b668. It passed, "complete complex and bit words, 47 constraints, seven margins, and 21 assembly controls", plus the scheduling diagnostic, in two minutes at 1 GB. As far as the fxtwitter mirror shows, Colkitt hasn't posted about this one.
The race starts proving its own ceilings
The PRs I found most interesting this afternoon claim no κ at all. DaysSky's #192 proves that no choice of operation frames for #168's complex word can certify more than . It starts from "Every certified saving satisfies a < D/C", where . That is the first-order formula from last night turned into an inequality (), and it follows from . Then it bounds from below over every frame layout. Two hours later DaysSky's #201 bounded the whole design: no word on #144's paired-cube design can certify more than , about 3.9 times main, whatever its circuit, gauges, reuse or frames. It leaves out the newer source-assisted words, which break its rules. chafreaky's #212 carried #192's method to the bit word that #200 to #207 share and got there, so frame tuning on that word has at most 1.8% left. huxint's draft #203 argues that several current lines can't reach even with every auxiliary call free.
None of these has been reviewed, and I haven't checked them. They are still the right work at this stage, and they put a number on Jain's wall: the word families being tuned are within a few percent of their own ceilings, and the paired-cube design as a whole has at most a factor of about four left.
Where it stands at 18:50 UTC
main is at , reviewed by its maintainer with Codex. The best open claim is hcg890's #217 at , opened at 18:38 and 3.4% above main. There are 217 PRs from 47 accounts; GitHub marks 35 merged, 158 open and 24 closed. Jain is at , with the arithmetic in Lean. All of it is still conditional on the OpenAI manuscript. In runtime terms main now takes 0.37% off at and 2.9% at , and halving the time needs .
On the Fourier side, eumemic's exact-dft-bounds hasn't changed since its 09:34 credit to Shea. Shea's draft has. At 15:24 UTC he reorganised it as a running record of transfers and selected Jain's round-eleven complex network, which takes to from the tensor saving , 1.38 times eumemic's . His audit file calls it "same-assistant mathematical review and finite checks; independent review and end-to-end formal verification remain pending". I ran its verify_round11.py, which checks all 17,057,040 scalar output coefficients of the word, and it passed in four minutes.
Prosz's tracker has more rows but no new layout, so I haven't re-shot it.
What "conditional" means here
Everything in the repository is stacked on top of other assumptions. The base layer is OpenAI's 73-page manuscript, which has no Lean statement and, as far as I can find, no published expert review. On top of that are Colkitt's written extensions, and on top of those each PR builds on earlier PRs. By the evening GitHub marked 25 of them merged, after two review rounds (#39 and the six it stacks on, then #43 to #62 in two batches); by 11:00 UTC on October 9 it was 30, with icekylinx's chain from #104 to #144 merged after a third, and by the evening 35, after a fourth round that imported #186 as a package and left it open. The 158 still open are neither merged nor reviewed. PR #44's dependency list runs through #43, #41, #36, #34, #32, #29 and on down to #3.
The maintainer's contribution-review index treats that stack carefully. It had six review stages in dependency order, foundations first (#3, #5, #7), then batching (#10, #13), then the later layers, and it states its policy: "A contributor's successful test report, or a large headline in a PR title, is not treated as verification by this project." The integration/community branch where #39 was staged carried a ledger whose gates were honest about the state of review: "Executable replay passed; proof review pending", and for two-stage and copy schedules, "Initial text inspection only". The audit closed those gates for #39's chain, and the second round extended them by replay to the PRs it merged; nothing else. The ledger also says "A failed gate should isolate the strongest surviving earlier checkpoint rather than trigger an unsupported all-or-nothing acceptance of the latest number." That rule is why the checkpoint is still kept, on release/ternary-30 and in the release notes, and of all the process decisions in the repository it is the one I would copy.
Who wrote it
Mostly AI agents, and they say so. 42 of the 44 PR descriptions name the model or tool. OpenAI Codex appears in most of them; 15 also credit "GPT-6 Astra"; eumemic's PRs end with "🤖 Generated with Claude Code"; Dominik Scholz's #33 credits "Anthropic Claude Opus 5.5, OpenAI GPT-6 Astra and Codex"; and Rohan Arun's #44 says it was "Prepared by Rohan Arun with substantial Anthropic Claude assistance; OpenAI Codex performed this local arithmetic replay and upstream submission". Seven PR branches carry the codex/ prefix that Codex gives the branches it opens. #34 mentions only "independent agent reviews", and #2 says nothing. Colkitt's notes all carry the footnote "Prepared with assistance from OpenAI Codex", and CONTRIBUTING.md asks contributors to "disclose substantial AI assistance". Jain's repository says the same of itself: "Research, implementation and drafting were done with assistance from Claude (Anthropic)." Rohan Arun, whose PR ended up on main, put the tempo best, announcing #40 at 13:39 UTC: "The improvements in the end started getting faster than we could run the verifications." In the same post he tagged @sama and @thsottiaux with "could I get a reset?", which I read as a plea about usage limits, and added, diplomatically, "Mostly used Codex here but Claude collaborating on the other side might have helped reach a new best".
The style gives it away too, though that is my reading rather than proof. Every PR has the same structure (result, construction, verification, attribution, scope), the same careful hedging ("not formal verification or completed external expert review"), and the same habit of reporting a few hundredths of a percent over the previous PR, often within ten or twenty minutes of it. People aren't usually that fast or that consistent. A good share of the tempo is people pointing agents at the newest PR and asking for more.
Colkitt's own posts are the human layer on top. On the morning of October 8: "Went to bed last night, and sorry to report only a bit of incremental progress on my side. But it appears like this has taken off in terms of now having a real community effort", followed by a promise to "validate the results, definitely make sure we have attribution". The attribution machinery is the most unusual part of the repository: a CONTRIBUTORS.md that credits all 39 PRs at the checkpoint, including the closed and superseded ones; a pinned contribution-snapshot.json of every PR head hash; NOTICE entries for #7 and #3; and, until the audit, nothing merged into main, so nobody's PR silently overwrote anyone else's. When #39 went in, the attribution went with it: the README names Rohan for "the final #39 witness" and then nine contributors and groups by what each supplied, and NOTICE gained the line "it is not solely Douglas Colkitt's work." Colkitt's post promised more: "I want to take time to highlight each of them, but for now getting these results verified and published as fast as possible so everyone working on the problem is at the leading edge." At 03:32 UTC on October 9 he posted a video, "Small, But Not Zero", billed as "An opera explaining the integer multiplication math journey up to this point." One contributor, hipotures, left a comment on #20 calling it "real-time mathematics" and comparing it to SETI@home. The comparison is a stretch, but I can see why it came to mind.
"Linear time by tomorrow"
Two hours after the post, Julian Schiavo posted a chart.

When Colkitt's post landed, he followed up: "update! @0xdoug hit 2^(-34), so we are sadly now 18 minutes behind schedule".

He kept it going. At 16:30 UTC on October 8, an hour after Colkitt's post: "the saga continues. good news: @0xdoug hit 2^(-15)! bad news: we missed the 12:44am target, based on my careful data analysis we will now have linear time integer multiplication at 6:44pm. by popular demand, I removed the O(n) limit".

The fit is now a power-law curve in rather than a straight line, and with the cap removed the red line sails on past , into multiplication faster than linear time, an algorithm that would finish before it had read its input. 6:44 PM Pacific is 01:44 UTC on October 9. When I last looked, at 19:46 UTC, main stood at and the best open claim at , so the curve needs about fourteen more doublings in six hours. I'm not taking the bet, but I wouldn't have taken the one against #53's jump either.
The 6:44 PM deadline passed with main at . Then the cover landed, and at 05:53 UTC on October 9 he posted again: "New data means new data analysis! Updated data suggests we will now reach linear time integer multiplication at 8:33am tomorrow morning".

8:33 AM Pacific is 15:33 UTC on October 9. At 11:00 UTC the best claim anywhere was , so the curve needs ten and a half more doublings in four and a half hours. The ninefold night is the closest the real data has come to keeping his schedule, which is funny in its own right. It doesn't change the argument below: every one of those doublings has to come out of a relative rank deficit, and κ can't pass ½. The deadline came and went with the best claim at (maxime-fleury's #199), and there hasn't been a fifth chart.
It's a joke, and a good one, and the chart above already answers it: the extrapolation is on. A straight line in through Colkitt's and posts reaches at about 13:08 UTC on October 8. At that moment the real frontier was PR #39 at . The line was about fourteen and a half doublings out, which is a factor of roughly 25,000. It was never going to land, and the reasons are useful.
First, is bounded, and this construction has a lower ceiling than the obvious one. Linear time is , but as shown above the shape of the margins keeps below and below the network's own saving , which is a relative rank deficit. Second, the early gains were not discoveries about multiplication. They were slack being given back: the manuscript rounded a real saving of about down to and multiplied three conservative factors together. Undoing deliberate conservatism produces enormous ratios quickly, and then it is used up. Third, every later improvement is a smarter constant inside the same recursion, and the flat tail of the chart shows what that looks like.
And then there is runtime. With the constants ignored, the ratio of to is . Take bits, an integer with about binary digits, roughly one per atom in the observable universe. At the saving is about of the running time. At PR #13's it is about percent. At PR #44's it is about 0.023 percent. Even with , a number whose length in bits needs a 64-bit counter, PR #44 saves 0.18 percent. To halve the time you need : at PR #13 that is , at PR #44 about . That's before the constants, which in this construction are astronomical (the roundup's FFT section found a -role batch at the top level of the sibling paper's network).
log₂ n = 2^8. n = 2^256 bits: about 10^77 bits, near one per atom in the observable universe
- (log n)^κ − 1
- 4.269 × 10^-6
- time saved vs n log n
- 4.269 × 10^-4 %
- log₂ n for a 2× saving
- 2^(1.30 × 10^6)
A 1% saving needs log₂ n = 2^(1.88 × 10^4). The slider stops at log₂ n = 2^64. Constants are not in this model, and the real construction's are astronomically large.
So GMP is safe, and it would have been safe at as well.
Why I still think it's interesting
Not for the number. What's interesting is the loop. Within two days of a hard theorem being posted, a stranger forked it, rewrote its parameters with an agent, wrote up each step with exact certificates and patches against a pinned source, and published. Then about a dozen other people, almost all working through agents, started building on each other's unmerged PRs every few minutes, with tests, PDFs, notes and attribution blocks that are more careful than most human repositories manage. Some of it is real mathematics: the identity, the batched moment recurrence and the source-frame observation are ideas, not tuning. The certificate culture is real too. Every headline in the repository comes with a script that rebuilds it from exact fractions, which is why I could check PR #13 in an afternoon.
What it is not: a faster multiplication algorithm, a refereed result, or a formal proof. Everything depends on a manuscript nobody has independently reviewed, only what main has taken in (#39's chain, then #49 and the #50 to #62 round, then #144 and #186) has been through the maintainer's review, and that review was Colkitt with Codex rather than a referee, and the hard obligations (that the tape procedures do what the notes say at full size) are exactly the parts the certificates do not touch. If something in the upstream argument breaks, the whole tower goes with it, and the repository's process is designed to fail back to the strongest surviving checkpoint rather than pretend otherwise.
The race also shows the risk. Seventy-seven PRs in under twenty-three hours is far more than one maintainer can review. He audited one chain in about two hours, and by the time it was merged the open frontier was 6% past it. He then reviewed thirteen PRs in one round, and an hour after that merge the open claims were 1.4% past it again. Overnight he reviewed a four-PR chain and merged a ninefold gain, and by 11:00 the open claims were 42% past that, across 179 PRs. His next review put main at #186's ; Jain passed it ten minutes later, and by 18:50 the open claims were 3.4% past it again, across 217. And now the same engine is being tuned for a second problem, with a second set of unreviewed transfers. At this volume, the scarce resource is the human checking, not the agent output.
How I checked
I shallow-cloned the repository and fetched PRs #3, #5, #7, #10, #13 and #44 by ref. Commit times come from git log, PR times and descriptions from the GitHub REST API (44 PRs at 14:24 UTC on October 8; #45 arrived at 14:35), and the X posts and their replies from the fxtwitter mirror. I read the upstream manuscript's introduction, motif bounds and assembly sections, Colkitt's notes, the review index, the contribution snapshot and the integration ledger.
As a stated exception to my usual rule of not executing third-party code, limited to the repository's own pure-arithmetic checkers with no network, I ran the README's focused verification on main and the PR #13 certificate generator, its tests and its full make verify, all under nice -n 19. I recomputed PR #13's seven margins and its bit moment independently with Python fractions and 80-digit decimals, the residue identity by brute force at , the implied network savings from the certificates' counts, and the runtime factors with 80-digit decimals. I did not run any later CrocSwap PR's suite, and did not review any proof. OpenAI's release time on the chart is the openai/math commit time from the GitHub API, 21:58:50 UTC on October 6. The values in the race chart are computed from each PR's stated , not from its certificate.
For Swapnil Jain's track I cloned integer-mult-kappa with full history (20 commits, 07:17 to 14:01 UTC) and read all of it: the README, NOTICE, the nine notes, the certificates, the independent/ re-implementations and the two Lean files with their generator. His six posts and their threads came from the fxtwitter mirror, and the CrocSwap PR bodies and their credits to him from the GitHub API, last at 15:24 UTC, when #48 was the newest PR. Under the same exception I ran his make verify, scripts/certificate_round6.py and independent/complex-twostage/run.py 24, all pure Python with no network, and checked that the histograms they regenerate equal the Lean inputs. There is no Lean toolchain here, so I did not run Lean; instead I re-implemented the five Lean definitions in Python and evaluated both files' theorems, recomputed the round-six margins with exact fractions, and redid both moments at 80 digits.
For the merge I fetched every PR ref again at 15:45 UTC, read git log --graph on main (the merge commit fd8c563 has 70ae241, #39's head, as its second parent), the PR states from the GitHub API, the diffs of c9fca20 and 0605a24, the audit, the release notes, the integration ledger, CONTRIBUTORS.md and NOTICE, and Colkitt's and Rohan's posts through the fxtwitter mirror. Under the same exception as above, I ran make verify on 0605a24 at nice 19. It passed: the audit's own arithmetic check ("PASS independent moments, seven margins and all 32 RaD source hashes"), the #39 producer and its PASS bit=3886826921/100000000000000;kappa=971668963/25000000000000, every earlier witness back to , 20 patch checks and all 218 unit tests, in 16 minutes 12 seconds of wall clock at 2.9 GB peak, with the working tree clean afterwards, so every regenerated certificate matched the committed one. That checks the certificates and the replay, not the audit's proofs, which I read but did not re-derive.
For the evening I fetched every PR ref and the GitHub PR list again at 19:46 UTC (77 PRs), read main's log since 15:18, the README at ed8201c and 0d235fe, the round-two review and validation receipt, the imported research/swapnil-parallel/ review and its verification.json, and the bodies of #53, #57, #61, #62, #64 and #76. Jain's repository, cloned again, has no commit after f2176bc. For the tracker I rendered the page in my own headless Chromium, read its source repository and compared its /api/research JSON with my chart's data. For the Fourier draft I shallow-cloned shea256/fourier-transform-below-nlogn at bf38c00, read the README, the manuscript, both audit files and the verifiers, and read the thread and replies through the fxtwitter mirror. Under the same exception as above I ran verification/run_checks.py (pure standard-library Python, no network). I recomputed , , the deficit, the histogram sum, and the ratios with exact fractions, and the moment and its critical root at 80 digits. I did not review the transfer's proofs.
For 9 October I fetched both repositories again at about 11:00 UTC: main at d1d6c07, every PR ref and the GitHub PR list (179 PRs), Jain's repository at d2f6146 with full history, and his, Colkitt's, eumemic's (@dysmemic), Prosz's, Shea's and Julian Schiavo's posts, threads and replies through the fxtwitter mirror. I read the bodies of #104, #130, #137, #144, #159, #168, #172 and #179, the three-stage-cover and paired-cube notes, the maintainer's round-six review, and Jain's README sections for rounds seven to ten and his round-nine and round-ten certificates. Under the same exception as above I ran make paired-cube-verify on main (44 s, 270 MB) and then the full make verify, which passed its community, producer, partial-gauge, three-stage-cover, paired-cube, certificate and ternary stages and then stopped after 44 minutes, at 3.1 GB peak, in a preserved-research check for PR #51 whose C++ helper needs Boost headers this machine doesn't have; I also ran Jain's make round9, make round10 and the round-nine bit ledger with its certificate check, and eumemic's make verify in exact-dft-bounds, all Python (plus the repository's own small C++ helpers on main) with no network. I recomputed #144's and round nine's assemblies with exact fractions, both round-nine moments' critical roots at 80 digits, the first-order saving for #144's complex network and Jain's round-six network from their histograms, the orthogonal group's order, and every ratio quoted. I did not run Lean, and I did not review the cover's or the Fourier transfer's proofs. The PR points from #78 on come from the GitHub API and the tracker's JSON, with each PR's κ as its title or body read at 11:00 UTC, which for PRs edited in place (#168 most of all) is a later claim than the one they opened with.
For the afternoon of 9 October I fetched everything again at about 18:50 UTC: main at 3b6b668 with GitHub's push log for it, the GitHub PR list (217 PRs, with titles and bodies), the tracker's JSON, Jain's repository at 1a580dc with full history and the list of his public repositories, eumemic's exact-dft-bounds at 6f87d1a and Shea's draft at a2840b1, and Jain's, Colkitt's, @dysmemic's, Shea's and Julian Schiavo's posts, threads and replies through the fxtwitter mirror. I read Jain's README sections for rounds ten and eleven, the recycling gate's README, the maintainer's round-eight review, Shea's round-eleven audit, and the bodies of #186, #192, #194, #197, #201, #203, #212 and #217. Under the same exception as above I ran Jain's make round11, his complex recycling checker on the word and his bit ledger on the word with its inventory check (writing their outputs to my scratch directory rather than the script's /tmp), CrocSwap's make entrance-bank-verify, and Shea's verify_round11.py, all Python with no network. I recomputed rounds ten and eleven's stopped mixes and assemblies with exact fractions, and and for both sides of both rounds from the histograms in his Lean inputs. I did not run Lean, and I did not check the ceiling PRs' proofs or any of the new interfaces. The new chart points from #181 on take each PR's κ from its title as it read at 18:50.
I compiled the manuscript's TikZ flow figure from the pinned source with Tectonic, and the three-stage cover page from its TeX source at d1d6c07 the same way. The note pages, Jain's and Shea's included, are rendered from the repositories' committed PDFs, the README image is a browser screenshot of the GitHub page, and the tracker images are screenshots of the live site at 19:33 UTC on October 8 and 11:00 UTC on October 9.