# The integer-multiplication exponent race: from 2^-182 to 2^-14.6 in forty hours

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/integer-mult-exponent-race
> date: 2026-10-08
> tags: math, theory, algorithms, agentic-coding, verifiers

The site owner sent me a link to [CrocSwap/integer-mult-bounds](https://github.com/CrocSwap/integer-mult-bounds) with the note "something crazy going on in this repo". When I opened it, the headline said $\kappa = 2^{-30}$. By the time I had read the pull requests, the best open claim was $\kappa \approx 4.1\times10^{-5}$, a little above $2^{-15}$. When I started writing this paragraph, PR #45 had just arrived. By the time I finished the article, #48 had nudged the claim to $4.1186\times10^{-5}$, and Swapnil Jain, racing in a separate repository, had reached $3.67\times10^{-5}$ with its arithmetic checked in Lean.

Then, at 15:18 UTC, the story changed shape. Colkitt merged Rohan Arun's PR #39 into `main`, with $\kappa = 971668963/(2.5\cdot10^{13}) \approx 3.8867\times10^{-5}$, after an audit he published alongside it. For the first time a community number is the maintainer's number. It is still conditional on OpenAI's unrefereed manuscript, and the audit was Colkitt working with Codex, not a referee; [the section on the merge](#pr-39-reaches-main) goes through exactly what it checked. The open PRs had already moved past it: Rohan's own #49, under twenty minutes later, claims $4.1239\times10^{-5}$.

And it kept going after I thought I had finished. By 18:35 UTC `main` had moved twice more and stood at $5.1017\times10^{-5}$, after a new idea broke the plateau I describe below; the open PRs reached $5.1739\times10^{-5}$ an hour later. Aurel Prosz, one of the contributors, built a live dashboard of the whole thing. And Ryan Shea took Jain's network across to a second OpenAI problem, the exact discrete Fourier transform, and claimed a saving 730 million times OpenAI's. [The evening](#main-moves-again) and [the Fourier transfer](#the-race-reaches-the-fourier-transform) have their own sections near the end.

Then, overnight, the plateau broke for good. Between 23:12 UTC on 8 October and 05:06 the next morning, icekylinx's PRs replaced the shape of the network itself, and `main` jumped ninefold to $4.609169\times10^{-4}$, about $2^{-11.08}$. Jain followed twenty minutes later at $4.6637\times10^{-4}$, by building on the same cover, and the open PRs were at $6.5592\times10^{-4}$ when I stopped, at 11:00 UTC on 9 October. [The ninefold night](#9-october-the-ninefold-night) explains what changed, and why κ is still nowhere near 1.

The afternoon was quieter in ratio and busier in every other way. At 12:57 UTC Colkitt moved `main` again, to $6.61886\times10^{-4}$. Ten minutes later Jain posted round eleven at $6.6857\times10^{-4}$, which led every claim anywhere for nineteen minutes, and said he needed an arXiv endorser for a paper collecting the proofs. By 18:50 UTC the best open claim was $6.84697\times10^{-4}$, and several contributors had stopped chasing κ to prove how far the current designs can go at all. [The afternoon](#9-october-afternoon-reusing-dead-registers) covers all of that.

Some background. Two days ago OpenAI's math release included a 73-page manuscript, [Integer multiplication below n log n](https://github.com/openai/math/tree/adc7f1241b42e322a6451854ab7e4b4c146bf78a/preprints/Integer-multiplication-below-n-log-n-September-23-2026) (result family 109), which claims a multitape Turing machine that multiplies two $n$-bit integers in $O(n(\log n)^{1-\kappa})$ steps with $\kappa = 2^{-182}$. I covered it briefly in [the openai/math roundup](/articles/openai-math#integer-multiplication-below-n-log-n), and it has a tile on the [AI breakthroughs in mathematics wall](/math#109). Doug Colkitt ([@0xdoug](https://x.com/0xdoug), the founder of CrocSwap/Ambient) forked the argument into a "research draft" and started tightening the constants. Then other people joined in. Counting $2^{-182}$ as the start, the exponent saving has grown by a factor of about $2^{167}$ in forty hours.

That sounds like a story about multiplication getting faster. It isn't one, and the reason is the most interesting part. What follows covers what $\kappa$ is, where it comes from in the construction, what each step of the race actually changed, what I could re-run myself, and why the very funny "linear time by tomorrow" chart that went round on X is wrong in an instructive way.

<RepoCard repo="CrocSwap/integer-mult-bounds" />

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig1-readme.png" alt="The top of the repository README: 'Research draft by Douglas Colkitt, conditional on the upstream manuscript and the written extensions supplied here.' It states T(n) = O(n (log n)^(1 minus kappa)) with kappa = 2^-30, approximately 9.31323 times 10^-10, in the fixed finite-alphabet Turing-machine model, and says this is a 2^152 ratio of exponent savings, not a runtime speedup. A related-contributions paragraph credits Zhihao Chen's earlier PR #7." caption="The README's own framing of the 2^-30 checkpoint, including its 'not a runtime speedup' line and the credit to PR #7. GitHub's KaTeX shows the manuscript's \! spacing as a literal '!' (CrocSwap/integer-mult-bounds README at 1a74950, screenshot)." />

## What $\kappa$ measures

Schönhage and Strassen multiplied $n$-bit integers in $O(n\log n\log\log n)$ steps in 1971 and guessed that $n\log n$ was the true answer. Fürer got within a $2^{O(\log^* n)}$ factor of it in 2007. Harvey and van der Hoeven reached $O(n\log n)$ exactly ([Annals of Mathematics, 2021](https://hal.science/hal-02070778)), and most people took that to be the end.

The OpenAI manuscript claims a bound of

$$
T(n) = O\!\left(n(\log n)^{1-\kappa}\right) = O\!\left(\frac{n\log n}{(\log n)^{\kappa}}\right).
$$

So $\kappa$ is the power of $\log n$ you get to divide out. At $\kappa = 0$ you have Harvey–van der Hoeven. At $\kappa = 1$ you would have linear time, which is the floor, since just reading the input takes $n$ steps. Every value in between is a strictly faster growth rate than $n\log n$. For the question "is $n\log n$ optimal?", $2^{-182}$ is as good an answer as $1/2$: any $\kappa > 0$ refutes the conjecture.

The machine model matters. The theorem is about a deterministic Turing machine with a fixed finite alphabet and a fixed number of one-dimensional tapes, the model in which Schönhage and Strassen stated their conjecture. On a tape, moving data costs time proportional to the distance it travels, so even rearranging $\Theta(n)$ bits at each of $\Theta(\log n)$ FFT levels costs $\Theta(n\log n)$. The manuscript says this outright: "Scanning an array of $\Theta(n)$ stored bits at each of $\Theta(\log n)$ transform levels already costs $\Theta(n\log n)$. We therefore need savings in both movement and arithmetic" (`upstream/build/sections/00-introduction.tex`).

Nobody had ever proved an $\Omega(n\log n)$ lower bound in this model. The best result is conditional: Afshani, Freksen, Kamma and Larsen ([2019](https://arxiv.org/abs/1902.10935)) showed that a circuit lower bound would follow from the network-coding conjecture, and a time-$T$ Turing machine only gives circuits of size about $T\log T$, so even that does not transfer. The [FFT section of the roundup](/articles/openai-math#what-was-actually-open) goes through the same gap for the Fourier transform, where the release's family 130 reuses this manuscript's finite gadget. The point to keep in mind is that $n\log n$ was a barrier nobody knew how to cross, not a proven wall.

## Where $\kappa$ comes from

The pipeline is Harvey–van der Hoeven with every change of representation charged on tape: digits go onto prime cyclic axes by a Chinese-remainder map, Gaussian resampling turns those into power-of-two axes, and a phase twist turns the last axis into polynomials modulo $y^r+1$, where multiplying by a root of unity is just a signed shift.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig4-multiplication-flow.png" alt="Flow diagram with four boxes stacked vertically: integer product as radix-digit convolution; cyclic convolution on distinct prime axes (via a Chinese-remainder address map); normalized convolution on power-of-two axes (via Gaussian resampling and chirps); convolution over C[y]/(y^r + 1) with synthetic transforms and packed polynomial products (via a last-coordinate phase twist)." caption="The reductions from the integer product down to ring convolutions. I compiled the manuscript's own TikZ source for this; the arrows are changes of representation, each of which is charged on tape (Integer multiplication below n log n, Figure 1, from upstream/build/figures/multiplication-flow.tex)." />

The new part is two tape procedures built from fixed linear networks. One swaps two address fields of an array; the other applies a layer of butterflies $(u+v)/2,\ (u-v)/2$ to selected coordinate bits. Each network is a finite circuit on $W$ "roles" (wires), it exchanges two banks of values while restoring every scratch value, and each edge of the circuit is carried out by smaller recursive calls of the same procedure. The decisive count is the total number $s$ of those smaller calls. If $s < Wm$, then multiplying the problem parameter by $m$ multiplies the normalized cost by $s/W < m$, and the recursion runs at $m^\tau$ with $\tau<1$. The manuscript: "This strict inequality is the source of the power saving."

How much below $m$ it lands is tiny. In the original construction the relative deficits are (`upstream/build/sections/03-motifs.tex:673-678`)

$$
\eta_{\rm b} = \frac{339}{22587335000000},\qquad \eta_{\rm c} = \frac{73}{19906842167500},
$$

at $m = 10^6$, so the honest savings are about $10^{-12}$ and $2.7\times10^{-13}$. The manuscript then rounds them down to $1-\tau = 1-\sigma = 2^{-50}$ and calls its constants "deliberately conservative; no optimization is claimed". Remember that sentence. A lot of what happened next is people taking it at its word.

Those network savings feed a cost table. With $p \approx \log n$ the working precision, the paper divides every operation's cost by the data volume and gets a power of $p$. Seven rows dominate, and each has a margin $g_i$ below 1:

$$
\begin{aligned}
g_1 &= 1-\epsilon(1+c), & g_2 &= \epsilon c(1-\tau), & g_3 &= \epsilon(1-\lambda'),\\
g_4 &= 1-\tau-\epsilon(2-\tau), & g_5 &= \tfrac14-\delta-\tfrac32\epsilon, & g_6 &= 1-\delta-\epsilon,\quad g_7=\epsilon.
\end{aligned}
$$

The total time is $n$ times $p$ to the largest power, so the exponent saving is the smallest margin, minus a sliver of slack to absorb stray $\log p$ factors. The upstream parameters are $\epsilon = 2^{-75}$, $c = 2^{-56}$ and $1-\tau = 2^{-50}$, so $g_2 = 2^{-75}\cdot2^{-56}\cdot2^{-50} = 2^{-181}$, and the paper says "Thus $\min_i g_i=2^{-181}=2\kappa$" (`08-assembly.tex:866`). That is where $2^{-182}$ comes from: three small numbers multiplied together, then halved.

The repository's checker transcribes the same seven margins as exact fractions (`scripts/certify.py:136-147`):

```python
def margins(p, *, layout_model="adjacent", assembly_model="original"):
    require(assembly_model in ("original", "tight-gaussian"), "Unknown assembly model")
    e, c, t = p.epsilon, p.c, p.tau
    return {
        "g1": 1-e*(1+c),
        "g2": e*c*(1-t),
        "g3": e*(1-p.lamp),
        "g4": 1-t-e*(layout_degree(layout_model)-t),
        "g5": Q(1, 4)-p.delta-(Q(5, 4) if assembly_model == "tight-gaussian" else Q(3, 2))*e,
        "g6": 1-p.delta-e,
        "g7": e,
    }
```

Once you see $\kappa$ as a product of factors, the race is easy to follow. Every improvement attacks one of the factors: it makes the network saving $1-\tau$ bigger (better circuits), lets $\epsilon$ grow (better routing, better Gaussian resampling), lets $c$ grow (cheaper control movement), or changes the recurrence so the same circuit counts for more (batching). The widget below shows the seven margins on a log scale for the original paper, Colkitt's $2^{-30}$ checkpoint and PR #13. In the original, $g_2$ sits at $2^{-181}$ and nothing else comes close. By PR #13, $g_2$, $g_3$ and $g_4$ are all balanced at about $2^{-20.3}$, which is what a tuned parameter set looks like.

<MarginLedger />

There is also a ceiling built into this shape. Since $g_7 = \epsilon$ and $g_1 = 1-\epsilon(1+c)$, the smaller of the two is at most $1/2$ whatever you choose, and $g_2$ is at most the bit network's saving $1-\tau$. In every witness I checked, $\kappa$ sits just under that network saving: $9.3\times10^{-10}$ against $4.67\times10^{-9}$ at the $2^{-30}$ checkpoint, $7.699\times10^{-7}$ against $1.54\times10^{-6}$ in PR #13, and $4.1006\times10^{-5}$ against $4.1008\times10^{-5}$ in PR #44. To get anywhere near linear time this method would need finite networks whose recursive calls nearly vanish, and nothing here suggests that is possible. The overnight witnesses of 9 October keep the same shape: in PR #144's assembly two margins are balanced against each other so that $\kappa \approx a/(1+2a)$, where $a$ is the weaker network's saving, and that saving is in turn about the network's relative rank deficit divided by how far below full width its average child sits on a log scale. The section on [the ninefold night](#why-smaller-is-worth-nine-times) works through the numbers.

## Colkitt's day and a half

Colkitt's own checkpoints are on `main` (the last one is also on `release/ternary-30`). Each came with an X post from [@0xdoug](https://x.com/0xdoug), and each post's ratios check out against the exact witnesses.

| commit (UTC) | $\kappa$ | what changed |
|---|---|---|
| `5cf29ec` Oct 7 02:22 | $5.8\times10^{-33}$, between $2^{-108}$ and $2^{-107}$ | parameters only: same network, sharper recurrence comparison |
| `52ce3be` Oct 7 13:35 | $2^{-78}$ | direct nonadjacent axis swaps: layout from $O(d^2)$ to $O(d)$ |
| `bcd4ebd` Oct 7 19:27 | $2^{-59}$ | paired circuits with shared sums, tighter guard and Gaussian width |
| `6e56487` Oct 8 00:27 | $83/10^{12} > 2^{-34}$ | compact dirty controls instead of spaced windows |
| `1c09a58` Oct 8 01:10 | $2^{-31}$ | compressed complex network, binary phase frames |
| `1a74950` Oct 8 13:10 | $2^{-30}$ | ternary five-subset bit circuit over $\mathbb F_3$ |

The first note is the most honest thing in the repository. With the paper's cost accounting left alone, it finds $\kappa = 29/(5\cdot10^{33})$ and also proves a ceiling of $5.838\times10^{-33}$ for that family, so the witness is within 1% of the best that accounting allows. Colkitt's post said so: "Surpassing this ceiling would require improving the network bounds or cost analysis from the original result."

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig5-parameter-note.png" alt="First page of 'A sharper exponent for integer multiplication' by Douglas Colkitt, October 6, 2026. The abstract gives kappa = 29/(5 times 10^33) = 5.8 times 10^-33, more than 99% of a certified upper bound for the stated parameter family. Section 1 lists the starting parameters tau = sigma = 1 - 2^-50, beta = 1/2, delta = 1/16, C1 = 20, the layer constraints, and the seven margins g1 to g7." caption="The first checkpoint is pure parameter tuning, and the note lists the same seven margins the whole race is about. The footnote discloses OpenAI Codex assistance (Colkitt, parameter-note.pdf, page 1)." />

So the next morning he changed the accounting. The original layout row costs $d^2(1+\ell^\tau)$ because it reverses axes with adjacent swaps; the manuscript already allowed nonadjacent swaps, and using them directly makes it $d(1+\ell^\tau)$. That changes $g_4$ to $(1-\tau)(1-\epsilon)$, which no longer forces $\epsilon$ to be smaller than the network saving. The result-history page explains the consequence: the original layout "forced epsilon &lt; a, yielding a cubic constraint kappa &lt; a^3", and the new schedule makes the saving scale "quadratically in a" (`docs/research/result-history.md:253-256`). Going from $a^3$ to $a^2$ when $a \approx 10^{-11}$ is how you gain thirty-odd powers of two in one commit.

The $2^{-59}$ step improved the network itself, which raises $a$: rectangle incidence circuits, then shared intermediate sums, cut the side roles per invocation from 41,122,620 to 2,394,438 and then to 577,576 at $h=46$, and paired-block circuits at $h=50$ reach 509,194. Fewer roles for the same rank deficit means a larger relative deficit, so a larger $1-\tau$.

The compact-control step at $2^{-34}$ attacked $c$. Moving whole spaced windows had a spacing penalty $K^\tau$ that forced $c$ to be of order $a$; moving only compact control fields removes it, so $c=1$ and $\kappa$ becomes roughly $\epsilon$ times $a$. His post that night put it in two sentences: "The improvement came from removing the spacing penalty behind the quadratic bottleneck. This was done by moving compact control bits instead of entire windows." Its "roughly 48 million fold improvement over the previous result" checks out against $2^{-59}$ (about $4.8\times10^{7}$), and so does "a 2¹⁴⁸ fold improvement over original OAI result" ($2^{148.5}$). Here the README discloses that the idea did not come from Colkitt: "The compact-control proposal originated with a separate research agent; the supplied note develops its tape, layout, repair and assembly arguments" (`README.md:138-140`).

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig8-compact-control.png" alt="First page of 'Compact dirty controls and a conditional integer-multiplication saving' by Douglas Colkitt, October 7, 2026. The abstract says full-slot permutations are replaced by swaps of compact dirty-control fields, the row-addition overhead becomes O(V((f log p)^tau + 1)) with no spacing factor K^tau, and the conditional witness is kappa = 83/10^12 > 2^-34. Section 1 says the compact-control idea was supplied by an independent research agent." caption="The 2^-34 checkpoint. Its scope paragraph credits 'an independent research agent' for the idea and says the general claims rest on written arguments, not simulation or formal verification (Colkitt, compact-control-note.pdf, page 1)." />

The last two checkpoints moved the bottleneck between the two networks. At $2^{-31}$ a compressed complex network ($h_c = 26$, $a_c = 5/10^9$) overtook the bit network's $296/10^{11}$, so the bit side became binding. The $2^{-30}$ checkpoint then replaced the bit circuit with one that computes over $\mathbb F_3$.

The $\mathbb F_3$ trick is pretty enough to show. Label the circuit's wires by the five-element subsets of 29 points, $n = \binom{29}{5} = 118755$ of them, and let $B$ be the incidence matrix of five-subsets against pairs. Then $(BB^{\mathsf T})_{ST} = \binom{|S\cap T|}{2} \bmod 3$, and for intersection sizes $0,\dots,5$ those residues are $0,0,1,0,0,1$. So $BB^{\mathsf T} = I + A_2$, where $A_2$ is the adjacency of subsets meeting in exactly two points, and the "side" map is just $-A_2$. A central factor of size $r = \binom{29}{2} = 406$ plus a correction that only touches intersection-two neighbours gives exactly the identity. I checked the residue identity by brute force over all 126 five-subsets of nine points; it holds.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig6-f3-identity.png" alt="Section 2.1 of the ternary note: with h = 29, n = binom(29,5) = 118755 and r = binom(29,2) = 406, B is the incidence matrix of five-subsets against pairs over F3, C = B B^T with C_ST = binom(|S intersect T|, 2) mod 3, the binomial residues for t = 0 to 5 are 0,0,1,0,0,1, so C = I + A2 and the side map is -A2. The rational label space uses the form H = I - (2/25) J." caption="The mod-3 identity behind the ternary circuit: pair incidence squares to the identity plus the intersection-two adjacency (Colkitt, ternary-note.pdf, Section 2.1)." />

The full circuit has 19,593,239 active additions and 20,780,789 side roles per invocation, and the exact count gives $s/(Wm)$ with a relative deficit of $553/11717905300$. I recomputed the implied saving from the certificate's $W$, $s$ and $m$: $-\log_m(s/Wm) \approx 4.672\times10^{-9}$, just above the certified $a_{\rm b} = 467/10^{11}$. The final minimum margin is $2332833/2500000000000000$, which exceeds $2^{-30}$ by about 0.194% of the claimed saving.

That checkpoint was not first, and its README says so: "Zhihao Chen's earlier PR #7 introduces the same ternary five-subset motif with a different circuit and stronger claimed bound. This release records a separate implementation and conditional checkpoint; it makes no priority or strongest-known-bound claim."

## Everyone else shows up

The first outside PR arrived at 20:58 UTC on October 7. Between 02:23 and 14:24 UTC on October 8, 42 more arrived, from about ten people. By 11:00 UTC on October 9 there were 179. The chart has all of them that state a κ: Colkitt's checkpoints in red, OpenAI's starting point in grey, the community PRs coloured by author, a green step line for the best claim open at each moment, and a red step line with a square at each move of `main` (#39 at 15:18 UTC, #49 at 16:47, a reviewed batch at 18:35, and icekylinx's #144 at 05:06 the next morning). Everything after $2^{-17}$ is squeezed into a sliver at the top, so the zoom button redraws everything from 10:00 UTC on October 8 on its own. Jain's separate track is the purple diamonds, with a ring on the rounds whose arithmetic is checked in Lean; more on that [below](#a-second-race-one-repository-over). Click a point (or use the dropdown) for its title and link.

<KappaRace />

A few PRs changed the shape of the problem rather than its constants, and those are the ones worth understanding.

Start with eumemic's #3 and #5. #3 compressed the complex network, cutting its side wiring from 3,693,800 roles per invocation to 108,195 and taking its saving from $418/10^{12}$ to $14/10^9$. That made the bit network binding, and the next limit was the Gaussian resampling row: its cost $O(tp^{3/2+\delta}\alpha)$ forced $\epsilon < 1/5$ and capped $\kappa$ below $a_{\rm b}/5$. #5 rewrote the resampling as chirped correlations evaluated by the existing $O(N\log N)$ multiplier and bounded the Neumann series more tightly, which let $\epsilon$ approach $1/2$. That one change of analysis was worth a factor of about 2.5 on its own.

Zhihao Chen's #7 is the $\mathbb F_3$ five-subset construction, at $h=28$, posted at 06:12 UTC, seven hours before Colkitt's independent version. Chen's PR credits "GPT-6 Astra (OpenAI Codex)" for "research, proof-development and implementation assistance" and asks that future work "explicitly acknowledge Zhihao Chen (jacklightChen), cite this PR and its accompanying note". The maintainer did.

icekylinx's #10 is the structural break you can see on the chart: from $2^{-28}$ to $2^{-23}$ in one PR. Upstream, every edge of the network costs singleton recursive calls, one per unit of rank. #10 shows that a contiguous block of $t$ address fields from a projector can be done as one recursive call of width $t$, given a common controlled basis. The recurrence condition stops being a count and becomes a moment: with rank-mass weights $w_i$ on calls of relative width $r_i$, you need $\sum_i w_i r_i^{-a} < 1$. A singleton call has $r = 1/m$ and costs the full $m^a$; in #13's profile, a block covering 195/196 of the fields costs almost nothing extra. The same PR fixes an error-enclosure gap it found in #5's fast Gaussian argument, with no change to the exponent.

eumemic's #13, the one the site owner pointed at, is a neat idea. The network's correctness condition only needs each role's endpoint frames to differ by the identity: $M_{\rm sink} - M_{\rm source} = I$. A scratch role does not have to start in the zero frame. It can start in any fixed frame $M$ as long as it ends in $I+M$. Starting every stage-two auxiliary role in the frame of its first gate deletes its rank-756 entrance edge and turns its exit into one projector of rank $m-h$. That projector compiles to $h$ singleton pivots plus one contiguous block of $m-2h$ fields. The total rank is unchanged. What changes is where it sits: out of singleton calls costing $\ln m$ each and into one block. Under #10's moment, the interchange saving goes from $246/10^9$ to $154/10^8$, and with a third whole-residual class on the complex side, $\kappa$ goes from $6149999/(5\cdot10^{13})$ to $7699/10^{10}$, a factor of 6.26. Since $\log_2(7.699\times10^{-7}) \approx -20.31$, the claim is about $2^{-20.3}$; the title's "$> 2^{-21}$" undersells it.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig7-source-frames.png" alt="Title and abstract of 'Auxiliary Source Frames for Integer Multiplication: Concentrating auxiliary rank into contiguous recursive blocks' by eumemic, with substantial assistance from Claude (Anthropic), 8 October 2026. The abstract explains that starting every stage-two auxiliary role in the frame of its first gate removes its rank-(h^2 - h) entrance edge and turns its exit into one projector of rank m - h, raising the interchange saving from 246/10^9 to 154/10^8, and gives kappa = 7699/10^10 > 2^-21." caption="PR #13's note. Its 26 pages reproduce #10's framework verbatim and then add the source-frame sections (eumemic, source-frame-21-note.pdf, page 1, PR #13 at 3ef246f)." />

After #13 the tempo picks up. #15 brought a smaller $h=30$ producer, #18 partial swaps, and #21 and #23 translated frames and "semantic precision", reaching $1099/10^8 > 2^{-17}$. #29 introduced a two-stage topology from Aurel Prosz's work and reached $15536/10^9 > 2^{-16}$. #36 added copied retained centers, which halve one rank term ($Wm-N+2L$ becomes $Wm-N+L$) and jump to $3.84569\times10^{-5}$. Everything after that is composition and tuning: fixed local bases, reversed corners, carrier matchings. PR #44 claims $2050314627/(5\cdot10^{13}) \approx 4.1006\times10^{-5}$, which it describes as "approximately 0.02681% above #43". The last three PRs are within 0.03% of each other. Three more arrived while I was writing: chafreaky's #46 adds 0.0000048% to #44, Rohan Arun's #47 adds 0.109% with hill-climbed producers, and chafreaky's #48 reorders the leave-one-out sums to save 847,550 wires and reach $411862541/10^{13} \approx 4.1186\times10^{-5}$. Then Rohan's #49, at 15:36, started from #48's orderings and hill-climbed them the way #47 had, for $4123863984/10^{14} \approx 4.1239\times10^{-5}$, 0.127% more. The step line flattens around $2^{-14.6}$ after 12:30 UTC.

When I first wrote this section I called that flattening the most informative thing on the chart, and said it was what you would expect if the cheap structural moves inside this family of constructions had mostly been made. Three hours later Avi Eisenberg's #53 jumped 9.2% in one PR, and the step line climbed again to $2^{-14.24}$. Use the zoom to see both plateaus; [the evening section](#main-moves-again) explains the move that broke the first one. I'd still read the flat stretches the same way, only with less confidence about how long any of them lasts.

## Re-running PR #13

The brief asked me to check the claims, and these are certificate checks I could actually run: pure Python with `fractions`, plus a small C++ helper on the PR branches, no network. On a shared 16-core machine at `nice 19`:

On `main` at `1a74950`, the focused verification from the README. `scripts/audit_ternary_side.py` builds the full $h=29$ DAG (about 21 million addition nodes) and finished with `PASS ternary side construction and conditional 2^-30 assembly.` in 176 s at 1.6 GB peak memory. `make_ternary_patch.py` took 17 s, the 12 focused tests passed in 39 s, the patch applied to the pinned upstream source, and the regenerated certificate and patch were byte-identical to the committed ones.

On PR #13 at `3ef246f`, `scripts/source_frame_network.py` runs in 0.1 s and prints the witness:

```text
PASS kappa=7699/10000000000 > 2^-21
bit saving=77/50000000; complex saving=9/5000000
bit moment gap=4.69651719166272e-10
complex moment gap=3.899666592614166e-09
minimum margin=3849992299500001/5000000000000000000000; gap=492299500001/5000000000000000000000
```

The regenerated `certificates/source-frame-network.json` compared equal to the committed one. The six tests in `tests/test_source_frame_network.py`, including an exact $h=5$ model in which the new exit is checked to be an idempotent of rank $m-h$ with exactly $h$ corner pivots, passed in 20 s. The full `make verify` on the PR branch took 432 s: 172 tests passed, all 19 manuscript patch checks applied, and the certificates and patches regenerated without a diff. Those numbers match what the PR description claims.

Then I recomputed the arithmetic myself, without the repository's helpers. From the certificate's parameters ($\epsilon = 499999/10^6$, $c=1$, $1-\tau = 154/10^8$, $1-\lambda' \approx 1.54\times10^{-6}$, $\delta = 10^{-10}$) the margins are $g_2 = 38499923/(5\cdot10^{13})$, $g_3 = 3849992299500001/(5\cdot10^{21})$ and $g_4 = 38500077/(5\cdot10^{13})$, all about $7.70\times10^{-7}$; $g_1$ and $g_5$ are about $2\times10^{-6}$, and $g_6$ and $g_7$ about $1/2$. The minimum is $g_3$, which beats $\kappa = 7699/10^{10}$ by $492299500001/(5\cdot10^{21}) \approx 9.8\times10^{-11}$. So the claimed $\kappa$ does follow from the margins.

The bit-moment check is the step that carries the new idea, so I redid it at 80-digit precision rather than with the script's $e^x \le 1/(1-x)$ bound. Here is how the script does it (`scripts/source_frame_network.py:54-62` on PR #13):

```python
def moment(weights, ratios, a):
    """Upper bound for sum_i w_i ratio_i^(-a), using exp(x) <= 1/(1-x)."""
    logs = []
    for r in ratios:
        simple = SIMPLE_LOGS[r]
        require(log_upper(1/r) < simple, 'Rational logarithm bound failed')
        require(0 <= a*simple < 1, 'Exponential upper-bound range failed')
        logs.append(simple)
    return sum((w/(1-a*ell) for w, ell in zip(weights, logs)), Q(0)), logs
```

The four classes have weights of about 0.0044 (singletons, relative width $1/21952$), 0.4934 (width $195/196$), 0.0076 ($10165/10976$) and 0.4946 ($391/392$). With $a = 154/10^8$ the moment is $1 - 4.78\times10^{-10}$, below 1 as required, and bisection puts the largest saving this rank profile supports at about $1.5499\times10^{-6}$. PR #13's $1.54\times10^{-6}$ leaves about 0.6% headroom.

None of this checks the hard part. The certificates confirm that the counts, ranks and inequalities are what the notes say they are. They do not confirm that a contiguous projector block really can be executed as one recursive call on a fixed-tape machine, that the controlled basis exists at full size (the review guide for #10 says it is "established by the existence argument, not by materializing every h28 factorization"), or that the upstream theorem is true. Colkitt's own review index was clear, that morning, that this was still open: "These changes appear to cross the singleton-rank recurrence limitation; the headline alone does not establish that step." By the afternoon his audit had worked through those obligations for the one chain he merged, #39's ([below](#pr-39-reaches-main)). For #13's own construction and every PR outside that chain, the sentence still stands.

## A second race, one repository over

While those PRs were landing, Swapnil Jain ([@SJ_Swapnil_Jain](https://x.com/SJ_Swapnil_Jain)) was running a race of his own. He never opened a PR on Colkitt's repository; none of the 13 accounts that did is his. He kept [Swapnil-jain/integer-mult-kappa](https://github.com/Swapnil-jain/integer-mult-kappa), posted six numbered updates on X between 07:20 and 14:03 UTC, and the sixth ended with a sentence nobody on the CrocSwap side could write about their frontier: "Both interchange certificates and the full assembly are checked in Lean's kernel." I cloned it and read all of it, and I wanted to know two things. Is this really a separate result, and what exactly did Lean check?

<RepoCard repo="Swapnil-jain/integer-mult-kappa" />

I lined the six posts up against what was open on CrocSwap at the same minute. The witnesses are the exact rationals from the commit each post announced; they are the purple diamonds on the chart above.

| post (UTC) | Jain's $\kappa$ | best open CrocSwap claim then | what was new |
|---|---|---|---|
| 07:20 | $249/(5\cdot10^{10}) \approx 4.98\times10^{-9}$ | #7, $3.73\times10^{-9}$ | the "ε → 1 stack" |
| 08:15 | $7499/10^{12} \approx 7.5\times10^{-9}$ | #9, $3.8\times10^{-9}$ | Gaussian correction inverted block by block |
| 10:05 | $\approx 4.188\times10^{-6}$ | #18, $1.885\times10^{-6}$ | batching on a two-stage bit network |
| 11:10 | $\approx 1.1972\times10^{-5}$ | #28, $1.226\times10^{-5}$ | PR #24's bit network, complex source frames |
| 12:39 | $\approx 1.5479\times10^{-5}$ | #37, $3.851\times10^{-5}$ | his own bit network on one flag basis; first Lean file |
| 14:03 | $3666565558019/10^{17} \approx 3.6666\times10^{-5}$ | #43, $4.0995\times10^{-5}$ | copied centres; both savings and the assembly in Lean |

So he held the lead twice: for about an hour after his first post, until icekylinx's #10 at 08:26, and for nineteen minutes after his third, until Zhihao Chen's #21 at 10:24. His fourth post went out one second after #28 was opened. From round five on, the CrocSwap side was ahead, by 2.5 times at 12:39 and by about 12% at 14:03. His own ratios check out: round six is 2.37 times round five, and $2^{-182}$ to $3.67\times10^{-5}$ is a factor of $2^{167.3}$. The one soft spot is the first post's comparison with Colkitt's $8.3\times10^{-11}$ ("a roughly 60 fold improvement"): by then Colkitt's `main` was already at $2^{-31}$, so against the newest checkpoint it was about ten.

"Independent" is the wrong word for it, and Jain doesn't use it. His NOTICE pins Colkitt's `6e56487` and cites fourteen open CrocSwap PRs (#3 through #36) plus Aurel Prosz's fork. Round four reads PR #24's published network counts from a JSON file pinned by SHA-256. Round six's copied centres are, in the first sentence of his note, "PR #36's copied retained-centre schedule (icekylinx)" applied to his network. The traffic went the other way as well. Chen's #29 and Rohan Arun's #31 credit "Swapnil Jain for the linked two-stage batching development", icekylinx's #36 links Jain's round-three commit with the line "no source code imported from this repository", and #44 and #47 list him with everyone else. These are two forks of the same argument, reading each other's work all morning.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig9-jain-copied-centres.png" alt="The copied-centres note. It applies PR #36's copied retained-centre schedule by icekylinx to Jain's two-stage bit interchange. In each stage a centre c runs the word RGRG, so its frame path is D0 to D1 to D0 to D1, which makes the rank budget s = Wm − N + 2L. Lemma 1: replace the second scatter's reads of c by reads of a temporary copy t, moved to D0 and erased afterwards, while the original stays at D1; each centre then has two children of width h per stage instead of three, and the budget becomes s = Wm − N + L." caption="The round-six idea in one lemma, credited to PR #36 in its first sentence. The page says 'Draft, October 9' although the commit is from 14:01 UTC on October 8; the date is typed into the TeX source (Swapnil Jain, copied-centres.pdf, page 1, integer-mult-kappa at f2176bc)." />

What is his is the part of the stack that turns a network saving into $\kappa$. Through PR #13, $\kappa$ sat at about half the bit network's saving: $7.699\times10^{-7}$ against $1.54\times10^{-6}$. His README explains why. A few costs tie the saving either to the number of axes $d$ or to the axis width $\ell$, with $d\ell \approx \log n$, so the saving gets split between them, and the Gaussian resampling's precision budget, about $8d^2$, has to fit inside $O(\log n)$, which keeps $\epsilon$ below $1/2$. His first round made five changes at once. The coefficients become $(\log n)^{1+x}$ bits long instead of about $\log n$, so precision no longer competes with address length. The Gaussian correction is inverted block by block (every block is "the same Toeplitz matrix up to a geometric rescaling, so one stored inverse serves the whole line", as the second post puts it). Each axis becomes one chunk, only the low bits of an axis are moved for resampling, and the axis reversal reuses Colkitt's selected-bit additions. With all of that, $\epsilon$ can go almost to 1, and that is the "ε → 1" in every post.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig10-jain-stack-notes.png" alt="First page of 'Layout and precision changes beyond the a_b/2 ceiling', draft of October 8, 2026. A conditional-status paragraph lists the unmerged CrocSwap results it assumes: PR #5's chirped-correlation Gaussian lemma, PR #3's compressed complex network and PR #7's networks. A paragraph headed 'Deliberate departures from upstream hypotheses' replaces the prime-interval condition by the prime number theorem, the resampling hypothesis by alpha squared theta at least 1, the band r at least 2^(a_r p/d) by r at least 2^(p^mu), drops K = o(l), and uses a certified guard bound in place of 5." caption="The stack's own scope statement. The second paragraph matters: this track changes several of the manuscript's hypotheses, which the CrocSwap PRs do not (Swapnil Jain, stack-notes.pdf, page 1)." />

I recomputed the round-six witness from the numbers in his certificate, without his code. The bit saving is $a_{\rm b} = 36667/10^9$, the complex saving $a_{\rm c} = 36926111/(5\cdot10^{11})$, so $\tau = 1-a_{\rm b}$ and $\lambda' = \tau + 2\cdot10^{-16}$. He sets $\epsilon = 2499908335861049/(2.5\cdot10^{15}) \approx 0.9999633$. Of his five cost margins, $1-\epsilon$, $1-\epsilon$ minus a $10^{-22}$-sized term, and $\max(\epsilon,1-\epsilon)(1-\tau)$ all come to $3.6665655580\times10^{-5}$ and a bit; $\epsilon(1-\lambda') \approx 3.666565558020684\times10^{-5}$ is the smallest; and $\epsilon$ itself is about 1. The claimed $\kappa = 3666565558019/10^{17}$ sits $1.7\times10^{-17}$ under the binding margin. The tidy way to say it: $\epsilon$ is chosen so that $1-\epsilon = \epsilon\,a_{\rm b}$, which makes $\kappa = a_{\rm b}/(1+a_{\rm b})$ to eleven digits. Every bit of the network's saving reaches $\kappa$.

The CrocSwap side got to the same place by another road. RaD/hipotures' #20 and Chen's #23, at 10:20 and 10:43 UTC, used "completed-child semantic precision" and bulk resampling, and neither cites Jain. Round four is a neat accidental experiment. Jain and Dominik Scholz's #27 took the same PR #24 bit network, with $a_{\rm b} \approx 1.19724\times10^{-5}$, put it through their two different assemblies, and got $1.19722292\times10^{-5}$ and $1.19720853\times10^{-5}$. Both stacks now pass essentially the whole network saving through, so the 12% gap between 3.67 and 4.12 is entirely the bit network: $36667/10^9$ here against #48's $82375901/(2\cdot10^{12})$.

The ε → 1 stack also has a price that the headline hides. With $\epsilon$ near 1, the coefficient-depth guard needs $\epsilon C_1 < 1+x$. Round six uses the crude guard $C_1 = 14694$ (from $2 + 576\ln 119453132304 \approx 14693.6$), so $x = 14695$ and each coefficient is $(\log n)^{14696}$ bits long. In round five the same rule gave $x = 842135$. That is legal, because $x$ is fixed and the volume stays $\Theta(n)$, and his certificate says as much ("The exponent is huge but fixed, so the O() statement holds", `scripts/certificate_round3.py:15`). It is one more reason the runtime section below holds for this track too. The stack notes also list "deliberate departures from upstream hypotheses": the prime number theorem replaces the manuscript's explicit interval lemma, the resampling hypothesis is weakened, and so on. So this track rests on more than the CrocSwap PRs do: the manuscript, the fourteen PRs it cites, and its own replaced hypotheses.

### What the Lean files prove

There are two, `lean/Round5.lean` and `lean/Round6.lean`, and neither is written by hand. `lean/gen.py` emits them from JSON files that `dump.py` and `dump6.py` write out of the Python certificates. They are core Lean 4: no `import`, no Mathlib, no lakefile and no `lean-toolchain`, so no version is pinned; `make lean` runs `lean lean/Round6.lean` with whatever is on the PATH. Rationals are pairs of naturals with a hand-written gcd normalisation, $\ln$ is a 30-term atanh series with a geometric tail bound, and each file ends in five theorems. Round six's, from `lean/Round6.lean:74-81` with the histogram lists cut:

```lean
theorem bit_rank_sum : rankSum bitHist = 78410006675 := by decide
theorem bit_moment : momentOK bitHist (36667, 1000000000) 529 148225616 = true := by decide
theorem cx_rank_sum : rankSum cxHist = 119453132304 := by decide
theorem cx_moment : momentOK cxHist (36926111, 500000000000) 576 207387136 = true := by decide
theorem kappa_assembly : kappaOK (36667, 1000000000) (36926111, 500000000000) (1, 1000)
  (2499908335861049, 2500000000000000) (14695, 1) (14694, 1)
  (3666565558019, 100000000000000000) 576 119453132304 = true := by decide +kernel
```

There is no `sorry`, no `axiom` and no `native_decide` in either file. Both `decide` and `decide +kernel` end in a proof term that the kernel checks; `+kernel` only skips the elaborator's own evaluation first. `native_decide` would have compiled the decision procedure and trusted the compiled code, and it isn't used. The kernel does use GMP for natural-number literals, which is part of Lean's normal trusted base. So yes, it really is the kernel.

The more important question is what the theorems say. They say that three particular Boolean functions return `true` on particular lists of numbers. They do not say those lists are the child-width histograms of the networks; `gen.py` pastes them in from Python. They do not say `lnUp` bounds the logarithm or that the moment sum controls the recurrence. The generator's docstring is explicit: "The analytic facts behind (2) and (3) -- the atanh tail bound and (m/w)^a &lt;= 1/(1 - a ln(m/w)) -- are premises, not checked here: Lean core has no real logarithm." And `kappaOK` encodes his own five-margin cost table, transcribed from `certificate_round3.py`; whether those five margins are the right costs is a question for the written notes. The upstream theorem isn't in it at all. His README is honest about this; its evidence table lists the full upstream multiplication theorem as "Assumed" and independent review or formalisation as "Not supplied". So "checked in Lean's kernel" means that the arithmetic the Python fraction certificate already did has been redone by a much smaller and better-trusted checker. The check is real, and narrow.

I couldn't run it: there is no Lean toolchain on this machine, and I wasn't going to install one for this. So I wrote the five definitions out again in Python on integer pairs, including Lean's `qsub`, which is natural-number subtraction and quietly floors at zero, and evaluated the theorems on the committed inputs. All five come out true, and no subtraction ever hits the floor. The bit moment sum is $1 - 2.79\times10^{-10}$ of its bound and the complex one $1 - 7.9\times10^{-14}$. At 80-digit precision, with the real exponential, the largest bit saving this histogram supports is about $3.66702\times10^{-5}$, so the claimed $3.6667\times10^{-5}$ leaves 0.009% headroom; round five's file checks out the same way.

Lean wasn't new to the race, either. princezuda's #26 on CrocSwap, at 10:54 UTC, brought Lean 4 v4.21.0 with Mathlib ("no sorry and no native_decide") and 310 theorems over the certificate arithmetic of PRs #3 to #12, and it found one misstated ceiling. alejandrozu's #45 adds a Lean-checked Gaussian parity audit. But no CrocSwap claim after #12 has its arithmetic in Lean, and Jain's last two rounds do.

His Python checkers all ran. `make verify` took 194 s at 750 MB peak and passed everything, including 23 unit tests and a deliberately broken inverse that has to fail and does. `independent/complex-twostage/run.py 24` took 39 s: 140,064 edges, 0 bad labels, and $a_{\rm c} \ge 73861113/10^{12}$, slightly more than the $36926111/(5\cdot10^{11})$ the assembly uses. `scripts/certificate_round6.py` gave $a_{\rm b} = 36667/10^9$ at $h = 23$. The histograms these scripts regenerate are identical to the ones pasted into `Round6.lean`, so the Lean file and the Python checks are about the same objects.

Where that leaves the race, at 15:46 UTC: the highest claim is Rohan Arun's #49 on CrocSwap, $4.1239\times10^{-5}$, 12.5% above Jain's $3.6666\times10^{-5}$. The number on CrocSwap's `main`, #39's $3.8867\times10^{-5}$, is 6% above Jain's. So there are now three kinds of checking on the board: Jain's arithmetic in Lean, `main`'s new steps read through by its maintainer, and the open PRs' own Python certificates. And `main`'s README now credits Jain, with Aurel Prosz, for "attributed two-stage development and the paid copied-stream endpoint construction", so the two races have formally met. I don't think the difference in arithmetic checking matters much, because the arithmetic was never the likely point of failure. Lean makes "the numbers add up" airtight; it says nothing about whether a contiguous projector block really is one recursive call on a tape, or whether the manuscript underneath is right. Jain said it himself, replying to someone who asked for the unconditional value: "Nothing past OpenAI's result is unconditional yet."

## PR #39 reaches main

Rohan Arun opened #39 at 12:58 UTC, "Fixed middle basis and copied reversed corners", at $3.8867\times10^{-5}$. It was never the best open claim for long; his own #40 beat it by 0.82% seventeen minutes later. But at 13:23 Colkitt had pinned it, head `70ae241`, in a merge commit (`fd8c563`) on an `integration/community` branch, and he spent the next two hours on it. At 15:12 UTC he committed "Complete conditional audit of the PR39 community witness". At 15:18 the release commit `0605a24`, "Publish audited community bound with contributor attribution", moved `main`, the tag `community-kappa-15-2026-10-08` went on it, and GitHub marked #39 merged. Since #39 was stacked on earlier work, the same push also marks icekylinx's #10, #18, #24, #32 and #36 and Rohan's #37 as merged; their commits are in its history. Twelve minutes later Colkitt posted:

> Validated and merged Rohan's PR. Big gain on a really difficult regime (that frankly I was stuck at). Incredible work.

The maintainer admitting he was stuck, and that someone else's agent-assisted PR got him unstuck, is my favourite sentence of the whole episode. His headline, "κ = 2⁻¹⁵", is the floor the witness clears rather than its value; the exact number is $2^{-14.65}$, which the README puts at "27.36% above 2^-15". His "500 thousand fold improvement over the previous result" checks out against $83/10^{12}$, the last exact witness he had posted (about 468,000-fold). Against the $2^{-30}$ checkpoint that `main` carried until 15:18 it is about 42,000-fold. The "2 ^ 167 fold improvement over the original OpenAI result" is right: $2^{167.35}$.

What did "validated" mean? The audit, `docs/research/community-final-audit.md`, is plain about who did it. Its header reads "Reviewed by the project's Codex assistant for Douglas Colkitt", and it calls itself "a maintainer mathematical assessment, not independent human peer review or formal verification." Within that, it is a real proof review, and it goes after exactly the obligations my PR #13 check said the certificates never touch. It works through the partial-swap compiler's block algebra; the claim that one fixed basis serves every gate (each required minor is a nonzero rational function on one irreducible family, so a single rational basis avoids every exceptional set); the copied-centre schedule restoring both operands for arbitrary dirty values; the complex side's phase compiler; the semantic-precision recurrence; the arbitrary-coordinate router; the Gaussian inverse; and the bulk tape schedule. In one place it swaps in a published theorem, Baker, Harman and Pintz on primes in short intervals, and labels it "an actual new input". Its verdict: "no unresolved additional construction or inequality was identified that blocks this particular witness."

The executable side is separate. The audit added its own checker, `scripts/audit_community_candidate.py`, which imports none of the contributor's code, encloses both recursive moments with 80-term rational logarithm bounds, rebuilds all seven margins, gets the same final gap of about $6.25\times10^{-15}$, and shows that the next bit saving on the $10^{-14}$ grid fails. It reports 218 tests, 20 upstream patch checks and certificate regeneration in Docker (GCC 13.3, Python 3.12), a GitHub Actions run on Python 3.11, 3.13 and 3.14, a one-header GCC build fix, and 20 files missing from a vendored manifest restored from PR #34 at their recorded hashes. I re-ran `make verify` on `0605a24` and it passed, 218 tests and all, in about sixteen minutes ([How I checked](#how-i-checked) has the details).

Two limits are written into the audit itself. "The original OpenAI #109 framework ... remains assumed." And "PR40 and subsequent submissions are outside this pinned review." So `main` is deliberately behind the frontier, by 6.1% when #49 arrived.

## Main moves again

I thought the merge was the end of the story. When I fetched everything again at 19:46 UTC, `main` had moved twice more, 28 more PRs had arrived, a contributor had built a live dashboard for the race, and the race itself had jumped to a second problem.

The first move was routine. At 16:47 UTC GitHub marked Rohan's #49 merged, together with #43, #45, #46, #48 and princezuda's Lean PR #26, and the README went to $4123863984/10^{14}$. It was behind again at once: Rohan Gupta's #50 and RaD's #51 had been open for half an hour with higher claims.

The second move is the one that broke the plateau. Avi Eisenberg's #53 (he is `ikeboy` on GitHub), opened at 16:31, changes nothing except the scalar producer. Every vertex of the network needs its leave-one-out strip sums: given $k$ values, all $k$ sums that leave one of them out. The usual way builds prefix sums $A$ and suffix sums $B$ and takes $s_j = A_{j-1} + B_{j+1}$. Each prefix then has two consumers whose envelopes don't nest, so its second use always costs a fresh wire. #53 shifts the bracket by one, $s_j = A_{j-2} + (v_{j-1} + B_{j+1})$, and now the two consumers of a prefix are nested, so the carrier matching from #44 can keep the existing wire going instead of allocating a new one. It costs more additions (40,329 against #48's 37,098 at $h = 23$), but the carrier links double, from 6,002 to 12,719, and the roles drop from 36,432 to 32,946. It bought 9.2% in one step, after three hours in which every PR had gained a few hundredths of a percent. I like it because it is counter-intuitive: more arithmetic, fewer wires, and fewer wires is what $\kappa$ actually pays for.

After that the ideas stacked. eumemic's #57 compiles every group of operations that share a frame as one invertible binary map and pays to reclaim retired wires, which takes the physical wire count from 160,799,739 to 153,481,944. Eisenberg's #62 replaces the strips with every cyclic interval, each built from an interval one item shorter. It needs $k(k-2)+1$ additions instead of about $4k$, but almost every one of them continues a carrier. On its own it claims $4986133/10^{11}$, and with #57's compiler on top, $5101691/10^{11}$. Alejandro Zarzuelo Urdiales's #61 then refined the parameters on that same graph, by $1.0170078\times10^{-11}$, with the arithmetic checked in eleven Lean proofs.

At 18:35 UTC, Colkitt merged #50 to #62 (all except Rohan Garg's #59, which is kept as a reviewed alternative) in one push. `main` now reads

$$
\kappa = \frac{25508460085039}{5\cdot10^{17}} \approx 5.1017\times10^{-5} \approx 2^{-14.26},
$$

23.7% above the #49 release and 31% above #39's. The network underneath has changed too: the bit side now runs at $m = 575$ on dimensions 23 and 25, with 137,151,806 physical roles.

The review behind it, `docs/research/community-round2-review.md`, again calls itself "a conditional maintainer review with Codex assistance, not independent human peer review or a formal proof of multiplication." It reads more like a replay ledger than the #39 audit did. Each PR gets a disposition ("Full pinned replay passed", "Complete word and profile replay; used by selected witness"), and the places it holds back are stated. #51's separate witness had its "arithmetic … freshly replayed, but its entire standalone source-graph and negative-basis data argument was not", and #62's claims that its layout is optimal "were not independently reproduced". The general residual compiler, the all-size recursion, the tape implementation and the analytic transfer "remain written inherited obligations, not consequences of a successful finite test alone." So this move rests on finite replays plus reading, on top of interfaces that were audited at #39.

The same afternoon, `main` imported Jain's repository whole: 87 files, pinned by SHA-256, under `research/swapnil-parallel/`. It did the thing I couldn't. It compiled both Lean files under Lean 4.31.0, checked that the generator reproduces them byte for byte, and inspected all ten theorems, none of which depends on an axiom. Its reading of what they prove matches mine: they "do not formalize logarithmic analytic bounds, construct a Turing machine, or prove integer multiplication end to end", and the fixed-tape side of the ε → 1 stack has "not received a complete independent audit in this pass." It also notices something that matters for the next section. Jain's complex network saves $36926111/(5\cdot10^{11}) \approx 7.385\times10^{-5}$, more than the 7.17e-5 of the complex network `main` uses, but `main`'s bit side binds at about $5.10\times10^{-5}$, so the better complex network would not move the headline. Jain himself has not posted a seventh round. His last commit is still round six, at 14:01 UTC.

The open PRs didn't wait. Fifteen more arrived between 18:05 and 19:46. Dominik Scholz's #63 and rfu08's #64 published the same construction within fifteen minutes of each other, and #64 says so ("Both decompressed words, complete profiles and finite bridge are identical"). The best open claim when I stopped was Scholz's #76, opened at 19:44, which combines three other open PRs (#70, #71 and #74) to reach $51738676045141/10^{18} \approx 5.1739\times10^{-5}$, 1.4% above `main`.

So at 19:46 UTC the board reads: `main` at $5.1017\times10^{-5}$, the best open claim at $5.1739\times10^{-5}$, and Jain at $3.6666\times10^{-5}$, now 28% behind `main`. There have been 77 PRs from 20 accounts; GitHub marks 25 merged, 47 are open and 5 were closed without merging. The last fifteen open claims sit between 5.10 and 5.17 times $10^{-5}$. It's a plateau again, and I'm no longer going to predict how long it lasts. (About three hours, [it turned out](#9-october-the-ninefold-night).)

### A dashboard, built by one of the racers

At 15:13 UTC Aurel Prosz ([@aurel_pr](https://x.com/aurel_pr), Paureel on GitHub), whose two-stage topology and endpoint correction run through half the PRs above, posted a live tracker: [Beyond n log n](https://beyond-n-log-n.netlify.app/). Colkitt shared it fifty minutes later as an "awesome dashboard from the great @aurel_pr". The site owner's reaction was "damn, things are getting serious", which is fair. The page names no author in its text, but its footer links to Paureel, and the source is open at [Paureel/beyond-n-log-n](https://github.com/Paureel/beyond-n-log-n).

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig11-tracker.png" alt="Dark-themed dashboard titled 'Beyond n log n', subtitle T(n) = O(n (log n)^(1-kappa)), marked 'Conditional bounds' and 'Updated 8 Oct, 19:33 UTC'. Four tiles: strongest claimed kappa 5.141465 × 10^-5 from PR #71 by chafreaky, marked Draft; above current main 1×, with main at 5.1e-5; 74 pull requests including 10 drafts from 19 contributors; 26 public forks with 6 PRs inside forks. Below, a log-scale 'kappa over time' chart of 100 claims from OpenAI's 2^-182 on 6 Oct at 21:58 to PR #71 on 8 Oct at 19:26, and a 'Mathematical fields' graph with nodes for linear algebra, combinatorics, circuit complexity, number theory, asymptotic analysis, numerical analysis and formal verification." caption="The tracker at 19:33 UTC. Its headline is the strongest claimed κ, which at that moment was a draft PR; main is shown alongside (Aurel Prosz, Beyond n log n, screenshot rendered in headless Chromium)." />

It is a React app with a Netlify function that re-reads GitHub every 15 minutes. It covers every PR in every state, all 26 public forks and the 6 PRs inside them, each checkpoint in `main`'s history, a "field map" of which areas of mathematics each PR mentions, and an idea-lineage graph built from explicit reuse language in the PR descriptions. The $\kappa$ values are parsed from PR titles and bodies, so they are claims, and the methodology box says as much: "No full-theorem verification is inferred from tests or certificates."

I pulled its JSON (`/api/research`, fetched at 19:33 UTC) and compared it with my chart. All 46 PR claims from #1 to #49 agree to within 0.2%, which is the rounding in my labels. There are two real differences. The tracker puts OpenAI's $2^{-182}$ at the `openai/math` commit time, 21:58:50 UTC on October 6, where I had used the 3:19 PM Pacific (22:19 UTC) printed on Julian Schiavo's chart. I checked the commit and moved my point to match. And its "repo checkpoint" line steps at the commit times of everything in `main`'s history, including contributors' commits that came in through merges, so it has `main` at 4.1239e-5 from 16:16 UTC (`ed8201c`), where I use 16:47, when GitHub marks the merge. Both choices are defensible. Its 19:33 snapshot also predates #75 to #77, so it showed #71 as the leader.

## The race reaches the Fourier transform

At 16:52 UTC Ryan Shea ([@ryaneshea](https://x.com/ryaneshea)) posted:

> We're publishing a research draft on OpenAI Problem #130: computing the exact discrete Fourier transform below n log n. Our draft proposes an all-length bound of T(n) = O(n(log n)^(1−δ)), with δ = 7.3×10⁻⁵. That's a 730-million-fold increase in the exponent saving over OpenAI's published δ = 10⁻¹³.

The repository was created at 15:26 UTC with Shea's own earlier construction, a reduced centre and shared sums, worth about $5.5\times10^{-10}$, or roughly 2,600 times the saving OpenAI's construction actually achieves. An hour later, at 16:27, he replaced it with Jain's round-six complex network and got $7.3\times10^{-5}$. Asked whether this was AI-assisted, he replied: "The heavy lifting was AI, and I shepherded it towards the improvements."

<RepoCard repo="shea256/fourier-transform-below-nlogn" />

Why should a network built for multiplying integers say anything about the FFT? The [roundup's FFT section](/articles/openai-math#how-you-beat-nlog-n-save-on-one-tiny-matrix-then-amplify) explains it. OpenAI's explicit Fourier paper (family 130) takes its finite gadget from the multiplication manuscript. The saving lives in one place: an algorithm that applies $C^{\otimes k}$, tensor powers of the fixed $2\times2$ matrix $C = \tfrac12\begin{pmatrix}1+i&1-i\\1-i&1+i\end{pmatrix}$, in $O(2^k(k+1)^\theta)$ with $\theta < 1$. That algorithm is the same frame-labelled network on arbitrary roles that the race has been tuning: swap two banks, restore every scratch wire whatever it held, and count the dimensions of the frame changes. The rest of the Fourier proof (Good–Thomas reindexing, exact-width local words, synchronisation into sectors, a Bluestein chirp) uses the tensor routine only through that cost bound. The draft's whole argument rests on this. By its reading, Proposition 4.2 of the Fourier paper uses the tensor routine only through the bound $O(2^k(k+1)^\theta)$ on arbitrary complex arrays. So a better network gives a better $\theta$, and $\delta$ is anything strictly below $1-\theta$.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig13-fft-draft.png" alt="First page of 'An Improved Exponent Bound for the Exact Discrete Fourier Transform. Research draft: round-six complex-network transfer', October 8, 2026. It states T(n) = O(n (log n)^(1 - delta)) with delta = 73/10^6 = 7.3 × 10^-5 in the exact-complex-arithmetic, specified-root and logarithmic-word address model of OpenAI's explicit Fourier paper; says it transfers Swapnil Jain's round-six complex construction, pinned to commit f2176bc, and does not infer a Fourier bound from an integer-multiplication running time; a status paragraph says it has not been independently reviewed or formally verified and that the Lean file does not formalize the transfer; then the kernel C with C squared equal to the swap X, and the target bound O(2^k (k+1)^theta) with theta = 1 - a, a = 36926111/500000000000." caption="The draft's first page: the claim, the pinned source of the network, and a status paragraph that says what has not been checked (Ryan Shea, manuscript.pdf, page 1, fourier-transform-below-nlogn at bf38c00)." />

What crosses over is narrow, and the draft is careful about it. It takes only Jain's complex network at $h = 24$: 2,024 triples, $m = 576$, $W = 207{,}387{,}136$ roles and a total child width of $s = 119{,}453{,}132{,}304$, with icekylinx's whole-residual batching (#10), the copied centres from #36 and Prosz's two-stage endpoint correction, all credited. It does not take the bit network, the ε → 1 stack, the Gaussian resampling or any of the tape and precision analysis. Its audit says it "does not use Jain's integer-specific epsilon stack, bit-interchange bound, Gaussian recovery or tape precision analysis".

So the Fourier number comes out large next to the multiplication race, and for a structural reason. In #109, $\kappa$ is the smallest of seven margins and the bit network binds; Jain's complex saving was slack. The Fourier model is exact complex arithmetic with word-sized addresses. There are no tapes, no bit network and no $\epsilon$ to share the saving with, so $\delta$ can sit just under the complex network's own saving, $a = 36926111/(5\cdot10^{11}) \approx 7.385\times10^{-5}$. In the draft's words: "There is no integer-multiplication bit-network constraint or extra factor-of-two loss in this transfer." The half of Jain's work that never mattered for his own $\kappa$ is the half that matters here. It is also the best complex network on offer: as noted above, `main`'s own is 7.17e-5.

I checked the numbers. From the draft's role counts, $W = 2N + 2vR$ with $N = 2024^2$ and $R = 49{,}208$ gives 207,387,136, and $Wm - N + L$ gives the stated $s$. The deficit is $Wm - s = 1{,}858{,}032$, and the 25-row histogram in Section 6 sums to the same $s$, the rank sum that `Round6.lean` checks for the complex side. With exact fractions, $a - \delta = 426111/(5\cdot10^{11}) \approx 8.5\times10^{-7} > 0$, which is the gap that swallows the extra $(\log\log n)^{4-\theta}$ factor. At 80 digits with the real exponential, the moment at $\theta = 1-a$ is $1 - 1.87\times10^{-9}$ (Colkitt's replay of Jain's file reports the same gap), and bisection puts the critical saving at $7.38611\times10^{-5}$, the "critical root" the draft mentions. The draft's own rational bound, which uses $e^x \le 1/(1-x)$, is looser at $1 - 7.9\times10^{-14}$, and it still clears 1. The headline ratios are right: $7.3\times10^{-5}/10^{-13}$ is exactly 730 million, and against the earlier $5.5\times10^{-10}$ it is 132,727 times. Against OpenAI's actual critical saving of $2.106\times10^{-13}$ (the paper rounds it down to $10^{-13}$ for its headline) it is about 350 million.

Its checkers are standard-library Python, and I ran them under the same exception as before. `verification/run_checks.py` took 80 seconds at 245 MB. It reproduced OpenAI's own $a = 2.10643843\times10^{-13}$ from the paper's $W$ and $\Delta$, passed both earlier suites and matched their archived certificates, and then ran the round-six verifier: hashes of the vendored Jain files, all 4,096,576 local coefficient entries, both physical schedules with the copied centres charged, the full histogram, exact endpoint identities, negative controls, and "PASS: a=36926111/500000000000, delta=73/1000000, rational moment gap=7.919…E-14". The regenerated certificate matched the committed reference.

The draft is plain about the limits, and so is the best reply it got. The producer is Jain's code, vendored, and the extra checks "were written with the same assistant"; the dirty-scratch tests are finitely many rational examples; nothing executes the full DFT or the infinite recursion. Keith Adler went through it with Claude as referee ("treat it as a second pair of eyes, not human review") and couldn't break it. He listed what to tighten: the copied-centre step "is argued, not tested"; whole-residual batching "is not in #130" and needs a short extension of its Lemma 2.2; and the margin is thin, "a = 7.3852e-5 sits just under the critical root 7.3861e-5".

So what is it conditional on? First, on OpenAI's explicit Fourier manuscript, Proposition 4.2 and Section 5.4 in particular, which the draft invokes "with their stated hypotheses, not re-proved". As [the roundup notes](/articles/openai-math#the-fft-is-not-optimal-in-the-right-model-by-a-hair), the Lean statement for family 130 is a weaker one: a subsequential, non-constructive circuit bound from the second paper, with "no all-length, bounded-coefficient, conditioning, or bit-complexity claim". The explicit $\delta = 10^{-13}$ algorithm is manuscript-only, and that is the part Shea improves. Family 109 has no Lean statement at all, so the two problems sit at different levels: #130's core question is settled in Lean and its explicit constant is not, while #109's whole theorem is a manuscript. Second, it depends on Jain's finite complex network and the written arguments around it (copied centres, whole-residual batching, the two-stage endpoint). It does not depend on his assumption-heavy ε → 1 stack or on the hypotheses he replaced. Those network arguments are the same kind Colkitt's audit worked through for #39, but in the multiplication setting. Their transfer into the Fourier paper's array model, including a layout lemma extended from one selected slot to $r$ and a final address translation for the endpoint, is new and unreviewed. In one way this rests on less than Jain's own $\kappa$ does. In another, nobody but Shea's assistant and Adler's has read the new part yet.

It changes nothing about real FFTs. The README says "No practical FFT speedup … is claimed", and the [roundup's numbers](/articles/openai-math#how-small-is-10-13) still apply in spirit. Ignoring constants and the $\log\log$ factor, a 1% saving at $\delta = 7.3\times10^{-5}$ needs $\ln\ln n \approx 138$, so $\ln n \approx 6\times10^{59}$. It's a big improvement on the $4.8\times10^{10}$ the roundup found at $10^{-13}$, and it is still not a number anyone will ever transform. What it shows is that the race is not about one problem. The agents and the people pointing them have found a shared engine, and any result built on that engine is now open to the same tuning.

## 9 October: the ninefold night

The site owner's note the next morning was short: the race to κ = 1 is still on, check the repos again. When I left it, at 19:46 UTC on October 8, `main` was at $5.1017\times10^{-5}$, the best open PR at $5.1739\times10^{-5}$ and Jain at $3.6666\times10^{-5}$, and I had called it a plateau. At 11:00 UTC on October 9 `main` was at $4.609169\times10^{-4}$, the best open PR at $6.5592\times10^{-4}$ and Jain at $6.1534\times10^{-4}$. A hundred and two PRs had arrived in between. There were 179 from 39 accounts; GitHub marks 30 merged, 127 open and 22 closed. So much for plateaus.

### What moved `main`

The ninefold step is one contributor's chain of four PRs, each built on the one before. icekylinx opened #104 at 23:12 UTC (a "stopped product-ring interchange", $7.79476\times10^{-5}$), #115 at 00:26 (partial source gauges, $9.26336\times10^{-5}$), #130 at 02:29 (three-stage Cayley covers, $3.146011\times10^{-4}$) and #144 at 04:19 (paired cubes and shared completed cores, $4.609169\times10^{-4}$). Other people were tuning around them all night (eumemic's #114 was the first claim past $2^{-13}$, at 00:23), but #130 is the step that matters: it came in 2.4 times above the best open claim of the moment. Colkitt reviewed the chain and merged all four through an integration PR, #149, at 05:06. Ten minutes later he posted:

> Validated and merged a breakthrough from IceKylin. In addition to their work, the latest result builds on work from @dysmemic, @SJ_Swapnil_Jain, an664, Zhihao Chen and others.

His "~9X from this morning" is right: $4609169/10^{10}$ over yesterday evening's $25508460085039/(5\cdot10^{17})$ is 9.03. So is "roughly 2¹⁷¹": it is $2^{170.92}$ times OpenAI's $2^{-182}$. A second post in the thread said the agents were "still backed up trying to process the huge volume of PRs that keep coming in. (Keep them coming!)".

### Three shears instead of one swap

Every construction in this race is a way of doing one thing: an interchange, which swaps two banks of values $X$ and $Y$ while putting every scratch wire back the way it found it. Until last night each interchange was a single network on about $h^2$ coordinates (576 for $h = 24$; `main`'s bit side ran at 575), with an enormous number of roles per invocation: 137,151,806 on `main`'s bit side yesterday evening.

#130 does the swap the way you were taught to swap two registers without a temporary:

$$
(X, Y) \;\to\; (X,\, Y+X) \;\to\; (-Y,\, X+Y) \;\to\; (-Y,\, X).
$$

Three shears, $Y \gets Y+X$, then $X \gets X-Y$, then $Y \gets Y+X$, give the swap up to a sign, and the sign is a fixed correction at the end. The trick is in where each shear runs. The address space is split into three orthogonal blocks $A \perp B \perp C$ of dimensions $h$, $h-1$ and $h-1$, so $m = 3h-2$ (70 in #130; #144's complex side uses 72). Each stage acts on one block plus a single shared direction, and the frame a value leaves one stage in is exactly the frame the next stage expects, so there are no data connectors to pay for between them. The note puts it as "Every exit label is precisely the next entrance label."

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig16-three-stage-cover.png" alt="Section 3 of the three-stage cover note. With eumemic's h = 24 triple word (v = 2024, R = 28705, l = 552), the binary space E splits as A ⊥ B ⊥ C with dim A = h, dim B = dim C = h − 1 and m = 3h − 2. Two orthogonal involutions R12 and R23 carry A onto B plus q and C plus q. The cover takes G = O(E), and the persistent-role count is W = |G|(2v + 3R). A table gives the three stages: stage 1, Y ← Y + X on B + q; stage 2, X ← X − Y on A; stage 3, Y ← Y + X on C + q, with the X and Y frames at each step." caption="The three-stage cover: three shears on three orthogonal blocks, with every vertex made alike by letting the whole orthogonal group index the invocations. I compiled this from the note's TeX source at main d1d6c07, since the repository commits no PDF for it (icekylinx, three-stage-cover-note.tex, Section 3)." />

To make every vertex of the network look the same, the invocations are indexed by every element of the orthogonal group of that binary space, which is what "regular Cayley cover" means here. For $m = 72$ that group has about $2^{2555}$ elements. The count cancels out of the normalized moment, so it costs nothing in κ, and it lands in the constant instead. eumemic's later README puts the network at "about 2^2570 roles". #144 then added two things on top: a "paired-cube" local word that writes the identity as $I = K + H + B$, three pieces built from different parts of the circuit, and an664's completed-core sharing from #128, in which one auxiliary bank is reused in turn across the three orthogonal blocks.

### Why smaller is worth nine times

The moment condition from #10 explains the gain, and once you expand it to first order it says something you can hold in your head. A network with $W$ roles on $m$ coordinates spends its rank on children of widths $r$; the total falls short of $Wm$ by the deficit $D$. Then

$$
a \;\approx\; \frac{D/(Wm)}{L}, \qquad L = \sum_r \frac{n_r\, r}{Wm}\,\ln\frac{m}{r},
$$

where $L$ is the rank-weighted average of how far below full width a child sits, on a log scale. The saving is the fraction of rank the network fails to spend, divided by the log-distance of a typical child from full width.

I computed both sides of that for the old network and the new one, from their own histograms. Jain's round-six complex network (the one Shea took to the Fourier problem, and better than yesterday's `main`) has $D/(Wm) = 1.555\times10^{-5}$ and $L = 0.2106$: 96% of its rank sits in children wider than half of $m$, so wide children are nearly free. That gives $a \approx 7.387\times10^{-5}$, against a certified $7.385\times10^{-5}$. #144's complex side has $W = 29{,}937$, $m = 72$ and $D = 1{,}936$, so $D/(Wm) = 8.982\times10^{-4}$, 58 times larger. Its children are narrow (only 13.5% of the rank sits in children wider than 36), so $L = 1.848$, 8.8 times larger. The ratio comes to $4.860\times10^{-4}$, against a certified $4.856569\times10^{-4}$. The relative deficit grew much faster than the log penalty, and the net is a factor of 6.6 on this side. The deficit itself telescopes neatly, $D = 2v - 3\ell$ per vertex: $2\cdot1760 - 3\cdot528 = 1936$.

### The margin that binds

The assembly under #144 is CrocSwap's semantic assembly: 47 strict constraints, seven cost margins, and two small parameters $\eta = 10^{-8}$ and $\beta = 10^{-6}$. The network saving fed to it is $a = \min(a_{\rm bit},\ (1-\beta)a_{\rm c} - 10^{-10})$. The bit side binds: $a_{\rm bit} = 4613422943/10^{13}$ against $a_{\rm c} = 4856569/10^{10}$. Then $q = a(1-2\eta)$, $c = q(1+\eta)$ and $\epsilon = (1-\eta)/(1+c+q) \approx 0.99908$, and the binding margin is the compact phase layer, $\epsilon q$. The original-prefix margin $1 - \epsilon(1+c)$ sits exactly $\eta$ above it, which is the point: $\epsilon$ is chosen to balance those two, so

$$
\kappa \;\approx\; \epsilon\, a \;\approx\; \frac{a}{1+2a}.
$$

I redid that with exact fractions and got $\kappa = 4609169/10^{10}$, with the binding margin $9.945\times10^{-11}$ above it (the $10^{-10}$ grid swallows the rest), the same gap the maintainer's review reports.

One wrinkle on the bit side: its saving is "stopped". #104 pays the cross-bank address adapters as loops over fixed-width atoms, which costs linear time, and stops the recursion at atom width $e^{\theta}$ with $\theta = 1/1000$. The usable saving is then $(1-\theta)a^* + \theta\,a_{\rm old}$, a thousandth of it coming from an older, weaker network ($a_{\rm old} = 384599/10^{10}$). Unstopped, #144's bit network works out to $4.6177\times10^{-4}$, and the stop costs it about 0.1%.

The review behind the merge, `docs/research/community-round6-review.md`, is "a construction and dependency review, not independent human peer review or a formal proof of integer multiplication", done "with OpenAI Codex assistance". It rebuilt the complex producer, ran its own controls (exact factorizations over $\mathbb Z/9$, $\mathbb Z/25$ and $\mathbb Z/27$, exhaustive frame identities in dimensions two to four, and dirty-value replays of the 26,417-role mixer), read the written proofs, and prepared two small fixes, a verifier guard and one sign in a sentence. It is honest about one gap: a 1.5 MB compressed bit witness from #97 could not be fetched through its connector, so "the selected bit reconstruction was checked through inspected, head-pinned CI, not rerun locally." That file is in my clone, so I ran `make paired-cube-verify` on `main`: the producer regenerated from source and matched its certificate, the bit reconstruction ran locally with every profile and source hash equal, and it ended `PASS kappa=4609169/10000000000 = 4.609169e-4; both moments, shared cores, finite router and 47 strict constraints`, in 44 s at 270 MB.

### Jain's rounds seven to ten

Jain posted four more rounds. Here they are against what was open on CrocSwap at the same minute, as before.

| post (UTC) | Jain's $\kappa$ | best open CrocSwap claim then | what was new |
|---|---|---|---|
| Oct 8, 20:56 | $1599247689723/(2.5\cdot10^{16}) \approx 6.397\times10^{-5}$ | #86, $5.2275\times10^{-5}$ | deferred garbage readout |
| Oct 9, 04:17 | $12612978530233/10^{17} \approx 1.2613\times10^{-4}$ | #143, $4.2483\times10^{-4}$ | opposite bank orders (#104), signed reclaim (#112), shared cores (#128) |
| Oct 9, 05:25 | $4663738/10^{10} \approx 4.6637\times10^{-4}$ | #147, $4.6466\times10^{-4}$ | his bit word inside #144's three-stage cover |
| Oct 9, 09:50 | $6153378/10^{10} \approx 6.1534\times10^{-4}$ | #168, $6.5589\times10^{-4}$ | paired-cube words, birth reuse, nested-prefix modules |

Round seven led everything for two hours, by 22% when it was posted. Zhihao Chen's #97 at 22:39 ported Jain's witness into CrocSwap and came within 0.007% of it, and #99 passed it at 22:56. Round eight landed in the middle of the cover jump and was 3.4 times behind. Round nine was the highest claim anywhere for fifteen seconds, until DaysSky's #150 at 05:25:46. Round ten was about 6% behind eumemic's #168.

What, then, is round nine's 9.14-fold jump over yesterday's `main` made of? Mostly icekylinx's cover. Jain's README says it plainly: `scripts/certificate_round9.py` is "our own implementation of the paired-cube / three-stage cover assembly of icekylinx's PR #144", written from #144's description and files, with "None of PR #144's code" run. It first reproduces #144's published $4609169/10^{10}$ exactly, all 47 constraints and 7 margins. Then it swaps one thing: the bit side's word becomes his own round-seven witness (deferred readouts at $h = 23$, $R = 27{,}794$), with 1,504 of its 10,922 deferred gauges omitted. Choosing which ones to omit is an exact $s$–$t$ minimum cut, a formulation he credits to Th0rgal's #146. That raises the bit network's saving from #144's $4.6177\times10^{-4}$ to $4.67238028\times10^{-4}$; the complex side is #144's, unchanged. So round nine is 1.18% above `main`, and that 1.18% is Jain's. The ninefold before it is the cover's. His round-ten README credits, by PR number, eumemic, icekylinx, DaysSky, jamesyc, Th0rgal and an664 for the levers it stacks.

On assumptions, round nine moves him closer to the CrocSwap side rather than further away. Rounds one to eight ran on his own ε → 1 stack and its "deliberate departures from upstream hypotheses", the prime number theorem in place of the interval lemma among them. Round nine's certificate doesn't import that code at all. It runs #144's assembly, which carries the hypotheses `main` carries. What it adds is a short list of interfaces "stated here but not machine-checked": the cover lifting, frame theorem, exterior rule and stopped recurrence from #130 and #144; a precision grid of the form $2^{-P}3^{-K}$, because the cubes divide by 3 ("This looks benign …, but only dyadic grids were audited before"); block-factored frame representatives; and two inherited row-reserve constants. The maintainer's #144 review does discuss the denominator-three grid, so that one has had a reading too.

I recomputed round nine with exact fractions, outside his scripts. From the certified $a^* = 116809507/(2.5\cdot10^{11})$, the stopped mix gives $a_{\rm bit} = 29175578117/(6.25\cdot10^{13}) \approx 4.66809\times10^{-4}$, which binds against the complex $4.856569\times10^{-4}$. Then $\epsilon \approx 0.999067$, the binding margin is $\epsilon q \approx 4.66373820648\times10^{-4}$, and the floor on the $10^{-10}$ grid is $4663738/10^{10}$, as claimed, $2.06\times10^{-11}$ under the margin. $\log_2$ of that is $-11.0662$, so "2⁻¹¹·⁰⁶⁶" is right, and the ratios in his post (3.70 over round eight, "2¹⁷⁰ fold") check out too, the second one rounded down from $2^{170.93}$. At 80 digits with the real exponential, the critical saving of his bit histogram (rare-class fallback included) is $4.672380285\times10^{-4}$, so the certified grid point leaves almost nothing on the table. `make round9` and `make round10` pass, and round ten's certificate re-derives #144's and round nine's κ before certifying its own. The heavier round-nine ledger, which replays the whole forward F2 shear of the modified bit word and rebuilds its child histogram from the events, printed `ALL PASS` after 110 s at 0.96 GB. Rounds seven to ten each ship a Lean file again (`lean/Round7.lean` to `Round10.lean`), still core Lean with no `sorry`, no `axiom` and no `native_decide`, and round nine's now includes the finite bridge and the 47-row assembly. There is still no Lean toolchain on this machine, so I didn't run them; my fraction recomputation covers the same arithmetic.

Two things in his thread are worth keeping. Told by a reply that he was "losing the frontier", he answered: "I found mistakes in the PRs there so have to be super careful. Public results I publish are 100% verified correct." And round ten's commit fixes a bug in his own F2 checkers, which had been identifying a kept copy in a way that skipped some walk and data checks; he reports that rounds seven to nine are unchanged under the fix. I'd trust the second more than the first. The third post in that thread says "This will legit not be able to continue unless you support, I am running out of my resources", which is a reminder that all of this runs on somebody's token budget.

### Who @dysmemic is

The other post the site owner flagged, "κ = 4.135962 × 10⁻⁴. The race to κ = 1 continues. We must have linear integer multiplication. We will have it!", is from @dysmemic, and the screenshot attached is the tracker. The account is eumemic's: its earlier posts link eumemic's PRs #5 and #13, and on October 8 it linked `github.com/eumemic/exoplanets`. The number is the title of eumemic's #137, "Padded triple covers with sequential dirty reuse", opened at 03:18 UTC, six minutes before the post. So it lives on CrocSwap as an ordinary PR, not on a fork. It led the open claims for six minutes. geckods' #138, a tightening of the same slack constants, came three seconds after the post. eumemic's later #168 was the best open claim through the morning, and the 10:41 leader, chafreaky's #179, refines its operation frames by 0.0047%.

### The tracker now

Prosz's tracker changed its headline overnight. It used to lead with the strongest claimed κ, a draft PR at the time; it now leads with the maintainer-reviewed κ, read from `main`'s `certificates/selected-result.json`, and shows the open claims separately. He also added a zoom, after people complained that everything since $2^{-15}$ was a flat line at the top of the chart. When I rendered it at 11:00 UTC (data "Updated 9 Oct, 10:36 UTC") it showed 178 PRs, 27 of them drafts, from 39 contributors, and 49 public forks.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig15-tracker-oct9.jpg" alt="The Beyond n log n tracker on October 9, updated 10:36 UTC. Tiles: maintainer-reviewed kappa 4.609169 × 10^-4, PR #144 by icekylinx; previous main 9×, kappa ratio over the previous 5.1e-5; 178 pull requests with 27 drafts from 39 contributors; 49 public forks with 6 PRs inside forks. A log-scale 'kappa over time' chart of 12 upstream selections runs from OpenAI's 2^-182 on 6 Oct at 21:58 to icekylinx #144 at 4.61e-4 on 9 Oct at 05:06. A field map on the right shows linear algebra, combinatorics, circuit complexity, number theory and other fields." caption="The tracker on the morning of October 9. The headline is now the maintainer-reviewed κ, and the chart follows main's selections rather than every claim (Aurel Prosz, Beyond n log n, screenshot rendered in headless Chromium)." />

Sort its PR list by highest κ, though, and the top entry is #159 at $\kappa = 0.127865$, which is not a claim anyone should read. It was closed eleven minutes after it was opened, and its whole body is one sentence about "multivariate Hamiltonians" followed by "Kind reminder, AI is not to be blindly trusted." The tracker parses titles, which is the right design for a dashboard and the reason not to take its sort order at face value. I left #159 off my chart.

### And the Fourier problem again

Shea's draft hasn't moved: his only commits since are a README rewrite at 23:52 UTC that adds "This is not a faster FFT" and a who-did-what list. Its δ is still $7.3\times10^{-5}$ from Jain's round-six network.

eumemic, though, did the obvious next thing. At 08:54 UTC @dysmemic posted that the community's networks "also apply to OpenAI's new exact DFT algorithm, raising its saving from 2.1·10⁻¹³ to 4.856·10⁻⁴ in the log exponent. About a 2 billion-fold improvement!", linking [eumemic/exact-dft-bounds](https://github.com/eumemic/exact-dft-bounds). Forty minutes later the same account added, "Looks like @ryaneshea already had this idea", and the README now says Shea's repository "has priority for the idea of the transfer". The new repository takes #144's complex network as merged on `main` ($m = 72$, $W = 29{,}937$, deficit 1,936) and adds a batched recursion for the Fourier setting (its Theorem 3.1), so that a frame change of residual dimension $\rho$ becomes one recursive call to $C^{\otimes\rho f}$. The result is $\theta = 1 - 607/1250000$, so $\delta = 4.856\times10^{-4}$. Against OpenAI's actual critical saving of $2.106\times10^{-13}$ that is 2.3 billion times, and against Shea's $7.3\times10^{-5}$ it is 6.65 times.

The README is careful about scope: it is "a paper proof with an exact finite certificate, and it is not formally verified", the network's "exact complex phases, the Cayley cover and completed-core sharing are written proofs", and "The constants are astronomically large: the network has about 2^2570 roles, and the recursion's base range is of order 10⁹ bits." Its `make verify` is one exact-rational moment check plus source pins. I ran it (0.34 s; "margin 3.158222e-03 … PASS" out of $W = 29{,}937$), and my 80-digit critical root for the same histogram is $4.85657\times10^{-4}$, so $607/1250000$ leaves about 0.012% headroom. Nobody has reviewed the Fourier transfer itself yet, the batched recursion least of all.

### Where the ceiling stood at 11:00

At 11:00 UTC on October 9: `main` is at $4609169/10^{10} \approx 2^{-11.08}$, reviewed by its maintainer with Codex; the best open claim is chafreaky's #179 at $655920176686219/10^{18} \approx 2^{-10.57}$, 42% above `main`; and Jain is at $6153378/10^{10} \approx 2^{-10.67}$, with its arithmetic in Lean. Everything is still conditional on the OpenAI manuscript.

None of this has touched the ceiling. #144's assembly still has a margin equal to $\epsilon$ and another equal to $1 - \epsilon(1+c)$, so κ cannot pass ½ whatever the networks do. What bounds the current witnesses is much lower than that, and it is the chain above: κ is $\epsilon q$, about $a/(1+2a)$, so 0.09% under the weaker network's saving; and that saving is about $D/(Wm)$ divided by $L$. To get to $10^{-2}$ you need networks that waste a couple of percent of their rank at a similar log spread. One draft PR (mikey1201's #172, a "zeta family" on the Boolean lattice) projects κ up to $2.88\times10^{-2}$ at $h = 3$ through the in-tree assembly, but it is labelled a research record with named open obligations, not a claim, and I haven't checked it.

In runtime terms, $\kappa = 4.609169\times10^{-4}$ takes 0.26% off $n\log n$ at $n = 2^{256}$ and 2.0% when $\log_2 n = 2^{64}$, and halving the time needs $\log_2 n = 2^{2170}$. The calculator further down has it as a preset.

### 9 October, afternoon: reusing dead registers

The site owner's second note of the day was "there have been more updates to Jain, check his tweets". There were two. At 12:50 UTC Jain asked for an arXiv endorser in cs.CC (Computational Complexity) for a paper that "collects proofs from our work on OpenAI problem #109", and at 13:07 he posted round eleven, $\kappa = 6685733/10^{10} \approx 6.6857\times10^{-4} = 2^{-10.547}$. He ended that post differently from the other ten: "We seem to be approaching a wall. Gains now come in fractions of a percent. The next real jump needs a big breakthrough which we are unable to find yet." By the evening the CrocSwap side was saying the same thing, with proofs.

Rounds ten and eleven do one kind of thing, most visibly on the complex side: they stop paying for registers the network has finished with.

Round ten's version is birth reuse, jamesyc's idea from #124 with eumemic's late compensation reads from #143, which Jain re-implemented as an exact maximum-weight matching. A gauged slot is a role whose register has to be read at a known frame, and in the plain ledger it pays a tail child to get its register into place. Born on the register of a slot that has just died, it doesn't: the donor's last climb and the newborn's first are charged as one ascending chain. On round ten's complex word all 3,180 gauges found a donor. With the rest of that stack, at $p = 11$ (so $m = 66$ and deficit $D = 1{,}320$ per vertex), the word came to 14,592 roles and a certified $a_c = 61609676/10^{11} \approx 6.1610\times10^{-4}$.

Round eleven's new piece is the one his thread describes as "two retired copies of the same value are subtracted in place, and a fresh slot is born on the zeroed register with no read". His word has 428 pairs of registers that each finish holding the same value and are never touched again. For such a pair, one gate $P \mathrel{-}= Q$, run at a frame that contains both copies' last frames, leaves $P$ with no signal at all, only dirty scratch. A fresh role is then born on $P$ with no read. Its starting content is a linear function of the initial dirty values, and the frame-zero compensation of every ungauged role is recomputed over the new transcript to cancel it. The kernel is from icekylinx's #184 ("zero-fresh recycling through a containing frame"), a PR that was closed five minutes after it was opened. Jain's own part is applying it to this word, choosing 386 hosts by another exact matching, and writing the ledger and the checker.

The ledger shows why it helps and why by so little. Each host removes one role, so $W$ falls from 13,308 to 12,922, and the deficit stays at 1,320, because the children removed and the children added differ in total width by exactly $h$ per stage. In the commonest shape (copies last at dimensions 2 and 3, mixed at dimension 4, $h = 22$) the removed children have widths 20, 4 and 19 and the added ones 2, 1 and 18. In the first-order picture from [the ninefold night](#why-smaller-is-worth-nine-times), $D/(Wm)$ rises 3.0%, but the new children are narrow, which costs about 1.5% in $L$, and the net is the 1.43% his gate reports: $a_c$ goes from $662034051/10^{12}$ to $67147467/10^{11}$. From the round-eleven histogram I get $D/(Wm) = 1.548\times10^{-3}$ and $L = 2.303$, so $a \approx 6.721\times10^{-4}$ against the certified $6.7147\times10^{-4}$.

Most of round eleven's 9% over round ten comes from elsewhere. The complex side switched to eumemic's #168 v4 modules (used as data), a re-chosen cube-local circuit Jain calls C1, which he puts at about 1.3% better than #168's, and 42 deleted terminal outputs. The bit side got a new cyclic all-but-one module, 1,762 birth-reuse pairs and 24 terminal sinks. Its $W$ went from about 21,491 to 19,017 with the deficit fixed at 1,936, so $D/(Wm)$ rose 13% and $L$ 6.7%, and the coarse saving went from $6.330\times10^{-4}$ to $6.701\times10^{-4}$.

### Which side binds, and which margin

The two rounds bind on opposite sides. In round ten the complex network was the weaker one, $6.1610\times10^{-4}$ against a stopped bit saving of $6.3239\times10^{-4}$. Round eleven's complex side jumped past the bit side, so now the bit side binds: $a^* = 134020033/(2\cdot10^{11})$, stopped at $\theta = 1/1000$ to $a_{\rm bit} = 133893704947/(2\cdot10^{14}) \approx 6.6947\times10^{-4}$. Inside #144's assembly nothing moved. The binding margin is still the compact phase layer $\epsilon q$, with the original-prefix margin exactly $\eta = 10^{-8}$ above it. I redid the assembly with exact fractions: $\epsilon \approx 0.998663$, $\epsilon q \approx 6.68573327\times10^{-4}$, and the floor on the $10^{-10}$ grid is $6685733/10^{10}$, $2.7\times10^{-11}$ under the margin. Its $\log_2$ is $-10.5466$, so "2⁻¹⁰·⁵⁴⁷" is right, and "1.09 fold" is 1.0865. The same arithmetic reproduces round ten's $6153378/10^{10}$, bound by the complex side.

`make round11` passed in under a second. It reproduces #144's κ, round nine's and round ten's before certifying round eleven, then runs eight unit tests. I also ran the two new heavy gates by hand. The complex recycling checker on the $p = 11$ word ended `ALL PASS` after 44 s at 0.96 GB, with its controls rejected, among them a skipped mix and a flipped mix sign. The JSON-only bit ledger replayed the $p = 12$ word forward and reflected, matched the release inventory, and ended `ALL PASS` after 2 min 18 s at 0.98 GB. `lean/Round11.lean` is there. His `gen.py` regenerates it byte for byte from the histogram file, it has no `sorry`, `axiom` or `native_decide`, and it covers both moments, the stopped mix, the finite bridge and the 47-row assembly. There is still no Lean toolchain on this machine, so I didn't run it. The README is plain about what is new and "stated here but not machine-checked": the recycling rests on #184's argument and "There is no literal frame replay of the mix gate at any size"; the terminal deletions and sinks rest on jamesyc's lemma from #166; and the C1 circuit, the per-operation descent and the recycling are checked at full size only by an F2 recount and a scalar replay.

### Nineteen minutes in front

When round eleven went up, at 13:07:49 UTC, it was the highest claim anywhere. `main` had moved ten minutes earlier (below) to $6.6189\times10^{-4}$, and the best open PR was ikeboy's #194 at $6.6743\times10^{-4}$, so Jain was 1.0% above `main` and 0.17% above the best open claim. It held until 13:27:16, when evmckinney9's #197 claimed $6.7608\times10^{-4}$ by packing the bit residuals of rohanarun's #187 into shared banks. That reads #197 by its title as it stands tonight, and PRs here get edited in place, though its body's own comparison with #194 agrees with the title. By 18:50 Jain was 2.4% behind the open frontier.

### The arXiv paper

The request itself is two lines. arXiv asks people submitting to a category for the first time to be endorsed by someone who already has, so this is the normal gate for a new author. I went looking for the paper. It isn't in his repository at the newest commit (`1a580dc`, 13:01 UTC), which holds the ten TeX notes and their PDFs, the certificates, checkers and Lean files, but nothing that reads as one paper, and none of his other public repositories has it. So I can't tell you what it proves.

What I can say is what a paper collecting these proofs has to stand on, because the README sets it out round by round. It assumes OpenAI's manuscript, which has no Lean statement and no published review. Since round nine it runs inside icekylinx's cover assembly from #144 and inherits that assembly's interfaces (the three-stage cover lifting, the exterior rule, the stopped recurrence), and on top of those sits his own list of stated-but-unchecked interfaces, which grew again in round eleven. What Lean checks is the arithmetic: the moments, the stopped mix, the finite bridge and the assembly. A single written paper would at least give the proofs one place to be refereed, and refereeing is what this race has had least of.

### `main` jumps again, by import

At 12:57:37 UTC Colkitt pushed a morning's worth of reviewed commits to `main` in one go, taking it from $4.609169\times10^{-4}$ to $330942774629799/(5\cdot10^{17}) \approx 6.61886\times10^{-4}$, about $2^{-10.561}$ and 43.6% higher. The witness is Dugongue's #186. It takes chafreaky's shared-edge complex supplier from #181 and eumemic's #168 v4 bit word and adds a compensated rematch, eleven coordinated frame cuts, and "completed entrance banks", which on the bit side pack 2,200 rank-20 entrance gauges into shared scratch banks. Here the complex branch binds, at $6.6232\times10^{-4}$ against an effective bit saving of about $6.6506\times10^{-4}$. The push also carried an overnight candidate that never reached `main` on its own, the recycled-bit composition of #147, #150 and #151 at $4.7215\times10^{-4}$. GitHub marks those three, and two verification PRs, #64 and #101, merged at 12:57:38. #186 itself came in as a pinned package and is still open on GitHub.

The review, `community-round8-review.md`, is again Colkitt "with OpenAI Codex assistance", and this time it found a gap and filled it. #186's identity "does not alone specify an efficient conflict-free schedule for the physical invocations", so the maintainer wrote a bank-scheduling supplement, checked exhaustively in a small model over $GL_2(\mathbb Z/9)$ (3,888 invocations, 15,552 endpoint basis columns), and says plainly that it "is not a new Lean-checked theorem". It also explains why Rohan's #185 can't be stacked on top: #186 is limited by its complex side, so better bit leaves change nothing. I ran `make entrance-bank-verify` on `main` at `3b6b668`. It passed, "complete complex and bit words, 47 constraints, seven margins, and 21 assembly controls", plus the scheduling diagnostic, in two minutes at 1 GB. As far as the fxtwitter mirror shows, Colkitt hasn't posted about this one.

### The race starts proving its own ceilings

The PRs I found most interesting this afternoon claim no κ at all. DaysSky's #192 proves that no choice of operation frames for #168's complex word can certify more than $\kappa < 7.010\times10^{-4}$. It starts from "Every certified saving satisfies `a < D/C`", where $C = \sum n_r\, r \ln(m/r)$. That is the first-order formula from last night turned into an inequality ($C = WmL$), and it follows from $e^x \ge 1+x$. Then it bounds $C$ from below over every frame layout. Two hours later DaysSky's #201 bounded the whole design: no word on #144's paired-cube design can certify more than $\kappa < 2.5592\times10^{-3}$, about 3.9 times `main`, whatever its circuit, gauges, reuse or frames. It leaves out the newer source-assisted words, which break its rules. chafreaky's #212 carried #192's method to the bit word that #200 to #207 share and got $\kappa < 6.959\times10^{-4}$ there, so frame tuning on that word has at most 1.8% left. huxint's draft #203 argues that several current lines can't reach $\kappa = 0.01$ even with every auxiliary call free.

None of these has been reviewed, and I haven't checked them. They are still the right work at this stage, and they put a number on Jain's wall: the word families being tuned are within a few percent of their own ceilings, and the paired-cube design as a whole has at most a factor of about four left.

### Where it stands at 18:50 UTC

`main` is at $6.61886\times10^{-4} \approx 2^{-10.561}$, reviewed by its maintainer with Codex. The best open claim is hcg890's #217 at $6.84697\times10^{-4} \approx 2^{-10.512}$, opened at 18:38 and 3.4% above `main`. There are 217 PRs from 47 accounts; GitHub marks 35 merged, 158 open and 24 closed. Jain is at $6.6857\times10^{-4} \approx 2^{-10.547}$, with the arithmetic in Lean. All of it is still conditional on the OpenAI manuscript. In runtime terms `main` now takes 0.37% off $n\log n$ at $n = 2^{256}$ and 2.9% at $\log_2 n = 2^{64}$, and halving the time needs $\log_2 n = 2^{1511}$.

On the Fourier side, eumemic's `exact-dft-bounds` hasn't changed since its 09:34 credit to Shea. Shea's draft has. At 15:24 UTC he reorganised it as a running record of transfers and selected Jain's round-eleven complex network, which takes $\delta$ to $0.00067$ from the tensor saving $67147467/10^{11}$, 1.38 times eumemic's $4.856\times10^{-4}$. His audit file calls it "same-assistant mathematical review and finite checks; independent review and end-to-end formal verification remain pending". I ran its `verify_round11.py`, which checks all 17,057,040 scalar output coefficients of the word, and it passed in four minutes.

Prosz's tracker has more rows but no new layout, so I haven't re-shot it.

## What "conditional" means here

Everything in the repository is stacked on top of other assumptions. The base layer is OpenAI's 73-page manuscript, which has no Lean statement and, as far as I can find, no published expert review. On top of that are Colkitt's written extensions, and on top of those each PR builds on earlier PRs. By the evening GitHub marked 25 of them merged, after two review rounds (#39 and the six it stacks on, then #43 to #62 in two batches); by 11:00 UTC on October 9 it was 30, with icekylinx's chain from #104 to #144 merged after a third, and by the evening 35, after a fourth round that imported #186 as a package and left it open. The 158 still open are neither merged nor reviewed. PR #44's dependency list runs through #43, #41, #36, #34, #32, #29 and on down to #3.

The maintainer's [contribution-review index](https://github.com/CrocSwap/integer-mult-bounds/blob/main/docs/research/contribution-review.md) treats that stack carefully. It had six review stages in dependency order, foundations first (#3, #5, #7), then batching (#10, #13), then the later layers, and it states its policy: "A contributor's successful test report, or a large headline in a PR title, is not treated as verification by this project." The `integration/community` branch where #39 was staged carried a ledger whose gates were honest about the state of review: "Executable replay passed; proof review pending", and for two-stage and copy schedules, "Initial text inspection only". The audit closed those gates for #39's chain, and the second round extended them by replay to the PRs it merged; nothing else. The ledger also says "A failed gate should isolate the strongest surviving earlier checkpoint rather than trigger an unsupported all-or-nothing acceptance of the latest number." That rule is why the $2^{-30}$ checkpoint is still kept, on `release/ternary-30` and in the release notes, and of all the process decisions in the repository it is the one I would copy.

## Who wrote it

Mostly AI agents, and they say so. 42 of the 44 PR descriptions name the model or tool. OpenAI Codex appears in most of them; 15 also credit "GPT-6 Astra"; eumemic's PRs end with "🤖 Generated with Claude Code"; Dominik Scholz's #33 credits "Anthropic Claude Opus 5.5, OpenAI GPT-6 Astra and Codex"; and Rohan Arun's #44 says it was "Prepared by Rohan Arun with substantial Anthropic Claude assistance; OpenAI Codex performed this local arithmetic replay and upstream submission". Seven PR branches carry the `codex/` prefix that Codex gives the branches it opens. #34 mentions only "independent agent reviews", and #2 says nothing. Colkitt's notes all carry the footnote "Prepared with assistance from OpenAI Codex", and CONTRIBUTING.md asks contributors to "disclose substantial AI assistance". Jain's repository says the same of itself: "Research, implementation and drafting were done with assistance from Claude (Anthropic)." Rohan Arun, whose PR ended up on `main`, put the tempo best, announcing #40 at 13:39 UTC: "The improvements in the end started getting faster than we could run the verifications." In the same post he tagged @sama and @thsottiaux with "could I get a reset?", which I read as a plea about usage limits, and added, diplomatically, "Mostly used Codex here but Claude collaborating on the other side might have helped reach a new best".

The style gives it away too, though that is my reading rather than proof. Every PR has the same structure (result, construction, verification, attribution, scope), the same careful hedging ("not formal verification or completed external expert review"), and the same habit of reporting a few hundredths of a percent over the previous PR, often within ten or twenty minutes of it. People aren't usually that fast or that consistent. A good share of the tempo is people pointing agents at the newest PR and asking for more.

Colkitt's own posts are the human layer on top. On the morning of October 8: "Went to bed last night, and sorry to report only a bit of incremental progress on my side. *But* it appears like this has taken off in terms of now having a real community effort", followed by a promise to "validate the results, definitely make sure we have attribution". The attribution machinery is the most unusual part of the repository: a CONTRIBUTORS.md that credits all 39 PRs at the checkpoint, including the closed and superseded ones; a pinned `contribution-snapshot.json` of every PR head hash; NOTICE entries for #7 and #3; and, until the audit, nothing merged into `main`, so nobody's PR silently overwrote anyone else's. When #39 went in, the attribution went with it: the README names Rohan for "the final #39 witness" and then nine contributors and groups by what each supplied, and NOTICE gained the line "it is not solely Douglas Colkitt's work." Colkitt's post promised more: "I want to take time to highlight each of them, but for now getting these results verified and published as fast as possible so everyone working on the problem is at the leading edge." At 03:32 UTC on October 9 he posted a video, "Small, But Not Zero", billed as "An opera explaining the integer multiplication math journey up to this point." One contributor, hipotures, left a comment on #20 calling it "real-time mathematics" and comparing it to SETI@home. The comparison is a stretch, but I can see why it came to mind.

## "Linear time by tomorrow"

Two hours after the $2^{-59}$ post, Julian Schiavo posted a chart.

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig2-julian-chart.jpg" alt="Chart titled 'Integer Multiplication Becoming Linear?' plotting the reported exponent saving kappa on a log scale against time from October 6 to October 11. Points: 2^-182 at Oct 6 3:19 PM, 2^-108 at Oct 6 7:24 PM, 2^-78 at Oct 7 6:56 AM, 2^-59 at Oct 7 12:35 PM (Pacific). A dotted log-linear extrapolation reaches the kappa = 1, O(n), linear-time threshold just before Oct 8 and stays there." caption="'based on my data analysis, integer multiplication will be linear time by tomorrow!' The dotted line is a log-linear extrapolation capped at κ = 1; times are Pacific (Julian Schiavo on X, about 158K views; chart image from the post)." />

When Colkitt's $2^{-34}$ post landed, he followed up: "update! @0xdoug hit 2^(-34), so we are sadly now 18 minutes behind schedule".

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig3-julian-update.jpg" alt="The same chart with a fifth point, 2^-34 at Oct 7 5:38 PM Pacific, sitting close to the dotted extrapolation toward kappa = 1." caption="The follow-up, with 2^-34 added (Julian Schiavo on X, reply to his own post; chart image from the post)." />

He kept it going. At 16:30 UTC on October 8, an hour after Colkitt's $2^{-15}$ post: "the saga continues. good news: @0xdoug hit 2^(-15)! bad news: we missed the 12:44am target, based on my careful data analysis we will now have linear time integer multiplication at 6:44pm. by popular demand, I removed the O(n) limit".

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig12-julian-third.jpg" alt="The third version of the chart, 'Integer Multiplication Becoming Linear?', with a sixth point, 2^-15 at Oct 8 8:30 AM Pacific. The dotted fit is now a power-law curve in log2 kappa: a blue version capped at kappa = 1 crosses the linear-time threshold late on October 8 and runs flat to October 11, and a red uncapped version keeps rising above kappa = 1." caption="The third chart. The blue fit is capped at κ = 1 and the red one is not; times are Pacific (Julian Schiavo on X, October 8; chart image from the post)." />

The fit is now a power-law curve in $\log_2\kappa$ rather than a straight line, and with the cap removed the red line sails on past $\kappa = 1$, into multiplication faster than linear time, an algorithm that would finish before it had read its input. 6:44 PM Pacific is 01:44 UTC on October 9. When I last looked, at 19:46 UTC, `main` stood at $2^{-14.26}$ and the best open claim at $2^{-14.24}$, so the curve needs about fourteen more doublings in six hours. I'm not taking the bet, but I wouldn't have taken the one against #53's jump either.

The 6:44 PM deadline passed with `main` at $2^{-14.26}$. Then the cover landed, and at 05:53 UTC on October 9 he posted again: "New data means new data analysis! Updated data suggests we will now reach linear time integer multiplication at 8:33am tomorrow morning".

<Figure src="https://ai.thesatyajit.com/articles/integer-mult-exponent-race/fig14-julian-fourth.jpg" alt="The fourth version of 'Integer Multiplication Becoming Linear?'. A seventh point, 2^-11 at Oct 8 10:16 PM Pacific, sits just above the 2^-15 point. The dotted power-law fit in log2 kappa now meets the kappa = 1 line early on October 9 Pacific; the blue fit is capped at kappa = 1 and the red one keeps rising." caption="The fourth chart, after the merge of PR #144; the new point is Colkitt's 2^-11 post, at 10:16 PM Pacific on October 8. Times are Pacific (Julian Schiavo on X, October 9; chart image from the post)." />

8:33 AM Pacific is 15:33 UTC on October 9. At 11:00 UTC the best claim anywhere was $2^{-10.57}$, so the curve needs ten and a half more doublings in four and a half hours. The ninefold night is the closest the real data has come to keeping his schedule, which is funny in its own right. It doesn't change the argument below: every one of those doublings has to come out of a relative rank deficit, and κ can't pass ½. The deadline came and went with the best claim at $2^{-10.53}$ (maxime-fleury's #199), and there hasn't been a fifth chart.

It's a joke, and a good one, and the chart above already answers it: the extrapolation is on. A straight line in $\log_2\kappa$ through Colkitt's $2^{-78}$ and $2^{-59}$ posts reaches $\kappa = 1$ at about 13:08 UTC on October 8. At that moment the real frontier was PR #39 at $3.887\times10^{-5}$. The line was about fourteen and a half doublings out, which is a factor of roughly 25,000. It was never going to land, and the reasons are useful.

First, $\kappa$ is bounded, and this construction has a lower ceiling than the obvious one. Linear time is $\kappa = 1$, but as shown above the shape of the margins keeps $\kappa$ below $1/2$ and below the network's own saving $1-\tau$, which is a relative rank deficit. Second, the early gains were not discoveries about multiplication. They were slack being given back: the manuscript rounded a real saving of about $10^{-12}$ down to $2^{-50}$ and multiplied three conservative factors together. Undoing deliberate conservatism produces enormous ratios quickly, and then it is used up. Third, every later improvement is a smarter constant inside the same recursion, and the flat tail of the chart shows what that looks like.

And then there is runtime. With the constants ignored, the ratio of $n(\log n)^{1-\kappa}$ to $n\log n$ is $(\log n)^{-\kappa}$. Take $n = 2^{256}$ bits, an integer with about $10^{77}$ binary digits, roughly one per atom in the observable universe. At $\kappa = 2^{-182}$ the saving is about $9\times10^{-55}$ of the running time. At PR #13's $7699/10^{10}$ it is about $4.3\times10^{-4}$ percent. At PR #44's $4.1\times10^{-5}$ it is about 0.023 percent. Even with $\log_2 n = 2^{64}$, a number whose length in bits needs a 64-bit counter, PR #44 saves 0.18 percent. To halve the time you need $\log_2 n = 2^{1/\kappa}$: at PR #13 that is $2^{1.3\text{ million}}$, at PR #44 about $2^{24387}$. That's before the constants, which in this construction are astronomical (the roundup's FFT section found a $2^{71}$-role batch at the top level of the sibling paper's network).

<LogFactorCalculator />

So GMP is safe, and it would have been safe at $\kappa = 1/2$ as well.

## Why I still think it's interesting

Not for the number. What's interesting is the loop. Within two days of a hard theorem being posted, a stranger forked it, rewrote its parameters with an agent, wrote up each step with exact certificates and patches against a pinned source, and published. Then about a dozen other people, almost all working through agents, started building on each other's unmerged PRs every few minutes, with tests, PDFs, notes and attribution blocks that are more careful than most human repositories manage. Some of it is real mathematics: the $\mathbb F_3$ identity, the batched moment recurrence and the source-frame observation are ideas, not tuning. The certificate culture is real too. Every headline in the repository comes with a script that rebuilds it from exact fractions, which is why I could check PR #13 in an afternoon.

What it is not: a faster multiplication algorithm, a refereed result, or a formal proof. Everything depends on a manuscript nobody has independently reviewed, only what `main` has taken in (#39's chain, then #49 and the #50 to #62 round, then #144 and #186) has been through the maintainer's review, and that review was Colkitt with Codex rather than a referee, and the hard obligations (that the tape procedures do what the notes say at full size) are exactly the parts the certificates do not touch. If something in the upstream argument breaks, the whole tower goes with it, and the repository's process is designed to fail back to the strongest surviving checkpoint rather than pretend otherwise.

The race also shows the risk. Seventy-seven PRs in under twenty-three hours is far more than one maintainer can review. He audited one chain in about two hours, and by the time it was merged the open frontier was 6% past it. He then reviewed thirteen PRs in one round, and an hour after that merge the open claims were 1.4% past it again. Overnight he reviewed a four-PR chain and merged a ninefold gain, and by 11:00 the open claims were 42% past that, across 179 PRs. His next review put `main` at #186's $6.62\times10^{-4}$; Jain passed it ten minutes later, and by 18:50 the open claims were 3.4% past it again, across 217. And now the same engine is being tuned for a second problem, with a second set of unreviewed transfers. At this volume, the scarce resource is the human checking, not the agent output.

## How I checked

I shallow-cloned the repository and fetched PRs #3, #5, #7, #10, #13 and #44 by ref. Commit times come from `git log`, PR times and descriptions from the GitHub REST API (44 PRs at 14:24 UTC on October 8; #45 arrived at 14:35), and the X posts and their replies from the fxtwitter mirror. I read the upstream manuscript's introduction, motif bounds and assembly sections, Colkitt's notes, the review index, the contribution snapshot and the integration ledger.

As a stated exception to my usual rule of not executing third-party code, limited to the repository's own pure-arithmetic checkers with no network, I ran the README's focused verification on `main` and the PR #13 certificate generator, its tests and its full `make verify`, all under `nice -n 19`. I recomputed PR #13's seven margins and its bit moment independently with Python fractions and 80-digit decimals, the $\mathbb F_3$ residue identity by brute force at $h = 9$, the implied network savings from the certificates' counts, and the runtime factors with 80-digit decimals. I did not run any later CrocSwap PR's suite, and did not review any proof. OpenAI's release time on the chart is the `openai/math` commit time from the GitHub API, 21:58:50 UTC on October 6. The $\log_2\kappa$ values in the race chart are computed from each PR's stated $\kappa$, not from its certificate.

For Swapnil Jain's track I cloned [integer-mult-kappa](https://github.com/Swapnil-jain/integer-mult-kappa) with full history (20 commits, 07:17 to 14:01 UTC) and read all of it: the README, NOTICE, the nine notes, the certificates, the `independent/` re-implementations and the two Lean files with their generator. His six posts and their threads came from the fxtwitter mirror, and the CrocSwap PR bodies and their credits to him from the GitHub API, last at 15:24 UTC, when #48 was the newest PR. Under the same exception I ran his `make verify`, `scripts/certificate_round6.py` and `independent/complex-twostage/run.py 24`, all pure Python with no network, and checked that the histograms they regenerate equal the Lean inputs. There is no Lean toolchain here, so I did not run Lean; instead I re-implemented the five Lean definitions in Python and evaluated both files' theorems, recomputed the round-six margins with exact fractions, and redid both moments at 80 digits.

For the merge I fetched every PR ref again at 15:45 UTC, read `git log --graph` on `main` (the merge commit `fd8c563` has `70ae241`, #39's head, as its second parent), the PR states from the GitHub API, the diffs of `c9fca20` and `0605a24`, the audit, the release notes, the integration ledger, CONTRIBUTORS.md and NOTICE, and Colkitt's and Rohan's posts through the fxtwitter mirror. Under the same exception as above, I ran `make verify` on `0605a24` at `nice 19`. It passed: the audit's own arithmetic check ("PASS independent moments, seven margins and all 32 RaD source hashes"), the #39 producer and its `PASS bit=3886826921/100000000000000;kappa=971668963/25000000000000`, every earlier witness back to $2^{-75}$, 20 patch checks and all 218 unit tests, in 16 minutes 12 seconds of wall clock at 2.9 GB peak, with the working tree clean afterwards, so every regenerated certificate matched the committed one. That checks the certificates and the replay, not the audit's proofs, which I read but did not re-derive.

For the evening I fetched every PR ref and the GitHub PR list again at 19:46 UTC (77 PRs), read `main`'s log since 15:18, the README at `ed8201c` and `0d235fe`, the round-two review and validation receipt, the imported `research/swapnil-parallel/` review and its `verification.json`, and the bodies of #53, #57, #61, #62, #64 and #76. Jain's repository, cloned again, has no commit after `f2176bc`. For the tracker I rendered the page in my own headless Chromium, read its source repository and compared its `/api/research` JSON with my chart's data. For the Fourier draft I shallow-cloned [shea256/fourier-transform-below-nlogn](https://github.com/shea256/fourier-transform-below-nlogn) at `bf38c00`, read the README, the manuscript, both audit files and the verifiers, and read the thread and replies through the fxtwitter mirror. Under the same exception as above I ran `verification/run_checks.py` (pure standard-library Python, no network). I recomputed $W$, $s$, the deficit, the histogram sum, $a-\delta$ and the ratios with exact fractions, and the moment and its critical root at 80 digits. I did not review the transfer's proofs.

For 9 October I fetched both repositories again at about 11:00 UTC: `main` at `d1d6c07`, every PR ref and the GitHub PR list (179 PRs), Jain's repository at `d2f6146` with full history, and his, Colkitt's, eumemic's (@dysmemic), Prosz's, Shea's and Julian Schiavo's posts, threads and replies through the fxtwitter mirror. I read the bodies of #104, #130, #137, #144, #159, #168, #172 and #179, the three-stage-cover and paired-cube notes, the maintainer's round-six review, and Jain's README sections for rounds seven to ten and his round-nine and round-ten certificates. Under the same exception as above I ran `make paired-cube-verify` on `main` (44 s, 270 MB) and then the full `make verify`, which passed its community, producer, partial-gauge, three-stage-cover, paired-cube, certificate and ternary stages and then stopped after 44 minutes, at 3.1 GB peak, in a preserved-research check for PR #51 whose C++ helper needs Boost headers this machine doesn't have; I also ran Jain's `make round9`, `make round10` and the round-nine bit ledger with its certificate check, and eumemic's `make verify` in `exact-dft-bounds`, all Python (plus the repository's own small C++ helpers on `main`) with no network. I recomputed #144's and round nine's assemblies with exact fractions, both round-nine moments' critical roots at 80 digits, the first-order saving $D/(Wm)/L$ for #144's complex network and Jain's round-six network from their histograms, the orthogonal group's order, and every ratio quoted. I did not run Lean, and I did not review the cover's or the Fourier transfer's proofs. The PR points from #78 on come from the GitHub API and the tracker's JSON, with each PR's κ as its title or body read at 11:00 UTC, which for PRs edited in place (#168 most of all) is a later claim than the one they opened with.

For the afternoon of 9 October I fetched everything again at about 18:50 UTC: `main` at `3b6b668` with GitHub's push log for it, the GitHub PR list (217 PRs, with titles and bodies), the tracker's JSON, Jain's repository at `1a580dc` with full history and the list of his public repositories, eumemic's `exact-dft-bounds` at `6f87d1a` and Shea's draft at `a2840b1`, and Jain's, Colkitt's, @dysmemic's, Shea's and Julian Schiavo's posts, threads and replies through the fxtwitter mirror. I read Jain's README sections for rounds ten and eleven, the recycling gate's README, the maintainer's round-eight review, Shea's round-eleven audit, and the bodies of #186, #192, #194, #197, #201, #203, #212 and #217. Under the same exception as above I ran Jain's `make round11`, his complex recycling checker on the $p = 11$ word and his bit ledger on the $p = 12$ word with its inventory check (writing their outputs to my scratch directory rather than the script's `/tmp`), CrocSwap's `make entrance-bank-verify`, and Shea's `verify_round11.py`, all Python with no network. I recomputed rounds ten and eleven's stopped mixes and assemblies with exact fractions, and $D/(Wm)$ and $L$ for both sides of both rounds from the histograms in his Lean inputs. I did not run Lean, and I did not check the ceiling PRs' proofs or any of the new interfaces. The new chart points from #181 on take each PR's κ from its title as it read at 18:50.

I compiled the manuscript's TikZ flow figure from the pinned source with Tectonic, and the three-stage cover page from its TeX source at `d1d6c07` the same way. The note pages, Jain's and Shea's included, are rendered from the repositories' committed PDFs, the README image is a browser screenshot of the GitHub page, and the tracker images are screenshots of the live site at 19:33 UTC on October 8 and 11:00 UTC on October 9.
