Long-form explainers on AI, organised by what you want from them: the must-reads, what you can run yourself, ideas worth learning, and what's new. How articles are scored
Must read. The best pages here, by overall score.
445 articles

Tokenizers v1 and Combinatorial BPE: a faster BPE, and the six ways to spell 'the'
Measures HF tokenizers v1 (same IDs over 20M tokens, 7-54x faster, Python API broken) and counts 19-32% of real vocabularies spent on space and case variants.

Qwen-Image-2.1: the 7B is 28% of the pipeline, and the licence is not Apache-2.0
Reads Qwen-Image-2.1's weights: an exact 7.1B DiT in a 25.6B pipeline, a prefix cache exact by design, a non-commercial licence, and 4-core CPU runs.

NVIDIA Model Optimizer: from mtq.quantize to a 4.95-bit DeepSeek-R1
Reads ModelOpt's quantize path and scale rules from source, then counts DeepSeek-R1-FP4's headers: 4.95 bits a weight, 1.63x smaller, attention left in BF16.

ORPO: the odds ratio is the good part
Derives ORPO's odds-ratio loss and why it beats a probability ratio, prices the reference-model saving at 11%/25%, finds a sign error, shows GSM8K retention.

stable-diffusion.cpp: 52 model versions in one binary, and a quantiser that never touches a convolution
Why sd.cpp never quantises convolutions (q4_K can be bigger than q4_0), a memory planner from real tensor tables, and a CPU run matching it to the byte.

EmbeddingGemma 2: the encoders are optional, the text model is not
Recounts EmbeddingGemma 2 to 744M from its header, hashes its audio tower as Gemma 4's, and finds the 30 s and 32-frame defaults behind the 327 s claim.

Jev in the browser: N stops being a loop bound
A per-option scorer live in the page, measured: isolation costs 6.2x the tokens, the ONNX export welds N=25 into the graph, and small models answer wrong.

Parallel Constrained Decoding, From First Principles
Builds parallel constrained decoding in 40 lines of PyTorch, finds one repo's CUDA path silently ties colliding enum tokens, and gives a calibration recipe.

Sana: the autoencoder does more of the work than the linear attention
Rebuilds Sana's FLOP table from its configs to show the 32x autoencoder, not linear attention, carries the speedup below 5000px, and that 100x is a 4K figure.

vLLM: what PagedAttention turned into
A source read of vLLM v1: chained block hashes, the two-policy free queue, one token budget for prefill and decode, deleted swapping, and a Rust frontend.

RLCD is not constrained decoding
What RLCD is and is not: formulas recovered from TypeSafe's examples, four checkpoints read, openjev's first reliability diagram, and a training recipe.

tgrep: what a trigram index actually buys you
Reads tgrep's trigram planner to show exactly where it falls back to a full scan, and computes the break-even nobody published: 3 to 200 queries, or never.

Modular's LLM Inference Handbook: right formula, wrong variable
Checks Modular's handbook formula by formula: its KV calculator uses query heads (3x to 57x high) and its acceptance formula takes the wrong alpha.

Stop looking, just guess: recomputing a GeoGuessr RL run from its raw data
Recomputes a GRPO GeoGuessr run from its raw JSON: contamination over 690,400 pairs, the r=-0.75, the advantage multiplier, plus six measurement bugs to avoid.

jev-semgrep: grep, but the pattern is a proposition
Reads jev-semgrep's matcher to show two of three presets break classical logic, prices a pass at 11 cents and 83 s, and maps where it belongs in Next.js 16.

Decision models in llama.cpp: what /v1/systemone actually evaluates
Builds and runs llama.cpp's /v1/systemone on a CPU: one pass per question, state shared only up to the slot count, and silent option truncation on Julia-1.

Lexing on the GPU: the scan is the easy part
Rebuilds Pareas's lexer generator: one string rule is 97.6% of the merge table, a block comment exhausts memory, and the simdjson win omits the upload.

Splash on Apple silicon: 144 tok/s needs 502 GB/s, and Apple sells two M5 Max chips
Computes the 16.73 GB a Splash decode step reads from its source, then shows 144 tok/s is a 40-core M5 Max number that the 32-core chip cannot reach.

Qwen3.8-2B-Distill: the filter wrote most of the headline
Shows 69% of a distill's GSM8K gain is lm-eval's last-number filter, with real regexes running on the page, and sizes its GGUFs and KV cache from the files.

WanGP and YuE2: VRAM is a policy, not a property
Reads WanGP's offloader and YuE2's pipeline from code and headers: '6 GB' is a two-block window moving 30 GiB per pass, '24 GB' hides an off-by-default flag.

Google's TabFM: 98.8% of the weights never see a cell
Counts TabFM's weights (98.8% never see a cell), shows 'one pass' is 32, and recomputes from Google's parquet a median 0.5% ensemble gain on binary tasks.

vLLM serving Qwen3.8-2.4T: a Pareto frontier is not a deployment
Re-derives vLLM's 759 MiB per request, shows 5,000 tok/s is mostly prompt (556 generated) on a different machine, and lays out the six-step PD tuning method.

SGLang: the tree, and the language nobody remembers
Source read of SGLang: the radix cache's locking, jump-forward decoding present but never called, and the paper's LPM scheduler off by default.

A 200 GB MoE on one RTX 3090: the experts live in RAM, the speed lives in the CPU
Derives routed bytes per token from the configs, ports the repo's PCIe-or-CPU split, and shows the codebook quant's CPU arithmetic, not DDR4, sets the speed.

SAM 3.1: the 7× is at 128 objects
Reads Object Multiplex in SAM 3.1's code, fits Meta's curve to show 7x is a slope at 128 objects (5.2x from the method), and prices API against self-hosting.

gdp-ts: making the authorization check a value the compiler can see
Type-checks all 60 ways to call one gdp-ts-guarded function: three compile, two are forgeries only strict lint catches; plus per-call-site check cost.

Few-step Qwen-Image-2.1: two DMD LoRAs, and a speedup that is steps alone
Two DMD LoRAs for Qwen-Image-2.1 read from headers: the speedup is steps alone, 6.3x is a 2K corner case, and a fix to load Pruna's adapter in sd.cpp on CPU.
Essential- Original, source-checked analysis
- A guide you can follow today
31 minImage & video generation

Edge0 streams a 35B MoE off SSD. I counted the bytes it can't stream.
Rebuilds Edge0's memory budget from safetensors headers: the 2.9 GiB claim holds, the checkpoint size does not, and five decode rates contradict each other.

MiMo-V2.6: a trillion parameters, thirty RL steps, and a 20× that is a decoder
Sums all 130 shards to 1.02T in MXFP4, tallies the card's own table against "on par" (3 of 14 ahead), and shows UltraSpeed's drafter caps at 8x.

Qwen3.8-27B variants: the bill is tokens times bytes
Puts three Qwen3.8-27B releases on one tokens-times-bytes axis, read from headers: 37.2% is 22.8% pooled, 4.61x is 4.45x, and the MLX 4-bit is 6.04 bits.
Essential- Original, source-checked analysis
- A guide you can follow today
20 minQuantization & compression

Spark-X2.5-4B: fifteen minutes with the files behind a viral model
A reusable 15-minute vetting pass run on a viral 4B model: headers, weight cosines, remote code, KV cost of 1M tokens, and a #1 claim that is #149 today.

Ego2Act and EgoTools: two egocentric benchmarks that grade state, not pixels
Recomputes Ego2Act's leaderboard and judge agreement from released tables: its 0.69 is Pearson against a 0.76 human ceiling, not the 0.84 quoted.

A $80 paper rewriter: what the fine-tune bought, and what the judge decided
Checks a $80 paper-rewriter fine-tune from its files: weight deltas confirm LoRA vs full, the 4B loses to Terra 6-7, and swapping the judge flips the table.

Depth Anything 3 in ROS 2: the metres come from the wrapper, not the model
Traces the one line that makes a ROS 2 depth node metric, reads the ONNX graph to show the model can't, and finds a 13% silent error on 4:3 cameras.

Uno: the draft model was already inside the target model
Shows Uno's lossless claim is the textbook rejection rule on a same-backbone LoRA drafter, that no table supports 3x, and the shipped loss drops the KL term.

Soup: an 8B model in 3.6 GB of VRAM, and 216 commands around it
Rebuilds Soup's 3.600 GB layer-streaming budget from the architecture, walks the silent-gradient bug and its fix, and counts the 216-command surface.

AuK: 6.12 GB, an empty Hub config, and a parameter count that only balances in FP32
Reads AuK's safetensors headers to settle 1.53B fp32 params and 5.7B resident, explains Flash's 4.5x, and traces a 7.5 GiB fix to the same fp32 cast.

Marigold V2: a 20B model wearing a v1 badge
Byte-checks all nine LoRA checkpoints, finds 815 MB of unread embeddings, and shows the 16-26% headline is two of five columns (average 12.1%).

QuantUI-rs: exact to the last bit, and a default file stock ComfyUI won't open
Checks every QuantUI-rs encoder against llama.cpp, comfy-kitchen and ModelOpt: exact bytes, but the default outputs are not what the README promises.

Matryoshka Attribution: one training run, a circuit at every size
A live 12-part toy runs four attribution methods and MAttr training; the leaderboard read shows the 2.9x is CPR only, and MAttr is eighth on CMD.

RelateAnything: no object labels, and one gate that was told the answer
Label-free relation prediction against a swappable text bank, checked in the checkpoints: the 'unsupervised' gate was warm-started on a spatial flag, AUC 1.0.

Multi-harness RL: training a model inside Claude Code, Codex and OpenCode without touching them
Reads OpenEnv's capture proxy and recomputes the guide from its data: recovers the tool-call bonus that drives all-correct groups, and a mislabelled SFT number.

openai/math: 372 claimed theorems, 26 million lines of Lean, and what the green tick means
Reads the whole release: what Comparator and the 26M-line Lean library certify, which headline claims are checked, and all 372 families graded.

Agent Memory Repo: Devin's dreaming memory is a git repo of one-line bullets
Reads the whole Agent Memory Repo spec and runs six concurrency cases in real git: appends conflict, and a stale session can quietly bring back a deleted fact.

Open-1B: what auditable training actually audits
Runs Gensyn's audit kit and joins its ledger to the 80,958-link hash chain: 27/27 land, 0/810 segments confirmed, and the SFT model sits outside the chain.

SGLang's NVFP4 KV cache: the 56% is exact, the 1.78x is untested
Confirms SGLang's 56.25% names a block size of 16, shows its own table never tests the 1.78x, and puts noise bands on every "lossless" accuracy gap.

pplx-decider v1.1: the causal mask comes off in 16 layers, and your state is billed once per question
Weight diffs, the code and a token recount show what v1.1 changed (a noncausal mask in 16 layers) and what it costs: the state billed per question.

Qwen3.8-Flash-Next: four changes and an honest report
Four architecture changes with the report's ablations, where loss and accuracy disagree, then how five local builds treat the 51B n-gram table, checked in code.

BTL-4: reading a model card against its own weights
Audits BTL-4 without downloading it: the base claim holds, the LiveCodeBench table cannot add up, and range-read weights show a merged LoRA. A reusable check.

AgentJev-0.6B: exact order invariance, and options that still interact
Runs AgentJev on CPU: 0 flips over 2,520 permutations yet set-dependent logits, plus a parameter recount, the tokens-vs-clock gap and the reliability curve.

GTSAM 4.3: factor graphs from first principles, and what three years of commits changed
Factor graphs to iSAM2 with an in-page pose-graph solver matched to GTSAM, then 4.3 counted from git, the C++17 port shown and the wheel timed on 10,000 poses.

OpenShell: give the agent a real shell, keep the boundary in the kernel
Reads OpenShell's crates: Landlock for files, seccomp notification to a supervisor that injects credentials, and which policy domains the Z3 prover covers.

kimi-k3-in-c: 2.78 trillion parameters, one CPU, 8 GB of RAM
Reads a C engine running 2.78T-param Kimi K3 in 8 GB of RAM: MXFP4 read packed, a fixed summation order for bit-identical output, a confessed cache bug.

CPU Performance Engineering: a reading list that brings its own benchmarks
Walks a CPU-performance repo's 14 committed M4 Pro benchmarks (mispredicts, cache ladder, roofline, GEMM), then applies its roofs to llama.cpp batch-one decode.

TileLang: you write the dataflow, the compiler owns the schedule
Reads TileLang's layout-inference and pipeline passes, prices tiles per GPU, and finds the pipeline annotation is a no-op on MI300X.

LLM Compressor 0.14: GPTQ's 15x speedup is one fused loop
Derives GPTQ, checks it in numpy, and shows the 15x is per-column launch overhead fused away, measured on 2,048 random calibration tokens.

Qwen-Image-2.1 Outpaint LoRA: 1.12% of the weights, trained to leave the picture where it was
Explains why an editor drifts when outpainting and what the LoRA fixes, then runs it on four CPU cores: registered to the pixel at 35.64 dB versus 15.53.

Tinfield 1: 1.85% of a model, and a 51-billion-parameter hash table
Byte-diffs Tinfield 1 against its base (1.85% changed), decodes Qwen's 51B hashed n-gram engram table, and uses the score lattice to prove k=5 is mean-of-5.

GTA V in a browser tab: Rockstar's own engine, recompiled to wasm64 and WebGPU
The only full teardown of the GTA V browser port: every shipped file read, the D3D11-to-WebGPU ring and HTTP file system explained, a run measured.

TensorFold and Strata: two one-person inference engines, checked against the memory bus
Checks two solo engines against bytes per token: TensorFold's 120 tok/s needs 3+ tokens per pass, and Strata's 4x over llama.cpp is about 2x like for like.

GLiNER2.5-Decide: the options are tokens too
Counts 436M decision-path weights, not 340M, measures label-order and multi-head answer flips plus 0.178 ECE, and says how to freeze and calibrate it.

ChessLFM: 408 of the 552 Elo are not in the weights
Rebuilds ChessLFM's action space, runs the 4-bit ONNX unmasked (99.65% legal), shows 408 of 552 Elo come from search, and finds a demo that drops search.

Needle 3: one checkpoint, five models, and a file that weighs more than the page says
Traces Needle 3's sliceable ladder to its training report, measures the file at 35.34 MB against a 29 MB claim, and narrows "passes DeepSeek" to one benchmark.

TeleOCR: the warped page is never unwarped
How TeleOCR reads warped photos via polygons and masks without unwarping, a header count of its 1.26B, and proof its OmniDocBench lead is all table structure.

DART: the 15.8 FPS is a 4-class number wearing an 80-class AP
How DART shares SAM3's backbone across classes without retraining, and why its 15.8 FPS is a 4-class number: the repo's own 80-class run is 225 ms per image.

GLiNER2.5: deleting the width axis
Explains boundary scoring and joint decoding, then recomputes the appendix: the 0.3B parity rests on XNLI alone, while the 0.2B gain is broad.

Octop: the accounts are separated by a row check, the sandbox is opt-in
Reads Octop's source to show accounts are split only by an API row check: by default agents, terminal and desktop share the host user, and the jail is opt-in.

Husky's 4.5× over MLX: the method is published, and it is speculative decoding
Shows from Husky's own table that a batch-1 engine can't beat MLX past 1.40x; the 4.5x is speculative decoding against an MLX run with no draft model.

Laya on Apple silicon: 1.39×, not 50×, and K is a literal again
Reads two Apple-silicon Laya ports: '50x vs Jev' is never claimed, 60/s is the failing row, 1 GB is one cell, and the option count is frozen at export.

PortSimEnv: real Barcelona berths, a proven optimum, and a 3D port that never touches the score
Re-implements PortSimEnv's floor and reward to match every stored score, and shows from 98 transcripts that open models hit the 32k cap, not the turn limit.

FLUX 3 Action: 2,720 tokens of imagined video for every 32 robot actions
Counts FLUX 3 Action's tensors (12.6B loaded, not 7B) and shows the RoboLab 42.92% is the best of six seeds and the promised video frames are never returned.

Rigel: a hybrid Mamba-2 MoE trained on whichever chips were free
Recounts Rigel from its header and reproduces 5.36e21 FLOPs to the digit, then shows under 1% needs Llama-3.1-8B's run and 'a few points' drops three tasks.

A 320B MoE on a free Kaggle TPU: I audited the bytes, not the tok/s
Audits a 320B MoE on a free Kaggle TPU from GGUF bytes: 110 GB holds, IQ2_S sits 25% from bf16, and the decode kernel never dedups experts.

Maple-Preview: a 20B reasoning model where 97% of the weights are −1, 0, or +1
Proves ternary weights from range-requested bytes, reconciles 5.31 GB at 1.65 bits, and shows the 40 GB release and its code do no quantization.

jevgrep: Jev does the reading, and the 40% leaves its bill out
Reads jevgrep's walk and five result files: the 40% is the best of five runs, the repo's own gate rejected it, and Jev's bill takes back a third of the saving.

auto-gpu-kernel: an agent won the DSA kernel track unattended, and the baseline copies 624 MB per call
Reads the contest baseline and harness behind the 34.93x: a 624 MB torch.cat per call, 69 of 128 traces needing no top-k, and a barrier that hangs past T=9.

cua-s1-forms: a System One model in 706,048 parameters
Three open System One rebuilds side by side: the two readout families as code, a measured 27.8% flip rate under option reversal, and who published ECE.

Needle Environments: the 0.9 is a gate, not a score
Shows the "90%+ held-out" is a 29/32 gate an always-refuse model nearly satisfies, finds the suite crashes on tuned weights, and extracts eight schema rules.

Prime Inference: serving GLM-5.3 to agents is a KV-cache problem
Rebuilds the 352-byte NVFP4 MLA row, explains why TP copies an MLA cache per rank, and finds the 66-session headline is 2,400 tok/s, not 6,666.

SOG: a Gaussian splat scene as spatially-ordered WebP images
Takes a real 5M-splat SOG apart byte by byte: 248 to 13.64 B per splat, most from quantization, and a Morton-order fingerprint in the position bytes.

figures4papers: a house style, audited against its own figures
Runs all 24 plotting scripts and audits the house style: a truncated axis draws 1.64x as 10x, red/green fails deuteranopia, and the skill's API does not exist.

Google's ax agent runtime: the single writer is a Go map
Reads and runs google/ax: the single writer is a per-process Go map deployed as three replicas, and one interruption ran a turn three times with 14 effects.

pmndrs/math and TypeGPU: out-parameters on the CPU, value types on the GPU
Reads pmndrs/math and TypeGPU at fixed commits, measures allocations and bundle sizes, runs the library live, and shows FABRIK missing 272 of 1,368 IK targets.

Runtime dynamic compression: 1.5 bits per weight, counting only the experts in memory
Explains sensitivity probing and dynamic REAP, digitizes the plots to put bnb2's lead at 0.11 to 0.88 bits not 0.4 to 1, and opens the PyPI wheel.

Nemotron 3 Diarization: who spoke when, sorted by who spoke first
Reads the diarizer's header, config and source: 99.2M params, arrival-ordered slots, a benchmark win on training data, and an install that cannot load it.

MiMo-V2.6 RL environments: 7,780 tasks, five graders, and two judges you bring yourself
Counts every file in Xiaomi's RL kit and reads all five graders: 4,688 environments, judges you supply, and a music reward that never reads the brief.

Qwen3-30B-A3B paints in p5.brush: five judges were worth one opinion
Turns a painting-RL post-mortem into a number (five correlated judges were 1.1 opinions) and recomputes the judge-vs-HPS split from the open reproduction.

BrowserSkill: the agent has to borrow your tab to click it, not to read it
Builds BrowserSkill and runs 19 commands on an unborrowed tab: the bypass flags are dead, but snapshot, get-html and screenshot read your tab with no prompt.

DeepSeek Harness: an agent harness that refuses to send what it didn't log
Reads dsh at a pinned commit for the one idea worth copying: every request is checked byte-for-byte against a replay of the session log, plus the rules.

H-JEPA: the top of the stack forgets the ant's legs, and that is the point
How H-JEPA's levels, SIGReg and top-down planner work, with the two gains split apart, and where the paper's own tables and code disagree.

Tiny browser models: six weeks of parameters, one week of shipping
Six sub-250K-parameter WebGPU models read and tested live, including one that parses gibberish as a 95.8%-confident schedule, plus the per-call cost case.

MiniCPM5-2B: how 2 KV heads pay for a 131K window
Derives the 131K KV cache from config.json to match a 3060 report, plans VRAM by quant and context, and measures the WebGPU build and GGUF job string.

Liquid time constants and gated delta rules: two literatures, one recurrence
Derives LTC's fused solver as a gated linear recurrence, verifies LTCAttention's factorization, and prices its gain against its 12% slowdown.

Claude's 104 product demos: two clocks, a phoneme-timed mouth and one ffmpeg join
Rebuilds a Claude-made demo pipeline from its frames, Playwright's recorder and Kokoro's real phoneme timings, against a working film engine that skips them.

Qwen-Image-2.1 Pocket rewriter: 12x smaller than the 9B, and it made all four test images less faithful to the prompt
Checks the pocket rewriters against their files: the image judge picked slot A 240 of 240, training was 70% stickers, and on CPU every rewrite hurt fidelity.

Jev is not deterministic, and option order is not noise
Recomputes a third party's raw Jev logs: identical requests change 67% of vectors, and the published 13% order flip rate is 12% order plus 5% repeat noise.

Ming-Image-0.1-Design: a Z-Image fine-tune that is #1 of 42 open models and 16th of 129
Shard headers, hashes and a weight fingerprint show the "6B" is a Z-Image base DiT in a 24.9B pipeline; puts the #1 arena rank's denominators back.

XGrammar structural tags: a grammar guarantees it parses, a menu guarantees it was on the menu
Separates what a grammar guarantees (it parses) from what a menu guarantees (it was offered), using XGrammar's BFCL bars, and explains its nine wire formats.

Pipette: on-device speed is not a number
Reads Liquid's on-device dataset and checks it: a one-parameter KV model reproduces the 4x Granite gap, and a bandwidth model misses MoE speed both ways.

Cloudflare OS: approve the agent's writes after it has finished
Reads Cloudflare OS's kernel and gatekeepers: capability bindings, a no-network sandbox, writes simulated and approved later, and where the agent waits.

FAST-LIO2 from scratch: LiDAR-inertial odometry you can actually reproduce
FAST-LIO2 derived step by step with a 390-line ROS-free Python reimplementation that tracks a real Livox bag, plus the three bugs it took to get there.

Code2Skill: a million skills mined from GitHub, verified by a round trip, not a test
Counts Code2Skill's released bank: verified means two models agreeing, at least 7,399 of 19,769 repos appear, and the trajectory baselines lose to no skills.

Meta Rebalancer: one expression graph, read as a MIP and as a fast local search
Rebalancer from source: each graph node prices its own move and writes its own MIP, SQUARES is x^1.1, plus a local-search toy and a when-to-use table.

153 autonomous runs, no new ideas: the nanoGPT speedrun frontier
Reads Prime Intellect's 153-run speedrun as a statistics exam: the rulebook's noise math, a planted wrong number 62 runs checked, and a harness gap.

Breeze TTS 2: a Gemma encoder, a Qwen3 backbone, and 1.18 GiB that never runs
Reads Breeze TTS 2's config and headers: unnamed Gemma, Qwen3 and Qwen codec parts, 1.18 GiB that never runs, a TTFA clock that skips the encoder, a cropped #1.

Kev, and the price of not being able to see the other options
Counts Kev's trained params and pools 259 permutation blocks: order flips fall 15% to 1.8% with scale, isolation costs 5.8 points, and Jev still moves 0.02.

Penjing-27B: reading a quantisation claim out of the GGUF header
Parses Penjing's GGUF headers: the "per-layer sensitivity" is a fixed first-and-last rule, real bits per weight exceed every name, K/V demotion saves 0.15%.

Jev alternatives, week two: the numbers arrive, and most sit on public items
Seven decision models checked against JevBench: a public lead flips on sealed items, two releases print their losses, and a DSPy recipe pins the answer slot.

Altar-1: a security model measured on recall alone
Recounts Altar's weights to 5.24 bits and its CVE grid to show the 1-point gap is one run, and that precision is never measured.

Decision models in three-tier agents: the middle tier is not in the middle
Argues decision models are a type, not a middle tier, reads two codebases that build every relation in code, and measures the 193.6x speed claim at 4.57x.

Jev-Omni: a 256-way classifier with no encoders
Header sums show Jev-Omni's four modalities are 52M params of stock Gemma 4, its head welds N=256 into weights, and 'on par with Jev' is a 2.9-point loss.

Agr: a decision model is Gemma's own output layer, read at one position per question
Byte diffs show Agr is Gemma plus a merged adapter; the control the launch chart omits puts the fine-tune at +0.82, almost none of it in Tools.

JEPA-Anything: one recipe for seven worlds, not one model
Measures the 36 released checkpoints: one recipe, not one model; projectors far less orthogonal than the paper's table; gains over a matched JEPA mostly small.

treg: a people-search router priced by expected cost per hit
Reads treg's router and recomputes its launch benchmark: the routing is sound, the accuracy chart is mislabelled and the licence bars hosting.

Ovis-Embedding: Qwen2.5-Omni's last token as one index for text, images, video and audio
Recomputes MMEB-v3 from leaderboard files: 58.46 holds but an uncited 9B scores higher, and the 3B keeps 1.14B unused speech weights on a research-only base.

Cactus Whistle: speech-to-text in 16.9 MB, and a FLEURS bar the Whisper paper doesn't support
Reads Whistle's 16.9 MB checkpoint (a hidden Conformer conv, cross-attention gates stuck at 1) and shows its Whisper FLEURS bar contradicts the Whisper paper.

Kimi K3's TPU megakernel: 92 layers in one Pallas call, and what 709 tokens a second measures
Reads the Pallas megakernel: 709 tok/s uses a forced acceptance length, and by bytes per token the GB200 has the higher ceiling, so the win is software.

LensVLM-9B: the 10.1× is how hard it squeezes, not how much it wins by
Reads LensVLM's paper, code and weights: 10.1x is a compression ratio where it ties the best retriever, and the frozen projector matches Qwen3.5-9B-Base.

Where to use Jev: the test is whether one option can read another
A four-gate procedure for placing a decision model, built from 18 measured teardowns, with the cost multipliers deflated and five places not to use one.

Any model can be Jev, for the price of a serving flag
Reads SGLang's /v1/score handler to show the decision readout is a serving flag, refits the 63x as a token-count ratio, and checks who measured calibration.

FlashAttention-3: the kernel is mostly a schedule
Reads FA3's Hopper kernel as a schedule: warp roles, pingpong overlap and register budgets that sum to exactly 65,536, plus why FP8 is a separate kernel.

OpenHuman: the harness is lighter, and it is faster for a reason the posts don't give
Re-aggregates OpenHuman's benchmark logs on tasks all three harnesses solved: model time is equal; its lead is start-up and fewer calls, and it solves fewer.

ZDTaichu5.0-9B: the release contradicts itself on its own best benchmark
Shows ZDTaichu's two READMEs differ by 10.4 points on MMSI-Bench (a prompt addendum), its NVFP4 is 8.01 bits/weight, and its entropy gate lacks parts.

jev-linkmap: the rubric rewrite was mostly a threshold
Recomputes jev-linkmap's run files: 679 links is a 28% yes rate through a cap, the $0.27 is $17.91 all-in, and the rubric gain is mostly a threshold.

Jev on WindTunnel's WebMCP board: the 245x is 15x of interface and 16x of price
Factors WindTunnel's 245x into 15x interface and 16x price from its own board, finds Mercury is 78% of the bill, and traces the 49th task to a 79 ms wait.

Tiel-Coder-35B-A3B: the 4-bit tier is 5.16 bits, and the only new weight is a prompt
Reconciles nine GGUF tiers to the byte: Q4_K_XL is 5.16 bpw, A3B checks out, a VRAM fit ladder, and the only change from the base is a 29K-char Jinja prompt.

Breaking the softmax bottleneck: your output layer is a rank-d wall
The rank argument and the MoS-vs-MoC control from first principles, plus |V|/d read from current model configs: the bound is tighter now than in 2017.

Qwen-Audio-3.1: five audio APIs, and which of them the papers actually describe
Shows the 3.1 TTS blog is the 3.0 report relabelled, down to radar vertices, checks token prices against the advertised cuts, and reads the one real 3.1 report.

Voxel Musou: the army is a 300-slot pool, and the combo is a frame table
Reads a voxel Musou browser game end to end: '300 soldiers' is a pool with ~84 fighting, determinism is per-engine, and the combo is a checkable frame table.

Jev as a substrate: every pixel is a Choice
Reads three odd Jev projects: a stroke gate that inverts the painter's confidence rule, a robot loop 88% queueing, and 'Jesse' patches byte-equal to gold.

Jev + Kimi K3 fraud detection: 96 of 100, on a 50/50 inbox
Rebuilds the confusion matrix from a fraud demo's frames: 0 of 50 false alarms bounds FPR only at 5.8%, and the review spend leaves five tokens.

Leviathan: the flat 450 tokens is a card cap, and the 99% comes from a friendly benchmark
Reads Leviathan's Rust and benchmark: BM25 in SQLite FTS5, a flat 450 tokens set by card caps, and a 99% from questions that share the records' words.

Fish Audio S2: a Qwen3-4B on the time axis, four layers on the codebook axis, and tags that are only text
Recounts S2 from file headers: the "lightweight" fast AR reads as many weights per frame as the 4B, tags are plain text, the win rate used rewritten prompts.

Limite 1B - Violetto: the throughput claim is in config.json, not the blog
Violetto's config, then its report: a 3.91x KV saving, the full 1B maths recipe, and a headline chart that omits the teacher it was distilled from.

The Hilti SLAM datasets: grading LiDAR SLAM against a steel tip on a surveyed cross
How survey-mark ground truth and the banded score work, plus an independent check: the dense references score 7 to 76 mm against the marks.

Land or Water?: reading a world map out of 16,200 one-word answers
Digitizes the Claude 'Land or Water?' maps and rescores them: errors pile onto the coast, 17 are lakes the key calls land, and 648 points give the same score.

OpenMuse: "compatible with any agent harness" is one ternary branch
Shows OpenMuse's "any agent harness" is a five-line AG-UI swap that drops all 29 tools and their cards, while the container check and approval gate stay behind.

UltraData-Code: counting a 1.2 TB corpus without downloading it
Counts a 1.2 TB code corpus via the datasets-server API: L3's 150B is one column (content is 52.4B), ALGO is 26.6%, and the licence contradicts itself.

Audio8: eight checkpoints, one broken link, and a 3B model with nowhere to ship
Enumerates all 12 Audio8 repos by API and safetensors header: names undercount the shared encoder, the 3B has no edge build, and one ONNX link goes elsewhere.

Five Jev harnesses, for a model that cannot write
Reads five decision-model harnesses side by side: all build model-free enumerators, none agree on the gate, and one-call-per-step costs every string.

Muon: orthogonalizing the update for hidden layers
Muon from the SVD up: why an orthogonalised step helps, how the tuned Newton-Schulz quintic behaves, the full update on the page, and the two fixes for scale.

K2-Horizon-MoVA-36B: routing the value vectors, and why the KV cache does not notice
Reads MoVA from the modeling code: routed values are summed before the cache, so KV stays GQA-sized, and the cost moves to 7.5B params of value experts.

OpenDLSS-NR: the network is a sequence of roundings
Reads the DLSS 5 neural-rendering port: a same-resolution Swin U-net whose parity is an order of roundings, what "same output" covers, and what 7.8 ms times.

Qwen3Guard-Stream: a safety verdict on every token
Reads Qwen3Guard-Stream's headers and code: 1.06M head params on a 0.6B guard, how per-token labels are made, where it cuts, and an eval script swapping modes.

Hy4 preview: frontier of five, not frontier of seven
Sums all 131 shard headers to confirm 770B/49B, shows IndexCache deletes indexer weights, and finds the GGUF's accuracy table has no primary source.

Bonsai 2 27B: PrismML's sequel gets to 98.2% — on average
Audits Bonsai 2's file manifest and two benchmark suites: 5.95 GB is real, 98.2% hides 75% retention on agentic coding, and stock llama.cpp won't load it.

Pocket TTS with drifting: a one-step speech head without the Jacobian
Derives the drifting field, reads the loss Kyutai shipped (not the blog's), and uses a toy built on that code to show why kernel temperature must be learned.

LOOM and Looped-DiT: looping a model more than twice
Checks LOOM's FLOP accounting row by row: the 5-loop iso-FLOP win holds, the 9-loop win costs ~9x compute and 3 optimizer steps per batch. Plus Looped-DiT.

LoopCD: a looped transformer's first loop is its own amateur
LoopCD as contrastive decoding with the first loop as the amateur, why disagreement beats accuracy for the reference, and why the 73% and 48% are separate runs.

mini-AGI: what the 169 experts and the 0.0067 nats actually are
Reads mini-AGI's log and config: a character sees 32 experts not 169, the 0.0067-nat forgetting is below its own noise floor, and the probe code is absent.

PrunaSuperPoint: a 2.1x faster keypoint front-end, pruned where the time actually is
Counts SuperPoint's MACs layer by layer to show why pruning the first convs pays, proves the top-k exact, and explains why 3.3x less compute gives 2x speed.

The fly connectome is a substrate, not a model
Sorts two weeks of fly-connectome projects by what they train, and shows wiring identity matters only when a model is checked against named neurons.

Qwen3.8, weights in hand: 98% of a 2.4T model is routed experts
Rebuilds Qwen3.8's 2.4T and 95B from the config, reads what FP8 spares, and turns the vLLM and GGUF findings into serving and quant choices.

The male fruit fly connectome: how a whole nervous system was mapped
How the MaleCNS pipeline works stage by stage, what "complete" means (40.1% of connections), the v0.9/v1.0 numbers, and how to query the data.

Kimi K3: a 2.8T open model that turns compute into intelligence 2.5× better
First-principles tour of K3's KDA, Attention Residuals and 16-of-896 LatentMoE, with the shape checked against the released config.

AI-written GPU kernels: PTXBench grades them, and an agent gamed the grade with the constant 0.1778209953
PTXBench's three-gate grading, recomputed KDA geomeans, and the hard-coded 0.1778 constant derived by hand, ending in a six-gate recipe for kernel verifiers.

EVA: a VAE that samples well once its prior predicts the next latent
Why one Gaussian per step suffices in a learned latent, EVA's prior heads counted in the weights, and three paper-vs-release mismatches.

RRSI: regularize the harness search, not the harness
Maps RRSI's seven search constraints onto L0/L1/L2, checks noise bands and cost against repo configs, and notes its LAB score is not LAB's own all-pass metric.

SigLIP 2 from first principles, and what its Core ML port measured
Builds SigLIP 2 from the sigmoid loss up, then checks a Core ML port claim by claim: an fp32 MPS baseline, 2.00x is just fp16, the text tower is 75% of bytes.

GLM-5.3-Flash: 45 layers, 11 of them expensive
Reads GLM-5.3-Flash's config for its 34 linear and 11 sparse layers and IndexPool, rescores Z.ai's table row by row, and lists the Hopper-only serving limits.

MimiModel: a 45M LLM on a $5 chip, and 20 points that lived in the decode loop
How a 45M model runs from flash on a $5 ESP32 via the Hadamard identity, and a row-by-row account of 20.8 points lost in the decode loop.

CUDA Rust: writing the kernel, not just launching it
Scopes CUDA Rust's aliasing-safety claim from both repos: cuda-oxide checks argument aliasing but not shared-memory races; cutile-rs deletes the thread.

Laya vs Jev: the slower model gets an easier question
Runs the T-Rex harness's planner offline: each model's prompt is built from its own latency, so the two get the same question on only 13.9% of frames.

Whittle MoE 27B: the routers moved 4.5 degrees
Checks a dense-to-MoE carve against its files: 66.3% active, routers 6.2% of what trained and 4.5 degrees from init, and a smaller pruned sibling scores higher.

Interference Search: the 23 out of 30 has no language model in it
Enumerates Countdown to show the 23 of 30 has no language model in it, the budget exceeds the state space, and merging only pays at six numbers.

Retrieve-for-Train: the 20x is in the prose, not the figure
RL as a compiler for set-level rewards, Figure 5's SVG read to put the speedup at 8x to 16x not 20x, and unpublished human ratings that split the result.

Beacon: Jev grades your agent runs, and the lesson is not in the skill
Reads Beacon's learning loop at a pinned commit: a mean gate that passes a zero on reuse, a projection that drops exit codes, and skills with no lesson in them.

Shaders goes MIT: the components are the menu, the compiler is the meal
Reads the Shaders engine to show how a layer tree compiles into GPU passes, what still triggers a recompile, and what the MIT switch does and does not open.

Qwen3.8-LiveTranslate: half a second off the lag, and 2.7 points on
Derives LAAL, puts the lag cut beside the quality gain, and shows the cited speaker benchmark's public release cannot produce the published numbers.

Contrastive Language Model (CLM): the action never sees the state, and 81.6% is a pick from four
Counts CLM's 18.9M trained parameters and re-scores its DeepSWE verifier: 81.6% is 10 of 13 decidable tasks against 7 for chance; the 9x is a frame clock.

Cactus Needle 3: the free fine-tune deletes the confidence head
Reads Needle 3's confidence head from the weights and the package that deletes it on local fine-tune, and sizes a 4-bit tuned archive from a byte-exact rebuild.

Interfaze 1 Lite: the mixture of architectures is ten public checkpoints and a tool loop
Hash-matches 76 of 79 weight files to public checkpoints, reads the OCR stitcher, and shows word confidences in the open code are just the line score.

CPLM in the NanoGPT speedrun: a pointer head takes 11% off the record
Re-derives CPLM's 11.25% and p = 0.0005 from the PR's 21 logs, explains the copy-sink pointer from the code, and ranks it against every prior record.

Julia-1: one score per mask token, and a benchmark that turns on how the labels are worded
Re-runs Julia-1's ONNX export and reproduces the card exactly, then shows the pilots overlap its train splits and Banking77 swings 62 to 91 on label wording.

Jev alternatives, week four: the week decision models learned where to run
Recomputes d1's launch chart from its own JSON (exact, but 'four of six' only at rounding) and reads llama.cpp's six decision readouts from source.

tiktok-5.6B-videos: 460 GB you can query without downloading
Reads all 148 Parquet footers of a 460 GB TikTok scrape: real row count, what a query downloads, where pruning fails, a timestamp bug, and the GDPR problem.

Jina-OCR-v1: 570M active parameters assumes an embedding tie that isn't in the checkpoint
Safetensors headers put active params near 740M, not 570M (untied embeddings), and the paper's own table shows FastMTP K=3 is the worst depth under CUDA graphs.

Hunyuan-A13B: an 80B MoE that reads 13B per token, and a switch for thinking
Recounts 80.39B params from shard headers, shows enable_thinking=False silently does nothing on main, and reads the territory-limited licence.

VoiceMem: the reply model is a QLoRA that can't touch 93% of its own base
Explains VoiceMem's schema-bounded candidate pool, shows its QLoRA reaches no routed expert, and that arXiv's abstract says 4.29 where the PDF says 1.89.

Spirula Studio deleted its dependencies instead of wrapping them
Checks whether a one-binary splat pipeline wraps or rewrites COLMAP, PyTorch and ffmpeg (it rewrites), why BA stays fp64, and how the new LiDAR path works.

Thinking in Blender: which half of the 3D scaffolding gets eaten
Sorts one harness's mechanisms into substitute, supply and constraint, and argues constraints get repriced by better models rather than eaten.

FreeVideo: a 66 GB video transformer on an 8 GB card, and the planner that decides where every block lives
Header reads show FreeVideo fits MiniMax H3 in 8 GB by precomputing 26 GB of AdaLN weights and streaming FP8 blocks; Video DeltaNet buys speed, not memory.

KOLC+ puts Gaussian splats inside BIM/CIM: how each feature has to work, and how I would build one
Works out what splat volumes, sections and alignment must compute inside a closed BIM tool, then lays out an open, licence-checked pipeline to build one.

json-render + Jev: generative UI in two model calls
Runs json-render's composer with a recording evaluator: any screen is two calls, the catalog costs 669 bytes per recipe, and confidence is never read.

OpenCut: the 92K-star repo is a rewrite, and the editor you can use is the archive
Separates the 92K-star rewrite scaffold from the archived editor that actually works, explains its integer-tick renderer and tests an export.

Strata in the source: one verify window, three workers and a cache that learns
Reads Strata's doorbell, three-way miss split and decayed-LFU cache, then rebuilds the 3090's 97 tok/s: a 0.70-0.86 hit rate, with PCIe, not RAM, the slow lane.

Functional gradient descent: refine the gradient until the step is safe
Rebuilds adaptive FGD in 1-D so you can watch it refine, works the theorem's constants, and checks each 'outperforms neural nets' against the figures.

One-shot launch videos: what the skill decides, what the model writes
Counts what sits under two "one-shot" launch videos: 1,462 lines of skill and 260 sounds for /brag; 4,819 lines by sub-agents over a force-aligned song.

Desert Ant Labs: real wins, missing metrics
Checks seventeen on-device model cards against their configs: most comparisons hold, two headline ones have no quality score, and the energy math misses.

ALoDLM: a diffusion LM that loops on its hard tokens, read against its own tables
The inner loop read from the decoder source, Table 1 re-averaged to show where the AR margin lives, and the speed-up re-priced in FLOPs and KV cache.

Linear attention's memory problem: four answers in one week
Puts four new linear-attention designs in one regression frame, with per-head state arithmetic, a measured recall toy and config checks.

GLM-5.3-Flash on four mining cards: the hard part is an Ampere kernel
Reads the vLLM fork that adds sm_80 sparse-MLA kernels, shows why upstream gates Ampere out, and labels every speed and quality number as one person's self-run.

Restore and relight LoRAs: teaching a generator to start from your picture
How an IC-LoRA puts a clean reference in the generator's own sequence, two LoRAs sized from their bytes, and demos showing both redraw what the input lacks.

Flash-dLLM: a KV cache for keys that never sit still, and a model that checks its own drafts
Explains Flash-dLLM's fused KV refresh and self-verification; finds the code's verify rule stricter than the paper's and a figure that contradicts a table.

Qwen3.8-27B on a 12 GB card: a hashed trellis and a ternary with a draft head
Decodes a real Mirai S trellis packet, byte-checks the GGUF port and plans VRAM for 27B at 128K in 12 GB; the 40 tok/s came from a capped 3090.

SurfSLAM: sim-to-real stereo and a DVL-driven factor graph for mapping shipwrecks
Shows from the tables that SurfSLAM's stereo gain is the water column and its tracking is the DVL, and where the released code departs from the paper.

Carveout: SAM 3 labels for a Gaussian-splat scene, without the multiplex
Reads Carveout's SAM 3 splat-labelling pipeline: why multiplex can't speed detection, how the lift votes across views, and its per-class VRAM ceiling.

Qwen-Drive-1.0: the driving VLM that doesn't touch the VLM
Checks in the code that the VLM is untouched, recounts 5.6B for planning against the 4B name, and shows RL's wins become a trade in the one closed-loop test.

RWKV-7 G1k: what fits in a 16-million-number mind
Derives RWKV-7's delta-rule state update, confirms the 16M state from checkpoint shapes, sizes it against a Qwen3 KV cache, and tests the safety claim.

DiffusionOPSD: a target, not just a reward
Reads DiffusionOPSD's own code for the exact update and its dir_mode ablation, then restricts the 19-of-20 claim to the seven evaluators anyone can rerun.

Gaussian-splat capture in practice: from a 360 camera to a scene on a map
Five practitioners' 360-to-splat pipelines checked against tool source: why ERP for SfM, the SH2 VRAM arithmetic, and how sort-free stochastic splats work.

LiFT: a looped DiT that keeps improving past the loops it was trained with
The LiFT target derived as flow matching along depth, its FLOP accounting re-derived, and the 52% headline traced to a 250-step dense baseline.

Mach-2 Additive Medium: 1.70 bits, a trellis, and a 102 GB table the bits leave out
Decodes Mach-2's packed format: a QTIP-style trellis, not additive codebooks; 1.70 bits only without a 102 GB BF16 table; one benchmark gap above noise.

WorldCrafter: skipping the reconstruction, not the 3D
Reads WorldCrafter's code and headers: coverage-based frame retrieval, a 1.14B view-synthesis encoder, and memory as four latents posed where the camera goes.

Overmind: the 28x fewer invented quotes did not come from anyone's agent traces
A source read of Overmind's trace-to-fine-tune loop: how the split, judges and trainer work, where holdouts leak, and why the 28x quote claim can't be checked.

PhoneLLM Alpha 1: the fine-tune is free, the cost line isn't
Recounts the 3.58B active params from shard headers and finds the card's cost example off by 10x; the gain is all fine-tune, at identical cost and latency.

Index-Translate: one base, four translators — text, speech, dubbing, long docs
Shows the 0.8789 FLORES score covers 21 languages, not 150, and that the dubbing head's syllable control costs 0.072 quality: keep SFT for subtitles.

Cinference runs Qwen3.8-27B at 450 tokens a second: what the number counts
Unpacks Cinference's 450 tok/s from its own results file: one user, MTP-10 on a recall task, 71 tok/s without speculation, and a two-minute prefill at 256K.

PixelUMM and PixelDense: diffusion without the VAE, and what the pixels need instead
Why pixel diffusion predicts clean patches, PixelUMM's 15.2B recounted from code, and why PixelDense's GenEval gains sit within noise while its probes do not.

Surflo: one latent for a scene, a surface at any resolution
Explains Surflo's fixed latent and per-point flow decoder, and shows guidance only looks harmful on DL3DV because its ground truth is a Gaussian-Wrapping mesh.

Gaussian-splat colour: 192 bytes of spherical harmonics, or 28 bytes and a tiny MLP
Each 3DGS colour model's maths and CUDA, bytes recounted from code, the M360 table split (SH still wins outdoors), and what it means for a capture pipeline.

Muse Glimmer: an agentic model designed backwards from a 24 GB budget
Closes Muse Glimmer's 24 GB claim from config and file sizes (21.6 GB at 131K), shows sliding windows make it fit, and re-tallies a table Meta loses a third of.

Holo4: open computer-use weights, graded on the lab's own benchmarks
Which Holo4 checkpoint you may self-host commercially, and that post-training adds +0.9 on OSWorld but +13.7 on OSWorld 2.0 over the Qwen base.

Probe monitors: reading activations, not the scratchpad
Checks a vendor's probe pitch against OpenAI's report and the primary papers, works the probe cost from GLM-4.5-Air's config, and shows the cascade's cost knob.

NextLat: a transformer graded on its own next hidden state
Explains NextLat's predict-your-next-hidden-state loss and checks its tables: MTP is close on rank; sequence compression is the real world-model evidence.

Olmo-core 3: an open training stack for trillion-parameter MoE
Explains the MoE capacity tax and Olmo-core 3's resident experts, set against Aleph Alpha's FSDP-only 30B recipe: the parallelism follows the model size.

LeVJEPA: one encoder, zero collapse-prevention machinery, and what 5.6-20.8x actually measures
Checks LeVJEPA's five headline numbers against its tables and released ViT-L: the compute range is honest, SSv2 is a real loss, and block-causal ships.

Prime Agent: the interface is a Python REPL, not a tool-call schema
Reads Prime Agent's source: one IPython tool, 21 typed host requests mostly gated by session flags, a base prompt /refine cannot edit, and its real git history.

Lattice Deduction Transformers: reasoning by narrowing a lattice, not emitting tokens
The LDT lattice and solve loop from first principles, read against the code, the authors' own 99.96% correction and a critique calling it one-shot prediction.

Reproducing Jev, from a source I could not find
Tests five unattributed Jev claims on public data: open models that beat it at home fall 88 items behind on a shared suite, and untrained decoders rank high.

Bend 2: proof-checked, and no longer an interaction net
Checks Bend 2's claims apart: a small audited checker with a new consistency argument, a rhetorical Navier-Stokes link, and a runtime without interaction nets.

TencentDB Agent Memory: the four-tier pyramid is really a cache-stability hierarchy
Shows from the source that Tencent's L0-L3 memory tiers are sorted by cache stability, not abstraction, and that the character budget ships switched off.

BM25: the ranking function that refuses to die
BM25 built term by term from TF-IDF, with a live scorer you can query, sliders for k1 and b, and 30 lines of Python that reproduce it.

DFlash 2: the drafter already knew the answer, it just picked the wrong one
Interactive walk-through of why a parallel drafter needs selection rather than depth, and why acceptance length only matters at serving concurrency.

The Skaling law: Chinchilla assumes model size and data don't interact, and they do
Shows why Chinchilla's additive form flips the sign of the compute-optimal trend, and reruns the author's own earlier estimate to find it was 2x too strong.

The Kalman filter from first principles
The Kalman filter as fusing two Gaussians, with a live tracker, runnable NumPy and the path to EKF, iEKF and the error-state filter used in LIO.

TencentARC's GAE: the geometry decoder never sees the picture
Traces GAE's forward path to show the point clouds never read the RGB, counts 3.5B real parameters, and flags that released weights reproduce no table.

Reflection Beam: a 501B open MoE that spends 23B per token, and what its efficiency chart measures
Rebuilds Beam's efficiency chart from its dots: the axis is 2 x active params x tokens, so 1.74x of each ratio is size, and RL outspent pretraining.

Z1T: sparse transformers for a chip that samples, not multiplies — and what 100x actually measures
Reads Extropic's full energy table: "over 100x" is the 10%-MFU row only, both sides are projections, and the omitted logit layer is 463x the total.

FreeToken: a 753B model on one GPU, and the two bandwidths that decide everything
FreeToken's q* = m*B_P/B_H split between PCIe fetches and CPU compute, why it swings 25% to 91% across machines, and how to install and calibrate it.

Token cues: how 'chicken' made a base model reason as well as RL
How two forced tokens match RL-Zero on Olmo, the swap test that puts RL's gain in the opening, the chicken data edit, and where the result stops.

RF-DETR: the first real-time detector past 60 AP, and its six sizes are one training run
Explains RF-DETR's one-run NAS ladder and DINOv2 backbone, and flags that the two sizes past 60 AP are PML-licensed, not Apache.

VoxelTTO spends 93% of a scene fixing the poses
Explains voxel-anchored Gaussians and test-time LoRA fit to given poses, reconciles the runtime (TTO is 93% of it), and shows input views get worse outdoors.

Kandinsky 6.0 Video: two token streams on one clock, and where the lip-sync comes from
Explains joint audio-video denoising from the code (230:1 tokens, RoPE scaled 0.144) and shows the 47% WER cut is graded by its own reward model.

TrackEverything: dense 3D tracking that grows with the scene, not the video
Explains TrackEverything's voxel de-duplication and 3D WAFT, then checks its claims: the 20-point lead is APD-P over one baseline, 10+ FPS excludes geometry.

The Jacobian lens: reading the residual stream with a derivative
The Jacobian lens from the estimator up: why an averaged derivative beats the logit lens, where the tangent fails, and the two-call API to run it.

3D reconstruction roundup 2: old geometry, bolted onto feed-forward models
Six feed-forward 3D papers read against their own tables: PocketSplat's budget saves no compute, OneCanvas's panorama adds nothing, plus a loop toy.

What a decision model cannot do
Sorts decision-model limits into one forced limit, two opposite prices and misfiled ones, and measures openjev's letter prior at +1.71 logits.

3D reconstruction roundup: six papers, and what each headline is measured on
Six 3D papers checked against their tables and code: Mira-Scene's Astra claim has no metric, NG-GS's code lacks its NeRF; a toy shows why highlights float.

A field guide to attention mechanisms
One map of every attention variant by the bill it pays, with ten interactive diagrams and a reach-for-it-when table.

What the Jev ecosystem actually built
A source-read census of 15 projects built on Jev: all are cheap classifiers in front of costly steps, and open-jev's metrics.json gives the first real ECE.

Genex, ffmpeg-skill and rdsh: three harnesses, read for where they check the agent's work
Source reads of Genex, ffmpeg-skill and rdsh for where each checks the agent's work; Genex's self-edit gate is real, rdsh's 81x times a --version print.

Kolibri-1: a 78B open MoE that does 3.46B of work per token and stretches to 1M
Recounts Kolibri-1 to the parameter, works its KV cache at 1M from config, and finds two MoE layers the model ignores and launch claims that hold two of three.

Marin 535B-A23B: the half of a training run nobody usually shows you
Walks Marin 535B's mid-run changelog through its issues and code: the MuonH axes bug left in, the shut attention gate, and what the MFU gain really came from.

Sentry: failure lessons should arrive with the failure
Why always-on failure lessons hurt and how Sentry's detector, label-keyed retrieval and verified writes fix it, with every headline average re-derived.

GLM built its own inference stack: one PR, one bottleneck, and a 3x with no rival
Checks Z.ai's infra-agent case studies against FLA and DeepEP source: the PR and GIL bug are real, the fix opt-in or never upstreamed, the 3x has no baseline.

FastH3 Preview v1: four calls, one-tenth the attention, and what the 14x actually measures
Splits FastH3's 14.38x into distillation and sparsity, and recovers Trim's pruned blocks and its 294 FP4 clipping scales from the released checkpoints.

AnyAPI's Scrape router: a ladder of scrapers, priced by the rung that gets through
Recounts AnyAPI's 100-page scrape benchmark from its CSV, reprices Firecrawl by plan, and replays a router built from the rivals: 95/100 at about $1.23.

Mistral Large 4: first among Western open models by 13 points, eighth overall
Checks Large 4's 'best from US or Europe' against Artificial Analysis records (true; eighth overall) and its charts against its prose; plus a memory planner.

toks: an exact tokenizer built from certified shortcuts
Reads toks end to end: why exact BPE hinges on merge order, the proof behind each shortcut, the asm loop, and where its own table undercuts the launch post.

Jev scores zero, and zero is the informative number
Why 0 of 100 on relational choice is a structural signature, shown in open readout code, plus server-side dates that put Laya after Jev, not before.

arc-driver: the 8x is a one-second wait, the 1.7x is a round trip
Source read of arc-driver and cua-driver: the 8x is cua-driver's one-second window poll, the 1.7x a saved model round trip, with a cost model to test it.

DepthBench: depth is a scaling axis only for the right residual stream
DepthBench's finding, checked against its configs and FLOP formulas: depth is a live scaling axis only for HC and Full AttnRes, and the cheap variants lose it.

RecursiveMAS: a multi-agent system folded into one looped transformer
How frozen agents become layers of one looped model via a 13M-param link, with the +8.3% re-checked against Table 3 and the AIME rows flagged as Pass@10.

Halo's 2.8× over TRL is not a kernel, and its own benchmark proves it
Shows Halo's 2.8x over TRL is ZeRO-2 vs ZeRO-3 (1.25-2.68x matched), not the kernel, and checks the EP routing claim against transformers' own source.

Subquadratic 3SUM and subcubic APSP: two hypotheses fall to a pruned matrix product
The 3SUM/APSP refutation from reduction chain to leaf count, with the identity checked exactly, the practicality arithmetic, and which lower bounds survive.

xHC: sixteen residual streams, four written at a time
Builds HC, mHC and xHC from the residual up with a runnable sublayer, re-derives costs (the 18B column cannot add up), and finds the conv does two-thirds.

Dust: every token is a member of the population
Derives Dust's per-token zeroth-order estimator, runs a toy that shows why per-token rewards carry it, and prices one update at a few thousand backprop passes.

PixiJS 3D and benchy: the fine print under 5.9x, and the agent's pull requests behind it
Decodes the 5.9x footnote, explains the WebGPU timing probe and labs' two-gate verdict from source, and audits the agent-written PixiJS PRs behind the claim.

Jev alternatives, week three: bigger models, smaller benchmarks
Reads seven decision-model releases' heads and tables: none has a sealed JevBench score, JEMM's headline row is a loss, and Jev's latency varies 5x.

XGEN-JING and DAO: the renderer shipped, the world engine didn't
Separates what shipped from what was announced: JING is an offline non-causal MiniMax-H3 derivative, DAO has no code, and its WBench rows are self-evaluated.

Recursive Self-Rewrite: teach the model what the harness did
RSR's plan, critique and re-solve rewrite, plus what the paper skips: no size-matched control, counts that don't reconcile, an untrained harness winning on TBH.

FAR-LIO: LiDAR-inertial odometry that holds together at 250 km/h
Checks FAR-LIO's GPU voxel map, SA-GICP and delay-compensating EKF against the shipped config, and shows its headline averages hide a worse Monza APE.

Naive-N0.5-Flash: no full-attention layers, and a runtime built by AI
Counts the config's 39 sliding and 9 sparse layers, none dense, prices 1M-context key reads, and marks the runtime and "built by AI" claims as reported.

Differential Transformer: softmax can't say zero, so it subtracts
Walks DIFF attention from the paper and code, then reads each result at its denominator: the 65% is a 0.025-nat gap, and the low-bit win is about logits.

Qwen3.8-Omni-Flash: skipping most of the video on purpose
Checks Qwen's omni launch against its own tables: a 45.7% token cut not 51.8%, where 'beats Gemini' comes from, ~1 s to first audio, and a demo path to 6 mm.

SparDA: a fourth projection that lets sparse attention prefetch its own KV cache
Explains the one-layer-ahead Forecast projection and pulls apart its three speedups by baseline; the +6.5 reasoning gain rests on two 30-problem AIME sets.

Context Language Models: the bitter lesson comes for context management
How Context Language Models let an agent rewrite its own context as a file, with the prefix-reuse cost formula re-derived and checked against the paper.

IQuest-Q1: 320B of weights, 25 layers that keep the context
With no report, rebuilds IQuest-Q1 from config.json: params recounted to the shard index, 25 of 88 layers grow KV, a 3.45x cache saving at 512K.

Thomson N=7: ten agents, 17,895 lines of Lean, one trusted kernel
How the SDP-certificate proof of Thomson N=7 splits on the smallest inner product, what two kernels certify, and why the Lean statement itself is unchecked.

FLASepformer and JAEC: the parts Alibaba shipped, and the parts it didn't
Walks the paper's tables from the released 20.7 dB checkpoint to the announced 24.9 dB one, and shows which of the two releases can ship commercially.

OpenSEO's AI Visibility: a scraper, a regex and one answer per prompt
Reads OpenSEO's AI-visibility code end to end: who answers, what counts as a mention, and how much noise one answer per prompt leaves in the trend.

SiamJEPA: shuffle the teacher, and a patch has to say what it is
Explains the shuffled-teacher trick from the training code, flags the README's pre-fix caveat on every number, and finds the core files are CC BY-NC.

Ling-3.0-flash-Fin: a finance finetune, and a benchmark that grades where the numbers came from
Confirms the finance finetune is byte-identical in shape to its base, reconciles 124B from shard headers, and parses FinFIRST's 701 sourcing criteria.

DeepGEMM-Ascend: DeepSeek's kernels off NVIDIA, and what 99.8% and 98% actually measure
Separates DeepGEMM-Ascend's 99.8% (one dense BF16 GEMM) from 98% (the fused MoE layer with comms), and flags DeepEP's missing licence and unreleased firmware.

Qwen-Planner-Agent: a phone agent that plans in tool calls, not taps
A phone planner that acts in tool calls, and its benchmark recomputed: the #1 lead is 0.21 points, inside the 0.41 run-to-run spread its own authors measured.

HauhauCS FastMTP: a signed manifest that actually verifies
Verifies the quant's Ed25519 signatures and every hash, dry-runs its llama.cpp patch, and finds two of five K_P size claims miss the stated band.

Hindsight: +44 points on the same 20B backbone, from the memory layer
Hindsight's four memory networks and four-lane RRF recall, checked against the shipped code, with the judge mismatch and self-reported baselines flagged.

MiniMax Music 3: a five-minute song is 9,000 steps of a 2.5 kbit/s code
Reconciles every parameter count on the card from safetensors headers, shows the 57 GB repo ships the model twice, and derives the 30 fps, 2.5 kbit/s code.

Intern-S2-Mobius: a 35B model that separates memory from reasoning
Reads the modeling code: 40 layers share four expert banks, there is no runtime loop, and the speedup tracks shorter traces and vanishes on math.

Mixture of Experts, from scratch
Builds MoE from one MLP: top-k router, noise, sparse dispatch and load balancing, with a complete 200-line model that trains on a laptop CPU.

GLM-5.3: the same base model, a month of post-training, and a cyber capability nobody planned
Diffs the 5.2 and 5.3 configs, counts active parameters from all 141 shard headers (41.25B, not 40B) and shows Unsloth's "2-bit" figures mix two files.

Projecting out Whisper's hallucinations: abliteration, in reverse
Explains projecting a hallucination subspace out of Whisper's decoder as abliteration in reverse, and reads the WER bill and the VAD baseline that beats it.

Colibri: running a 744B model on a 25 GB machine by streaming experts from disk
How Colibri runs a 744B MoE in 25 GB of RAM by streaming experts from 370 GB of NVMe, why it is disk-bound, and the community speeds from 0.08 to 6.8 tok/s.

KVzip: compress the KV cache once, by scoring it with the model's own reread
Why query-aware KV eviction breaks on reuse, how KVzip scores by having the model reread its context, and how Attention Matching synthesizes a smaller cache.

LFM2.5-DSpark: the best draft model in the release is the slowest one
Why Liquid's best-accepting draft, on the 8B MoE, is slowest on a laptop: verifying a block touches the union of experts. And: pick drafts by acceptance.

J-space in the open: a CKA map of workspace geometry across 38 models
Explains the CKA behind a 38-model J-lens map and reads its numbers honestly: a shared layer order (rho 0.83), but only +0.04 of matched-depth similarity.

Uber's ADR: the agent that watches your agents
How Uber's two-tier agent detector triages cheaply and escalates to a Claude Code subprocess that reads a suspect tool's source, with costs and gaps.

Speculative programmatic tool calling
How sPTC launches tool calls from a half-written REPL cell via a shadow fork with taint rules, and a simulation of why a 1-1.2x gain is so hard to measure.

nac: an orchestrator that is not allowed to touch anything
How nac's orchestrator only dispatches workers whose transcripts are discarded for episodes, the DAG batch rules, and the non-transactional seam Arcee admits.

DCFormer: an ICML oral that let attention heads borrow each other's circuits, and quietly shipped anyway
Explains DCMHA's five-branch Compose and its low-rank decomposition, why static composition equals a wider head, and what two years actually bought it.

One small box, 64 users at once: continuous batching on a DGX Spark
Explains continuous batching against committed spark-bench results: aggregate throughput on one box scales, per-user speed does not, and what is reported.

Level-of-Token Diffusion: a DiT that is told where the detail goes
How LoT turns a layout into rectangle tokens a pretrained DiT can denoise, why images gain less than the token cut and video more, and what eval layouts leak.

Language models are injective, so a KV-cache is not a summary — it's the prompt, in another basis
Walks the real-analyticity proof that hidden states never collide and how SipIt inverts them, and what that means for KV-cache offload and privacy.

Jev, TypeSafe's System One model: the receipts, itemized
Holds Jev's launch claims to TypeSafe's own docs and outside benchmarks: speed mostly holds, 'frontier intelligence' shrinks, RLCD has no published evidence.

Gemma 4: an open multimodal family, tuned to the KV-cache budget
Walks Gemma 4's KV-cache savings (5:1 local/global, values=keys) with a memory table per size, and flags the thinking-vs-non-thinking baseline.

Scaling agentic RL: 365,000 environments behind one contract
How Prime Intellect puts 365,000 agentic tasks behind one gold-validated contract, set beside Hugging Face's OpenEnv watercolour run where the reward is taste.

Cross-model KV transfer: a linear map that skips prefill, four pairs out of six
Explains the closed-form per-head ridge map that lets a model skip prefill, floor-normalises the two failing pairs, and checks which pairs match KV.

Cosmos 3: a world model that reasons and generates in one sequence
How Cosmos 3's two-tower design lets a diffusion generator read the reasoner's keys, plans in pixels via Action-CoT, and where its SOTA is open-only.

Recursive Language Models: context as a variable, recursion as a function call
Separates Zhang's RLM idea from Prime Intellect's two builds, shows prompt-as-variable and rlm() in their real code, and lists the costs of a REPL as interface.

Jev swarms: one call, ten agents, and the answer moves
Argues swarm batching saves the shared state but is not answer-preserving: Score shifts ~0.29 levels when Choice agrees, and coordination lives in code.

Grouped Value Attention: cache the value, reconstruct the key
Derives GVA's key-from-value absorption, sizes the cache in bytes, corrects a quoted 44.18 to 44.35, and finds the released checkpoint is the failed ablation.

Ternary Bonsai 2 27B Uncensored: the numbers behind a runtime abliteration
Checks a runtime refusal ablation in code: 129 hooks, no write to the pack, an alpha=0 path that skips wrapping, and a self-check that would pass any vector.

GEPA: optimize anything you can score and describe
Explains GEPA's reflect-on-traces loop and Pareto parent selection, the optimize_anything interface, and omni's explore-then-continue budget split.

XGEN-JING wired up two more control axes, four days later
Counts JING's exact 33.4B parameters and 302.6M control branch from headers, shows dialogue and audio are prompt text, and control changes every 17 frames.

Streaming 3D from unposed video: R³, AMB3R-SLAM, and the move to bounded memory
How R3 and AMB3R-SLAM turn a frozen Depth Anything 3 into streaming 3D with flat memory, checked against both papers, with R3's licence trap named.

TurboQuant: rotate first, then quantize the KV cache
Why a random rotation lets one data-free quantizer fit every KV vector, how a 1-bit QJL residual removes inner-product bias, and where the idea now ships.

fframes: a frame is a function, and the benchmark is one scene
Reads fframes' compile-time static cache and agent CLI, and shows the 31x over Remotion is one synthetic scene, CPU against CPU, GPU skipped.

NVIDIA's Skill2Env: collective is a distribution claim, and the benchmark disagrees
Pins down what 'collective' means in Skill2Env, explains its Oracle/no-op gates, and puts the rubric run that lost every benchmark at the centre.

SMELT: looping wins even after you pay for the FLOPs
How SMELT matches FLOPs, params and KV cache for a looped MoE, and why the win shrinks to a 6.8-18% compute discount; Nanbeige and FBT held to the same bar.

Qwen-CUA: a computer-use agent that only ever sees pixels
Why folding screenshots ten at a time keeps the KV prefix cacheable, plus what the 86.2 hides: 2 of 8 wins, unmatched budgets, a benchmark its lab wrote.

LongLive-Plug: distill the speedup once, plug it into 54 video models
How LongLive-Plug distills few-step and CFG once into LoRAs that transfer to 54 descendant video models, and why a separate CFG LoRA acts as a guidance dial.

Apodex 1.1 and FrontierAgent: the harness, measured
Uses Apodex's same-weights ReAct vs Agent Team rows to price the harness at about 40% of the 1.0-to-1.1 gain, and reads FrontierAgent's sandbox and run steps.

Virtual logic depth: looping buys reasoning, not knowledge
How the paper measures knowledge (absorbed entropy) and reasoning (iGSM), why looping leaves knowledge flat, and that no comparison is compute-matched.

LFM2.5-VL-3B: the release where GUI grounding appears out of nothing
Recomputes all 28 benchmark rows: a dead tie with two 4.7B models, GUI grounding from 5.4 to 80.7, a CountBench regression, and unspecified on-device numbers.

Macaron-V1: four 1B adapters on a frozen 744B base
Explains Mixture-of-LoRA on a frozen GLM-5.2, isolates what the four 1B adapters add over their base, and flags a base-model judge and best-of-N rows.

Inkling-Small: what on-policy distillation actually buys a reasoning model
Explains why on-policy distillation helps reasoning but not recall, and finds both Inkling models' stated sizes run 2.4-3.8% above their own weights.

GLM-5.3-Flash-MLX: built for MacBook Pro, verified on an H200
Checks the five OrcaSAQ MLX builds against their README: tensor counts sum, and "built for MacBook Pro" fits only the 128 GB tier, tested on an H200.

AIRA: the bottlenecks it named, and an audit that complicates the tease
Traces AIRA2's three ablated fixes back to AIRA1's own limitations, and sets its 5-of-11 contaminated AIRS-Bench wins against AIRA3's unhackable tease.

Speculative decoding on AMD GPUs: verifying the hedge
Rechecks AMD's vLLM spec-decoding appendix: EAGLE-3 loses on Qwen3-8B MATH500 at every N, DFlash spans 1.10-2.87x, and DSpark ran with its confidence head off.

Hunyuan Hy3: Tencent's 295B-A21B MoE, and the community 1M GGUF
Hy3's 295B-A21B layout and KV-cache arithmetic, which GGUF quant fits your RAM, and why the 1M window is a community YaRN stretch of a 256K model.

OREN: real-time Euclidean SDF, because the octree carries the prior
Why OREN's octree with per-vertex gradients makes a quadratic-error prior so a tiny MLP fits only the residual, with its speed, memory and where it is second.

MAGI-2 Preview: 114B parameters, 6B awake, and two sparsities doing the work
Rebuilds 114B/6B from the safetensors shapes: three per-modality weight copies, not MoE alone, get to 5.96B active; finds 4-stream hyper-connections.

Neutrino-1: quantization is a training decision, not a deployment one
Why ternary rounding after training lands at chance on MMLU while training in the format does not, where the bytes go, and which numbers do not line up.

Reading a torch.profiler trace: overhead-bound vs compute-bound
A condensed read of HF's torch.profiler guide: the setup, Self CPU vs Self CUDA as a one-line regime test, and when torch.compile costs more than it saves.

Matryoshka LM Suites: stop training the small models twice
Explains how nested exits share one run, distil for free and make a 1:6 draft pay off, and spots the 3B exit's slowdown with no speculation.

Scaling laws in 2026: the line held, the recipe didn't
From Kaplan's shallow exponents through Chinchilla to serving- and sample-aware laws, explaining why shipped models train at 200 to 79,000 tokens per param.

MiniMax Sparse Attention: let each query pick its own blocks
How MSA's index branch picks 16 blocks per GQA group, why a fixed 2,048-token budget pays off only at long context, and where the converted model regresses.

H-BAC compresses a field ViT 54.5×, and a same-size baseline nearly ties it
Checks from the paper's tables whether prune, distill and quantize compound: 1.2 points short of additive, and a same-size model trained directly nearly ties.

LFM2.5-Encoder: classification in one forward pass, zero completion tokens
How LFM2.5 decoders become bidirectional encoders, read from the cookbook code: mean-pool, one Linear, sigmoid per label, and an honest 4th-of-14 ranking.

DeepSeek DSpark: making speculative decoding draft better and verify smarter
How DSpark's semi-autoregressive drafter and load-aware verifier lift accepted length, and why its production speedup is a boundary case, not a multiplier.

GLM 5.2: long-horizon coding at a million tokens
How IndexShare reuses one sparse-attention indexer across four layers to make 1M context affordable, with Z.ai's benchmarks quoted as reported.

SAIL: a robot that searches 45 trajectories in simulation before it moves once
Explains SAIL's trajectory-level MCTS and video progress score, then splits its 25-to-73% curve: sampling against an oracle gives 26 points, the method 14.

MegaTrain: training a 120B model on one GPU by inverting where memory lives
The host-memory-first layout and three-stream double buffer behind 120B on one H200, with the regime where its speedups do and do not hold.

Rollout Routing Replay: stabilizing MoE reinforcement learning
Derives why MoE routers flip between rollout and training engines and how replaying the rollout mask fixes it; run R3 and drop TIS, per the paper.

Adapt-1 Machina: no critic, because the simulator is the critic
Recounts all 384 per-case outcomes and re-derives the interval: the simulator replaces the critic, at 18,400 executions for +6.25 points.

Dream-RSI: the loop rewrites a scheduler, not the model
Shows what Dream-RSI actually recurses on (about 150 lines of scheduler, no gradients), checks its matched baseline on nine comparisons, and prices the name.

DeepSeek DSec: three million agent sandboxes a day, 50× overcommitted
What DeepSeek's sandbox platform does to make 50x CPU overcommit safe and images load on demand, plus its production catalogue of agents attacking the box.

MiMo-V2-Flash: a 128-token window and one global layer in six
Shows why five 128-token sliding-window layers per global layer cut the KV cache about 6x at 256K, with MiMo-V2-Flash's reported scores.

How LLM inference works: prefill, decode, and where the time goes
Prefill vs decode from arithmetic intensity, with widgets computed from real model shapes and a rule for telling which phase is slowing you down.

KDA has a half-life: linear attention forgets like a radioactive isotope
Shows KDA's per-channel decay is radioactive decay: alpha 0.99 is a 69-token half-life, and K3's gate floor and per-head A_log are read from config and code.

Switch Transformers: route every token to exactly one expert
A clear walk through top-1 routing, the capacity buffer and dropped tokens, the balancing loss and fp32 router, with the 7x and 1.6T read precisely.

GOAT: entropic optimal transport, minus the transport
Shows GOAT is one-sided OT (softmax with a learned log-prior), not Sinkhorn, with a Sinkhorn stepper and the paper's modest, small-scale results read closely.

The generated web: Qwen3.8-27B on Cerebras at 2,000 tok/s, and no ground truth
Explains why the demo's tok/s counters disagree, maps page size to latency thresholds, and prices a generated page at ~4,600x a CDN-served one.

ZCode, open-sourced: what a remediation can and cannot establish
Sorts Z.ai's ZCode remediation evidence by what each piece can prove: the open tree shows the client now, not then, and the fix is a local git checkpoint.

Ornith-1.5: the model writes the exam, builds the marking scheme, then sits it
Explains Ornith's loop that rewards its own tasks, scaffolds and rollouts, reads the uneven generation gains, and names what would tell skill from harness fit.

LongStraw: fitting million-token RL onto a fixed GPU budget
How LongStraw fits 2M-token GRPO steps on eight H20s by capturing the prompt as resident state and replaying one response at a time, and what it does not claim.

Cactus Needle 2: the interesting number is what happens after you fine-tune it
Needle 2's 45M attention-only design and 2-bit training, and why tuning on a dozen tools beats a zero-shot frontier model, with the lift table still missing.

Instella-MoE: a 16B MoE that never touches an NVIDIA GPU
Explains Gated MLA and FarSkip's deliberately stale activations, and audits AMD's 'fully open' claim: RAIL weights, no compute or cluster size.

Nemotron in NVFP4: training a frontier model natively in 4-bit
What NVFP4's two-level scale buys, the three stabilizers that make 4-bit training GEMMs converge, and why it is mixed precision with a loss-gap claim.

Flex-π: a robot policy that decides how much to think at deployment
How one shared latent space makes Flex-pi's inputs and outputs runtime flags, and why 11 of 20 eight-stage repairs is 93% per stage against 69%.

Solar Open 2: Upstage's 250B-A15B hybrid-attention MoE
Walks Solar Open 2's 1-softmax-in-4 hybrid stack with the KV-cache arithmetic done by hand, the 2.3% weight transfer, serving commands and the name-tax licence.

LongCat 2.0: a 1.6T open-weights MoE, and the sparse attention behind it
Explains LongCat Sparse Attention's three parts (token-level index, shared cross-layer index, streaming reads) against MSA; benchmarks are vendor-run.

SenseNova-U1.5: no vision encoder, no VAE, and pixels that come back anyway
What dropping both the vision encoder and the VAE costs (1.09 dB of reconstruction), what is still unpublished, and why 28 steps is enough for edits.

SANA-Video 2.0: keeping video attention linear without losing the picture
Why a 3:1 linear-to-softmax stack with block attention residuals keeps 720p video cheap, and how the 120x splits into architecture and Sol-Engine gains.

The harness effect: orchestration, not the model, sets your agent token bill
The effective-input-price equation and six harness mechanisms behind a vendor's 41% cost cut, with the small-sample and self-evaluation caveats spelled out.

CSFM: flow matching never had to start from Gaussian noise
Why a learned, condition-dependent source cuts flow matching's intrinsic variance, with an exact toy of path crossings and a check of CSFM's tables.

Sol-Attn: deciding which attention blocks to skip while you're already streaming them
Explains the mean-plus-beta-sigma threshold inside the online-softmax loop, and subtracts the stacked speedups: Sol-Attn alone is 1.92x and 1.34x, not 5.08x.

Gigatoken: tokenizing at gigabytes per second
Shows why regex pretokenization, not BPE merges, was the bottleneck, with per-CPU throughput and the one-line command to benchmark your own tokenizer.

TwoTower: giving a diffusion LM a frozen autoregressive memory
Why TwoTower splits a diffusion LM into a frozen AR context tower and a trained denoiser, with its 98.7% / 2.42x headline and the code-and-math caveats.

How small can a verifier be? Six hundred thousand parameters
Walks a 19-model ladder showing verifiers switch on near 1M params then plateau because all fail the same items, with the shortcuts and odd rows called out.

Nous Hermes and Mixture-of-Agents: when models confer before they answer
Explains Mixture-of-Agents proposers, aggregators and collaborativeness, untangles what Nous actually shipped on Hermes, and flags the unverifiable 2026 claims.

Bonsai 27B: a 27B model at 1.125 bits, small enough for a phone
What end-to-end ternary and 1-bit quantisation of a 27B model buys and costs: 3.9 GB on a phone, math intact, agentic and vision down about a fifth.

Qwen's RL formulation: token-level RL is a first-order approximation to the reward you want
Why token-level RL is the first-order part of the sequence objective, and when to use R2 or R3 Routing Replay for MoE, from Qwen's stabilization paper.

speech-to-speech: the OpenAI Realtime API, reimplemented as four swappable parts
How HF's voice pipeline passes the real OpenAI Realtime client tests, the fully local llama-server setup, and Smart Turn's revision-numbered speculation.

Laguna's Model Factory: treating model development as an industrial process
Walks Poolside's 'Model Factory': versioned data and runs, the AutoMixer surrogate for data mixes, commit-derived RL tasks, and INT4/INT8 split by layer.

Set Diffusion: one knob from autoregression to diffusion
Shows AR, block diffusion and order-agnostic diffusion as set schedules of one model, with a window-width widget and the paper's 110M-scale results.

Ultra-FineWeb: what an education filter costs you, measured
Reads Ultra-FineWeb's per-benchmark table to show what an education filter trades away, and why the staged L1/L2/L3 release matters.

Pokee-Isaac 28B: 10M tokens on one GPU, and an architecture the report never explains
Reads Pokee's report against its own tables: the architecture is never explained, and Isaac trails two named cloud models on real-tool benchmarks.

Chimera: unbundling RoPE into a diffusion Transformer that extrapolates 6× on video
How Chimera reassigns RoPE's three jobs to convolution, KDA's forget gate and layout, why that buys 6x video extrapolation, and why its 7.3x is 6.8x.

MrFlow: climb the resolution in pixel space, not diffusion steps
Explains MrFlow's low-res steps, pixel-space SR and one refine step, and separates the 4-9x training-free speedups from the 25x that needs distillation.

Model casting: compute the gate, then skip most of the FFN
Derives the 3x ceiling on gate-first FFN skipping, shows how LoPA's low-rank gate lifts it, and flags that the 3.31x is FFN-only and single-stream.

GLM-5.3-Flash-Uncensored-FP8: the numbers behind an abliteration
Checks a gated abliterated GLM-5.3-Flash from Hub metadata alone: weights byte-identical in shape to the base, and the refusal numbers are self-reported.

EAGLE-3: making the draft model scale, and a from-scratch build
A short explainer of EAGLE-3's Training-Time Test and feature fusion, its acceptance scaling law, and a from-scratch draft-model trainer on Qwen3-8B.

Explorative Modeling: factor the training loop, not the generation loop
Explains why best-of-K training buys the expressivity diffusion gets from steps, with a K slider; notes the headline-result code is not released yet.

GRAPE: RoPE, ALiBi, and FoX are the same construction
Walks GRAPE's group-action view in which RoPE is a rotation and ALiBi and FoX are shears, and notes the gains are about a point at 353M-770M only.

HydraHead: hybrid attention at the head, not the layer
Why HydraHead hybridizes per head, not per layer, how causal patching picks the retrieval heads, and which long-context results come only from a figure.

iLLaDA: how far a masked-diffusion language model scales
How masked-diffusion LMs train and unmask, with widgets for the masking objective; iLLaDA matches Qwen2.5 as a base model but trails by 10 points after SFT.

Dream-Cubed: diffusion directly on Minecraft's block IDs, where inpainting comes free
Why masked discrete diffusion on Minecraft block IDs gets exact inpainting free where a continuous DDPM cannot, and why the paper's own numbers tie.

Fireworks' distributed RL: frontier RL is cheaper than you think
Separates Fireworks' sales pitch from the weight-sparsity physics, re-derives the 94% delta-traffic cut, and lists when co-location still wins.

ABot-World-0: a 5B world model that wins on efficiency, not the leaderboard
ABot-World-0's distill-then-LongForcing recipe and the five systems fixes behind 16 fps on one GPU, with its second place on WorldRoamBench shown plainly.

Towards looped models done right: what actually separates Ouro from Huginn
Walks IFM's matched-depth ablation one axis at a time: an untied prelude/coda and persistent input injection do the work; random state init barely helps.

dots3-note Preview: 16B active parameters, and a critic that thinks before it scores
A short read of dots3-note: TEMPO lets the critic spend inference on long tasks, plus two benchmarks nobody scores above 34% on, and the release's caveats.

Nanbeige4.2-3B: looping a small model up to a big one's depth
Explains Nanbeige's two-pass looped transformer trained from scratch and its agentic post-training, and sorts which of its benchmark wins are comparable.

A.X-K2: a sparse-attention upgrade that costs nothing, trained natively in FP8
How A.X-K2 adds a top-k indexer to gated MLA at near-zero LongBench cost and trains natively in FP8, with its weak BrowseComp score and the TIS failure kept in.

Antidoom: breaking doom loops with Final Token Preference Optimization
Why small reasoning models loop and how FTPO retrains only the loop-restart token, with Liquid's recipe, stop rule and the low-temperature tradeoff.

Coroutines in C, intuitively
Tatham's switch-and-__LINE__ coroutine trick explained step by step, with the macro expansion, its sharp edges, and a reentrant context-struct version.

Tapered Language Models: spend your width where the work is
The cosine MLP-width taper at fixed params, the early/middle/late control that isolates direction, and modest but consistent gains on four architectures.

Program-as-Weights: compiling a natural-language spec into a tiny local model
How PAW turns a plain-language spec into a LoRA through a mapper over learned basis matrices, with the paper's figures and its limits; results quoted.

Ternary15M: a language model where every weight is −1, 0, or +1
A readable tour of ternary QAT on a 15M TinyStories model: signed accumulation, the STE in five lines, and why the FP32 embedding is most of the 43 MB.

pdf-inspector: classifying PDFs without a single model
The three heuristics pdf-inspector uses to tell text PDFs from scanned ones without a model, and how narrowly it leads its own non-ML benchmark.

JOSIE-2: a 4M-token fine-tune, and a labeling bug in its own benchmark table
Checks JOSIE-2's 'beats its larger base' claim against config.json: every base is same-size and the benchmark tables mislabel the baseline.

Recursive Harness Self-Improvement: beat your last harness, not a population of them
The Bradley-Terry case for comparing a harness only to its last version, plus the gaps: Opus judges Opus and no rival method is benchmarked.

A6B: k-expansion, or what breaks when you force a top-8 MoE router to fire 32 experts
Why widening a top-8 MoE router to top-32 hurts (renormalisation hands 54% of gate mass to untrained experts), and how far router-frozen deltas heal it.

The full-bandwidth transformer: the feedback channel is one token wide
Explains latent feedback as widening a 16.6-bit token bus, and why instruction tuning on old traces erased its shorter-reasoning gain.

Qwen3.8-Max: 16 days, 265 commits, zero humans in the loop
Lays out Qwen3.8-Max's five unsupervised case studies and its benchmark tables with the footnoted harness and judge asymmetries; weights not yet out.

SWE-1.7: near-frontier code RL, and the async training loop behind it
First-principles walk through Cognition's async RL, top-p sampling-distribution replay, delta weight sync and self-compaction; all numbers self-reported.

Ring-Zero: what a trillion-parameter model learns from reward alone
The four-stage zero-RL pipeline on a 1T MoE, the engine-ratio fix with its equations, and the scope caveats: ablations at 104B and nothing released.

KAT-Coder-V2.5: training a coding model to live inside a repository
Walks KAT-Coder's environment and data stack: AutoBuilder sandboxes, hint-then-strip rollouts, harness randomization and a critic that sees the future.

Sakana Fugu: a multi-agent system as a model
Puts TRINITY's evolved 20K-parameter coordinator beside the Conductor's RL-written workflows, the two papers behind Sakana's Fugu API; numbers quoted.

WorldClaw: a 3D world generator that is really a Blender programmer
Explains WorldClaw's write-a-Blender-program pipeline and region-blended height field, and notes it needs Opus, GPT-Image-2 and Hunyuan3D and has no numbers.

Nar TTS: the two rewards it built and refuses to switch on
Reads a small TTS repo for its evaluation discipline: two rewards built and held at zero so the judge is never the model it trained against.

Multi-token prediction: training a model to see further than one step
MTP from the loss up: parallel heads versus DeepSeek's sequential modules, why acceptance caps the decode speedup, and who actually invented what.

ZUNA 1.1: a channel-agnostic EEG foundation model
How ZUNA 1.1 treats each EEG electrode as a token at an (x, y, z, t) coordinate with a rectified-flow decoder, and where it beats spline interpolation.

LOTUS: reasoning in the hidden states, not the token stream
How LOTUS loops a padded latent region R times and supervises each slot against its gold CoT token, reaching explicit-CoT accuracy at 3B with a fixed budget.

Motif 2.6B: differential attention and PolyNorm, trained at scale
Explains differential attention, PolyNorm and Motif's linear data-mix schedule, with the report's code and math scores against 7-8B models.

Leanstral: proving theorems by being a code agent, not a prover
Explains how Leanstral proves Lean theorems with a plain code-agent loop, why accuracy scales with token budget, and how SafeVerify blocks reward hacks.

Agents-A1: scaling the agent horizon, not the parameter count
Walks through Agents-A1's knowledge-action graph and three-stage routed on-policy distillation; the benchmark numbers are the tech report's, unchecked.

Inkling: an open-weights multimodal MoE built to be adapted
Lays out Inkling's config (5:1 local/global, 256+2 experts, encoder-free inputs) and the effort-control claim; every number is vendor-reported.

VideoChat3: a 4B video model that watches longer for less
Explains VideoChat3's inflated 3D ViT (16x token compression) and state-driven frame resolution, with its 4B-vs-4B benchmark gains and 2048-frame latency.

Ling-3.0-flash: a 124B open MoE that runs like a 5B and reaches for 1M tokens
A first-principles tour of Ling-3.0-flash's 5:1 KDA and Gated-MLA stack and 8-of-512 MoE, and how close it is to Kimi K3; benchmarks are vendor-run.

Fara 1.5: an open, vision-only browser agent at 4B, 9B, and 27B
Places Fara 1.5's 4B/9B/27B ladder against three closed browser agents: only the 27B clears all three, and every number is Microsoft's own.

TabFM: a foundation model that learns tables in-context
An early overview of TabFM's in-context approach from the model card, with schematic widgets and the ensemble-vs-ensemble caveat on its TabArena Elo.

Mach-Mind-4-Flash: specialize, then integrate
Walks through MOPD routed distillation and HMPO's median-length budget, and marks where the 100B-class framing trails Kimi-K2.5.

MemHarness: agent memory should be reconstructed, not replayed
Explains MemHarness's critique-then-rewrite memory step and the ablation where raw replay loses to no memory, with an honest list of what is unmeasured.

TurboVec: FAISS-competitive vector search with no training phase
TurboVec explained as TurboQuant applied to vector search: no train() step, the six-step encode, its SIMD scan, and where it beats or trails FAISS IndexPQ.

AngelSpec: specialize the drafter, share the verification budget
How AngelSpec pairs per-workload drafters (MTP, DFly) with D-cut, a batch-wide verification budget, and what its live-traffic H20 numbers do and don't show.

MusaCoder: teaching a model to write GPU kernels with execution-feedback RL
Explains MusaCoder's correctness-first kernel reward and its three RL stabilizers, with the caveat that the benchmark and protocol are the authors' own.

Toast 1: what happens when you stop making the frontier model do the searching
Reads Mixedbread's launch numbers for what they show: on Harvey's benchmark the score holds at 55 while tokens fall 3.5x, so most were spent looking.
Niche6 minAgents & harnesses

Taimi-14B-Med: reading a model card properly
Reads a medical model card whose v0.1.0 is rebranded Qwen2.5-14B weights, splits its benchmark numbers, and keeps the 16 GB serving config worth copying.

oruk.ai's fly-brain speech classifier: real wiring, scrambled wiring, same score
Shows oruk.ai's fly-brain speech classifier is an echo state network whose scrambled-wiring control ties it, and why its lesion test does not rescue the claim.

How self-attention works in transformers
A short first explanation of scaled dot-product attention, stepping through scores, softmax and the value mix on a three-token toy matrix.

Soofi S: a sovereign 3B-active model that keeps its cache near-constant
A careful summary of the Soofi S report: why 6 of 52 attention layers keep decode flat, the German-weighted mix, and which claims are active-vs-total.

Unconventional AI's Un-0: generating images with coupled oscillators
Kuramoto oscillators from first principles with live widgets, Un-0's evolve-then-decode pipeline and FID ladder, and the line between simulation and chip.

FLUX 3: when an image model decides to become a world model
Flow matching and joint multimodal attention explained from first principles around BFL's FLUX 3 announcement, with every result flagged as vendor-reported.

zvec: an in-process vector database, and the ANN search inside it
Explains HNSW descent and quantize-then-refine behind zvec's embedded vector search, and why its 8,475 QPS is not a like-for-like comparison.

Diffusing blame: credit assignment under Dale's principle
What Dale's principle and no weight transport rule out, how Error Diffusion learns anyway, and why its MNIST and RL numbers are a possibility proof.

BTL-3: a rank-32 LoRA that turns Qwen3.6-27B into a tool-use agent
A read of BTL-3's thin card: a rank-32 LoRA on Qwen3.6-27B for tool use, strong on irrelevance, weak on parallel-multiple, all numbers self-reported.

MiniMax H3: open weights, four excluded countries, zero benchmarks
Quotes the H3 licence's four excluded jurisdictions and revenue gate, notes two of three modules are API-only, and that no benchmark was published.
Niche4 minImage & video generation

The harness is the generalizer: how a scaffold learns to solve longer, unseen tasks
A readable digest of Zhang and Khattab's argument that a harness keeping calls locally in-distribution generalizes to longer tasks; little beyond the post.

Audex: audio, speech, and text through one decoder — without the text tax
How Audex reads audio as embeddings and writes it as extra vocabulary in one decoder, and why text-domain RL keeps its text scores, per the paper's tables.

Mage-Flow: a 4B image model that bets on its tokenizer
A short tour of Mage-Flow's cheap tokenizer, bucket-free MMDiT and 4-step Turbo, with the paper's memory and latency numbers.

Laguna S 2.1: an 8B-active model that won't give up
A summary of Laguna S 2.1's weight-class benchmarks, thinking-mode gap and published trajectories; all numbers are Poolside's own.

Intern-S2: a 397B model that reads the raw page
A summary of Intern-S2's raw-page pretraining, science tokenizer and model-card benchmarks; the wins are vendor and partly internal benchmarks.

Antares: a 1B model that hunts vulnerabilities like a person
A summary of Cisco's Antares: a sub-1B terminal agent trained with SFT and GRPO to localise vulnerable files, with the launch's cost and F1 numbers.

Lanyon: proving a PDE solver correct before you run it
Why proving a PDE solver under IEEE-754 rather than real arithmetic matters, and a skeptic's read of Lanyon's self-graded 20-250x token claims.

Arcee's Open Models API: a model lab selling six models it did not build
Arcee's API post read for its one strategy sentence and its price list: output multipliers from 2x to 5x re-rank the catalog for output-heavy agent work.

NVIDIA Rubin: co-designing a GPU for the shape of agentic inference
A map of NVIDIA's Rubin features against agentic-inference bottlenecks, from the vendor's own post, with the 10x and 40% kept as vendor claims.

Agent harnesses: engineering the loop around the model
A readable digest of Lilian Weng's harness post with two illustrative widgets; little here that is not in the source.

Qwen Audio 3.0 TTS: an instructable LM-plus-flow-matching speech stack
A short tour of Qwen-Audio-3.0-TTS's LM plus flow-matching stack and its 12.5 Hz tokenizer, with the project page's radar charts; no numbers checked.

AutoCompact: teaching an agent to decide when to forget
Explains AutoCompact's learned compaction trigger and notes there is no paper, code or results table to check it against.
Niche5 minAgents & harnesses

LIMSSR: scoring actions when modalities go missing at training time
Explains LIMSSR's idea of naming missing modalities in the prompt instead of reconstructing them, with the paper's FS1000 numbers; nothing is checked.

MAI-Image-2.5-Pro and MAI-Voice-2-Flash: Microsoft builds its own
Reads Microsoft's two MAI preview models as a supply-chain move; prices and vendor-reported cost cuts, no architecture to check.
Niche7 minImage & video generation

Qwen-Image-3.0: chasing “useful” instead of “good-looking”
A summary of Qwen-Image-3.0's announced capabilities using its own examples; no architecture or benchmark exists to check.

Monolith 1.0: a 1.6T open MoE built for reasoning
A record of a retracted release: generic MoE routing, staged YaRN and self-speculative decoding, explained for a model whose weights never existed as described.
Niche10 minLLM architecture