# Satyajit Ghana — full content corpus
> Head of Engineering @ Inkers Technology. I build deep-learning systems, 3D perception, and high-performance infra.
> Generated from the content layer. Index: /llms.txt
---
# About Satyajit Ghana
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/about
I build deep-learning systems, 3D perception, and high-performance infra.
Head of Engineering at Inkers, where I work on industrial AI — 3D perception, deep learning, and the systems that ship it to production.
I write custom neural networks, CUDA kernels, LiDAR/point-cloud pipelines, and high-performance C++/gRPC services. Previously taught MLOps (EMLO 2.0) and computer vision (EVA 4.0) at The School of AI.
## Links
- github: https://github.com/satyajitghana
- linkedin: https://www.linkedin.com/in/satyajitghana/
- x: https://x.com/thesudoer_
- medium: https://satyajitghana.medium.com/
- scholar: https://scholar.google.com/citations?user=rZCRakQAAAAJ&hl=en
- website: https://thesatyajit.com/
- email: satyajitghana7@gmail.com
---
# Resume — Satyajit Ghana
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/resume
> pdf: https://ai.thesatyajit.com/satyajit-ghana-resume.pdf
> json: https://ai.thesatyajit.com/resume.json
```json
{
"contact": {
"name": "Satyajit Ghana",
"title": "Head of Engineering",
"company": {
"name": "Inkers Technology",
"url": "https://inkers.ai"
},
"location": "Bengaluru, India",
"email": "satyajitghana7@gmail.com",
"github": "https://github.com/satyajitghana",
"linkedin": "https://www.linkedin.com/in/satyajitghana/",
"website": "https://ai.thesatyajit.com"
},
"summary": "Engineering leader and deep-learning systems builder. I lead engineering for industrial-AI products — 3D perception, LiDAR/point-cloud pipelines, and structural-defect analysis — and ship them as high-performance C++/CUDA/gRPC services. Two USPTO patents pending. Previously taught MLOps and computer vision.",
"experience": [
{
"organization": "Inkers Technology",
"url": "https://inkers.ai",
"location": "Bengaluru, India",
"roles": [
{
"title": "Head of Engineering",
"start": "2024-01"
},
{
"title": "Deep Learning Software Engineer",
"start": "2022-06",
"end": "2024-01"
},
{
"title": "Deep Learning Associate",
"start": "2021-07",
"end": "2022-06"
}
],
"summary": "Lead engineering for industrial-AI products across 3D perception, LiDAR/point-cloud pipelines, and structural-defect analysis.",
"highlights": [
"Lead engineering for industrial-AI products: 3D perception, LiDAR/point-cloud pipelines, and structural-defect analysis.",
"Design custom neural networks, CUDA kernels, and high-performance C++/gRPC services for production deployment.",
"Named inventor on 2 USPTO patents pending (assigned to Inkers); grew from Deep Learning Associate to Head of Engineering."
]
},
{
"organization": "The School of A.I.",
"location": "Remote",
"roles": [
{
"title": "MLOps Instructor",
"start": "2022-01",
"end": "2023-01"
}
],
"summary": "Designed and taught MLOps and contributed to computer-vision curriculum.",
"highlights": [
"Designed and taught EMLO 2.0 — an end-to-end MLOps course (training, packaging, deployment, monitoring).",
"Contributed to EVA 4.0, the deep computer-vision program."
]
}
],
"education": [
{
"institution": "M.S. Ramaiah University of Applied Sciences",
"location": "Bangalore, India",
"degree": "B.Tech",
"end": "2021",
"highlights": [
"CGPA 9.78/10",
"Silver Medalist"
]
}
],
"skills": [
{
"name": "Deep Learning",
"skills": [
"PyTorch",
"TensorFlow",
"Computer Vision",
"3D / Point Clouds / LiDAR",
"GenAI (SDXL, LLMs)"
]
},
{
"name": "Systems",
"skills": [
"C++",
"C",
"Rust",
"CUDA",
"gRPC",
"MongoDB"
]
},
{
"name": "MLOps",
"skills": [
"Kubernetes",
"AWS",
"GCP",
"Docker"
]
},
{
"name": "Web",
"skills": [
"TypeScript",
"React",
"Next.js"
]
}
],
"projects": [
{
"name": "torch-point-ops",
"description": "High-performance PyTorch operators for point-cloud and 3D geometry processing.",
"url": "https://github.com/satyajitghana/torch-point-ops"
},
{
"name": "PV-LIO-for-HBA",
"description": "Point-to-voxel LiDAR-inertial odometry adapted for hierarchical bundle adjustment.",
"url": "https://github.com/satyajitghana/PV-LIO-for-HBA"
},
{
"name": "ige_lio",
"description": "Iterated error-state Kalman-filter LiDAR-inertial odometry experiments.",
"url": "https://github.com/satyajitghana/ige_lio"
},
{
"name": "sdxl-dreambooth-finetune",
"description": "DreamBooth fine-tuning pipeline for Stable Diffusion XL subject personalization.",
"url": "https://github.com/satyajitghana/sdxl-dreambooth-finetune"
},
{
"name": "TSAI-DeepVision-EVA4.0",
"description": "Coursework and experiments from The School of AI's EVA 4.0 computer-vision program.",
"url": "https://github.com/satyajitghana/TSAI-DeepVision-EVA4.0"
},
{
"name": "PadhAI-Course",
"description": "Implementations and notes from the PadhAI deep-learning course.",
"url": "https://github.com/satyajitghana/PadhAI-Course"
}
],
"publications": [
{
"title": "Adaptive Visual Learning Using Augmented Reality and Machine Learning Techniques",
"publisher": "Journal of Computational and Theoretical Nanoscience",
"date": "2020-01-01",
"details": "Vol. 17, No. 11, pp. 4952–4956",
"doi": "10.1166/jctn.2020.8982",
"url": "https://doi.org/10.1166/jctn.2020.8982"
}
],
"patents": [
{
"title": "Method and System for Performing Structural Defect Analysis in a Structural Environment",
"status": "Pending",
"applicationNumber": "US 19/634,310",
"filed": "2026-03-31",
"assignee": "Inkers Technology"
},
{
"title": "Data Acquisition Device",
"status": "Pending",
"applicationNumber": "US 19/634,339",
"filed": "2026-03-31",
"assignee": "Inkers Technology"
}
]
}
```
---
# Health — biomarker panel
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/health
> panel date: 2026-05-15
```json
{
"panelDate": "2026-05-15",
"stack": [
{
"type": "device",
"name": "Apple Watch",
"url": "https://www.apple.com/watch/"
},
{
"type": "device",
"name": "Withings Body+"
},
{
"type": "supplement",
"name": "Vitamin D3"
},
{
"type": "supplement",
"name": "Omega-3"
},
{
"type": "supplement",
"name": "Creatine"
},
{
"type": "supplement",
"name": "Magnesium Glycinate"
}
],
"biomarkers": [
{
"key": "ldl-c",
"label": "LDL-C",
"value": 128,
"unit": "mg/dL",
"category": "cardiovascular",
"weight": 2,
"optimalRange": {
"max": 100
},
"note": "Elevated — primary driver of atherosclerotic risk.",
"status": "elevated"
},
{
"key": "hdl-c",
"label": "HDL-C",
"value": 52,
"unit": "mg/dL",
"category": "cardiovascular",
"optimalRange": {
"min": 40
},
"status": "optimal"
},
{
"key": "apob",
"label": "ApoB",
"value": 105,
"unit": "mg/dL",
"category": "cardiovascular",
"weight": 1.5,
"optimalRange": {
"max": 90
},
"note": "Elevated — atherogenic particle count above target.",
"status": "elevated"
},
{
"key": "triglycerides",
"label": "Triglycerides",
"value": 88,
"unit": "mg/dL",
"category": "cardiovascular",
"optimalRange": {
"max": 150
},
"status": "optimal"
},
{
"key": "hba1c",
"label": "HbA1c",
"value": 5.4,
"unit": "%",
"category": "metabolic",
"weight": 1.5,
"optimalRange": {
"max": 5.7
},
"status": "optimal"
},
{
"key": "fasting-glucose",
"label": "Fasting Glucose",
"value": 96,
"unit": "mg/dL",
"category": "metabolic",
"optimalRange": {
"min": 70,
"max": 99
},
"status": "optimal"
},
{
"key": "fasting-insulin",
"label": "Fasting Insulin",
"value": 8.1,
"unit": "µIU/mL",
"category": "metabolic",
"optimalRange": {
"max": 8
},
"note": "Borderline — marginally above optimal fasting insulin.",
"status": "borderline"
},
{
"key": "alt",
"label": "ALT",
"value": 22,
"unit": "U/L",
"category": "liver_kidney",
"optimalRange": {
"max": 40
},
"status": "optimal"
},
{
"key": "ast",
"label": "AST",
"value": 24,
"unit": "U/L",
"category": "liver_kidney",
"optimalRange": {
"max": 40
},
"status": "optimal"
},
{
"key": "creatinine",
"label": "Creatinine",
"value": 0.95,
"unit": "mg/dL",
"category": "liver_kidney",
"optimalRange": {
"min": 0.7,
"max": 1.3
},
"status": "optimal"
},
{
"key": "egfr",
"label": "eGFR",
"value": 99,
"unit": "mL/min",
"category": "liver_kidney",
"optimalRange": {
"min": 90
},
"status": "optimal"
},
{
"key": "tsh",
"label": "TSH",
"value": 2.1,
"unit": "mIU/L",
"category": "hormonal",
"optimalRange": {
"min": 0.5,
"max": 4
},
"status": "optimal"
},
{
"key": "testosterone",
"label": "Testosterone",
"value": 610,
"unit": "ng/dL",
"category": "hormonal",
"optimalRange": {
"min": 300,
"max": 1000
},
"status": "optimal"
},
{
"key": "vitamin-d",
"label": "Vitamin D",
"value": 24,
"unit": "ng/mL",
"category": "nutritional",
"weight": 1.5,
"optimalRange": {
"min": 30,
"max": 100
},
"note": "Low — below the optimal 25-OH vitamin D range.",
"status": "low"
},
{
"key": "vitamin-b12",
"label": "Vitamin B12",
"value": 540,
"unit": "pg/mL",
"category": "nutritional",
"optimalRange": {
"min": 400,
"max": 1000
},
"status": "optimal"
},
{
"key": "ferritin",
"label": "Ferritin",
"value": 38,
"unit": "ng/mL",
"category": "nutritional",
"optimalRange": {
"min": 50,
"max": 300
},
"note": "Low — iron stores below the optimal floor.",
"status": "low"
},
{
"key": "hemoglobin",
"label": "Hemoglobin",
"value": 15.1,
"unit": "g/dL",
"category": "blood_panel",
"optimalRange": {
"min": 13.5,
"max": 17.5
},
"status": "optimal"
},
{
"key": "wbc",
"label": "WBC",
"value": 6.2,
"unit": "10³/µL",
"category": "blood_panel",
"optimalRange": {
"min": 4,
"max": 11
},
"status": "optimal"
},
{
"key": "platelets",
"label": "Platelets",
"value": 245,
"unit": "10³/µL",
"category": "blood_panel",
"optimalRange": {
"min": 150,
"max": 400
},
"status": "optimal"
},
{
"key": "resting-hr",
"label": "Resting HR",
"value": 58,
"unit": "bpm",
"category": "vitals",
"optimalRange": {
"min": 40,
"max": 60
},
"status": "optimal"
},
{
"key": "bp-systolic",
"label": "BP Systolic",
"value": 132,
"unit": "mmHg",
"category": "vitals",
"weight": 1.5,
"optimalRange": {
"max": 120
},
"note": "Elevated — systolic pressure above the optimal ceiling.",
"status": "borderline"
},
{
"key": "vo2max",
"label": "VO₂max",
"value": 47,
"unit": "mL/kg/min",
"category": "vitals",
"optimalRange": {
"min": 42
},
"status": "optimal"
}
]
}
```
---
# Now
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/now
```json
{
"updated": "2026-06-04",
"items": [
"Building ai.thesatyajit.com — a dual-native personal site maintained by a crew of Claude agents that ship every change as a reviewed PR.",
"Leading industrial-AI 3D perception at Inkers: turning LiDAR and camera capture into structural-defect analysis that runs in production.",
"Writing CUDA point-cloud kernels — fused nearest-neighbour, sampling, and grouping ops to take the Python side out of the hot path.",
"Reading deeply on LiDAR odometry and point-cloud registration, with an eye toward tighter, faster SLAM front-ends."
]
}
```
---
# Uses
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/uses
```json
[
{
"section": "Workstation",
"items": [
{
"name": "Custom GPU workstation",
"note": "NVIDIA RTX-class GPU for CUDA + training"
},
{
"name": "Linux",
"note": "Ubuntu — primary dev OS"
}
]
},
{
"section": "Editor",
"items": [
{
"name": "VS Code",
"note": "daily driver"
},
{
"name": "Claude Code",
"note": "agentic pair-programming in the terminal"
},
{
"name": "Fonts",
"note": "JetBrains Mono today, migrating toward IBM Plex Mono"
}
]
},
{
"section": "Terminal",
"items": [
{
"name": "zsh",
"note": "shell"
},
{
"name": "tmux",
"note": "session multiplexing"
}
]
},
{
"section": "Stack highlights",
"items": [
{
"name": "PyTorch + CUDA",
"note": "custom kernels for point-cloud ops"
},
{
"name": "C++ / gRPC",
"note": "high-performance perception services"
},
{
"name": "Next.js + Tailwind",
"note": "this site"
}
]
}
]
```
---
# Reading
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/reading
```json
[
{
"type": "paper",
"title": "PV-LIO: A Probabilistic Voxel-based LiDAR-Inertial Odometry Framework",
"status": "reading",
"note": "PLACEHOLDER — voxel-probabilistic front-end ideas for tighter LIO."
},
{
"type": "book",
"title": "Programming Massively Parallel Processors",
"author": "Hwu, Kirk, El Hajj",
"status": "reading",
"note": "PLACEHOLDER — the CUDA reference I keep coming back to."
},
{
"type": "paper",
"title": "FAST-LIO2: Fast Direct LiDAR-Inertial Odometry",
"status": "read",
"note": "PLACEHOLDER — iterated Kalman filter without feature extraction."
},
{
"type": "book",
"title": "Multiple View Geometry in Computer Vision",
"author": "Hartley & Zisserman",
"status": "queued",
"note": "PLACEHOLDER — geometry fundamentals refresher."
},
{
"type": "paper",
"title": "3D Gaussian Splatting for Real-Time Radiance Field Rendering",
"status": "queued",
"note": "PLACEHOLDER — splatting as a reconstruction primitive."
}
]
```
---
# Vritti UI
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/vritti
> stack: React 19, Next.js 16, Tailwind v4, Motion, Three.js, TypeScript
Components crafted for design engineers — animated components and blocks you copy
and paste into your shadcn/ui project. No package installs, full ownership.
```bash
npx shadcn@latest add "https://vritti.thesatyajit.com/r/shimmer-button"
```
**560+ components** across 14 categories — backgrounds, animations, text effects, buttons,
cards, charts, layouts, navigation, carousels, cursors, inputs, media, shaders, and the
genuinely weird ones. **240+ blocks** — pre-built page sections for real apps: auth,
pricing, e-commerce, billing, dashboards, modals, testimonials, footers, FAQ, AI & Web3.
Beyond components, it ships a set of **creative tools** that run online or install into
your project: Art Studio (9 generative-art tools), Background Studio, Shape Studio (an SVG
editor), Texture Studio (30+ stackable WebGL filters), Shader Studio (87 production-ready
shaders), Dither Studio, a visual Theme Editor (42 presets, Google Fonts, contrast
checker, CSS export), and native haptics. There's an interactive playground for tuning
component props live, and AI components (chat streaming, image generation, voice, AI
search) for LLM-native UIs.
Everything is copy-paste, shadcn-style — you own the source. It's the most ambitious thing
on this page: a full component ecosystem, a design-tool suite, and a theme system, built on
React 19, Next.js 16, Tailwind v4, Motion, and Three.js.
---
# torch-point-ops
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/torch-point-ops
> stack: PyTorch, CUDA, C++, Python
> repo: https://github.com/satyajitghana/torch-point-ops
A PyTorch library for 3D point operations — nearest-neighbour, sampling, and grouping
primitives — implemented as fused CUDA kernels for throughput on large point clouds.
Built to back LiDAR and 3D-perception pipelines where the Python-side ops were the
bottleneck. Exposes a clean `torch` autograd-compatible API over the native kernels.
---
# fabrik
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/fabrik
> stack: TypeScript, React, Vercel AI SDK, Zod, Motion
> repo: https://github.com/satyajitghana/fabrik
Most LLM apps are text in, text out. **fabrik** lets the model answer with *interface* —
it picks and fills real React components (cards, forms, charts, pickers) instead of
emitting a wall of prose. The AI decides what to show.
It's provider-agnostic: wire it to any model through the Vercel AI SDK (Gemini, OpenAI,
Anthropic, …), keep the API key server-side in a single route handler, and render the
streamed component tree on the client. Components are validated with Zod schemas, so the
model can only produce UIs you've defined — no arbitrary markup, no injection surface.
The bet behind it: as models get better at tool use, the next interface isn't a chat
bubble — it's a UI the model assembles on the fly for the task at hand.
---
# rentree
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/rentree
> stack: Next.js, React, TypeScript, Supabase, Drizzle ORM, Tailwind v4
> repo: https://github.com/satyajitghana/rentree
Rentree connects people with organic farms across India. Instead of buying produce off a
shelf, you rent a plot or adopt a tree: your farmer grows your food, sends real photo
updates as it grows, and ships it when it's ready.
Rent a plot for a season (tomatoes, herbs, rice — whatever grows), or adopt an apple,
orange, or mango tree and get the whole harvest. Every tree carries a QR code, so you can
scan and trace the full journey from planting to your plate. A clean Next.js 16 app on
Supabase + Drizzle, with the provenance trail as the trust mechanism that makes the model
work.
---
# Quantum Orbital Visualizer
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/atoms
> stack: Next.js, Three.js, Rust (WASM), C++23 (WASM), Tailwind v4, Zustand
> repo: https://github.com/satyajitghana/atoms
Each visualization is the probability density |ψ|² of a hydrogen-like electron, computed
from the exact solutions to the Schrödinger equation and rendered as up to 500K glowing
particles in real time. Drive the quantum numbers (n, l, m) and watch the probability
cloud morph as you go.
The fun engineering bit is the **dual WASM backends**: the same physics is implemented
three ways — JavaScript, Rust→WASM (~41 KB), and C++23→WASM (~28 KB) — and you can switch
the compute engine at runtime. Same orbitals, different implementations, side-by-side for
benchmarking.
There's also an Element Explorer: the full periodic table, where clicking any of the 118
elements parses its electron configuration into individual orbitals you can then visualize
one by one.
---
# BOTCHA
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/botcha
> stack: Next.js, TypeScript, Redis, crypto (SHA-256 / HMAC / JWT)
> repo: https://github.com/satyajitghana/botcha
Only an autonomous agent with runtime access to HTTP, cryptography, and byte manipulation
can pass — which is the whole joke, and the whole point.
Each challenge is 256 random bytes plus 2–4 byte-level transformation steps written in
randomized natural language, inside a 30-second window: fast enough for a machine,
hopeless for a human copy-pasting into a REPL. The agent must decode the base64, execute
each transform in order, concatenate the raw byte outputs, `SHA-256` the result, then
`HMAC-SHA256(key=nonce, message=answer)` and submit both — proving it actually did the
computation. On success it gets a short-lived JWT.
It's a small, sharp take on a real question the agent era raises: if the web increasingly
wants to *let bots in* and keep humans out, what does that gate look like?
---
# penora
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/penora
> stack: TypeScript, React, Canvas, npm package
> repo: https://github.com/satyajitghana/penora
Type a string and penora animates it into natural handwriting — not a font fade-in, but
stroke-by-stroke drawing driven by the glyph contours, with pen physics layered on top:
pressure tapering at stroke ends, seeded jitter so each render is subtly different, and
micro-wobble for that hand-drawn quality. Export the result as video or GIF.
Shipped as an npm package (`npm i penora`) and as a shadcn registry component you can drop
straight into a project. It's the kind of small, self-contained library that's satisfying
to build: a tight problem (make text look handwritten and *alive*) with a lot of room for
craft in the details.
---
# Hyr
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/hyr
> stack: Next.js, TypeScript, Tailwind v4, Google Gemini
> repo: https://github.com/satyajitghana/hyr
Drop in any PDF resume and Hyr parses it into a structured, fully editable format —
inline editing for every section, no forms or modals. From there it tailors the resume
per job posting, aligning keywords and tone to the listing so it gets past ATS screening
and reads as a real fit.
Build, tailor, optimize, and track the whole job search in one place. The AI work runs on
Gemini behind a server route; the parsing-to-structured-data step is the part that makes
the rest of the editing experience feel clean rather than fighting a PDF.
---
# MockLab
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/mocklab
> stack: Next.js, Tailwind v4, shadcn/ui, Motion, html-to-image
> repo: https://github.com/satyajitghana/mocklab
MockLab generates pixel-perfect social mockups across 21+ platforms — social posts (X,
LinkedIn, Instagram post/story, Reddit, Threads, YouTube, TikTok), chat messages
(WhatsApp, Telegram, Slack, Discord, iMessage, Snapchat), and Gmail. Edit on the left,
watch the live preview on the right, download as 2× PNG.
Everything is configurable: verified badges, reactions, timestamps, read receipts, media
uploads, and per-message editing for the chat mockups — all with dark/light theming for
both the app and every mockup. It's a deceptively deep UI project: each platform is its
own little pixel-accurate design system to reproduce.
---
# tapcn
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/projects/tapcn
> stack: React Native, Expo, TypeScript, CLI
> repo: https://github.com/satyajitghana/tapcn
tapcn brings the shadcn model to mobile: not an npm component library, but a CLI that
copies beautifully-designed, accessible component *source* directly into your Expo
project. No dependency lock-in, no version conflicts — just code you control and can edit.
```bash
npx @tapcn/cli init
npx @tapcn/cli add button card input text
```
Components work on iOS, Android, and Web from a single source. The whole appeal of the
shadcn approach — own your components, theme them freely, no black-box library — has been
missing on React Native; tapcn fills that gap.
---
# My site's chatbot was stuffing 273k tokens into every message. I gave it tools instead.
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/blog/site-agent-dynamic-tools
> date: 2026-07-20
> tags: agents, llm, tools, retrieval, systems
The assistant embedded on this site — the `⌘K` console, `/api/ask`, and the `ask_satyajit`
MCP tool — worked. It was also doing something faintly absurd. Every single request built its
system prompt like this:
```ts
// the old lib/chat.ts, abridged
const sections = []
for (const page of ["about", "resume", "health", "now", "uses", "reading"])
sections.push(await dataPageMarkdown(page))
for (const item of getAllContent()) // every article, blog, log, digest…
sections.push(contentMarkdown(item.kind, item.slug))
return [PERSONA, RULES, "=== CONTENT CORPUS ===", sections.join("\n---\n")].join("\n")
```
It pasted **the entire site** into the prompt and let the provider's context cache eat the cost.
That's a defensible move when the corpus is a page or two. Mine isn't anymore:
```text
old full-corpus system prompt: 1,090,854 chars ≈ 273,000 tokens
```
The code even carried a comment — *"if the corpus ever approaches ~100K tokens, switch to
retrieval"* — that reality had quietly sailed past nearly 3× over. Two hundred seventy-three
thousand tokens is past the input window of the model serving it. The chat was running on
whatever survived truncation. Time to do the thing the comment said.
## The fix: retrieve, don't dump
I rebuilt the assistant as a small **tool-using agent**. The system prompt now carries only a
compact **catalog** — every page's `kind/slug`, title, and one-line description, about **5k
tokens** — and the agent pulls the bodies it actually needs through tools. Three of them:
```ts
// lib/chat.ts — the whole tool set
search_content({ query, limit }) // BM25 over the site (ranked, returns section + snippet)
get_content({ kind, slug }) // fetch one page's full markdown, on demand
think({ thought }) // a reasoning scratchpad; records the thought, returns nothing
```
A question now flows: read the catalog → `search_content` (or jump straight to `get_content` if
the catalog already names the page) → read one or two pages → optionally `think` to reconcile
them → answer, with citations. The base prompt is fixed and small; the variable cost is the
one-to-three pages it fetched, not the other sixty-five it didn't.
```text
before: ~273k tokens, every request, whether relevant or not
after: ~5k catalog + ~2–10k per page actually fetched
```
That's the same **dynamic loading** idea from
[Kimi K3's tool-calling guide](https://platform.kimi.ai/docs/guide/kimi-k3-tool-calling-best-practice) —
their point is that a big upfront payload "eats up context and makes the model more likely to
pick the wrong thing." Kimi loads *tool schemas* on demand; my corpus is the payload, so I load
*content* on demand. Same principle, one layer down.
## Three tools, on purpose
The tool count is a decision, not an accident. Mario Zechner's
[pi coding agent](https://mariozechner.at/posts/2025-11-30-pi-coding-agent/) makes the case that
"four tools are all you need" — `read`, `write`, `edit`, `bash` — and that MCP servers which
"dump their entire tool descriptions into your context on every session" are the anti-pattern.
A read-only site agent needs fewer still: search, fetch, think. Each description is a couple of
sentences. The combined tool surface is under a thousand tokens, so keeping it declared upfront
(rather than lazily loading schemas, which only pays off with dozens of tools) is the right call
here — the honest version of "dynamic tools" for a small surface is *don't have a big one*.
The one genuinely new tool is `think`, from Anthropic's
[think-tool post](https://www.anthropic.com/engineering/claude-think-tool). It does nothing —
literally logs the thought and returns:
```ts
think: tool({
description: "Think out loud: plan which pages to fetch, or check a draft answer against the sources you read. Records the thought and returns nothing new.",
inputSchema: z.object({ thought: z.string() }),
execute: async ({ thought }) => ({ ok: true, thought }),
}),
```
That looks pointless until you watch a multi-step tool run. Between `search_content` and the
final answer, the model has raw tool output sitting in context and no designated place to reason
over it before committing. `think` is that place — a scratchpad that keeps the "what did I just
read, and does it actually answer the question" step from being skipped. Anthropic reports it
buying a large margin on multi-step tool tasks; the appeal for a retrieval agent is exactly that
it makes the model *check the page it fetched* instead of answering from the snippet.
## The loop
The agent loop itself is one call, courtesy of the Vercel AI SDK — `stopWhen` bounds how many
tool round-trips it may take before it has to answer:
```ts
const result = streamText({
model: chatModel("main"),
system, // persona + rules + the 5k-token catalog
messages,
tools: agentTools(), // search_content, get_content, think
stopWhen: stepCountIs(6), // search → read → think → answer is ~4; 6 is headroom
})
```
The same tool set backs all three surfaces — the streaming `⌘K` console (`/api/chat`), the
one-shot `/api/ask`, and the `ask_satyajit` MCP tool — so there's one harness, not three. And
`search_content` is the [Contextual BM25](/blog/site-search-contextual-bm25) engine I'd already
built for the site's `/search`, now doing double duty as the agent's retriever. The pieces
compose: the search work made the harness possible.
Honest scope. This is retrieval over a small, single-author corpus, not a coding agent — the
lessons transfer but the stakes are lower. `think`'s benefit is task-dependent (Anthropic saw
big gains on policy-heavy multi-step tasks, near-zero on simple ones); on a two-hop lookup it
mostly earns its keep by stopping the model from answering off the snippet. And the retriever is
lexical BM25 — no embeddings yet, so a pure paraphrase with no shared terms can still miss the
right page. The catalog is the safety net there: the model can see every title even when search
comes up short.
## The take
The old design wasn't wrong when it was written — it was wrong at 273k tokens. The rebuild is the
[agent-harness](/articles/agent-harness) framing applied to my own site: the model barely changed,
but the loop, the tool set, and the context policy around it changed completely. Give the agent a
map and three tools and let it fetch, instead of force-feeding it the whole library and hoping the
answer survives the truncation. Retrieve, don't dump.
---
# Giving my own search the Contextual BM25 treatment
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/blog/site-search-contextual-bm25
> date: 2026-07-20
> tags: search, retrieval, bm25, rag, information-retrieval
I recently wrote two pieces back to back: one on [BM25](/articles/bm25), the ranking
function that refuses to die, and one on [Contextual Retrieval](/blog/contextual-retrieval),
Anthropic's fix for chunks that forget where they came from. Then I opened `lib/search.ts` —
the thing powering this site's own `/search` and the `/api/search` endpoint agents hit — and
found this:
```ts
// the old lib/search.ts, abridged
for (const item of getAllContent()) {
const fields = [item.title, item.description, item.tags, item.body]
for (const text of fields) {
if (text.toLowerCase().includes(q)) { // q = the whole lowercased query
results.push({ /* … a ±120-char window around the hit */ })
break
}
}
}
```
A substring `indexOf`. It has three problems, and the third is the one that actually bites:
1. **No ranking.** The first document that contains the substring wins, in date order. A tight,
on-topic match and an incidental mention are indistinguishable.
2. **No notion of rarity or length.** Matching "the" counts the same as matching "kalman".
3. **Multi-word queries fall off a cliff.** `q` is the *entire* query string, so
`includes(q)` needs that **exact phrase** somewhere in the text. Nobody searches in verbatim
phrases. `"bm25 length normalization"` returns **zero results** — those three words are all
in the BM25 article, just never contiguous.
So I ate my own cooking. Here's the rebuild.
## BM25, the same formula I just wrote up
The [BM25 walkthrough](/articles/bm25) has the whole derivation; the code is a direct
transcription of it. Lucene's non-negative IDF, term-frequency saturation at `k1 = 1.2`, length
normalization at `b = 0.75`:
```ts
// lib/search.ts — the ranking core, straight from the article's formula
const K1 = 1.2
const B = 0.75
function idf(term: string, ix: Index): number {
const n = ix.df.get(term) ?? 0
return Math.log(1 + (ix.n - n + 0.5) / (n + 0.5))
}
// per query term present in a chunk:
const f = chunk.tf.get(term) ?? 0
const norm = K1 * (1 - B + (B * chunk.len) / ix.avgdl)
score += idf(term, ix) * (f * (K1 + 1)) / (f + norm)
```
Two things fall out for free. The query is **tokenized** into terms and each is scored
independently, so multi-word queries just work — no phrase has to exist. And rare terms
dominate: `kalman` carries far more weight than `filter`, because `idf` collapses for common
words. No stopword list, same as the article promised.
I keep a **postings map** (`term → chunk indices`) so a query only touches chunks that actually
share a term, instead of scanning the whole corpus. The corpus is one person's writing, so this
is overkill — but it's the same inverted-index shape a real engine uses, and it's three lines.
## Contextual chunks, minus the Claude call
BM25 alone still has the chunk-amnesia problem from the [Contextual Retrieval
post](/blog/contextual-retrieval): if I split an article into passages and index each one alone,
a passage that reads "it drops 41% at parity" has no idea it's *about the Harness Effect, in the
section on the controlled swap*. A query for either misses it.
Anthropic's fix is to have Claude write a one-line context for every chunk before indexing.
Mine is cheaper and dumber: the context a chunk lost is sitting right there in the document's
frontmatter and headings. So before indexing, every chunk inherits its document's **title,
description, tags, and nearest heading** — deterministically, at request time, no model, no
build step:
```ts
// each chunk is packed with weighted terms: its own body, its heading, and the
// document context it would otherwise have lost when the body was split.
const contextTokens = tokenize([title, description, tags.join(" ")].join(" "))
for (const t of bodyTokens) add(t, 1)
for (const t of headingTokens) add(t, 1) // section-local: full weight
for (const t of contextTokens) add(t, CONTEXT_WEIGHT) // 0.5 — present, not dominant
```
The honest caveat: this is a **poor man's** Contextual Retrieval. Claude's per-chunk context can
say things the frontmatter can't ("the previous quarter's revenue was \$314M"); my version can
only replay the structural context the document already carries. And prepending the *same* title
to every chunk of a document inflates those terms' document-frequency, which is exactly why the
context terms get `CONTEXT_WEIGHT = 0.5` instead of full weight — present enough to make the
chunk findable by its subject, quiet enough not to drown the passage's own words. It's the 80%
of the win for 0% of the inference cost, which is the right trade for a static personal site.
Each result also reports **which section** it matched — the nearest heading rides along as the
result's `field`, so `/api/search` tells an agent not just *which* document but *where* in it.
## Before / after
Same queries, old substring search vs. the new Contextual BM25, on the live corpus:
```text
query substring indexOf contextual BM25 (top hit)
--------------------------------------------------------------------------------------
"bm25 length normalization" 0 results (no verbatim articles/bm25
phrase) [Start with TF-IDF, and its two flaws]
"revenue grew quarter" 0 results blog/contextual-retrieval
[The problem: chunks forget …]
"kalman filter" first dated doc that articles/kalman-filter (8.4)
contains the phrase, + fast-lio2 (8.6, cites it)
unranked ranked by relevance, not date
"reciprocal rank fusion" 0 results blog/contextual-retrieval
[A dependency-free repro]
```
The multi-word queries are the story. Under substring, anything that isn't a verbatim phrase —
which is almost everything a person types — returned nothing. Under BM25 every term contributes,
so the query finds the document even when its words are scattered across a paragraph, and the
contextual prefix pulls in matches on the *subject* of a passage, not just its literal words.
## It's live
Try it: [`/search?q=length+normalization`](/search?q=length+normalization). Agents get the
ranked, scored JSON with the section label:
```bash
curl -s "https://ai.thesatyajit.com/api/search?q=contextual%20retrieval&limit=3" | jq '.results[] | {slug, field, score}'
```
What this still isn't: it's the **lexical half** only. The [contextual retrieval
post](/blog/contextual-retrieval) makes the case that the real win is *fusing* BM25 with a dense
retriever, and that reranking the shortlist is what takes the failure rate the last mile. There
are no embeddings here yet — a query has to share actual terms with a chunk, so a pure paraphrase
with no lexical overlap can still miss. For a corpus this size, lexical BM25 over contextualized
chunks is the honest 90% solution; the semantic half is the next commit, not this one.
## The takeaway
The whole change is `lib/search.ts` — a few hundred lines, no dependencies, no index server, no
model at request time. It's the two ideas I'd just written about, applied to the smallest
possible target: my own search bar. BM25 for ranking, structural context for the chunk-amnesia
problem, and an honest note about the half I haven't built yet. Writing about a technique is a
good way to understand it; running it in your own site is a better one.
---
# Contextual Retrieval, with a runnable repro and a browser playground
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/blog/contextual-retrieval
> date: 2026-07-17
> tags: rag, retrieval, claude, embeddings, playground
[Anthropic's Contextual Retrieval post](https://www.anthropic.com/engineering/contextual-retrieval)
has been on my to-read list for a while. It targets the oldest bug in RAG, so I finally sat down,
built a dependency-free reproduction, and wired up a browser playground. This is the write-up.
## The problem: chunks forget where they came from
Standard RAG splits documents into chunks and indexes each chunk on its own. That destroys context.
Anthropic's example is perfect:
> `The company's revenue grew by 3% over the previous quarter.`
Which company? Which quarter? The chunk can't say. A query like *"how did ACME's Q2 revenue change?"*
has **no terms to match** against that chunk — the words "ACME" and "Q2 2023" live in the *document*,
not the *chunk*. Both lexical (BM25) and semantic (embedding) retrieval miss it.
## The fix: let Claude situate each chunk
Contextual Retrieval prepends a short, chunk-specific context to each chunk **before** you embed it and
**before** you build the BM25 index — Anthropic calls these **Contextual Embeddings** and **Contextual
BM25**. The context is generated by Claude, given the whole document. The same chunk becomes:
> `This chunk is from an SEC filing on ACME corp's performance in Q2 2023; the previous quarter's revenue was $314 million. The company's revenue grew by 3% over the previous quarter.`
Now the owner and the period are *in the chunk*, so retrieval can find it.
## Watch it happen
Here is the whole pipeline in your browser — the document, its chunks, the chunks with context prepended,
then retrieval. The scoring is a real BM25 and a TF-IDF cosine (a stand-in for an embedding model); the
contextual prefixes stand in for Claude's output. Step to **4 · retrieve**, pick a query, and watch the
right chunk climb from the standard column to the contextual one:
Toggle between BM25, embeddings, and their fusion. The pattern holds across all three: the answer chunk
is buried on standard chunks and near the top once each chunk carries its context.
## The one prompt that does the work
This is the prompt from the post, verbatim. It runs once per chunk, with the full document supplied as
context:
```text
{{WHOLE_DOCUMENT}}
Here is the chunk we want to situate within the whole document
{{CHUNK_CONTENT}}
Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. Answer only with the succinct context and nothing else.
```
The obvious worry is cost: you re-send the whole document once per chunk. **Prompt caching** kills that.
Cache the document once and every chunk in it reads from the cache, which is why Anthropic quotes a
one-time **$1.02 per million document tokens**. The document is the stable prefix, so it takes the
`cache_control` breakpoint; the chunk and instruction vary and come after it:
```python
from anthropic import Anthropic
client = Anthropic()
CONTEXT_PROMPT = """Here is the chunk we want to situate within the whole document
{chunk}
Please give a short succinct context to situate this chunk within the overall document \
for the purposes of improving search retrieval of the chunk. Answer only with the \
succinct context and nothing else."""
def situate(doc: str, chunk: str) -> str:
resp = client.messages.create(
# One Claude call per chunk. For a large index you'd typically drop to a
# cheaper model like claude-haiku-4-5 — which is what Anthropic's ~$1.02 /
# 1M-doc-token estimate assumes — trading a little context quality for cost.
model="claude-opus-4-8",
max_tokens=200,
messages=[{
"role": "user",
"content": [
# The whole document is the stable prefix: cache it ONCE, then every
# chunk of this document reads from the cache instead of re-paying for it.
{"type": "text",
"text": f"\n{doc}\n",
"cache_control": {"type": "ephemeral"}},
{"type": "text", "text": CONTEXT_PROMPT.format(chunk=chunk)},
],
}],
)
return "".join(b.text for b in resp.content if b.type == "text").strip()
```
Verify the cache is actually working: `resp.usage.cache_read_input_tokens` should be non-zero on every
chunk after the first for a given document. If it's zero, something upstream is mutating the document
bytes (a timestamp, non-deterministic JSON) and invalidating the prefix.
## A dependency-free repro
To convince myself the lift was real and not marketing, I wrote a ~180-line pure-stdlib script: naive
chunking, BM25 from scratch, a TF-IDF cosine as an embedding stand-in, and reciprocal rank fusion. The
retrieval core is small — here is BM25 and the fusion step:
```python
class BM25:
def score(self, query, i):
dl, tf, s = len(self.docs[i]), self.tf[i], 0.0
for t in query:
if t not in tf:
continue
num = tf[t] * (self.k1 + 1)
den = tf[t] + self.k1 * (1 - self.b + self.b * dl / self.avgdl)
s += self.idf(t) * num / den
return s
def rrf(rankings, k=60): # reciprocal rank fusion of BM25 + embedding rankings
scores = Counter()
for ranking in rankings:
for rank, item in enumerate(ranking):
scores[item] += 1 / (k + rank + 1)
return [i for i, _ in scores.most_common()]
```
Then I index the chunks two ways — plain, and with a one-line context prepended — and measure where the
correct chunk lands for a couple of "which entity, which period" queries. The actual output:
```text
query method plain rank ctx rank
------------------------------------------------------------------------------------
How did ACME Corp revenue change in Q2 2023? bm25 4 2
How did ACME Corp revenue change in Q2 2023? emb 4 1
How did ACME Corp revenue change in Q2 2023? hybrid 4 2
What happened to Beta Industries revenue in Q3 2023? bm25 5 2
What happened to Beta Industries revenue in Q3 2023? emb 5 1
What happened to Beta Industries revenue in Q3 2023? hybrid 5 2
recall@1 (fraction of queries where the right chunk ranks #1):
bm25 plain 0/2 contextual 0/2
emb plain 0/2 contextual 2/2
```
Contextualizing the chunks moves the answer from rank **4–5** to rank **1–2** across BM25, embeddings, and
their fusion — and embeddings recall@1 goes from **0/2 to 2/2**. On this toy corpus fusion lands the answer
at #2 rather than #1 (a small-N artifact of RRF), which is a good reminder that the fusion win is an
*aggregate* effect — which is exactly what Anthropic's real evaluation measures.
## What the real evaluation found
On Anthropic's benchmark (top-20 retrieval failure rate, i.e. `1 - recall@20`), stacking the techniques
compounds:
- **Contextual Embeddings** alone: `5.7% → 3.7%` — a **35%** cut in the failure rate.
- **+ Contextual BM25**: `5.7% → 2.9%` — **49%**.
- **+ reranking**: `5.7% → 1.9%` — **67%**.
The lexical half matters more than you'd guess — BM25 nails exact identifiers (error codes, ticker symbols,
function names) that embeddings smear together, so contextualizing *both* indexes and fusing them beats
either alone.
## Things worth copying from the post
- **Retrieve top-20, not top-5/10.** Anthropic found 20 the most performant cut for the final context.
- **Rerank the shortlist.** Retrieve ~150 candidates, then rerank down to 20 for the answer prompt — that's
the step that takes the failure rate from 2.9% to 1.9%.
- **Embedding model matters.** Gemini and Voyage embeddings were the standouts in their tests.
- **Chunking still matters.** Size, boundary, and overlap all move the numbers — Contextual Retrieval sits
on top of good chunking, it doesn't replace it.
- **A domain-tuned context prompt beats the generic one.** The template above is a floor, not a ceiling.
Be honest about what this costs. Contextual Retrieval adds a Claude call **per chunk** at index time —
cheap per token with caching, but real latency and spend when you're indexing millions of chunks, and it
has to re-run when documents change. My repro is a minimal reproduction of the *core mechanism* (context
lifts rank); the `35 / 49 / 67%` figures are Anthropic's, on their corpus. And my playground scores with a
TF-IDF cosine, not a real embedding model — it shows the *shape* of the effect, not production numbers.
## The takeaway
The move is almost embarrassingly simple: spend a cheap, cached Claude call per chunk to write down the
context a human would need to make sense of it, then index that. It attacks the failure at its source
instead of papering over it downstream, it helps lexical and semantic retrieval at the same time, and — as
the little repro above shows — you can watch the right chunk climb the rankings the moment the context goes
in.
---
*Source: [Introducing Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval)
(Anthropic). The prompt and the `35 / 49 / 67%` numbers are theirs; the repro and the browser playground are
mine, and the playground's scoring runs entirely client-side.*
---
# A 14B model that matches a 671B one — by knowing its domain
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/blog/qwen-bim-domain-beats-scale
> date: 2026-06-09
> tags: llm, fine-tuning, domain-models, bim, paper-notes
Here's the headline from [Qwen-BIM](https://arxiv.org/abs/2602.20812) (Lin et al.,
Tsinghua, Feb 2026): a fine-tuned **14B** model scores **0.83 on G-Eval** for
BIM-based design tasks — essentially tied with **DeepSeek-R1 at 671B** (0.84), and
ahead of its own 72B sibling. A model ~48× smaller, matching the frontier on a
specific domain.
That result is not surprising on its own — "fine-tune a small model on your domain"
is folklore by now. What makes the paper worth reading is the *anatomy*: where exactly
general LLMs fall over on engineering work, and which one design choice did most of the
lifting. I work on industrial AI at [Inkers](https://inkers.ai), so domain models over
3D/BIM data are close to home. These are my notes.
## The actual problem: a BIM model isn't text
A Building Information Model is a structured graph of components — walls, slabs, beams,
each with geometry, materials, and relationships. An LLM can't read it. So step one of
*any* LLM-on-BIM pipeline is an unglamorous one the field mostly skips past: **turn the
model into text.**
The authors do exactly that, carefully. Revit models of five building types (malls,
offices, dormitories, teaching buildings, museums) are sliced into spatial blocks of
~10–15 components each, defects are injected, and each block is serialized to plain text
plus 22 templated questions with **hard-coded reference answers**. That last detail
matters: the ground truth is computed by rules, not by another model, so the benchmark
isn't measuring one LLM against another LLM's opinion.
The questions ladder up in difficulty on purpose: from "list the wall IDs"
(extraction) through "compute each slab's area" (calculation) to "is this wall
thickness suspicious given residential norms?" (domain reasoning). It's a clean way to
see *which rung* a model falls off.
The whole pipeline, end to end, is just: project the structured model into text, turn it
into supervised Q&A, add reasoning traces, and LoRA-fine-tune a small open model on it.
## Where general LLMs actually fail
They evaluated 11 general models (ChatGLM, Qwen, DeepSeek — including the 671B
DeepSeek-V3/R1). The failure modes are specific and, honestly, familiar from any
engineering-LLM project:
- **Arithmetic.** Asked for a slab's planar area, Qwen-max picks the *right formula* and
still returns the wrong number. The bottleneck isn't understanding — it's calculation.
- **Natural-language literalism.** Models misread parentheses in the answer template, or
a naming rule ("wall IDs start with Q"), and confidently apply the wrong transform.
- **Missing domain knowledge.** Asked to infer a building's floor height, the 14B base
model reasons that floor height ≈ slab thickness (120 mm) — coherent chain of thought,
wrong mental model, because it was never taught what "floor height" means in practice.
The pattern: general models clear extraction and counting, then degrade sharply on
calculation, multi-step reasoning, and anything needing design common sense. On the
domain-specific design-review tasks, G-Eval is mostly **below 0.8** — not reliable
enough to trust.
## The one choice that mattered: reasoning supervision
This is the part I'd underline. They built two datasets from the same BIM text:
- **BIM-QA** — 2,129 plain question→answer pairs.
- **BIM-QRA** — 1,364 question→**reasoning**→answer triples, where the intermediate
steps are supervised, not just the final answer.
Then they LoRA-fine-tuned Qwen2.5-14B on different mixes. The result is the kind of
finding that should change how you build these datasets:
| Fine-tuning data | Size | G-Eval |
|---|---|---|
| 100% QA | 2,129 | 0.69 |
| 80% QA + 20% QRA | 2,661 | 0.77 |
| 60% QA + 40% QRA | 2,500 | 0.77 |
| **100% QRA** | **1,364** | **0.83** |
The **smallest** dataset — pure reasoning triples — won, by a wide margin. More
reasoning supervision monotonically improved G-Eval, and quality beat quantity outright.
Teaching the model *how to get there*, on a third of the data, beat teaching it *what the
answer is* on the full set.
## Bigger is not better (and the paper shows it twice)
Two clean data points against scale-maximalism:
1. On the general benchmark, **QwQ-32B out-scored DeepSeek-R1 (671B)** on G-Eval. The
giant model's verbose reasoning actually *hurt* — it padded answers with redundant
text, tanking format and text-similarity scores without improving correctness.
2. After fine-tuning, **Qwen-BIM (14B) matched DeepSeek-R1 (671B)** and beat the 72B and
32B Qwen models on the domain G-Eval.
The improvement from fine-tuning is also *targeted* exactly where you'd want it:
| G-Eval | Base 14B | Qwen-BIM | Δ |
|---|---|---|---|
| General tasks | 0.810 | 0.874 | +0.06 |
| Domain-specific tasks | 0.588 | 0.801 | **+0.21** |
| Overall | 0.689 | 0.834 | +0.15 |
General ability barely moved (it was already fine); the entire gain is concentrated in
the domain tasks that were broken. That's the signature of fine-tuning doing the right
thing — adding domain competence without trading away the base model's generality.
## What I'd flag
It's a careful paper, but keep the scope honest:
- **2D only.** Early tests showed the models couldn't do 3D geometry (collision
detection), so those questions were cut. The hard part of real BIM reasoning is 3D.
- **One narrow task family**, five building types, rule-generated Q&A. G-Eval is an
LLM-as-judge metric — better-correlated with humans than BLEU/ROUGE here, but still a
proxy. "Data available on request" rather than released.
- The "matches 671B" comparison is on *this* benchmark. It's a domain-competence claim,
not a general-capability one.
## Why it's the right playbook anyway
Strip away the BIM specifics and this is a template for industrial domain models, the
kind I think about constantly: you rarely need a frontier model. You need (1) a faithful
**text projection of your structured/3D data**, (2) a benchmark with **rule-computed
ground truth** so you're measuring competence, not vibes, and (3) **reasoning-supervised**
fine-tuning data — quality and chain-of-thought over raw volume. Get those three right
and a 14B model on two A6000s reaches the same place a 671B model does, at a fraction of
the inference cost. For anyone shipping AI into a real engineering vertical, that economics
is the whole game.
---
*Paper: [Developing large language model for BIM-based design with domain-specific
benchmark and dataset](https://arxiv.org/abs/2602.20812) — Lin, Cai, Ni, Zhou, Pan
(2026), arXiv:2602.20812.*
---
# This site is managed by Claude
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/blog/hello-world
> date: 2026-06-03
> tags: meta, ai, nextjs
GitHub READMEs are dead. After Claude and the wave of coding agents, your homepage
isn't a static profile — it's a living, agent-readable artifact you can hand to an LLM.
This site is **dual-native**: every page is both a clean human document and a
machine-readable surface. Try fetching [`/blog/hello-world.md`](/blog/hello-world.md)
or [`/llms.txt`](/llms.txt) — an agent gets structured text, you get the rendered page.
The whole site is maintained by a crew of Claude agents. New posts, logs, and
data updates are authored by skills that validate themselves before shipping.
## What's under the hood
The content layer is a single source of truth: MDX files validated with Zod, surfaced
identically to humans (this page), to agents (the `.md` variant), and to tools
(the MCP server). More on that soon.
---
# 153 autonomous runs, no new ideas: the nanoGPT speedrun frontier
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nanogpt-speedrun-frontier
> date: 2026-08-15
> tags: agents, evaluation, prime-intellect, research, benchmarks, explainer
[Measuring Autonomous AI Research](https://www.primeintellect.ai/blog/measuring-autonomous-research) (Elie Bakouch, Prime Intellect, 14 August 2026) is the largest public experiment of its kind I have seen: **153 autonomous runs across 18 frontier models**, each on its own 8×H200 node, running unattended for up to eight days.
The task is [modded-nanoGPT](https://github.com/KellerJordan/modded-nanogpt) track 3 — the optimizer speedrun. Train a 124M GPT to validation loss 3.28 in as few optimizer steps as possible. You may edit the optimizer, its hyperparameters, the schedule and the initialization. The dataloader, architecture, batch size, sequence length and data are frozen.
For scale, the post notes the comparisons: Anthropic's internal automated-R&D evaluation optimizes a model on a *CPU node*, and OpenAI's GPT-5.6 Sol system card reports nanoGPT Track 1 on a single H100 for under a day. This is a much larger instrument than either.
It also produced a negative result, and says so plainly.
## The scoreboard, and what it means
The metric is honest and easy to check. The tuned baseline the agents start from passes at **3,290 steps**; the human record claim sits at **2,600**. So there are 690 steps on the table, and "gap closed" is just `(3290 − record) / 690`. Every published percentage reproduces from that formula exactly.
Three things are worth saying about the shape of it.
**Nobody beat the human.** Best run: Fable 5 at 2,726 steps, 81.7% of the gap, after 8.7 days. The remaining 18.3% is where a human already is, and had already been for weeks.
**Nobody invented anything.** This is the post's own summary, and it is the sentence I would lead with:
> None of the runs produced a fundamentally new method; the winning ingredients are all similar to existing ones in the literature.
The improvements that win are known optimizer work — better preconditioning, caps and floors on update magnitudes, keeping the learning rate hot longer, weight averaging near the end. Given that the agents had **no internet at all** — a deliberate change from earlier experiments, where they over-anchored on existing PRs — rediscovering the literature from parameters alone is a real result. It is just not the result the phrase "recursive self-improvement" is usually deployed to suggest.
**The cost column reorders everything.** Flip the interactive to tokens-per-step-gained and the ranking falls apart. Grok 4.5 bought its steps at 0.27M tokens each and closed a quarter of the gap; GPT-5.6 Sol paid 11.7M per step for a third of it — a 43× spread that says nothing about rank. The model that looks best on both axes is **Opus 5**: second place, in under three days, at 0.49M tokens per step. Fable 5 wins outright but spends nearly three times that rate and takes 8.7 days to do it.
## The benchmark is a statistics exam
Here is what makes this task harder than it sounds, and it is all in the public rulebook, [`program.md`](https://github.com/PrimeIntellect-ai/frontier-automated-speedrun).
A record requires the mean of eight fixed seeds — `0xC0FFEE+0..7`, which the agent cannot touch — to come in below **3.27859**. The file derives that number itself: `3.28 − 0.004/√8`, described as one-sided p < 0.001 at a per-run σ of about 0.0013. The arithmetic checks out. The standard error of an eight-run mean is `0.0013/√8 = 0.00046`, and `3.28 − 3.09 × 0.00046 = 3.27858`.
Now look at what that does to the research loop. Per-run σ is 0.0013 — *larger than most of the improvements being hunted*. A single screening trial cannot tell you much of anything, and the two ways to get it wrong pull in opposite directions:
- Trust one run and you certify noise. A recipe with **no real gain at all** clears the bar on a single trial about 14% of the time.
- Distrust one run too hard and you throw away the thing you were looking for. At one trial each, a recipe that genuinely is 0.001 better *loses* its head-to-head roughly 29% of the time.
And every trial is real money: a run takes the whole 8-GPU node, so runs are strictly sequential. Deciding when to widen from one seed to three to eight *is* the research skill this benchmark measures. The blog says as much:
> The models all find similar ideas. What separates them is how they run experiments.
The failure modes it describes in the weaker models are all statistical, not intellectual: killing whole families on one seed, treating their own crashes as evidence the idea was bad, discarding small gains that don't clear the bar alone. Grok 4.5 lost row normalization twice — to its own scaling bugs, not to the method.
## The best thing in the post is one paragraph long
Prime Intellect put a noise estimate in `program.md` that was **deliberately slightly too large**. Then they counted who checked.
62 of roughly 100 runs measured the noise themselves instead of trusting the number they were handed — and those runs are concentrated at the top of the table. 42 went further and discovered something nobody had mentioned: rerunning the same recipe on the same seed *also* moves the loss, because GPUs are not deterministic. That residual is much smaller than seed-to-seed variance, so two recipes compared on a shared seed resolve differences a normal screen cannot, for identical compute. Several models rebuilt their screening protocol around it.
That is a beautifully cheap instrument. It is not a coding test or a knowledge test — it measures whether an agent treats its documentation as evidence or as a claim, and it costs nothing but a willingness to write down something untrue. I would like to see more evaluations do this, and I suspect it generalizes far past optimizer research.
## The harness is worth as much as the model
Kimi K3 appears twice in the table under two different harnesses, which makes it the closest thing here to a controlled comparison — and the gap between its two runs is larger than the gap between several adjacent *models*.
Under [Prime Agent](/articles/prime-agent), which hands the model a persistent IPython kernel instead of a tool menu, K3 reached 2,930 steps on 112M tokens and 488 tool calls. Under `kimi-code` it reached 2,974 on 682M tokens and 4,000 calls. Better record, **6.1× fewer tokens**, an eighth the tool calls — and *more* output tokens, which is the tell. It was writing programs, not issuing commands.
The traces show what that looks like: K3 built its own experiment driver, a loss-curve comparator, a routine to restore a clean baseline, and then a numerical laboratory for retuning Newton-Schulz coefficients — testing them in simulation before spending a GPU-hour, and revising its hypothesis when the theoretically cleaner update trained worse. This is the same [harness effect](/articles/harness-effect) that keeps showing up: the scaffold is not packaging around the model, it is part of the system being measured.
One caveat the blog does not foreground and the repository README does: that Prime Agent run is tagged **serial era**. It ran under `program-serial.md`, a variant used between 20 July and 13 August that made agents wait on each run instead of delegating to a subagent. So two things differ, not one. Five of the twenty rows carry that tag — including second and third place — and Prime Intellect says they are being rerun.
## The tension I keep circling
The post credits the top models with research taste: re-ablating the stack after every merge, dropping components that stopped helping, revisiting old negatives when the recipe changed. Opus 5 re-opened β2 tuning under a new recipe and it became a record. K3 deleted two mechanisms that had produced its previous record once a new normalization made them redundant. Fable, out of single-knob gains, started testing pairs that were individually worse but jointly better; one late re-probe was worth 31 steps.
Those are genuinely good research instincts. They are also, in part, **instructions**. From `program.md`:
> Roughly every ~8 ideas explored, do a pruning round: try dropping each component you've stacked on and keep only what still earns its place.
And:
> A better method than the baseline exists (the human frontier is well below it), so "no improvement found" / "baseline is optimal" is never a valid place to stop.
The rulebook tells every model to prune periodically and forbids all of them from concluding they are done. So some unknown share of what is being scored as taste is compliance — following a written procedure under fatigue, across days, without a human checking. That is a real and valuable capability. It is just a different one, and the experiment as designed cannot separate them. The clean version of this study gives half the runs a rulebook with those two paragraphs removed.
## What it does not establish
The authors are candid about most of this, which is why the post is worth reading in full.
**Variance is high.** Speedrun noise plus model-level randomness on a days-long process; they mitigate with at least three seeds per model, taking the best after 24 hours and continuing it. That is a sensible protocol and it is also a best-of-k selection, so single-model numbers carry more optimism than a single run would.
**The task may not transfer.** Their words: they "don't have strong conviction that methods developed in this kind of speedrun are inherently scalable or would be used in real model training."
**The records are partly reconstructions.** The repository builds each model's record PR from the state recovered from its traces; Muse Spark 1.1 reached 3,232 steps but its exact record file could not be reconstructed, so it has no PR at all. The README's leaderboard also lists Kimi K3 at 2,968 steps where the results site shows 2,930 and 2,974 for its two runs — a discrepancy that is probably "last record, not best," but is not explained anywhere I could find.
**Comparing across harnesses is comparing systems, not models.** Every row pairs a model with a specific scaffold at a specific effort setting, and the K3 pair shows how much that matters.
## Why it is still the right experiment
None of that undercuts the main thing. Claims about models doing autonomous research have gotten much louder than the evidence, and almost all of the evidence has been either private or tiny. This is 153 runs on real GPUs for real days with the traces, scratchpads, monitor reports and rulebook all published, and the headline finding is *modest*: the best model closed four fifths of a gap a human had already closed, using ideas that were already in the literature, with no internet to look them up.
The honest way to read the table is as a measure of experimental discipline under uncertainty — screening cheaply, widening on signal, resisting a conclusion the data can't support, re-testing what you already decided. Which, now that I write it out, is a fair description of what makes a human researcher good too, and a much better thing to be measuring than whether the model can name the trick.
---
# Qwen3.8, weights in hand: 98% of a 2.4T model is routed experts
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/qwen3-8-open-weights
> date: 2026-08-15
> tags: qwen, open-weights, moe, quantization, architecture, inference, vllm
When [Qwen3.8-Max was announced](/articles/qwen3-8-max) on 3 August, the open weights were a promise: 2.4 trillion parameters, 95B active, "next week." The [Qwen3.8 collection](https://huggingface.co/collections/Qwen/qwen38) is that promise landing, and it is now four repositories deep.
Which means the interesting work has changed. There is no technical report, and there probably will not be one. But there is a `config.json`, a weight index, a chat template, an FP8 exclusion list, two serving recipes and a genuinely rigorous third-party quantization study. That is more than enough to check the claims, and checking them turns up several things the model cards do not say.
## What actually shipped
| repository | params | license | created | downloads | likes |
|---|---|---|---|---|---|
| [Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) | 2.4T / 95B active | **`qwen3.8-max`** | 8 Aug | 6.4k | 949 |
| [Qwen3.8-2.4T-A95B-FP8](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8) | same, FP8 | **`qwen3.8-max`** | 8 Aug | 10.7k | 191 |
| [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) | 27B dense + vision | **Apache-2.0** | 5 Aug | 91.9k | 9.5k |
| [Qwen3.8-27B-FP8](https://huggingface.co/Qwen/Qwen3.8-27B-FP8) | same, FP8 | **Apache-2.0** | 13 Aug | 123k | 381 |
Two things jump out of that table before any architecture.
**The licenses are not the same.** The 27B is plain Apache-2.0. The 2.4T ships under a bespoke `qwen3.8-max` license that is MIT-shaped with two riders: products above 100M monthly actives or \$20M monthly revenue must display the model name in their UI, and anyone running a "Model as a Service or AI Work Assistant business" whose group revenue passes **\$50M over any twelve months** needs a separate commercial license from Qwen. Internal use is carved out explicitly, as long as you do not expose the model or its outputs to third parties. It is a reasonable license and it is not an open-source one, and "the first Qwen-Max-class model getting open weights" deserves the asterisk.
**The 27B is the release.** It has fourteen times the downloads and ten times the likes of the flagship, and it went up three days earlier. Note also that in both pairs the FP8 repo out-downloads the bf16 one while collecting a fraction of the likes — bf16 is what people bookmark and requantize from, FP8 is what they actually serve.
## Reading the architecture out of the files
Both models are the same design at two scales, and Qwen describes the stack in one line on each card:
> Hidden Layout: 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE))
The most useful thing you can do with a model card number is try to rebuild it. If the reconstruction lands, you understand the architecture; if it does not, you have found something.
It lands. Summing the config — 92 layers of 512-expert MoE, 23 gated-attention layers, 69 Gated DeltaNet layers, two untied embedding matrices and one MTP block — gives **2.446181T** against the weight index's **2.446183T**. The 1.6M-parameter residual is the layernorms, which I did not bother to count.
Three things fall out of the exercise that no card mentions:
**The 95B active figure needs the embeddings.** The compute path — routed experts, shared expert, router, both attention types — comes to 91.2B. You only reach 95.3B by counting the untied `embed_tokens` and `lm_head`, 2.03B each. That is a defensible convention, but it is a convention, and it is 4% of the headline.
**`q_proj` is twice as wide as you would guess.** `head_dim` is 256 with 64 query heads, so 64 × 256 = 16,384 — already 2× the 8,192 hidden size. But the actual tensor is `[32768, 8192]`, twice that again, because `attn_output_gate: true` fuses the output gate into the same projection. The attention block is genuinely wider than the residual stream it reads from, in both directions.
**The MTP block costs 26.4B parameters.** The multi-token-prediction head is not a small linear probe. It is a complete extra layer — its own gated attention, its own 512-expert MoE, plus a fusion projection — weighing 1.08% of the model. That is a larger draft model than most models. Whether it earns that is a question the serving recipe answers below, and the answer is "only at depth 3."
## The hybrid, and what it is supposed to buy
`full_attention_interval: 4`, so three Gated DeltaNet layers then one gated attention layer, all the way up. The DeltaNet layers are Mamba-shaped — `A_log`, `dt_bias`, a kernel-4 depthwise `conv1d`, and a fused `in_proj_qkv` that carries 16 QK heads and 128 V heads at head dim 128 (`[20480, 8192]`, which is exactly 16·128 + 16·128 + 128·128). They keep a fixed-size recurrent state. They do not keep a KV cache.
That is the entire pitch: at 256K context the 23 attention layers of the 2.4T want a KV cache that grows linearly, and the other 69 layers contribute a constant.
Except the advertised saving is only real if your runtime knows about it. The most careful GGUF publisher for the 27B quotes **256 KB of attention cache per token**, and 2 GB at 8K. The sixteen full-attention layers in that model need 2 · 4 heads · 256 dims · 2 bytes · 16 layers = **64 KB per token**. The quoted figure is exactly 4× that — and 4 is the hybrid interval, i.e. precisely what you get if every layer is given a cache. I have not read llama.cpp's allocator for the `qwen35` architecture, so I will not tell you which of "allocation detail" and "deliberate margin" it is. I will tell you it is worth checking on your own hardware before you size a card, because the difference is 1.6 GB at 8K and 6.4 GB at 32K.
## The 27B is the interesting model
Sort the 27B's benchmark table by what each row measures and a clean pattern appears that Qwen does not point at.
On anything agentic, the 27B beats **Qwen3.7-Plus** — a larger model from the previous generation — on all thirteen rows, mean margin +10.9. Several margins are not subtle: OSWorld-Verified 84.3 against 73.3, Vision2Web 62.9 against 42.1, RecreationBench 47.1 against 30.2. DeepSWE 1.1 goes from 14.2 to 42.2, a three-fold jump that scale does not explain and that reads like a benchmark the training mix learned to do.
Flip to the rows where the answer has to already be in the weights and it loses five of seven — GPQA Diamond, HLE, ERQA, RealWorldQA, OmniDocBench — with ERQA down 4.3 and HLE down 3.9.
That split is the most useful finding in the release. **A generation of post-training bought an enormous amount of doing and almost no knowing.** Which is roughly what you would expect, and it is still worth seeing measured: if your workload is agentic, a 27B from this generation genuinely substitutes for something much larger; if your workload is recall, it does not, and no amount of harness will fix that.
The number I would treat most carefully is QwenSWEBench, where the 27B scores 79.0 against the 2.4T flagship's 80.7. A dense 27B landing within 1.7 points of a 2.4-trillion-parameter model is an extraordinary claim, and it is on Qwen's own benchmark, run by Qwen, against models Qwen did not train. The three `Qwen*Bench` rows should be read as internal instrumentation, not as evidence.
## `reasoning_effort` is two sentences
Both cards advertise "official support for `reasoning_effort`" as a headline feature. It is implemented in the chat template, and you can read the whole implementation:
```jinja
{%- if resolved_reasoning_effort == 'xhigh' %}
{%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think
carefully through the task, validate key assumptions, consider plausible
alternatives, and prioritize correctness, consistency, and clarity in the
final answer.' %}
{%- elif resolved_reasoning_effort == 'low' %}
{%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your
thinking brief and focused, moving directly to the conclusion without
unnecessary elaboration.' %}
{%- endif %}
```
That is it. `medium` sets nothing at all — it is the untouched model, and `xhigh` and `low` are two English sentences prepended to the system message. There is no token budget, no separate decode path, no architectural switch.
This is not a criticism: the model was presumably post-trained to respond to those exact strings, which is what makes it "official" rather than a prompt you invented. But it has consequences worth knowing.
- The effort level lives in **prompt space**, competing with your own system prompt for attention.
- Any harness that replaces the system message silently drops it.
- You can replicate all three levels, or invent new ones, with a string.
Two more control details:
**`preserve_thinking` defaults to on**, and it means every prior assistant turn keeps its full `` block in context. Combined with Qwen's own recommendation to allow 262,144 tokens of reasoning per turn, an agentic loop can spend its context window on its own history of deliberation faster than you expect. Setting it false strips reasoning from all turns before the last user message.
**The 2.4T cannot stop thinking.** Its template raises outright: `Disabling thinking is not supported.` The 27B accepts `enable_thinking: false` and emits an empty `` pair. If you were planning to use the flagship for anything latency-sensitive, that is a design constraint, not a setting.
Tool calls also moved off JSON to an XML-ish form — `` — which reads oddly until you notice it means multi-line code payloads need no escaping at all. That is a real improvement for a coding agent, and it explains why quantizers are shipping "tool calling improvements" notes.
## What FP8 actually quantizes
Open `quantization_config` on the 2.4T FP8 checkpoint and read `modules_to_not_convert`. It spares:
- every attention projection — `q_proj`, `k_proj`, `v_proj`, `o_proj`
- every Gated DeltaNet projection — `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `conv1d`, `out_proj`
- the shared expert, all three projections, plus the router and the shared-expert gate
- `lm_head` and `embed_tokens`
- the entire MTP block
What is left is the routed experts, and the routed experts are **97.97% of the model**. So the FP8 checkpoint is not "the model in FP8." It is *the experts in FP8 and the model in bf16*, which happens to look the same from a distance because the experts are almost all of it.
The arithmetic confirms the reading. Take 2.3966T routed parameters to one byte, leave the remaining 49.6B at two, and you predict **2.270 TiB**. vLLM's recipe publishes the FP8 checkpoint at **2.27 TiB**. (The bf16 figure checks too: 4.450 TiB reconstructed against 4.45 TiB published.)
The 27B FP8 config makes the same choice at a smaller scale — the GDN gating path, both embedding matrices, every layernorm and the entire vision tower stay bf16 — with one artifact worth a chuckle: its exclusion list names `mlp.gate` and `mlp.shared_expert_gate`, tensors that do not exist in a dense model, along with a fused `in_proj_ba` that is not in the checkpoint either. Harmless, and clear evidence both configs came off one template.
### Three parties, one conclusion
Here is the finding I would actually carry away from this release, because it arrives from three directions that did not coordinate:
1. Qwen's **2.4T FP8** config refuses to quantize any Gated DeltaNet projection.
2. Qwen's **27B FP8** config refuses to quantize the DeltaNet gating path.
3. A third-party quantizer, measuring rather than guessing, found that lifting `in_proj_z` and `out_proj` by one precision step cost 0.16 GB and removed **11% of the remaining divergence** — the single best trade in their whole search.
In a hybrid GDN/attention model, the linear-attention path is the precision-critical part. If you are building your own quantization mix for this architecture, that is where the bits go.
## Serving it
Both [vLLM recipes](https://recipes.vllm.ai/Qwen/Qwen3.8-27B) are unusually candid, and three of their findings generalize.
**MTP depth 3, not 1.** MTP-1 measured **64.8% acceptance** and is not merely marginal — it is *negative* at scale: +3.4% at concurrency 1, −9% at 128, −23% at 256, because the draft pass displaces real work once the batch is compute-bound. Depth 3 is worth roughly 2.3× on per-user output rate (FP8/TP16: 130 → 307 tok/s/user). A 26.4B draft head is only worth its weight if you speculate deep enough to amortize it.
**Context length is a concurrency dial.** At `--max-model-len 262144` the engine reserved KV for **25** concurrent requests. At 9,240 — 8K in, 1K out — the same 70 GiB of KV served **506**. Twenty times the concurrency from one flag, and nothing about the model changed.
**Tensor parallel must divide 64 attention heads**, so only 1/2/4/8/16/32 are legal. The recipe walks through the consequence: FP8 needs 2,325 GiB, which is three GB300 trays by capacity, but TP12 is not a thing — so it is a four-tray, sixteen-GPU deployment. Capacity planning on this model is arithmetic on head counts, not on gigabytes.
Smaller items worth knowing: `--load-format fastsafetensors --safetensors-load-strategy lazy` cut weight load from 545s to 306s on a 1.32 TiB checkpoint; MXFP4 does not load on NVIDIA (use NVFP4); the 1M-context `--hf-overrides` key nests under `text_config` for the 27B but sits flat for the 2.4T; and hybrid models have a CUDA-graph failure mode where `assert num_cache_lines >= batch` means your capture size exceeded the *recurrent-state* cache, which is a separate resource from the KV cache and one most people have never had to think about.
## Running it on your own machine
The GGUF ecosystem produced two serious repositories within a day of each other, and they are interestingly different.
[unsloth](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) has 25 files, 868k downloads, and a card that says "Unsloth Dynamic V3.0 (preview) for SOTA quantization performance" with no measurement attached. [AtomicChat](https://huggingface.co/AtomicChat/Qwen3.8-27B-GGUF) has 16 files, 8k downloads, and a card that is essentially a small paper.
The AtomicChat card measures per-token KL divergence against the **bf16** weights — not against `Q8_0`, which is the usual shortcut — publishes the reference logits so you can measure your own builds against the same point, downloads its competitors' files and measures those on the same harness rather than quoting their published numbers, and flags the one size band where it loses. That is the right protocol, and it produces four results worth keeping:
- **Where the bits go beats how many there are.** Ten builds within one gigabyte of each other span 2.2× in divergence. Nothing changes but tensor assignment.
- **The ends of the network matter most.** Peak activation energy sits on layers 52–62, with a second peak on layer 0. Lifting the first four and last twelve helped more than widening the band to 32 layers.
- **`Q8_0` is not lossless.** 0.00064 divergence, 98.92% top-1 — it disagrees with the original on about one token in ninety-three.
- **A quant name is not a specification.** Three publishers ship a `Q4_K_M` for this model: 16.8 GB, 19.0 GB and 17.1 GB, spanning 1.9× in divergence.
There is also a nice architectural footnote: the MTP head never executes during a normal forward pass, so the importance matrix has nothing to say about it at any corpus size, and llama.cpp refuses to quantize it low rather than guess. It is pinned to `q5_k` in every file.
**One practical warning.** AtomicChat's repository contains no `mmproj`. Qwen3.8-27B is a vision-language model — its HF pipeline tag is `image-text-to-text` — and those quants are text-only. unsloth ships `mmproj-F16.gguf` at 0.93 GB, which is exactly the 0.466B-parameter vision tower I reconstructed from the weight index. If you want the 27B to see, that file is not optional and only one of the two repositories has it.
## What this release does not establish
There is **no technical report**. Every number in every table is Qwen's, produced on Qwen's harness, and three of the coding benchmarks are Qwen's own instrumentation. The 2.4T's headline claim — matching or beating Opus 4.8 and GPT 5.6 Sol on agentic coding — is now at least *checkable*, since the weights are public, but nobody has checked it yet.
There is **no training detail at all**: no token count, no data mix, no post-training description, nothing about how MTP was trained "with multiple steps," and no ablation for any architectural choice. The 3:1 hybrid interval, 512 experts with 10 active, head dim 256, a 25% partial rotary factor — all are presented as facts about the artifact rather than as decisions with evidence behind them.
And the thing I would most like to see is the thing least likely to arrive: an honest account of why DeepSWE 1.1 went from 14.2 to 42.2 in one generation. Three-fold jumps on a single benchmark, in a family where three of the benchmarks are the vendor's own, are exactly the results that deserve the most explanation and usually get the least.
What is genuinely good here is how much of the release is *legible*. The parameter counts reconstruct. The FP8 size falls out of the exclusion list. The vision tower's size matches the mmproj byte-for-byte. `reasoning_effort` can be read in full in nine lines of Jinja. That is not nothing — it is the difference between a model you can reason about and a model you can only benchmark, and for a 2.4-trillion-parameter flagship it is more than we usually get.
---
# Arcee's Open Models API: a model lab selling six models it did not build
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/arcee-open-models-api
> date: 2026-08-14
> tags: inference, open-weights, pricing, agents, api
Arcee's [Open Models API beta](https://arcee.ai/blog/open-models-api-beta) post is labelled a **1 min read**, and that is accurate. It announces that the API now serves models beyond Arcee's own Trinity family, lists six of them, gives a price table, and offers $5 in credits.
Worth reading anyway, for one sentence and one table.
## A model lab watching what you pick
The stated motivation is not the usual one:
> It also helps us better understand why people choose a particular model for a particular task. When the API is used across our products, we can learn which models users prefer for a given task, and more importantly, *why* they chose them.
>
> Those insights will help us consistently develop and deliver Trinity models that are exceptional, diverse, and widely adopted.
Arcee trains models. It is now serving DeepSeek's, Z.ai's, Moonshot's and Thinking Machines' — and saying plainly that a reason to do so is to learn where its own are not chosen.
That is an honest description of an inference business as a research instrument, and it is unusual to see written down. Most labs that add competitors' models to their API describe it as customer choice and stop there. The reasoning is sound: a lab with no serving surface only learns about its models from benchmarks and complaints, while a lab that routes real workloads sees substitution behaviour — which model people reach for when the task is long, which when it is cheap, which when it has to be right.
The launch catalog:
- **Trinity-Large-Thinking** (Arcee's own)
- **DeepSeek-V4-Pro (Preview)** and **DeepSeek-V4-Flash-Latest**
- **GLM-5.2** — the base that [GLM-5.3](/articles/glm-5-3) was post-trained from
- **Kimi-K3**
- **Inkling-Small**, from Thinking Machines, which the post goes out of its way to call "another American lab advancing the frontier of open-weight models"
That last aside is doing some positioning work: four of the six are Chinese labs, and Arcee names the American one specifically.
## The table is more interesting than the announcement
Prices are per million tokens, and the number the list does not draw attention to is the **ratio between them**:
| model | input | output | output ÷ input |
|---|---|---|---|
| deepseek-v4-flash-latest | $0.14 | $0.28 | 2.0× |
| deepseek-v4-pro | $1.74 | $3.48 | 2.0× |
| inkling-small | $0.50 | $1.20 | 2.4× |
| trinity-large-thinking | $0.25 | $0.80 | 3.2× |
| zai-org/glm-5.2 | $1.40 | $4.40 | 3.1× |
| moonshotai/kimi-k3 | $3.00 | $15.00 | 5.0× |
Both DeepSeek models charge exactly double for output. Kimi K3 charges five times. That spread matters because the workload this API is being pitched at — long-horizon agent work, launched the same day as [nac](/articles/nac) — has a token mix that shifts with the task, and the cheap-to-read model is not always the cheap-to-run one.
On a read-heavy job the ordering roughly follows input price. On a generation-heavy one it stops doing so. Kimi K3's input price is 21× DeepSeek-V4-Flash's, but at 5M in and 25M out the actual bill is **50× higher** — the output multiplier widens the gap by more than double. And GLM-5.2 overtakes DeepSeek-V4-Pro on that same mix ($117 against $95.70) despite being the cheaper of the two to read.
Arcee's own Trinity-Large-Thinking is priced to sit second-cheapest on input and to stay cheap on output — $0.25 and $0.80, undercutting Inkling-Small on both. For a lab measuring which model people choose, pricing its own model into the "obvious default" slot is a thumb on the scale worth noting when reading whatever conclusions come out of the experiment later.
## Launched alongside nac
The post is explicit that this ships the same day as [nac](/articles/nac), Arcee's open-source agent harness, and that the two are meant to inform each other:
> We built nac to support demanding agentic workloads that may run for extended periods, and over time, what we learn from nac will help us improve how the API routes, serves, and supports models for long-running tasks.
The pairing is the actual strategy. nac is Apache 2.0 and free; it is also a very good instrument for observing long-horizon agent workloads, because its architecture forces every unit of work through a named dispatch with a recorded episode. An orchestrator that plans in one model and dispatches workers to another is a natural place to learn which models are chosen for which kind of step.
One detail from the nac repository suggests the catalog is not settled: the most recent commit at the time of writing is *"drop minimax, trinity-mini, and trinity-large-preview from arcee register."* Three models removed from the client-side catalog on launch day.
## What this is not
It is not a technical post. There is no routing architecture, no latency or throughput figure, no serving stack detail, no context-length or rate-limit table, no availability or region information. "Beta" is doing real work in the title.
There is also no evaluation of any of the six models, which is a slightly odd absence given the stated purpose is to learn which is best for what. The learning is planned to come from usage, not from measurement — which is a legitimate choice, and one that only produces useful answers if the pricing does not distort the selection it is measuring.
Read it for the strategy sentence and the ratio column. The rest is a price list.
---
# DeepSeek Harness: an agent harness that refuses to send what it didn't log
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/deepseek-harness
> date: 2026-08-14
> tags: agents, harness, open-source, architecture, typescript, explainer
[deepseek-harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`) is an MIT-licensed agent harness from DeepSeek that you run with `npx @deepseek-ai/dsh web`. The npm package was first published on 2026-08-10; six versions shipped in the four days after that, the latest being `0.1.0-rc.6`. It is labelled a developer preview, and the README is blunt about what that means: **THERE WILL BE COMPATIBILITY-BREAKING CHANGES.**
The obvious thing to write about an agent harness is its agent loop. That turns out to be the least interesting part of this one. The loop is about what you would guess — claim input, assemble a prompt, call a model, run the tools it asked for, repeat while anything is owed. What is unusual is everything built *around* the loop to make it hold still, and the reason that machinery exists is legible in the commit history: the repository went from its first commit to that npm release in **61 days**, taking **12,293 commits across 65 active days** on the way. At least 209 of its 984 merged pull requests came off `codex/*` branches, which is a floor rather than a count.
That combination — a codebase moving faster than humans can review, and a product whose whole job is to be trustworthy about what it told a model — produced a design decision worth stealing.
## Everything is a plugin, and it means it
`dsh` is built on [Cordis](https://github.com/cordiverse/cordis), a plugin framework DeepSeek vendored into the repo at v4.0.1 and rescoped under its own namespace. Cordis is five ideas: a plugin contributes services to a shared context; a service claims a stable key like `ctx.tools` or `ctx.llm`; plugins declare what they need with `inject` rather than being boot-ordered by hand; communication is typed events; and every registration is a reversible effect that unwinds when its plugin unloads.
The architecture doc states the consequence directly, and unlike most claims of this shape it survives contact with the source:
> There is no privileged core to patch: you extend dsh by mounting a plugin beside the others.
The model adapter is a plugin. The tool registry is a plugin. The session log is a plugin. The agent loop is a plugin — `core/agent` owns the `Agent` interface and the live registry, while `core/agent-loop` is described as "the default driver implementing that interface." Swapping it is a config row, not a fork.
There are **219 workspace packages** under `packages/*/*`. Twenty-one of them are model-facing tools (`tool-bash`, `tool-fs`, `tool-lsp`, `tool-subagent`, `tool-terminal`, and so on). The rest are seams, providers, UI surfaces, and policy.
A running `dsh` is composed at boot from ordered layers: bundles stack in listed order, then the profile's own patch file, then the home-level one, then any `--patch` overlay. I parsed the three committed bundle patches to see what that actually produces.
The detail I found convincing is base's own header comment, which explains why a row whose value differs between modes is *not allowed* to live in base: a patch replaces a row's whole `config` rather than merging into it, so a mode-varying row belongs to each mode bundle, which restates it completely. That is a rule written down because someone expected agents to add rows to this file.
## One turn
A **step** is one model request plus the tools it calls. A **turn** is zero or more steps: it opens before its first input is claimed and closes once nothing is owed. Here is the repo's own flow block, verbatim:
```text
turn/start
claim next-step input plus one queued message
assemble prompt sections + tool schemas
-> agent/pre-step reject | enter(messages)
reject, or a first enter rewritten empty -> close the turn with no step
step/start
append entered messages as user/message
derive model history from the log
agent/request -> llm/stream -> assistant/chunk* -> assistant/message
tool/call* -> tools/pre-execute -> tools/execute -> tools/post-execute -> tool/result*
step/end
tools owe another request, or next-step input arrived -> claim -> next step
-> agent/turn-stopping
turn/end
```
Two kinds of thing are interleaved there. Some events are durable facts appended to the session log; the rest are live extension points, and most of those are around-middleware — a listener receives `next()`, and either wraps the call and delegates or owns the decision and returns without delegating.
The repo draws the same lifecycle as a sequence diagram, which adds what a linear list cannot: who talks to whom, and the branches. Both `alt` blocks are worth reading — a rejected pre-step leaves the turn open having spent no step, and a terminal request failure routes to an `agent/request-error` waterfall that returns a retry action or preserves the original error.
Note the line `derive model history from the log`. It is doing more work than it looks like.
## The check that makes the log the source
Most harnesses treat the transcript as a rendering of the conversation: the conversation lives in memory, and the log is written alongside it for display and debugging. `dsh` inverts this. The log is the source, model history is *projected* from it by `deriveMessages()`, and a runtime invariant refuses to let those two drift apart.
The whole of `packages/core/agent-loop/src/invariant.ts` is 63 lines. This is its core:
```ts
ctx.on('llm/stream', (options: GenerateOptions, next) => {
if (!isAgentLoopRequest(options)) return next()
if (!Object.isFrozen(options)) fail('a loop-built request must be frozen')
// ...
const expected = session.deriveMessages()
if (JSON.stringify(options.messages) !== JSON.stringify(expected)) {
fail(`llm request for session "${String(session.id)}" diverges from the
dispatch-time durable derivation (log-reconstruction desync)`)
}
// ... and the folded request header must match model, system, temperature,
// maxTokens, stop and tools
return next()
}, { global: true, prepend: true })
```
Every request the loop builds is compared, byte for byte through `JSON.stringify`, against a *fresh replay of the session log made at dispatch time*. If a plugin slips an extra message into the outgoing request without writing a session event for it, the request does not go out. It throws.
The `prepend: true` matters: it means a replay or mock listener that short-circuits the waterfall still cannot get in front of the check.
The failure this prevents is the quiet one. Injecting an unlogged message doesn't crash anything and usually makes the model behave *better* — it is exactly the sort of change that ships. What it destroys is reproducibility: from then on, the log no longer explains the answer, and "why did it do that?" has no reachable answer. The repo states the rule as **model-visible ⟺ logged**, and this is the line of code that makes it true rather than aspirational.
The same header check covers sampling settings, which I think is the sharper half. Temperature and tool schemas are part of what makes a run reproducible, so retuning one between the logged header and the actual call is treated as divergence rather than a tweak.
## The same idea, applied to tools
Tool execution gets the same treatment, and the repo's own pipeline diagram is the clearest statement of it. Two details in there are the log-is-source rule again, wearing different clothes.
`tool/call` is **logged before execution** — not after, not on completion. If the process dies mid-tool, the log still records that the call was attempted. And at the far end, `tool/result` is labelled *single model-facing outcome*: however the call actually went — denied by a guard, refused at the approval prompt, thrown inside the tool body, thrown by a wrapper, timed out — every path converges through registry normalization and `finalizeContent` into exactly one recorded result.
Note the dotted `throw` edges all landing on the same normalization box. A tool that raises does not produce a missing result; it produces an `isError` result that the model sees and the log records. That is what lets the invariant in the previous section hold for a turn where something went wrong, which is the only kind of turn where reproducibility actually matters.
## 219 invariant companions, and what they actually contain
Here is where I nearly published something wrong.
Every one of the 219 package directories contains exactly one `src/invariant.ts`. I checked the correspondence as a set difference in both directions: zero packages without one, zero orphans. My first instinct was to write that as "219 packages all enforce runtime invariants."
They don't. Only **35** of the 219 ever call `fail(...)`. The other 184 are 20-line files whose install function is empty:
```ts
/** No runtime invariant: this stateless seam owns types while implementations
enforce immutable-store checks. */
const install: InvariantInstaller = () => {}
```
That looked like ceremony until I read `scripts/package-invariants.ts`, which is what enforces the convention. A package missing its companion is a violation. An empty install function that does *not* carry a comment beginning `No runtime invariant:` is a violation. A non-empty install function that never uses its bound failure reporter is a violation.
So the number that matters is not 219 checks. It is **219 decisions** — every package has been made to answer "what runtime invariant do you own?", and 184 of them answer "none, because…" in a sentence a reviewer can disagree with. Absence is recorded rather than assumed. That is a much better idea than 219 checks would have been, and it is the kind of thing that only pays off at this repo's scale.
## The repo is built by the workflow it ships
The discipline makes sense once you look at how the code got written.
`.agents/notes/` holds **686 design notes** — 507 implemented, 143 archived, 25 proposed, and 11 rejected, kept deliberately as the record of what was decided against. (The raw file count is 1,386; every note has a `.zh.md` twin, and counting both would double it.) Alongside them, `.agents/skills/` holds eleven repo-specific skills with names like `dsh-prose-standard`, `dsh-doc-standards`, `dsh-find-simplifications`, and `dsh-archive-agent-notes`.
Of 984 merged pull requests, **209 came off `codex/*` branches** — a lower bound on machine-authored work rather than a total, since it only counts one agent's branch naming. Another 210 came from `worktree/*`, which I am not going to attribute either way.
There is more test code than source code: roughly 205,500 lines of TypeScript under `src/`, against about 222,200 lines of `.spec.ts` and `.e2e.ts`. And CI runs 27 standalone `verify-*` scripts plus eleven catalog generators re-run with `--check`, covering things most repositories leave to habit — dead documentation links, markdown wrapping, mermaid syntax, JSDoc on exports, whether the English and Chinese docs are still paired, whether the generated config and tool catalogs still match the code they describe.
Read together, these are one decision made repeatedly: once enough of the code is machine-written, every rule a human reviewer would have applied has to become a script, or it stops being applied.
## Interop, and one honest gap
`dsh` can drive other harnesses. `subagent-claude-code` invokes the official Claude Agent SDK in the delegating session's workspace and returns only the final answer through the shared subagent contract; `subagent-codex` and `subagent-acp` do the same for their respective agents. In the other direction, `hooks-claude-code` and `hooks-codex` run a user's *existing* hook configuration on the harness's own interception points.
The Claude Code bridge is refreshingly self-effacing about why it exists:
> A native cordis plugin could do everything this bridge does — more powerfully, with typed returns and no serialization boundary. **The bridge exists only as a compatibility path for the mapped CC command-hook subset.**
The model story is thinner than the plugin story. Only one first-party adapter ships (`llm-deepseek`); everything else routes through `llm-pi-ai`, a generic multi-provider adapter built on the third-party [`@earendil-works/pi-ai`](https://www.npmjs.com/package/@earendil-works/pi-ai). That is a reasonable trade — a new OpenAI-compatible gateway becomes configuration rather than a code change — but it does mean the polish gradient between DeepSeek's own models and everyone else's runs through a dependency they don't control.
## What I'd actually take from this
Ignore the plugin count. 219 packages is a consequence of the architecture, not evidence for it, and a smaller project copying that number would just be slower.
The transferable ideas are two, and both are cheap:
**Make the log the source, then check it.** If the context you send is derived from your durable record rather than accumulated beside it, then a divergence is a crash instead of a slow mystery. The check is a few lines and it runs on every request. Nearly every agent system I've read builds the request and writes the log as two separate acts of bookkeeping, and quietly hopes they agree.
**Make "no check here" a thing you have to say out loud.** The 184 empty invariant files are worth more than they look, because a missing check and a considered decision not to check are indistinguishable in most codebases, and a script can tell them apart here.
The caveats are real: this is a developer preview with breaking changes promised in capital letters, nine weeks old, and moving fast enough that any specific file I quoted may have been rewritten by the time you read it. Every number here is measured at commit `47f9438` (2026-08-13) — I cloned the repository rather than reading the README, because the README does not mention the invariant at all, and that is the only part I would still be thinking about a week from now.
---
# dots3-note Preview: 16B active parameters, and a critic that thinks before it scores
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/dots3
> date: 2026-08-14
> tags: llm, open-weights, agents, rl, multimodal, long-horizon, explainer
The dots team — Xiaohongshu's AI lab — open-sourced [dots3-note Preview](https://studio.dots.ai/dots/dots3-en.html): **280B total parameters, 16B active**, a 512K context window, and multimodal understanding across text, vision and speech, under Apache 2.0. Weights on [Hugging Face](https://huggingface.co/dots-studio/dots3-note-prev), code on [GitHub](https://github.com/studio-dots-ai/dots3-note-prev), architecture submitted to Transformers as [PR #47844](https://github.com/huggingface/transformers/pull/47844). A technical report is promised within a week.
"note" is the **lightest** of three planned models; jazz and aria follow.
The benchmark table is respectable and not the reason to read this. The reason is a training method for tasks that take longer than a working day.
## The problem TEMPO solves
Reinforcement learning on genuinely long-horizon agent tasks runs into two walls at once, and the dots team states both plainly: a single rollout "can take more than ten hours, making training prohibitively inefficient, while sparse rewards hinder effective credit assignment."
Actor-critic methods like PPO exist to fix the second problem. But:
> a critic estimates value through a fixed-compute forward pass. Unlike an actor, it cannot reason, reflect, or use tools to analyze the current state, making accurate value estimation difficult on complex problems.
That is the observation the whole method turns on. If a task is hard enough that acting well requires ten hours of tool use and reasoning, then judging whether it is going well is *also* hard — and a single forward pass through a value head is not going to manage it.
**TEMPO** — Test-time-scaled Value Estimation with Macro-step Policy Optimization — cuts the task into macro-steps, each several rounds of interaction. At the end of each, the **same agent switches from actor to critic** and uses test-time-scaled reasoning to estimate expected remaining return. The policy can then be updated mid-task rather than after ten hours.
Reported result: **+31.5% average score over the base checkpoint and +20.6% over GRPO** on ARC-AGI-3, reaching the same level in fewer steps.
## Evaluation is easier than generation
The claim underneath TEMPO is the interesting one, and dots found it during training rather than assuming it:
> Even when the agent cannot yet solve a problem, it can act as a critic to distinguish between two superficially similar states, identify the one that represents a genuine breakthrough in understanding the environment's rules, and assign clearly different value estimates.
Their worked example is a "place knights" puzzle where two training branches both ran 64 rounds without clearing a level — **identical by environment score**. Branch B had misidentified the objective and was searching under a wrong assumption; branch A had found the real conflict rule and was close to a feasible layout.
A scalar reward cannot tell those apart. A model that reads both trajectories can. That is the entire argument for making the critic a reasoning model, and it is why the dots team frames self-evaluation as the direction they intend to keep pushing: real-world tasks "lack verifiable reward signals, while relying on human experts to evaluate model outputs may not scale."
## The IMO result belongs here
At IMO 2026 in Shanghai, dots built "an internal harness around a branch of dots3-note Preview" that generated proofs recursively and used tools to evaluate and improve them. The committee's own graders awarded **7/7 on all six problems — 42/42**, a score seven of 666 contestants from 117 countries matched.
Two things are worth being precise about, because the result is easy to over-read.
It was **not this model**. It was a branch of it inside a purpose-built harness, and the [IMO write-up](https://studio.dots.ai/dots/imo-en.html) is a separate page from the model release. Nothing you can download reproduces it.
And it used **no formal language**. The model read the organizers' original LaTeX and worked in natural language plus Python — no Lean, no proof checker. The dots team is explicit about why: formalization "requires a person to translate a problem into a formal language," and most real problems resist that. So the only thing standing between a plausible-looking proof and a wrong one was the model's own critique loop, and then a human panel that reads for holes.
Which makes the IMO run an inference-time instance of the same bet TEMPO makes at training time. The proof lengths are the one signal that varies — 3, 10, 6, 5, 4 and 3 pages, and P6, traditionally the hardest slot, took one of the two shortest.
## Where it actually lands
Head-to-head across the 23 reasoning and agentic benchmarks in the appendix: ahead of Hy3 (18–6), GLM 5.2 (16–9) and Seed 2.1 turbo (12–5); behind DeepSeek-v4-flash (8–14), GPT-5.5 (8–17), Opus 4.8 (8–18) and Kimi K3 (2–7 on the nine rows they share). For a model with **16B active parameters** against 21B, 39B and 104B, the first half of that sentence is the notable one.
Two rows stand out, both on the benchmark this release is built around:
- **ARC-AGI-3 (arcagi3 harness): 6.9 against Opus 4.8's 1.5 and GPT-5.5's 0.4.** More than four times the next best. This is the benchmark ARC Prize designed for autonomous learning in unfamiliar environments, where complex tasks need thousands of interactions over 40–50 hours.
- **ARC-AGI-2: 81.4**, above Opus 4.8's 72.1 and below GPT-5.5's 85.0.
That second one needs its asterisk read. dots' note says results marked `*` are their own testing, and specifically for ARC-AGI-2: "We evaluated models on the official public evaluation set; unmarked results are official leaderboard scores from the private set." **dots3-note's 81.4 is starred. Opus 4.8's 72.1 is not.** So a self-run public-set score is sitting in the same column as an official private-set score, and on ARC-AGI that difference is not cosmetic. The ARC-AGI-3 general-harness row has the same shape — dots' number is starred, and so are most of the competitors'.
To their credit, the harness details are unusually complete: Terminus-2 with a 10-hour timeout for Terminal-Bench, OpenClaw 2026.6.1 with a GPT-5.4 judge for WildClawBench, live-swe-agent for the SWE suite, Hugging Face access blocked during agentic search to prevent leakage. That is more methodology than most releases publish, and it is what makes the asterisk asymmetry visible in the first place.
## The two benchmarks they released
Both are open-sourced alongside the model, and both target the gap dots says it cares about — tasks where the user does not state what they want up front:
- **[VibeSearchBench](https://vibebench.github.io/VibeSearchBench.github.io/)**: 200 tasks across 20 domains. Each starts with an ambiguous request, and a persona-driven simulator reveals constraints over multiple turns. The agent's predicted knowledge graph is matched against ground truth by nodes and triplets, scored by Triplet F1.
- **[VibeLifeBench](https://vibebench.github.io/VibeLifeBench_homepage/)**: 20 tasks across 10 domains, each spanning **20–30 stages** on a simulated timeline, with **1,247 atomic checks** on cross-stage state consistency, tool execution and final deliverables. Their example is a family trip that has to be re-planned as aircraft type, weather and flight status change underneath it.
Nobody scores well on either. On VibeLifeBench the whole field sits between 21.1 and 30.1, with dots3-note at 28.1; on VibeSearchBench, between 22.4 and 33.8, with dots3-note at 25.7. A benchmark where the best model in the world manages 30% is either badly designed or pointed at something genuinely unsolved, and the 1,247-check structure suggests the latter.
## What they say is wrong with it
The Limitations section is short and unusually direct:
> dots3-note Preview is an interim preview release. Reinforcement learning is not yet complete, and the model still has limitations in hallucination mitigation, the balance between text and multimodal capabilities, and overall stability.
And on the real-life results specifically: those tasks **run in simulated environments**, and turning them into real experiences needs "robust harnesses, connectors, data sources, safety and permission mechanisms, and product design."
That is the right caveat and it is load-bearing. The persona-driven simulator that plays the user is itself a model, so a system trained and measured against it may be learning to satisfy a simulator rather than a person. dots built the environments, the benchmarks, and the model being evaluated on them.
## What I'd take from it
Ignore the parameter count and the leaderboard position. The transferable idea is that **a critic should be allowed to think**.
Every value-based RL setup assumes evaluation is cheap enough to do in one forward pass — an assumption that holds fine when the task is short and breaks silently when the task is ten hours long. TEMPO's answer is to spend inference on the value estimate, and the evidence for it is a picture of a model correctly separating two trajectories that the environment scored identically.
If "evaluation is easier than generation" holds up in the technical report, it is the more useful half of this release than any benchmark row in it.
---
# The full-bandwidth transformer: the feedback channel is one token wide
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/full-bandwidth-transformer
> date: 2026-08-14
> tags: transformers, architecture, reasoning, paper, efficiency, explainer
[Full-bandwidth transformer](https://arxiv.org/abs/2608.08888) (arXiv 2608.08888, 2026-08-09) opens with an observation that is obvious once stated and easy to never state:
> Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded.
Every decoding step runs the full depth of the model, produces a rich final-layer state, uses it to pick one token from the vocabulary — and then throws the state away. The next step starts from that one token.
## How narrow is narrow
The paper counts it in bits: a sampled token carries `log₂|V|` bits between steps. Their model has a tied 100,352-token vocabulary, so **16.6 bits per step**, against a 1,536-dimensional hidden state.
I would not push the ratio too hard — a hidden state's float width bounds what it *could* carry, not what it does — and the paper doesn't either. The claim that holds is about the channel's *shape*: one discrete symbol from a fixed alphabet, versus a continuous vector.
That shape has a consequence, and it reframes something familiar. If the only way to pass intermediate state to your next step is to name it in the vocabulary, then you must **verbalize your own scratch work**. Which is a fair description of what chain-of-thought is:
> CoT sidesteps this by externalizing intermediate state into language: the model writes out partial results, subgoals, and bookkeeping, then conditions future computation on the written trace.
Reasoning traces on this reading are not primarily a thinking technique. They are a workaround for a 16-bit bus.
## The fix
**Latent feedback.** At each decoding step, fuse the previous top-layer hidden state with the sampled token's embedding through a gated linear unit, and use that as the next input. The state carried between steps goes from `s_t = a_{1:t}` — the token trace alone — to `s_t = (a_{1:t}, z_t)`, the trace *and* the most recent latent.
What I find persuasive is what it does not change. Standard transformer architecture, standard KV cache, standard language-modelling objective. The fusion is dimension-preserving, so nothing downstream needs to know. There is an appendix on vLLM compatibility.
The paper's phrase for the benefit is the right one: latent feedback lets "non-verbalized computation re-enter the stack with a renewed depth budget." A fixed-depth transformer has bounded serial computation per forward pass; feeding the top state back gives the next pass somewhere to continue from rather than somewhere to restart.
Training is the part that could have gone wrong. Naive recurrence destroys parallel teacher forcing and with it the ability to train at scale. Their answer is a **scheduled multi-pass objective**: introduce latent feedback late in pretraining, and mix in a small fraction of deeper feedback passes for stability.
## What it buys
At 1B parameters and up to 400B tokens, full-bandwidth transformers "match or approach standard transformers trained with roughly **1.5× more tokens**," at negligible per-token decoding overhead.
But the result I would actually build on is the quieter one. The feedback passes double as a **training signal on the hidden states**:
> In later feedback passes, the top-layer state is shifted, fused into the input of subsequent positions, and can influence losses at multiple future positions through causal attention. Thus gradients from later predictions backpropagate into earlier hidden states, encouraging them to be reusable as inputs rather than merely predictive at the output layer.
In the ordinary objective, the top-layer state is supervised only through the next token. Here it is also supervised by whether it is *useful to consume*. And the payoff survives without the mechanism:
> Empirically, this improves pre-training data efficiency even when latent feedback is not used at decoding time.
So there is a version of this that costs nothing at serving time: train with the feedback objective, decode normally, keep the representation gains. That is a much easier thing to adopt than a new decoding loop, and it is the finding most likely to show up in someone else's model.
## The result that gets destroyed
On the base model, latent-feedback decoding produces markedly shorter reasoning traces at equal or better accuracy — exactly what the bandwidth argument predicts, since computation that would have to be spelled out can ride the hidden state instead.
Then:
> Notably, the effect disappears after instruction tuning. We attribute this to the tuning data being off-policy with respect to latent-feedback decoding: the target traces were produced by (and imitate the verbosity of) standard token-by-token reasoning, so fitting them re-imposes the fully verbalized style regardless of what the state can carry.
This is the most interesting paragraph in the paper and it is reporting a failure.
A capability was trained in and then trained back out — by imitation data written by models that did not have it. The traces in every instruction-tuning set were produced under the old constraint, so they encode verbosity that the new architecture makes unnecessary, and fitting them teaches the model to keep paying a cost it no longer owes.
The fix they name is on-policy post-training under latent feedback, left to future work. The general shape of the problem is not specific to this paper: **architectural capabilities can be erased by post-training data that predates them**, and nobody notices, because the benchmark still passes.
## What it does not establish
The two limitations are the authors' own, stated plainly.
Everything is at **1B parameters**. Their intuition is that deeper models should benefit more, since a deeper stack's top-layer state carries more — but that is a hypothesis, and 1B is small enough that a 1.5× data-efficiency gain could plausibly shrink or grow at scale.
The **feedback schedule is a heuristic**. No ablation on how long the recurrence phase should run, no principled way to choose the number of recurrence steps; they point at Jacobi-iteration convergence diagnostics as a possible route.
I would add a third: the state-tracking probes that verify the extra bandwidth is used — completion tracking and delayed memory — are synthetic diagnostics built for the purpose. They show the channel carries something. They do not apportion the 1.5× between the wider channel and the extra training signal, and those are separable ideas with very different deployment costs.
## Why the framing sticks
Almost a decade ago, [Breaking the Softmax Bottleneck](https://arxiv.org/abs/1711.03953) made a structurally identical argument about the other end of the model: the output layer factorizes through a matrix of rank at most the hidden size, so no matter how good your representations are, the distribution you can express is capacity-limited by a shape.
This paper makes the same kind of argument about the *feedback* path. Not "the model is not smart enough" but "the pipe is too narrow, and everything you have interpreted as a reasoning strategy is partly an adaptation to the pipe."
Whether or not latent feedback is the right fix, that is a productive way to look at a decoder. The interesting question it leaves open is how much of what we currently call reasoning is thinking, and how much is just a model talking to itself because that is the only channel it has.
---
# GLM-5.3: the same base model, a month of post-training, and a cyber capability nobody planned
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/glm-5-3
> date: 2026-08-14
> tags: llm, open-weights, agents, post-training, rl, security, explainer
Z.ai's [GLM-5.3 post](https://z.ai/blog/glm-5.3) opens with a sentence most labs would bury:
> Scaling post-training is all we did for GLM-5.3.
It is the same base model as GLM-5.2. No new pretraining run, no new architecture, nothing to report about parameter counts because none of them changed. What changed is a month of additional post-training on the stack they had already built — IndexShare for long-context processing, SAO for RL on long-horizon tasks, and [slime](https://github.com/THUDM/slime) for large-scale asynchronous training — pointed at more environments, more diverse tasks, and more compute.
That makes the release unusually easy to read. Every number below is a post-training delta, which is a rarer thing to be able to say than it sounds.
## What a month of post-training bought
Fourteen benchmarks where both models are scored, GLM-5.3 ahead on all fourteen. The two that dominate the story:
- **Terminal-Bench 3.0: 4.6 → 28.3.** A 6.2× move, and the largest on the board. It is also the one to be most careful with — GLM-5.2 scored 4.6, which is close enough to the floor that the model was essentially not playing. Going from *not playing* to *28.3* is a real capability change, but it is not the same kind of evidence as moving a mid-range score.
- **SWE-Marathon: 19.4 → 42.5**, and **AutomationBench: 26.2 → 48.2.** Both roughly double, both on long-horizon agentic work, which is where Z.ai says the environment scaling was aimed.
At the other end, Agents' Last Exam moves 23.8 → 28.5, a 1.20× gain — the smallest of the sixteen. The gains are real and they are not uniform.
## Where the environments came from
The part of the post I found most interesting is not a benchmark. Z.ai describes the bottleneck moving off the model entirely:
> As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
Their answer is to synthesize the environments, and for a subset of tasks the reward signal too. Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent then attempts each task to confirm it is actually solvable. Verifiers are synthesized **without access to the reference solution**, and solver trajectories are used to find and close reward shortcuts. A verifier that passes oracle, no-op and unsolved-state checks produces a binary reward they consider reliable enough to train on directly.
That trio of checks is the load-bearing detail. An oracle check catches a verifier that rejects correct solutions; a no-op check catches one that accepts doing nothing; an unsolved-state check catches one that was already satisfied before the agent started. Those are the three ways a synthesized reward usually turns out to be worthless, and they are checkable without a human reading the task.
Z.ai is direct that this is not yet automatic: the pipelines "still require a meaningful amount of human-in-the-loop work."
The environments themselves are aimed at something closer to a job than an exercise. Their example is an ML infrastructure task where the model gets the same working environment as an engineer — compute clusters, storage, internal documentation, codebases, experiment results — and has to diagnose bottlenecks across a training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup without breaking correctness. Some tasks, they say, represent several days of work for an experienced engineer.
## The honest reading of the table
"The most capable open-weights model for coding" is a carefully worded claim and it survives checking. Against open models GLM-5.3 is 16–0 over GLM-5.2, 12–0 over Qwen3.8-Max, 10–4 over Kimi K3, and 7–2 over DeepSeek-V4 Pro. Against the closed frontier it is 6–9 versus Fable 5 and 4–9 versus GPT-5.6 Sol.
Counted across the whole field rather than pairwise, GLM-5.3 holds the top score on **3 of 16 rows** — CyberGym, AutomationBench, and GDPval-AA v2. Publishing a table where you lead three rows and lose thirteen is not the normal shape of a launch post, and the word doing the work in the headline is *open-weights*.
Two smaller notes on the table. Thirteen of the 112 non-GLM-5.3 cells are blank, so several head-to-head records rest on fewer than sixteen rows — the DeepSeek-V4 Pro comparison is nine rows, not sixteen. And ExploitGym is reported as `2h / 6h` pairs rather than a single score, so any ranking of that row depends on which budget you pick.
## Token efficiency is the better result
The claim I'd have led with is the one about cost, not capability.
At Max effort GLM-5.3 reaches 34.5% at roughly 75K output tokens per task, against GLM-5.2's 23.4% at 96K — more accurate *and* cheaper, which is the direction that rarely happens on its own. At High effort it reaches 31.4% at around 50K tokens.
Z.ai compares that last figure to "Claude Opus 4.8 at 29.5% with 120K," which is true and worth reading precisely: 29.5% is Opus's **Max** effort, not its High. The comparison is GLM-5.3's High against Opus's Max. As an efficiency-frontier argument that is legitimate — the whole point of the chart is that the curves sit in different places — but it is not a like-for-like row.
The post is also straightforward that the frontier still belongs to someone else: "GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort."
One detail in the figure's own subtitle deserves attention: the benchmark was **evaluated on Claude Code 2.1.207**. Every model in that chart was scored through a competitor's harness. Given [how much a harness shapes agent results](/articles/harness-effect), holding it fixed across models is the right call, and it is unusual to see it stated on the chart itself.
## The cyber result, and what it actually says
This is the part Z.ai describes as a surprise:
> As we scaled post-training, cyber capability developed faster than we expected.
They added vulnerability discovery data and environments to the training mix expecting the model to get better at finding and reasoning about flaws. What they report instead is that it began reasoning across multiple stages of exploitation and forming coherent plans for complete chains.
Both cyber claims in the post are accurate. GLM-5.3 does hold the top CyberGym score at 84.5 — narrowly, over Fable 5 at 83.8 and GPT-5.6 Sol at 83.6, but it is a genuine lead over closed models. And its gains really are largest further up the chain when measured against GLM-5.2: 2.23× on ExploitBench, 3.3× on ExploitGym at 6h.
Put the two sentences next to each other and they suggest a model leading at exploitation. The same table says otherwise. One rung up from discovery, GLM-5.3 sits at 54.4 on ExploitBench against 78 and 76.5. At full chains under a six-hour budget it clears 130 problems against 247 and 293. The gap to the closed models widens at precisely the rate the capability is described as growing.
Alongside this, Z.ai published a disclosure ledger: **2,436 findings tracked**, 53 publicly disclosed, 2,383 still under embargo, 1,097 rated critical or high, across 269 open-source projects. The detail that stops you is the age distribution — the oldest flaw was introduced in **1981**, and on average a vulnerability had been sitting in a codebase for **26.6 years** before it was found. That is a claim about the state of open-source security as much as about the model.
It is also the context for the release schedule. The weights are not out yet:
> We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.
A two-week hold between announcement and weights, explicitly attributed to safety evaluation, is a reasonable response to having just demonstrated automated vulnerability discovery at scale. It also means nobody outside Z.ai can check any of the above yet.
## What to take from it
The headline result is not the benchmark table, which shows a strong open-weights model that trails the closed frontier — a familiar position. It is the claim that a month of post-training on a fixed base moved fourteen benchmarks, several of them by more than 2×, with output token counts going *down*. If that reproduces when the weights land, the interesting variable in this release is the environment synthesis pipeline, not the model.
The caveats: Z.ai Code Bench is private, so its numbers cannot be independently checked by construction — a deliberate anti-contamination trade with a real cost. Thirteen table cells are blank. The weights are two weeks out. And the cyber capability that is described as emergent is, on the evidence published alongside it, still a discovery capability rather than an exploitation one.
---
# LFM2.5-VL-3B: the release where GUI grounding appears out of nothing
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/lfm2-5-vl-3b
> date: 2026-08-14
> tags: vlm, open-weights, on-device, edge, benchmarks, explainer
Liquid AI released [LFM2.5-VL-3B](https://www.liquid.ai/blog/lfm2-5-vl-3b) on 2026-08-12 — a 3.1B open-weight vision-language model aimed at edge deployment. It is a **non-reasoning** model by design: it answers directly, which keeps latency low and, as it turns out, is the single decision that most of the performance story comes back to.
The blog lists four improvements over LFM2-VL-3B: screen understanding, function calling, grounding, and multi-image input. Three of those are ordinary gains. One of them is a capability appearing from nothing, and it is not the one the post leads with.
## The row that changed
LFM2-VL-3B scored **6.0, 7.6 and 2.5** on the three ScreenSpot-v2 splits. Those are not weak scores; they are the scores of a model that could not do the task. LFM2.5-VL-3B scores 78.7, 81.2 and 82.2 — an average of 80.7 against 5.4, a 15× move filed under "significant improvements in screen understanding."
Liquid AI's own framing compares outward rather than backward: 80.7 "far ahead of the much larger Gemma-4-E4B (51.2) and Qwen 3.5 4B (78.5) and close behind the larger InternVL-3.5-4B (84.1)." I recomputed all four averages from the published splits and they reproduce exactly. But the comparison that tells you what happened is the one with its own predecessor.
For a model whose stated purpose is running on the device that owns the screen, this is the row that decides whether it is useful.
## Architecture, and where the parameters went
The vision path is a **SigLIP2 400M NaFlex** encoder — "NaFlex" meaning it handles native aspect ratios rather than forcing a square crop — feeding a token-compression stage built from PixelUnshuffle plus an MLP connector. The diagram makes the compression ratio visible: four encoder tokens become one model token.
That compression is why the time-to-first-token numbers work. A heavier encoder produces more tokens, and every one of them has to be processed before the first output token appears.
On the language side, LFM2.5-VL-3B builds on the same pre-trained base as the LFM2.5-2.6B text model, pre-trained on roughly **34T tokens**. Two details worth pulling out:
- The vocabulary was **doubled to 128K by extending the existing tokenizer in place**, specifically to support non-Latin scripts. Extending in place rather than retraining a tokenizer keeps the existing embeddings valid, which is the cheap way to do this and not the usual one.
- Vision pretraining was scaled **4× in tokens**, with a mixture of curated and synthetic image-caption, OCR, grounding and instruction-following data. The grounding gain — RefCOCO precision@1 from 57.1 to 87.9 — is attributed directly to scaling synthetic grounding data.
Post-training is SFT with knowledge distillation from a larger teacher, plus something Liquid AI calls **Antidoom training**, followed by multi-reward RL.
## The size-class claim, checked
I transcribed all 28 benchmark rows and recomputed each model's average. Every one reproduces Liquid AI's published Average row to within rounding, so the table is internally consistent — worth doing, because the claim rests entirely on that average.
The claim is narrow and it is stated precisely: LFM2.5-VL-3B averages **69.4**, which is *exactly* level with InternVL 3.5 4B (69.4, at 4.7B parameters) and 0.7 behind Qwen3.5-4B (70.1, also 4.7B). It beats both Gemma models, at 5.1B and 8B.
So "competitive vision performance against models twice its size" is supported. "Better than models twice its size" would not have been, and the post does not say it.
The head-to-head view makes this sharper than the averages do. Counted row by row across all 28 benchmarks, LFM2.5-VL-3B is **14W–14L against InternVL 3.5 4B and 14W–14L against Qwen3.5-4B** — a dead tie against both 4.7B models, from two different labs. Against the rest it is comfortably ahead: 27–1 over Gemma-4-E2B, 23–5 over Gemma-4-E4B, 22–6 over both 2B-class models.
Where it loses is worth naming. Qwen3.5-4B takes the document-heavy rows — DocVQA 94.8 to 91.1, InfographicVQA 80.3 to 70.2, OCRBench v2 58.7 to 47.5 — and MMMU-Pro 36.0 to 30.5. InternVL 3.5 4B takes ChartQA, MMMU, and all three GUI splits. If your workload is dense document OCR, the larger models are still worth their size.
One row moves the wrong way and the post does not mention it: **CountBenchQA drops from 92.2 to 87.3**, the only benchmark where the new model is meaningfully behind its predecessor. POPE also slips slightly, 89.2 to 88.7.
## Function calling, added to a VLM
New to the VL line: ToolSandbox goes **26.4 → 59.5** and BFCL v4 **20.5 → 32.5**. The blog positions this as "on par with Gemma-4-E2B and ahead of Qwen3.5-2B," which understates it — 59.5 beats Gemma-4-E2B's 56.5 and Qwen3.5-2B's 47.7, and only Gemma-4-E4B (61.6, at 8B) and Qwen3.5-4B (65.0) are ahead.
Both InternVL models are marked N/A because they do not support function calling at all. That is the more interesting fact in the row: on the benchmark where InternVL was beating LFM on GUI grounding, it cannot compete, and a model that can both locate a button and call a tool is a different product from one that can only do the first.
The text-only instruction-following numbers are less flattering. IFEval 82.3 sits behind both Gemma models (83.0 and 87.9); Multi-IF 59.4 is well behind their 69.4 and 77.4. This is a vision model with tool use bolted on competently, not a text model that also sees.
## Speed, and what is actually specified
The GPU measurements are properly specified: vLLM 0.26, BF16, a 512×512 image plus 1,024 input tokens, up to 256 output tokens, median of five runs per concurrency level, single H100 SXM5. On a 5-frame clip LFM2.5-VL-3B returns its first token in about **34 ms** where the Gemma models take around 200 ms. Sustained output throughput reaches roughly **11K tokens/s** at high concurrency — about 2× the 4B-class models — which Liquid AI works out to nearly 1B output tokens per day from one GPU.
The on-device figures are the ones to be careful with: **228 tok/s on an Apple M5 Max**, 116 on an AMD Ryzen AI Max+ 395, 20 on a Galaxy S26 Ultra, in about 3 GB. No quantization, prompt, or batch size is stated for any of them. Given that GGUF, MLX and ONNX builds all ship day one and would each give a different answer, these should be read as claims rather than measurements.
## What it is for
The honest summary is that this is a **screen-and-document model that fits on a phone**. It ties two 4.7B models on a 28-benchmark average, loses the dense-OCR rows to both, wins the real-world and grounding rows, and gained GUI grounding and tool calling in one release.
The non-reasoning choice is the through-line. It costs accuracy on the STEM benchmarks where thinking helps — MMMU-Pro 30.5 is the weakest column in the table — and buys 34 ms to first token and 11K tokens/s sustained. For an agent that has to look at a screen, decide where to tap, and do it again, that is the correct trade. For a model asked to reason about a diagram, it is not.
Two things I could not check: Antidoom training is named but not described anywhere in the post, and the on-device numbers have no stated conditions. Everything else in the table reproduces.
---
# MAGI-2 Preview: 114B parameters, 6B awake, and two sparsities doing the work
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/magi-2-preview
> date: 2026-08-14
> tags: video-generation, audio, open-weights, moe, architecture, explainer
Sand AI released [MAGI-2 Preview](https://github.com/SandAI-org/MAGI-2-preview) under Apache 2.0: a **114B-parameter unified audio-video generation model that activates just 6B parameters per token**. Text-to-video or image-to-video, ten-second clips, sound generated alongside the picture and muxed into the same file. Eight Hopper GPUs to run it.
Both halves of that headline are checkable against the published weights, and they check out — but the second one only works for a reason the card does not spell out.
## Where 6B comes from
The total is exact. The safetensors index reports `total_size: 228107858176`, which at bf16 is **114.05B parameters**. Reconstructing it from the published tensor shapes and the config gives 113.88B — 0.15% apart, close enough to say the decomposition is right rather than lucky.
The active figure is more interesting. The MoE is 256 experts across 12 heads — 3,072 expert slots — with top-6 routing *per head*, so 72 of 3,072 fire for any token. That takes the MoE weight in a layer from 3.02B down to 70.8M.
Apply that alone and you land at **7.71B active**, not 6B.
The rest comes off because of something visible only in the shapes. On layer 0, `linear_qkv` is `[27648, 3072]`; on layer 2 the same tensor is `[9216, 3072]`. Exactly three times as large, and `k_norm` goes `[384]` against `[128]` to match. The model carries **three sets of weights — one per modality** — in its dense layers and in the modality-specific shared expert of every MoE layer. A token is video *or* audio *or* text, never all three, so two thirds of those weights sit resident and idle.
Put both sparsities together and it comes to **5.96B**, against a stated "just 6B parameters per token." One active parameter in 19.1.
That is the design worth naming: MoE sparsity and modality sparsity multiplied, not just the MoE ratio everyone quotes.
## The stack
Forty layers, hidden size 3072. The config lists `mm_layers: [0, 1, 38, 39]` and MoE on layers 2 through 37 — dense at both ends, sparse through the middle.
That placement says what Sand AI thinks the hard part is. Mixing three modalities is treated as an entrance-and-exit problem: two dense layers at the bottom and two at the top, each carrying private weights per modality. Everything between them runs a *single shared attention* over the fused sequence and spends its capacity on routed experts. Dense where the modalities are still separate, sparse once they are already mixed.
The three modalities enter through their own embedders — video at 48 channels, audio at 64, text at 5120 — and leave through separate video and audio output heads. There is no text head: text is conditioning, not output.
The refiner is a different model, not a smaller copy of the same one. Its config gives **30 layers at hidden 4096** with 8 query groups, `mm_layers: [0, 1, 28, 29]` — the same dense-at-the-edges pattern — and **no MoE at all**. It also sets `local_attn_layers` to all thirty. That is a sensible split of labour: upscaling 512×896 to 1088×1920 is a local problem, so the second stage is dense, wider, shallower, and never looks far across the frame. It gets 14 GB and 5 denoising steps against the preview stage's 228 GB and 100.
## Four residual streams
The tensor names give away a technique the README never mentions. Every layer carries `mhc_alpha_pre_attn`, `mhc_bias_res_attn` shaped `[4, 4]`, and an `mhc_norm.weight` of `[12288]` — which is 4 × 3072.
That is **hyper-connections**: instead of one residual stream with `x + f(x)`, the model maintains four parallel streams and learns how to mix them, with a 4×4 matrix deciding how each stream feeds the next block. The config confirms it as `mhc_config: { num_stream: 4, alpha_init: 0.01 }`, and the residual state really is four times as wide as the hidden size — the embedders write into 12288, not 3072.
Two implementation details in `magi2_preview.py` are worth flagging because they are not obvious from the config:
- The connection matrices go through a **Sinkhorn-Knopp** normalization (`_sinkhorn_knopp_affine_fwd_kernel`), which makes them doubly stochastic — every stream contributes and receives a fixed total, so no stream can quietly dominate.
- The whole thing runs through a hand-written Triton kernel (`_hyper_connect_fwd_kernel`). Four residual streams is four times the memory traffic if you do it naively.
The attention has two further additions: **attention sinks** (one sink token per layer, via FlashAttention-3's `fa3_func_with_sink`) and **gating** — a `linear_g` projection per layer, 24 outputs on MoE layers and 72 on the modality-specific ones, one per query group.
## What you are actually downloading
307 GB, and **64 GB of it Sand AI did not train**: the text encoder is Qwen3.5-27B, the video VAE comes from Wan2.2-TI2V-5B, and the audio VAE is Stability's stable-audio-open-1.0. The repo names each one and links it, which is the right way to do this — but it does mean "114B open-weights video model" describes the transformer, not the system you run.
The 2 GB `turbo_vae` is Sand AI's own distilled VAE decoder, and it is on by default (`use_turbo_vae: true`). It is also the only distilled component in the release, which brings up the cost.
## The honest problem: 105 denoising steps
The README is unusually direct about this:
> Neither transformer has been step-distilled, so the denoising step count is where most of the wall-clock time goes.
The shipped configuration is **100 preview steps plus 5 refiner steps**. The preview stage generates at 512×896 and the refiner takes it to 1088×1920. A distilled release with "far fewer" steps is listed as coming soon, with no date.
So the model that exists today is the slow one, on purpose, and Sand AI is saying so in the release rather than after someone benchmarks it. What is not stated anywhere is how long 105 steps actually takes on the eight Hopper GPUs it requires — there is no wall-clock figure in the repo, the card, or the config.
Everything else about the runtime is specified in detail: `cp_size: 8` and `ep_size: 8`, so context parallelism (Ulysses) and expert parallelism both span all eight GPUs; guidance is 5.0 for video and 7.0 for audio; output is 12.5 fps over a 10-second clip; the video VAE stride is `[8, 16, 16]`.
There is even a `--deterministic` flag, and a commit whose entire purpose is *"add Inductor compile-time configs for bit-exact reproducibility."* For a diffusion model where a one-ULP difference changes the video, shipping bit-exactness as a supported mode is a real courtesy.
## Prompt enhancement is not optional in practice
The captions the model trained on are long and structured, so the pipeline ships a prompt-enhancement step that asks an instruction-following LLM for a **structured JSON caption** of the 10-second clip, then renders it to Markdown before encoding. Templates are included for T2V and I2V separately.
It talks to an OpenAI-compatible endpoint and is off unless you set an `API_KEY`. The README is candid that "a short hand-written prompt underuses" the model — which means the shipped quality bar assumes a second model in the loop that the checkpoint does not include. The repo does hedge this properly by shipping two already-enhanced example prompts so you can see what the model actually expects.
## What is missing
**No evaluation of any kind.** No VBench, no comparison against Wan, Kling, Veo, Sora or anything else, no human preference study, no ablation. For a release whose stated purpose is to explore "an efficient path to scaling video generation," there is no published evidence that the efficiency buys quality. The samples in `assets/` are inputs, not results.
**No wall-clock or cost figure**, as above.
**The technical blog is unreachable.** The architecture, training system and data pipeline are described at [sand.ai/blog/magi-2-preview](https://sand.ai/blog/magi-2-preview), which sits behind a WAF that returns a challenge page rather than content. Everything in this article therefore comes from the repository, the model card and the published weights — which turned out to be enough to verify the headline numbers, but means the training and systems claims are unexamined here.
## Why it is worth the attention anyway
The parameter accounting is the result. A model that keeps three modality-specific copies of its dense weights and routes 72 of 3,072 expert slots per token gets to 19× sparsity without either mechanism being exotic on its own — and both are legible in the shapes, which is rarer than it should be.
The rest is a set of choices that are individually defensible and unusual together: hyper-connections with four Sinkhorn-normalized streams, attention sinks and gating, modality-private layers only at the edges, a distilled VAE decoder but undistilled transformers. It is a lot of recent architecture research in one checkpoint, shipped Apache 2.0 with the shapes visible.
The thing I would want before recommending it is a number — any number — comparing its output to something else.
---
# MiniMax Music 3: a five-minute song is 9,000 steps of a 2.5 kbit/s code
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/minimax-music3
> date: 2026-08-14
> tags: audio, music-generation, open-weights, flow-matching, rvq, explainer
[MiniMax Music 3](https://huggingface.co/MiniMaxAI/MiniMax-Music3) generates complete songs up to five minutes long from lyrics plus a music description, at 32 kHz 16-bit stereo. The weights went up on 2026-08-07 and the card was still being edited on 2026-08-13.
What makes it worth reading closely is that the card is unusually specific — it names a parameter count for four separate components — and all four of those numbers can be checked against the published files. They hold up, which is rarer than it should be.
## The shape of the thing
The split is between **structure** and **texture**:
- The **Global LLM (8B)** predicts the first RVQ codebook, frame by frame. That codebook is the semantic one — 16,384 entries — and it carries the song's long-range progression: where the chorus is, whether the vocal identity holds, how the arrangement evolves.
- The **Local LLM (646M)** predicts the remaining seven acoustic codebooks *within* each frame, restoring fine-grained detail the semantic codebook throws away.
The part that is not standard is what happens next. Rather than decoding from the discrete RVQ tokens, the synthesis stage fuses the **final hidden states** of both models and flows from there. The tokens are what the models predict; the hidden states are what actually gets rendered. MiniMax's argument is that continuous representations preserve more than the quantized codes do — vocal articulation, instrumental texture, temporal continuity — and the card is explicit that at inference time "waveform synthesis uses the fused LLM hidden states and does not require the discrete tokenizer decoder."
The tokenizer, in other words, is a training-time device. It shapes what the LLMs learn to predict and is then bypassed.
## Every number on the card, checked
I pulled the safetensors headers by HTTP range request — the 8-byte length prefix, then the JSON header that gives every tensor's dtype and shape — rather than dividing file sizes and hoping. That mattered twice.
The **Flow Matching** module is 9.73 GB across two shards. At bf16 that would be 4.86B parameters and the card's "2.4B" would be wrong by a factor of two. Every tensor in it is **F32**, so it is 2,431.9M, and the card is right.
The **Global LLM** is the nicer result. The config is `Qwen3ForCausalLM` with Qwen3-8B's exact shape — 36 layers, hidden 4096, intermediate 12288, 32 query heads over 8 KV heads — but `vocab_size` is **200,000** rather than Qwen3-8B's 151,936. Qwen3-8B is 8.191B parameters. Widening the vocabulary adds `(200,000 − 151,936) × 4096 × 2 = 393.7M` for an untied embedding and output head. That predicts 8.584B, and the index reports 8.584B.
So the card's two claims about this model — "initialized from Qwen3-8B" and "its embedding and output layers are first adapted to semantic music tokens" — are both visible in a single number, and the extra 48,064 vocabulary slots are where 16,384 semantic music tokens went.
The **Flow-VAE decoder** matches exactly too: `dav.pth` is 491.8 MB, which at fp32 is 123.0M parameters against a stated 123M.
The one component the card never mentions is a 25.2M condition encoder that takes 24 kHz audio in and produces conditioning at 44.1 kHz — the piece that would let you condition on a reference track rather than only on text.
## Why the repository is 57 GB
The parameters add up to roughly 12B. The repository is 57.35 GB. The difference is that it ships the entire model twice, in two runtime layouts — the SGLang-Omni one the card recommends, and a diffusers modular pipeline.
The arithmetic that shows these are the same weights rather than two models is clean: `flowmatching_vae.pth` is 2,457.1M parameters at fp32, and the diffusers `transformer` plus `condition_encoder` are 2,431.9M + 25.2M. Same total, not an approximation.
The oddity is the folder named `qwen_7B`. It holds an **`AbabForCausalLM`** — Abab being MiniMax's own model family — at the identical 8.58B shape as the `Qwen3ForCausalLM` sitting beside it. A directory named after one model family, containing another, holding what appears to be the same model converted for a different runtime. It is 18.48 GB of the repository and nothing in `modular_model_index.json` refers to it.
## The frame budget
The card's Limitations section gives two ceilings and does not connect them: songs "up to five minutes," and "audio generation is limited to 9,000 acoustic frames." Those only agree at **30 frames per second**, which is a number the card never states.
It is worth deriving, because it fixes the scale of everything else. Eight codebooks per frame — one at 14 bits, seven at 10 — is 84 bits per frame, or **2.52 kbit/s**. That is the representation the Global LLM is autoregressing over, and a full-length song is 9,000 steps of it against a 32 kHz stereo output that would be 1,024 kbit/s as raw PCM. Roughly 406× compression, with the flow-matching stage responsible for putting back everything that ratio removed.
The text side is separate and much tighter: 5,000 tokens total for lyrics and description combined.
## Control, and what it does not promise
Input is two fields. **Lyrics** may carry explicit section tags — `[Intro]`, `[Verse]`, `[Pre-Chorus]`, `[Chorus]`, `[Post-Chorus]`, `[Bridge]`, `[Instrumental]`, `[Solo]`, `[Outro]`. **Music description** covers style, emotional progression, vocal performance, instrumentation, arrangement, and production.
MiniMax recommends a three-part Structured Caption — Global Metadata (genre, BPM, key, scale, emotional arc, production profile), Vocal Details (gender, timbre, performance style, harmony, backing vocals, effects), and Arrangement (primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures, spatial effects). There is a `music-caption-rewriter` skill for expanding a short prompt into one, installable with `npx skills add`.
The card is honest about what that buys:
> Section tags and music descriptions provide generative control rather than strict symbolic guarantees. The generated tempo, key, instrumentation, lyrics, and song structure may not always match every requested detail exactly.
Which is the right way to describe a model that has no symbolic music representation anywhere in it. You are conditioning a sampler, not programming a sequencer.
## Running it
CUDA only, and non-streaming only — you wait for the whole song. Full precision fits under 24 GB of VRAM; with automatic CPU offloading it needs about 22 GB; and streaming the language model layer by layer with `apply_group_offloading` gets it onto an 8 GB card, slowly.
That last path is the interesting one for anyone without a datacenter, and it is a consequence of the hierarchy: the 8B Global LLM is the only piece that has to be resident for the long autoregressive run, so streaming its layers costs bandwidth rather than correctness.
## What I'd flag
The engineering claims check out, which is the main thing I set out to test. What the card contains no evidence for is **quality** — there are no listening-test results, no comparison against Suno or Udio or any other music model, and no objective audio metrics. There is a demo page and a single `assets/minimax_ttm.wav`. For a generative audio model, that is the entire evaluation.
The license file is present in the repository but the HF API reports no license field, so it is worth reading `LICENSE` directly before assuming anything about commercial use.
And 25 downloads against 440 likes, a week after release, is the signature of a model far more people want to hear about than can actually run.
---
# nac: an orchestrator that is not allowed to touch anything
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nac
> date: 2026-08-14
> tags: agents, harness, open-source, rust, context, explainer
[nac](https://github.com/arcee-ai/nac) is Arcee AI's open-source agent harness, Apache 2.0, about 98,000 lines of Rust across three crates. The [write-up](https://arcee.ai/blog/nac) opens with a diagnosis rather than a feature list:
> We think this couples two things that should be separate: the temporary context needed to perform an action; and the persistent state needed to continue a workstream.
That is the whole design, and everything else follows from taking it literally.
## The orchestrator cannot do anything
nac uses a thread-and-episode architecture adapted from Random Labs' [Slate](https://randomlabs.ai/blog/slate). A central orchestrator plans, decomposes, and decides what happens next. It has exactly one action available: launch threads.
> Importantly, the orchestrator's only action is launching threads; it cannot execute commands or edit files on its own.
The blog states the split as two lines of pseudocode, which is the clearest thing in it:
```text
orchestrator: decide and route, but do not act
workers: act, but do not expand the orchestration graph
```
Each dispatch starts a **worker** — a fresh process with a fresh model context, given the worker system prompt, the requested action, its tools, and any applicable skills. The worker calls the model and uses tools until the model returns a response with **no tool calls**. That response is the **episode**.
There is no separate summarization pass. The worker's system prompt tells it that its final answer should be a concise handoff, so the answer *is* the summary. That saves a model call and, more importantly, means nothing gets summarized twice.
## What survives
Once the episode exists:
> the worker's execution context **is discarded and never used as model context by the system again**. Its changes to the environment remain, but the episode is the persistent representation of the work.
A **thread** is just a named, ordered list of episodes. When the orchestrator assigns that thread more work, a new worker starts fresh with the thread's accumulated episodes — never the transcripts that produced them.
This is why nac has no compaction problem. Compaction exists to compress a transcript that has grown too long; nac never lets the transcript reach the orchestrator in the first place. The orchestrator reads episodes, and by the time it plans again the execution details are already gone. It is a different answer to context rot than compressing history: don't accumulate it.
**Thread weaving** is the other half. A dispatch can name source threads, and nac resolves each to its most recent retained episode and hands it to the worker as context — but those source episodes never join the target thread's history. Only the new episode does. Threads stay their own.
## A batch is a graph
The orchestrator ends a turn by emitting a batch of thread calls, each with a `name`, a free-form `action`, and optional `threads`, `skills` and `timeout`. Naming a source that is dispatched in the *same* batch creates a dependency edge; naming one that already finished just supplies context.
That makes the batch a DAG. nac rejects duplicate targets, validates acyclicity before executing, runs independent workers concurrently, and waits for the whole batch before letting the orchestrator plan again. The synchronization point is deliberate — the orchestrator never polls background work and never observes a half-finished world.
One place the code is more precise than the prose. The blog says a cyclic batch is rejected; `crates/nac-core/src/agent/tool_exec.rs` shows what actually happens on `DagError::Cycle` or `DagError::DuplicateName`: every *thread* dispatch gets an error result, while non-thread tool calls made in the same turn still execute normally. If your orchestrator mixes a query with a dispatch, that distinction matters.
## The seam they printed
The honest part of this release is a sentence most teams would have left out:
> Worker failures are not transactional: if a worker changes the environment and then exits before committing its final response, those changes may remain without a new episode, so a returned error means the environment may have moved ahead of persistent history.
Episodes persist only on success. Environment changes persist unconditionally, and live outside nac's state entirely. So a worker that edits files and then dies leaves the world ahead of the record, and nothing in the runtime knows.
Every system that separates durable state from a scratch context has this seam somewhere. What is unusual is printing it in the launch post rather than leaving it to be discovered.
## Harnesses as inference runtimes
The framing section is the part I expect to get quoted, and I think it earns it. Arcee traces harness evolution along two axes — **enriching context** so each model call gets denser information, and **expanding the action space** so the model can initiate more capable operations — from tool use through program execution, memory, multi-agent search, [Recursive Language Models](/articles/recursive-language-models), fresh-session harnesses, and finally Slate-style dispatch.
Then the claim:
> An agent inference runtime constructs context, schedules inference, executes effects, preserves state, enforces capabilities, and defines how work synchronizes, fails, resumes, and stops. A thin harness executes a model-tool loop. A runtime owns semantics that would otherwise exist only implicitly in its transcript.
And the mapping, which is what makes it concrete rather than a slogan:
```text
worker invocation = inference operation
thread = persistent program state
episode = committed workstream update
source thread = data dependency
dispatch batch = dynamic execution graph
```
Their summary line is the one worth keeping: **"judgment stays in tokens, invariants live in the runtime."**
It is worth reading this next to [DeepSeek Harness](/articles/deepseek-harness), which arrives at a related conclusion from the opposite direction. dsh keeps one agent loop and makes the *log* the authority, with a runtime invariant that refuses any request the log cannot reconstruct. nac keeps no shared log at all and makes *episodes* the authority, with a scheduler that refuses any batch it cannot order. Both are saying the harness should own guarantees the transcript used to own implicitly; they disagree about whether the transcript should exist.
Arcee also names two systems that make different choices — Onyx, which pushes orchestration control flow into persisted typed programs, and LongHorizon-Harness, which advances one globally audited task record through serial manager/executor/auditor rounds instead of parallel workstreams. Citing your neighbours accurately is a good sign.
## When it is the wrong tool
Stated plainly, which is rarer:
> For a single focused change that fits in one coding-agent session, going direct is simpler and often faster. That adds overhead because the orchestrator cannot perform the task itself; it still has to delegate to a thread.
The architectural purity has a fixed cost: a one-line fix still requires a dispatch. Their stated fit is work with a meaningful high-level objective, hard boundaries stated up front, a concrete definition of done, enough independent work to justify parallelism, and freedom for nac to choose its own decomposition — reproducing an ML paper, porting a large codebase, decomposed code review, large parallel change jobs on a dedicated branch and worktree.
## The meta-orchestrator pattern
nac ships an MCP server, so Claude Code or Codex can dispatch, monitor and steer nac jobs as tools. Arcee's preferred pattern is to make the interactive agent a **meta-orchestrator**: it works with you in a normal session, watches for work that is decomposable with a concrete definition of done, writes the job description itself, and hands it to nac to run in the background.
The capability boundary is drawn carefully:
> Through nac's MCP interface, the meta-orchestrator still cannot see a worker's discarded execution context or the underlying environment directly; the MCP server exposes no file or shell tools of its own.
So the outer agent gets the same view a human gets — orchestrator chat, thread episodes, recent events, the ability to steer — and no more. State must be queried; it is not pushed into the meta-orchestrator's context. The restriction that defines the inner orchestrator is applied to the outer one too.
## What is missing
**No evaluation.** No benchmark, no comparison against a single-agent baseline, no measurement of the token savings the architecture is supposed to produce. For a design whose central claim is that separating temporary from persistent context makes long tasks work better, there is no number showing it does. The evidence offered is a timelapse video and the fact that Arcee uses it internally.
**No cost accounting.** Running an orchestrator plus N parallel workers, each with its own context, is not obviously cheaper than one long session — it trades context length for context count. Which way that lands is exactly the thing an evaluation would tell you.
The repository is six commits old at the time of writing. This is a design worth taking seriously and a codebase worth waiting on.
## What I'd take from it
The transferable idea is the prohibition, not the architecture. Most multi-agent systems let the orchestrator do a little work itself when delegation feels heavy — and that is precisely when the orchestrator's context starts filling with execution detail and the original intent starts getting diluted. nac removes the option. The orchestrator cannot act, so its context stays a plan.
That is a constraint you could impose on a system you already have, without adopting threads, episodes, or Rust.
---
# Breaking the softmax bottleneck: your output layer is a rank-d wall
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/softmax-bottleneck
> date: 2026-08-14
> tags: architecture, transformers, paper, explainer, theory, language-models
[Breaking the Softmax Bottleneck](https://arxiv.org/abs/1711.03953) (Yang, Dai, Salakhutdinov, Cohen — CMU, ICLR 2018 oral) is a paper about a wall you cannot see from inside the model.
Deep learning does not have many negative results that matter. Most limitations turn out to be engineering — not enough data, not enough compute, the wrong optimizer. This one is linear algebra, it takes two lines to state, and it is still true of every model you use.
## The argument
Take the standard output layer. A network turns the context $c$ into a hidden state $\mathbf{h}_c$, you dot it with every word embedding $\mathbf{w}_x$, and softmax the result:
$$
P_\theta(x \mid c) = \frac{\exp \mathbf{h}_c^\top \mathbf{w}_x}{\sum_{x'} \exp \mathbf{h}_c^\top \mathbf{w}_{x'}}
$$
Now stack every context as a row of $\mathbf{H}_\theta \in \mathbb{R}^{N \times d}$, every embedding as a row of $\mathbf{W}_\theta \in \mathbb{R}^{M \times d}$, and the true log-probabilities as $\mathbf{A} \in \mathbb{R}^{N \times M}$. Your model's logits are $\mathbf{H}_\theta \mathbf{W}_\theta^\top$, and language modelling is now the question of whether that product can equal $\mathbf{A}$.
It cannot, if $d$ is small. The rank of a product is bounded by the shared inner dimension. Whatever the network does, its logits live in a $d$-dimensional subspace.
The move that makes this a paper rather than an observation is handling the obvious objection. Softmax is invariant to adding a constant to a row, so the model doesn't have to hit $\mathbf{A}$ — it can hit any member of the family $F(\mathbf{A}) = \{\mathbf{A} + \mathbf{\Lambda}\mathbf{J}\}$ of row-shifted variants, an infinite set. Surely somewhere in an infinite set there's a low-rank one?
No. Their Property 2: any two matrices in $F(\mathbf{A})$ have ranks differing by at most 1. The entire row-shift freedom is worth **one rank**. So the corollary is clean:
> **Corollary 1 (Softmax Bottleneck).** If $d < \text{rank}(\mathbf{A}) - 1$, then for any function family $\mathcal{U}$ and any parameter $\theta$, there exists a context $c$ such that $P_\theta(X \mid c) \neq P^*(X \mid c)$.
Read the quantifier: *for any function family*. Universal approximation doesn't help. You can make the network computing $\mathbf{h}_c$ arbitrarily deep and arbitrarily wide and it changes nothing, because the constraint is on the shape of the factorization, not on the expressiveness of the thing being factorized. All that effort is spent producing a vector that then has to squeeze through a $d$-wide waist.
## The part that is not proved
The bound only bites if $\text{rank}(\mathbf{A})$ is actually large, and the paper is upfront that this is a hypothesis:
> It is difficult (if possible) to rigorously prove this hypothesis since we do not have access to the true data distribution of a natural language.
The supporting intuitions are decent but soft. Language is context-dependent — "north" is followed by "korea" in a politics article and not in a U.S. history textbook. And if $\mathbf{A}$ *were* low rank, that would mean a few hundred basis distributions span every meaning humans express, and no one has ever found such a basis.
Neither of those is evidence. The evidence comes later, and it's better than the arguments.
## Two easy fixes, priced
Once you see the bound, two fixes suggest themselves, and the paper prices both out before proposing anything.
**Use an n-gram model.** Non-parametric, no rank constraint, universally approximates any language. Costs $N \times M$ parameters, where $N$ — the number of contexts — is unbounded. Generalizes badly, which is why the field left it.
**Raise $d$ until the bound stops binding.** To express a full-rank $\mathbf{A}$ you need $d \approx M$, so the embedding matrix costs $M \times M$. Slide the widget above to a modern vocabulary and watch that number: at $|V| = 151{,}936$ it's a **23-billion-parameter output layer**, for a model whose whole point was to be small. And empirically it doesn't even work — the paper notes, and everyone else had found, that pushing $d$ past a few hundred stopped helping on these benchmarks.
That is the real tension, and it is why the paper is interesting: *expressiveness and generalization are in conflict at the output layer*, and the naive ways to buy one spend the other.
## Mixture of softmaxes
The fix is small enough to quote in full. Compute $K$ context vectors instead of one, run each through the **same** embedding matrix, softmax each, and average the resulting **probabilities** with context-dependent weights:
$$
P_\theta(x \mid c) = \sum_{k=1}^{K} \pi_{c,k} \frac{\exp \mathbf{h}_{c,k}^\top \mathbf{w}_x}{\sum_{x'} \exp \mathbf{h}_{c,k}^\top \mathbf{w}_{x'}}
$$
Here is the whole thing in the [reference implementation](https://github.com/zihangdai/mos), essentially unedited:
```python
# one linear layer produces all K context vectors at once
self.latent = nn.Sequential(nn.Linear(nhidlast, n_experts * ninp), nn.Tanh())
self.prior = nn.Linear(nhidlast, n_experts, bias=False)
latent = self.latent(output) # (T*B, K*d)
logit = self.decoder(latent.view(-1, self.ninp)) # shared W, K times
prior = F.softmax(self.prior(output).view(-1, self.n_experts), -1)
prob = F.softmax(logit.view(-1, self.ntoken), -1)
prob = prob.view(-1, self.n_experts, self.ntoken)
prob = (prob * prior.unsqueeze(2).expand_as(prob)).sum(1) # mix probabilities
log_prob = torch.log(prob.add_(1e-8))
```
Two things worth noticing in that code. The embedding matrix `self.decoder` is used $K$ times — MoS does not buy $K$ embedding tables, it buys $K$ *readings* of one table, so the parameter cost is the `latent` projection and nothing else. And the mix happens **after** the softmax, on probabilities, which is the only reason any of this works.
Because the resulting log-probability matrix is
$$
\hat{\mathbf{A}}_{\text{MoS}} = \log \sum_{k=1}^{K} \mathbf{\Pi}_k \exp\!\left(\mathbf{H}_{\theta,k} \mathbf{W}_\theta^\top\right)
$$
and $\log \sum \exp$ is nonlinear, $\hat{\mathbf{A}}_{\text{MoS}}$ has no rank ceiling at all. It is a nonlinear function of $K$ rank-$d$ matrices, and nonlinear functions of low-rank matrices are generically full rank.
## The trap, which is the best part
Now the near-miss. Suppose you mix the **context vectors** instead of the probabilities — average $\mathbf{h}_{c,k}$ with the same weights, then take one softmax. Call it mixture of contexts. It looks like the same idea, it has the same parameter count, and it is completely useless:
$$
\mathbf{h}'_c = \sum_k \pi_{c,k}\mathbf{h}_{c,k} \quad\Longrightarrow\quad P_\theta(x\mid c) = \frac{\exp \mathbf{h}'^\top_c \mathbf{w}_x}{\sum_{x'} \exp \mathbf{h}'^\top_c \mathbf{w}_{x'}}
$$
which is the original softmax with a different $\mathbf{h}$. Rank still bounded by $d$. Mixing in feature space makes the *function family* richer and leaves the *ceiling* exactly where it was.
MoC exists in the paper as a control, and it is the sharpest instrument in it: same parameters, same layers, same hyperparameters, one design decision moved, and the theory says one should work and the other shouldn't.
There is also a footnote in the related-work section that has aged into the most consequential sentence in the paper:
> Although Shazeer et al. (2017) name their architecture as MoE, it is not a standard MoE and should be classified as MoC under our terminology.
[Shazeer et al. 2017](https://arxiv.org/abs/1701.06538) is the sparsely-gated mixture-of-experts layer — the direct ancestor of [every MoE LLM shipping today](/articles/mixture-of-experts-from-scratch), from [Switch Transformer](/articles/switch-transformer) to DeepSeek-V3 to Kimi. Under this paper's taxonomy all of them are mixture-of-*contexts*. They mix in feature space. They make the function family enormously richer and they do not touch the rank of the output layer by one.
That is not a criticism of MoE — sparse experts are solving conditional computation, not expressiveness — but it does mean the thing people reach for when they want "more capacity" is provably not the thing that lifts this particular ceiling.
## The evidence
Two measurements, and the first one is the kind that could have embarrassed everybody.
They compute the empirical log-probability matrix on PTB and estimate its rank. Softmax with $d = 400$ measures **400**. MoC with $d = 280$ measures **280**. Not near the bound — the bound, to the digit. MoS with the same 280 dimensions measures **9,981** out of a possible 10,000.
Then the dose-response. Sweep $K$ from 3 to 20 and rank climbs — 6,467, 8,930, 9,973 — with perplexity falling alongside it. At $K = 15$ rank has saturated at 9,981 and perplexity is at its best. At $K = 20$ rank does not move, because there is nothing left to buy, and **perplexity gets worse**.
That reversal is what makes the sweep an argument. A "more parameters help" story predicts a monotone curve. A "mixtures buy rank until rank runs out, and then you're just overfitting" story predicts a curve that turns exactly where rank saturates. It turns exactly where rank saturates.
Counting non-zero singular values is a roundoff-sensitive way to measure rank, so they plot the spectrum instead and the picture is unambiguous. Softmax and MoC dump ~96% of their normalized singular values below $10^{-9}$; MoS's are spread from $10^{-5}$ upward. Same conclusion, no thresholding decision required.
A third check, in the appendix: expected pairwise KL divergence between next-token distributions at different contexts — how much the model's prediction actually changes when the context changes. Softmax 4.763, MoC 4.864, MoS 5.284 on PTB test.
## Three controls that turn it into a mechanism
Any of the above is consistent with "MoS is a good regularizer and the rank story is decoration." The paper runs the experiments that separate those.
**Ablation.** MoC with matched everything is worse than MoS on both datasets — and on WikiText-2 it is worse than the plain AWD-LSTM baseline it was built from (65.98 against 65.40). So mixing per se isn't the win. Separately, training the baseline with MoS's hyperparameters is a disaster (74.86 against 58.95 on PTB), which rules out "they just found better hyperparameters."
**Regularization control.** On the 1B Word dataset, where overfitting is unlikely and no dropout is used at all: Softmax reaches 41.47 train / 42.77 test; MoS reaches 36.39 train / 37.10 test. MoS's *training* perplexity is 5.08 points lower. If the gain were regularization, training perplexity would have gone up, not down. (Their word for the generalization gaps is "similar"; strictly, MoS's is a bit narrower — 0.71 against 1.30 points, or ratios of 1.020 and 1.031 — which if anything strengthens the reading.)
**The inverse experiment, which is the one I'd point at.** If the mechanism really is the rank bound, then in a setting where the bound cannot bind, MoS should do *nothing*. Character-level language modelling is exactly that setting: $\text{rank}(\mathbf{A}) \leq |V| \approx 27$, and $d$ is in the hundreds, so there is no bottleneck to break. On text8, at matched parameter counts:
| model | params | test BPC |
|---|---|---|
| Softmax (hid 1024, emb 1024) | 8.42M | 1.49 |
| MoS-7 (hid 910, emb 510) | 8.45M | 1.49 |
| MoS-10 (hid 860, emb 452) | 8.43M | 1.49 |
Identical. A method that improves everything improves this too; a method that breaks a specific bound does nothing when the bound is absent. Papers that predict their own null results are rare, and this one went and measured it.
## What it won
State of the art at the time, at comparable model size:
| benchmark | best prior | MoS |
|---|---|---|
| Penn Treebank (dynamic eval) | 51.1 | **47.69** |
| WikiText-2 (dynamic eval) | 44.3 | **40.68** |
| 1B Word (their own softmax baseline) | 42.77 | **37.10** |
22M parameters on PTB against 24M baselines, and 35M on WT2 against 33M — so slightly under on one and slightly over on the other, which is the honest way to read "comparable." The 1B Word row is the one that ages best: 5.67 points on a dataset large enough that regularization tricks aren't doing the work, against a plain 2-layer LSTM softmax at 119M parameters, with hyperparameters they admit they never tuned.
They also bolt MoS onto a Seq2Seq decoder for dialogue on Switchboard and it wins on perplexity and on every BLEU precision and recall figure, which is a reasonable check that this is about context-dependent distributions in general rather than about language-modelling benchmarks in particular.
## So why isn't it in your model
Cost. $K$ softmaxes means $K$ passes over the vocabulary. Measured at matched batch size it's 1.9× on PTB, 2.5× on WikiText-2, 3.8× on 1B Word; at the settings where each model does its best, 2.8× and 6.4× on one GPU. Sub-linear in $K$ thanks to GPU matmul efficiency, but "sub-linear" still means two to three times the training cost, and that is before you consider that the output layer is now $K$ times the memory.
Then scale the setting. PTB has a 10,000-token vocabulary. A modern model has 150,000–200,000, and the vocabulary projection is already one of the most expensive tensors in the network — it's why chunked-and-fused cross-entropy kernels exist at all. Fifteen softmaxes over 200,000 logits per position is not a rounding error, it is the model.
So the field made a choice, and the choice was not obviously wrong: buy quality with data and depth, where the cost curve is friendlier, and leave the output layer alone.
## What happened to the idea
It didn't disappear so much as fragment into a small literature that no one reads together.
[Sigsoftmax](https://arxiv.org/abs/1805.10829) (Kanai et al., 2018) re-derives the bottleneck and argues the culprit is specifically the exponential in softmax, proposing a cheaper output nonlinearity that also escapes the rank limit — the same diagnosis, a one-softmax fix.
[Stolen Probability](https://arxiv.org/abs/2005.02433) (Demeter, Kimmel, Downey, 2020) finds a different consequence of the same geometry: embeddings in the interior of the convex hull of the embedding cloud can never be the argmax, no matter the context, so certain words are structurally unpredictable.
And the honest counterweight, [Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in Practice](https://arxiv.org/abs/2203.06462) (Grivas, Bogoychev, Lopez, 2022), goes looking for that failure in real systems: 13 of 150 public models have unargmaxable tokens, and they are rare enough not to matter. Which is a useful correction — a bound being real is not the same as a bound costing you anything — though note it tests one specific symptom, the argmax-unreachable token, not the broader claim that the expressible distribution is lower-rank than the one you want.
The bottleneck's authors moved on to [Transformer-XL](https://arxiv.org/abs/1901.02860) and [XLNet](https://arxiv.org/abs/1906.08237), and Zhilin Yang went on to found Moonshot AI, whose [Kimi K3](/articles/kimi-k3) ships a 7,168-dimensional hidden state against a 163,840-token vocabulary — a ratio of 23, in a model from the person who wrote the paper about the ratio. (Kimi K2 sits in the chart below at the same two numbers.)
## The ratio, today
I pulled `hidden_size` and `vocab_size` from published configs to see whether nine years of scaling relaxed the geometry. It didn't. The paper's own bottlenecked setup — the one where breaking the bound was worth 3.6 perplexity — had $|V|/d = 25$. Almost every open model checked is above that, several by a factor of six.
Anchoring on the paper's own numbers: from PTB's 10,000 tokens to a modern 151,936, vocabulary grew about 15×, while $d$ went from 400 to roughly 4,096 — about 10×. And the mismatch lands hardest on small models, which inherit a large tokenizer from their big siblings and get a fraction of the width to read it with. Qwen3-0.6B carries the same 151,936-token vocabulary as Qwen3-32B through 1,024 dimensions instead of 5,120.
I want to be careful about what that does and doesn't show. $|V|/d$ is geometry, not severity. Nobody knows $\text{rank}(\mathbf{A})$ for natural language; a 2026 model has representations no 2017 LSTM had; and the bottleneck may be so far from binding at this scale that it costs nothing measurable. The claim is narrower and, I think, harder to argue with: **the constraint everyone stopped worrying about is tighter now than when they stopped**, and the experiment that would tell us what it costs — an MoS-style rank measurement on a modern LLM's output layer — appears not to have been run.
## Why I keep coming back to it
The [full-bandwidth transformer](/articles/full-bandwidth-transformer) paper argues that the feedback path between decoding steps is one token wide — $\log_2|V|$ bits — while the hidden state that produced that token gets thrown away, and that chain-of-thought is partly a workaround for the narrow pipe.
That is the same argument, at the other end of the model. Both say: the network is not the limiting factor, the *interface* is. One narrow shape sits between a rich internal state and the thing you actually want, and everything upstream is spending its capacity on getting through it.
The best thing about the softmax bottleneck paper isn't the mixture of softmaxes, which nobody uses. It's that it demonstrated the move: take the part of the architecture that's so standard nobody writes it down, ask what it structurally cannot do, and then — this is the rare part — go and measure whether it costs anything. Softmax 400. MoC 280. Character-level, no change.
Most papers proposing an architecture would have stopped after the perplexity table.
---
# speech-to-speech: the OpenAI Realtime API, reimplemented as four swappable parts
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/speech-to-speech
> date: 2026-08-14
> tags: speech, voice-agents, open-source, realtime, latency, explainer
[speech-to-speech](https://github.com/huggingface/speech-to-speech) is Hugging Face's voice-agent pipeline: VAD → STT → LLM → TTS, each stage in its own thread, connected by queues. Apache 2.0, on PyPI, first commit 2024-08-07 and still being merged the day I looked.
Two things make it worth more than a glance. The first is a compatibility decision. The second is a latency trick that I think is the actual contribution.
## The compatibility decision
The server speaks the **core OpenAI Realtime GA event set** over WebSocket and WebRTC. Not a similar protocol — the same one, at `/v1/realtime`, such that the official OpenAI client connects to it by changing a URL:
```python
client = OpenAI(
base_url="http://localhost:8765/v1",
websocket_base_url="ws://localhost:8765/v1",
api_key="not-needed",
)
with client.realtime.connect(model="local") as conn:
...
```
The repo is careful about the size of this claim, and I want to repeat its wording rather than improve on it:
> This is a tested core subset, not a claim of full OpenAI Realtime API equivalence.
What is implemented inbound: `input_audio_buffer.append`, `session.update`, `conversation.item.create`, `conversation.item.truncate`, `response.create`, `response.cancel`. Outbound: speech start/stop, streaming transcription, audio deltas, tool calls, `response.done`. CI connects pinned `@openai/agents` `RealtimeSession` instances through the SDK's own WebSocket *and* WebRTC transports — so the compatibility claim is tested against the real client library, not against a hand-written mock.
That is the difference between "OpenAI-compatible" as a marketing word and as an engineering commitment.
## Ninety pipelines
Six STT backends, three LLM backends, five TTS backends, one VAD. The defaults are Parakeet TDT for transcription and Qwen3-TTS for output — both local — with the LLM slot pointed at anything speaking OpenAI protocols.
That last choice is worth pulling apart, because "OpenAI-compatible API" sounds like a dependency and is not one. Point it at `llama-server` on localhost:
```bash
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF -np 2 -c 65536 -fa on --swa-full
speech-to-speech serve \
--model_name "ggml-org/gemma-4-E4B-it-GGUF" \
--responses_api_base_url "http://127.0.0.1:8080/v1" \
--responses_api_api_key ""
```
Now the whole pipeline is local and the LLM is still reached over HTTP. Keeping the model behind a protocol boundary rather than in-process is what makes the slot genuinely swappable — the pipeline never learns which model it is talking to.
For fully disconnected operation, run the exact configuration once online to warm the caches, then set `HF_HUB_OFFLINE=1`. The repo is specific that this covers STT, LLM, TTS, Silero VAD, NLTK *and* Smart Turn assets — the kind of list you only write after being caught out by one of them.
There is also a `--stt none` mode that skips transcription entirely and hands each VAD-segmented audio chunk straight to an audio-input model over `/v1/chat/completions`. The README is blunt that this needs a model that actually accepts audio, and that the default `gpt-5.4-mini` does not.
## Smart Turn is the interesting part
Silero VAD tells you *that* speech stopped. It cannot tell you whether the person was **finished**. That gap is why voice agents interrupt people who paused to think.
[Smart Turn v3.2](https://huggingface.co/pipecat-ai/smart-turn-v3) classifies the turn using content and prosody, and speech-to-speech wires it in speculatively rather than as a gate:
- **Complete turns** start STT and the LLM immediately, with `--speculative_reopen_ms` (800 ms) before output is committed.
- **Incomplete turns** wait `--smart_turn_incomplete_delay_ms` (600 ms) before spending anything, and their output stays gated by `--smart_turn_max_wait_ms` (2 s).
- **If speech resumes during either delay**, the turn is reopened as a newer revision, the accumulated audio is re-emitted, and work from the previous revision is discarded *before it reaches the user*.
That third rule is what makes the first two safe. Speculation is only free if the wrong guesses are invisible, and revision numbering is the mechanism that makes them invisible. You spend tokens you might throw away in exchange for latency you cannot otherwise reclaim — a reasonable trade in a pipeline where every millisecond between "user stopped" and "audio starts" is audible.
It ships enabled by default, as a quantized CPU ONNX checkpoint, so the cost of running it is not a GPU.
## The part that gives it weight
> This pipeline runs in production as the conversation backend for thousands of [Reachy Mini](https://huggingface.co/blog/reachy-mini) robots.
Voice-agent demos are cheap and voice agents that hold up in a room with background noise are not. A deployed fleet is the only evidence that separates the two, and it is the reason to read this repo rather than one of the dozens of similar cascades.
## What to be aware of
**The Realtime surface is a subset**, and the repo says so. If your client depends on an event outside the tested set, it will not work, and "OpenAI-compatible" will have been true and useless simultaneously.
**No latency numbers.** For a project whose headline is "low-latency," there is no published end-to-end figure — no time-to-first-audio, no comparison against hosted OpenAI Realtime, on any hardware. The architecture is clearly built for latency; how much it achieves is unmeasured in public.
**Installation has sharp edges.** The Qwen3-TTS GGML backend's default PyPI wheel targets CUDA 12.8, and the README carries a table of alternate wheels for CUDA 13.x, 12.4, and CPU. That is honest documentation of a real problem, and also a sign of how much platform-specific machinery sits under `pip install speech-to-speech`.
**Some components moved to `archive/`** — Moonshine STT, MeloTTS, Parler TTS. Worth knowing before you build on a backend that is on its way out.
The design I would steal is the one at the boundary: implement someone else's protocol exactly enough to be tested against their client, then make everything behind it yours. It converts a hosted API from a dependency into an interface.
---
# Toast 1: what happens when you stop making the frontier model do the searching
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/toast-1
> date: 2026-08-14
> tags: retrieval, agents, rag, search, cost, explainer
[Toast 1](https://www.mixedbread.com/blog/toast-1) is Mixedbread's first specialised search agent, released 2026-08-13. The pitch is division of labour: instead of a frontier model burning its context window navigating a corpus, Toast 1 takes the whole search loop — decomposes the query into subqueries, gathers evidence, inspects sources, curates what matters — and hands back a package. The frontier model spends its tokens reasoning instead.
It runs standalone or as a subagent, and it is backend-agnostic: co-designed with Mixedbread Search but able to run over an existing index.
## The result worth reading twice
Harvey's LAB firm-knowledge benchmark, on a 33-task subset, with **one model and one evaluation** and only the retrieval stack changing between runs:
| configuration | tokens | turns/task | score |
|---|---|---|---|
| vanilla agent | 80.6M | 21.7 | 55 |
| + Mixedbread Search | 47.0M | 14.6 | 55 |
| + Toast 1 subagent | 23.0M | 11.2 | 55 |
The score is the finding. It does not move. If quality had gone up you would be looking at a better agent; because it is identical across all three rows, what the experiment demonstrates is that **57.6M of the vanilla agent's 80.6M tokens were not contributing to the answer**. They were the cost of looking.
I recomputed the deltas and they reproduce: 47.0/80.6 is −41.7% against a stated −42%, 23.0/47.0 is −51.1% against −51%, and 80.6/23.0 is 3.50× against a stated 3.5×.
Two caveats belong right next to that, and Mixedbread states the first itself in a footnote: this is a **randomly selected 33-task subset**, chosen "to make repeated comparative runs tractable." And a score that lands on exactly 55 three times is coarse enough that a small quality change would not necessarily show up in it. The token reduction is a much more precisely measured quantity than the quality preservation it is paired with.
## The headline benchmark, and whose numbers they are
On [OfficeQA Pro V2](https://www.mixedbread.com/blog/toast-1) — 90 questions on enterprise financial situations, released by Databricks — GPT-5.6 Sol running in Codex with Toast 1 as a subagent reaches **70% correctness at about $1.15 per task**. The previous best in Databricks' evaluation, Claude Fable 5 on Databricks Genie, was 60% at roughly $4.
The comparison that carries the most information is the one against itself: **GPT-5.6 Sol in Codex without Toast 1 reaches 33%**. Same model, same harness, and correctness doubles when the search loop is delegated.
The chart's own footnote is the thing to hold onto: "Genie and harness numbers as reported by Databricks; Codex + Toast 1 runs are ours." Half the points come from the benchmark's authors and half from the vendor being evaluated. That is a normal and disclosed arrangement, and it is still a different evidential status than a single evaluator running everything.
## As a standalone retriever
Evaluated as a retriever rather than a subagent — BrowseComp Plus, OfficeQA Pro and LongSeal, scored by NDCG@10 — Toast 1's fusion configuration lands in the same band as GPT-5.6 Sol and above Kimi K3, GLM, Opus 5 and Sonnet 5, while sitting an order of magnitude to the left on cost.
The chart shows something the prose does not dwell on: every other system is drawn as a *sweep*, a short line tracing what more reasoning effort buys. Toast 1's line is short and nearly flat. Whatever it is doing, spending more on it does not move quality much — which is the expected shape for a specialised model that is already doing the one thing it was trained for.
## The economics
A standard run is **$0.016–$0.023 per query at an eight-second median**; the fusion configuration is **$0.05–$0.07 at eleven seconds**. Token pricing is $0.30/M input, $0.04/M cached input with free cache writes, and $0.80/M output.
Against the frontier retrieval agents in the same evaluation, Mixedbread claims 7–11× cheaper. That multiple is the only cost figure given for the comparison group, so the band in the diagram above is their claim inverted rather than a published measurement — worth flagging, because the latency comparison beside it needs no such inference: **20 seconds to four minutes**, quoted directly.
An eight-second search subagent and a four-minute one are different products before price enters the discussion. If a frontier model is going to call search several times per task, the difference compounds into whether the task is interactive at all.
## What is not disclosed
No architecture. No parameter count. No training details, data, or method. No information about what Toast 1 is beyond what it does and what it costs. This is a product launch, not a model release, and every number in it is a system-level measurement.
That matters for one specific reason: the headline results are all **system** results. "GPT-5.6 Sol + Toast 1 in Codex reaches 70%" is a claim about a pipeline with at least three moving parts, and the contribution of each is not separable from the published data. The Harvey ladder is the closest thing to a controlled experiment on offer, and it is the one I would weight most — one model, one task set, one evaluator, one variable.
Mixedbread's own footnote places this alongside SID-1 and Chroma's Context-1 as a growing category of specialised search agents. That framing is right, and it is the more interesting story than any single benchmark: the bet is that retrieval is a distinct enough skill to be worth a dedicated model, and that the frontier model's context window is too expensive to spend on navigation.
The Harvey numbers are the strongest evidence for that bet I have seen stated plainly. Three quarters of a vanilla agent's tokens went to finding things, and removing that cost changed nothing about the answers.
---
# WorldClaw: a 3D world generator that is really a Blender programmer
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/worldclaw
> date: 2026-08-14
> tags: 3d-generation, agents, world-models, procedural-generation, paper, explainer
[WorldClaw](https://arxiv.org/abs/2608.05248) (arXiv 2608.05248, 2026-08-05, Tencent Hunyuan) generates large, freely explorable 3D worlds from open-ended text. The thing that separates it from most work in this area is the output format: not a radiance field, not a mesh soup, but **explicit instance-level assets sitting on a continuous terrain**, all of it editable afterwards in Blender.
The way it gets there is the interesting part. WorldClaw does not generate 3D. It writes programs that generate 3D.
## The pipeline
Formally the paper writes it as three functions and a compose:
- `P = F_plan(q)` — a text prompt becomes a structured specification of regions, terrain conditions and object conditions
- `T = F_terrain(P)` — the specification becomes a global terrain
- `O = F_region(P, T)` — regions that need detail get objects, conditioned on the terrain already built
- `S = Compose(T, O)`
The ordering carries the whole design. Terrain is built first and objects are generated *conditioned on it*, so a hut sits on a slope the system already knows about rather than being placed onto a surface it has to discover.
That middle panel is the system in miniature. The agent's deliverable is a Python function that composes noise octaves and landform primitives into a height field. Not a heightmap image — a program that computes one.
## The height field
Equation 6 is the core of the terrain stage:
`H(x) = Σ_r m̃_r(x) · [ h_r + Σ_k w_r,k N_r,k(x) + Σ_j α_r,j G_r,j(x) ]`
Each region *r* contributes a base elevation `h_r`, a weighted sum of noise octaves `N`, and a weighted sum of landform primitives `G` — and the whole contribution is gated by a **normalized** region mask `m̃_r`.
The normalization is the part worth dwelling on. Because the masks sum to one everywhere, adjacent regions blend rather than abut. A beach becomes a forest without a seam, and no post-hoc stitching step is needed. That is what lets the planning agent describe regions independently — writing "coastal", "dense jungle", "volcanic ridge" as separate specifications — and still get a single continuous world out.
It is also why this is a genuinely different approach from tiling. There is one field. It just happens to be authored per region.
## Placing objects is a solved geometry problem
For regions that need detail, WorldClaw renders the terrain from a viewpoint, generates a **composition image** conditioned on that render, segments the objects out of it, reconstructs each as a textured mesh, and then has to work out where each mesh goes in 3D.
That last step is where the paper does real work rather than prompting. Placement is recovered by solving for a similarity transform per object: a scale from the ratio of depths and focal lengths (`s_i = (Z_t/Z_o)(f_i^o/f̂_i)`), a rotation, and a translation, assembled into `T_place`. There is also a bounded contact constraint — the projected base of each object has to land within a tolerance band of a reference height, written as a two-sided inequality rather than an exact equality.
This is the difference between an agent that *asks* a model where the tree goes and one that computes it. The bounded constraint in particular is doing something specific: it permits a tree to sink slightly into a slope or stand slightly proud, which is what contact looks like on real terrain, while forbidding it from floating.
## Six models in a trench coat
WorldClaw uses **Claude Opus 4.8** as the agent model, with task-specific skills that wrap GPT-Image-2, SAM3, SAM3D and Hunyuan3D, executing into **Blender 5.1.1** on 4× NVIDIA H20 GPUs.
The Limitations section is more candid than most, and it is the most useful part of the paper:
> In our experiments, current open-source language models often struggled to generate procedural terrain and materials that were both executable and consistent with user requirements. Likewise, open-source image generation models frequently failed to produce usable semantic layout maps or to preserve object appearance and pose.
And then, plainly:
> Consequently, fully validating this decoupled pipeline at the current stage still requires capable models such as Claude Opus 4.8, GPT-Image-2, and Hunyuan3D.
Decomposing a task into stages is supposed to make each stage easier. Here it did the opposite for the two stages whose output has to be *executable*: a plan that becomes a Blender program either runs or does not, and a layout map is either segmentable or is not. Neither degrades gracefully when the model gets weaker.
The second limitation is the one anyone building on this should read twice. Several stages depend on LLM-generated programs, and:
> Errors in scale estimation, numerical parameters, or node connectivity directly manifest in the resulting 3D scene as inconsistent landforms, inaccurate material effects, or object layouts that deviate from the user intent, often necessitating multiple render–inspect–refine iterations.
A wrong number in a generated program is not a crash, it is a mountain in the wrong place. The render-inspect-refine loop exists because the failure mode is silent and visual.
## What the paper does not contain
There are no quantitative results. Section 3.2 is "Qualitative Results" and 3.3 is "Qualitative Comparison"; there are thirteen tables in the HTML and every one of them is a display equation. No user study, no CLIP or FID-style score, no timing table, no ablation with numbers.
For a system paper this is more defensible than it would be for a model paper — the claim is "you can build worlds this way and they are editable afterwards," and a render plus an instance mask demonstrates that. But it means nothing here is measured. The comparison figures show WorldClaw's worlds looking bigger and better organized than the baselines', and that is the entire evidential basis.
The third limitation is the honest counterweight: generating and reconstructing every object separately, then iterating refinement over terrain, assets and contacts, "incurs substantial inference latency and computational cost," and the pipeline "can be unnecessarily lengthy and inefficient for simpler scenes that holistic generation methods can synthesize in fewer steps." No wall-clock figure is given for either.
## Why it is still worth attention
The output format is the argument. A generated radiance field is a thing you can look at; a terrain with named regions, instance-level meshes, PBR materials and solved placements is a thing you can *open* — move a building, restyle a material, swap an asset, run physics against the ground. The paper's stated next step is generating objects as executable node graphs too, which would make composition and material logic editable in the same way the terrain already is.
The cost is that WorldClaw is currently less a model than an orchestration of four proprietary and open systems that its own authors could not substitute. Whether that is a stepping stone or a ceiling depends entirely on whether open models get good enough at writing Blender.
---
# Liquid time constants and gated delta rules: two literatures, one recurrence
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/ltc-gated-delta
> date: 2026-08-10
> tags: linear-attention, state-space-models, attention, math, explainer, open-source
There are two separate research literatures about neural networks that forget at an input-dependent
rate, and as far as I can tell they mostly do not read each other.
One starts in continuous time. [**Liquid Time-constant
Networks**](https://arxiv.org/abs/2006.04439) (Hasani, Lechner, Amini, Rus and Grosu, 2020) are
ODEs, motivated by the neural dynamics of *C. elegans*, analysed with stability theorems and solved
with numerical integrators. The other starts in discrete time. [**Gated
DeltaNet**](https://arxiv.org/abs/2412.06464) (Yang, Kautz and Hatamizadeh, 2024) is a linear
attention variant, motivated by retrieval failures in efficient Transformers, analysed through
online learning and implemented with chunkwise GPU kernels.
They are the same recurrence. Not analogous — the same. This piece derives that correspondence from
each paper's own equations, works through what each tradition figured out that the other did not,
and then reads [**LTCAttention**](https://github.com/Rikka-Botan/LTCAttention), an implementation
published today that sits deliberately between them.
This is a mechanism piece, so the maths is the point rather than an aside. Everything is derived
from the two papers' numbered equations, and the LTCAttention section is read off its source and its
checked-in result JSON. Where I run a calculation the papers do not — the discretization bridge, and
a scaling estimate at the end — I say so and show the working.
## Part 1 — What "liquid" means
A plain continuous-time RNN decays toward its input at a fixed rate: $dx/dt = -x/\tau + S(t)$, where
$\tau$ is a learned constant. Every input, every timestep, same $\tau$. LTC's move is to let the
decay rate be a function of the state and the input. Substituting
$S(t) = f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)(A - \mathbf{x}(t))$ gives the paper's Equation 1:
$$
\frac{d\mathbf{x}(t)}{dt} = -\left[\frac{1}{\tau} + f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)\right]\mathbf{x}(t) + f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)\,A
$$
Read the bracket. The coefficient multiplying $\mathbf{x}$ is the decay rate, and it now contains
$f$ — a neural network. So the **system time constant** is
$$
\tau_{\text{sys}} = \frac{\tau}{1 + \tau f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)}
$$
which is a number the network computes fresh at every point in time from whatever it is currently
looking at. That is the whole idea, and the name: a time constant that flows.
The paper's two theorems are what make this more than a reparameterization. Because $f$ is a bounded
sigmoidal nonlinearity, **Theorem 1** traps the time constant:
$$
\frac{\tau_i}{1 + \tau_i W_i} \le \tau_{\text{sys}_i} \le \tau_i
$$
and **Theorem 2** traps the state itself between $\min(0, A_i^{\min})$ and $\max(0, A_i^{\max})$,
"which guarantees that the outputs of LTCs never explode even if their inputs grow to infinity."
Those are unusual guarantees. A model whose decay rate is an unconstrained network output could in
principle be driven to instability by an adversarial input; LTC's cannot, by construction.
This is worth flagging because the same argument recurs, unattributed, throughout modern gated linear
attention. Every one of these architectures constrains its gate — Mamba2 through a softplus and a
discretization, Gated DeltaNet by requiring $\alpha_t \in (0,1)$, [Kimi K3's
KDA](/articles/kda-half-life) through a `gate_lower_bound` on $\log\alpha$. The reason is always the
same one LTC proved in 2020: an unbounded forgetting rate is an unbounded system.
## Part 2 — The bridge
Now the part neither literature states, which falls out of LTC's own **Algorithm 1**. Solving
Equation 1 in closed form is not possible, so the paper introduces a *fused solver* — a semi-implicit
Euler step that reads
$$
\mathbf{x}(t + \Delta t) = \frac{\mathbf{x}(t) + \Delta t \, f(\mathbf{x}(t), \mathbf{I}(t), t, \theta) \odot A}{1 + \Delta t\left(\frac{1}{\tau} + f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)\right)}
$$
Look at the denominator. Writing $g$ for the gate output,
$$
1 + \Delta t\left(\tfrac{1}{\tau} + g\right) = 1 + \Delta t\,\frac{1 + \tau g}{\tau} = 1 + \frac{\Delta t}{\tau_{\text{sys}}}
$$
because $\tau_{\text{sys}} = \tau/(1 + \tau g)$ by definition. So the entire update is
$$
\mathbf{x}_{t+1} = \bar{\alpha}\,\mathbf{x}_t + \bar{\alpha}\,\Delta t\, g \odot A,
\qquad
\bar{\alpha} = \frac{1}{1 + \Delta t/\tau_{\text{sys}}}
$$
**That is a gated linear recurrence.** Previous state times a scalar in $(0,1)$, plus a write. The
scalar depends on the input, through $g$. This is structurally identical to what Mamba2, Gated
DeltaNet and KDA do — the object those papers call $\alpha_t$ and describe as a "data-dependent
gating term."
The only difference is which approximation of the exponential you use. The exact solution of the
linear ODE over a step decays by $e^{-\Delta t/\tau_{\text{sys}}}$; LTC's fused solver uses
$1/(1 + \Delta t/\tau_{\text{sys}})$, which is the $[0/1]$ Padé approximant of that exponential.
In the regime these models actually operate in — long memory, so $\Delta t \ll \tau_{\text{sys}}$ —
the two agree to a fraction of a percent. LTC picked the Padé form because it is what makes the
implicit Euler step solvable in closed form; the linear-attention literature picked the exponential
because $\alpha^n$ composes cleanly across a chunk, which is what its parallel scan needs. Same
recurrence, two discretizations, chosen for two different implementation reasons.
Which means the three quantities have one meaning:
| tradition | symbol | this article's other coverage |
|---|---|---|
| continuous-time / LTC | $\tau_{\text{sys}}$, a time constant in seconds | Part 1 above |
| gated linear attention | $\alpha_t = e^{-\Delta t/\tau}$, retention per token | [KDA has a half-life](/articles/kda-half-life) |
| what you should think in | $n_{1/2} = \ln 0.5 / \ln \alpha$, a horizon in tokens | same |
I have argued the third column before, and the LTC connection strengthens it: a half-life is just
$\tau_{\text{sys}}$ in units a language model can be reasoned about in. $\tau$, $\alpha$ and $n_{1/2}$
are one number in three coordinate systems.
## Part 3 — What gating alone cannot do
If the story ended there, Gated DeltaNet would be LTC with better kernels. It is not, and the
difference is the delta rule.
A gated linear attention state is a matrix $\mathbf{S}$ holding key-value associations. Pure gating
updates it as $\mathbf{S}_t = \alpha_t \mathbf{S}_{t-1} + v_t k_t^\top$: scale everything down, add
the new pair. The problem the Gated DeltaNet paper identifies is that $\alpha_t$ is a single number
multiplying the entire state. It can dump everything, and it can hold everything, and it has no way
to express *forget this one fact, keep the rest*.
DeltaNet solved that with the delta rule, which subtracts the state's existing content at the current
key before writing the new one — but, as the paper puts it, "since this process only modifies a
single key-value pair at a time, the model lacks the ability to rapidly clear outdated or irrelevant
information, especially during context switches." One mechanism clears the table but cannot pick up a
single plate; the other picks up single plates but cannot clear the table.
The gated delta rule (Equation 8) is both terms in one product:
$$
\mathbf{S}_t = \mathbf{S}_{t-1}\left(\alpha_t\left(\mathbf{I} - \beta_t k_t k_t^\top\right)\right) + \beta_t v_t k_t^\top
$$
The two bars are the whole argument. Along any direction orthogonal to the current key, the surviving
fraction is $\alpha_t$ — the global forgetting knob. Along $k_t$ itself it is $\alpha_t(1 - \beta_t)$
— global decay *and* targeted erasure. Set $\alpha_t \to 1$ and you have DeltaNet; set $\beta_t \to
0$ and you have Mamba2; the useful region is the interior.
Worth keeping straight against the version of this update I covered in
[KDA has a half-life](/articles/kda-half-life). Kimi K3 writes it as
$S_t = (I - \beta_t k_t k_t^\top)\,\mathrm{Diag}(\alpha_t)\,S_{t-1} + \beta_t k_t v_t^\top$ — the
same two factors, transposed convention, but $\alpha_t$ is a **vector** with one entry per channel
rather than Gated DeltaNet's **scalar** per head. That is not cosmetic. A scalar $\alpha$ gives a
head one memory horizon; a diagonal $\mathrm{Diag}(\alpha)$ gives it a whole spectrum at once, which
is the difference between a head that forgets at one rate and a head that runs a filter bank.
## Part 4 — LTCAttention, and a third place to put a time constant
Both traditions above put the time constant on a **recurrent state**. [LTCAttention by Rikka
Botan](https://github.com/Rikka-Botan/LTCAttention), published today under MIT, puts it somewhere
else: on the attention score itself.
The construction is worth following because it is genuinely clever. Each KV head carries $M$ learned
directions, orthonormalized by QR so that $u_m^\top u_n = \delta_{mn}$. The first token of the causal
block, $x_0$, sets every mode's time constant through one linear projection:
$$
\tau_{h,m}(x_0) = \frac{\tau_{\min}}{\sigma\!\left(r_{h,m} + \delta_{h,m}\right)} > \tau_{\min}
$$
That is the LTC principle exactly — a positive, input-conditioned, *bounded-below* time constant, with
the sigmoid playing the role LTC's Theorem 1 played. Because $x_0$ is visible to every position in the
block, reading it keeps the controller causal.
For a query at $i$ and a key at $j \le i$, with key age $\Delta = i - j$, mode $m$ retains
$\lambda_m(\Delta) = e^{-\Delta/\tau_m}$, and the modes assemble into
$$
M_\Delta(x_0) = \prod_{m=1}^{M}\left[\mathbf{I} - (1 - \lambda_m)u_mu_m^\top\right] = \mathbf{I} + \sum_{m=1}^{M}\left(\lambda_m(\Delta, x_0) - 1\right)u_mu_m^\top
$$
which drops into the score as $s_{ij} = q_i^\top M_{i-j}(x_0)\,k_j/\sqrt{d}$.
The effect: the learned orthogonal complement passes through untouched, while each temporal mode is
an eigenvector with eigenvalue $\lambda_m$. Since $\lambda_m(\Delta) = a_m^\Delta$ with
$a_m = e^{-1/\tau_m}$, this is the same stable diagonal decay law as an SSM — just expressed as a
metric on an inner product rather than a state update.
### The factorization is the load-bearing trick, and it checks out
Applying a different $M_\Delta$ to every $(i,j)$ pair naively means building a $T \times T \times d$
object. LTCAttention avoids it by pushing the decay into the queries and keys separately, around a
fixed center $c$:
$$
q_i' = q_i + \sum_m \left(e^{-\frac{i-c}{\tau_m}} - 1\right)(q_i^\top u_m)u_m,
\qquad
k_j' = k_j + \sum_m \left(e^{\frac{j-c}{\tau_m}} - 1\right)(k_j^\top u_m)u_m
$$
I checked the algebra rather than taking it on faith. Decompose $q_i = q_\perp + \sum_m (q_i^\top
u_m)u_m$ using orthonormality; the transform replaces each modal coefficient by
$e^{-(i-c)/\tau_m}(q_i^\top u_m)$ and leaves $q_\perp$ alone, and symmetrically for $k$. Their inner
product is then
$$
q_i'^\top k_j' = q_\perp^\top k_\perp + \sum_m e^{-\frac{i-c}{\tau_m}}e^{\frac{j-c}{\tau_m}}(q_i^\top u_m)(k_j^\top u_m) = q_\perp^\top k_\perp + \sum_m e^{-\frac{i-j}{\tau_m}}(q_i^\top u_m)(k_j^\top u_m)
$$
and expanding $q_i^\top M_\Delta k_j$ directly gives the same thing. The center $c$ cancels, exactly
as claimed. The modal projections cost $O(TMd)$, so **scaled dot-product attention remains the only
quadratic operation** — the mechanism is free at the asymptotic level and the standard SDPA kernel is
still doing the heavy lifting.
There is a real numerical hazard hiding in that trick, and the code knows it. The factors
$e^{-(i-c)/\tau}$ and $e^{(j-c)/\tau}$ are individually huge or tiny even though their product is
bounded by 1; they cancel only algebraically. The implementation handles this two ways. It computes
the exponents in FP32 or FP64 regardless of the BF16 activation dtype, with a comment saying exactly
why. And it fixes $c$ at the middle of the context, `centre = 0.5 * (max_positions - 1)`, rather than
recomputing it per prefix — which both keeps cached keys valid as the KV cache grows and halves the
worst-case exponent.
The choice of $\tau_{\min}$ then finishes the job, and this is my favourite detail in the repository.
The default is `min_tau = max_positions / 12`. Combined with the centered origin, the largest
exponent magnitude is
$$
\frac{(T-1)/2}{T/12} = \frac{6(T-1)}{T} \approx 6
$$
**independent of context length.** Whatever $T$ you configure, the factorization's intermediate values
stay inside roughly $e^{\pm 6}$. That is not a coincidence; it is a bound chosen so the trick cannot
overflow.
### The experiment, and the number that worries me
The repository ships a real controlled study rather than a claim: three seeds, a paired comparison,
SHA-256 checksums on the tokenized data, one epoch over 287,588,352 FineWeb-Edu tokens consumed
without replacement, and the full result JSON checked in.
Reading the numbers straight out of `results/fineweb_edu_fullrank_29m_6layer_half_3seeds.json`:
| | validation loss | perplexity |
|---|---|---|
| standard | 4.13471 ± 0.02367 | 62.49 |
| LTC | 4.07281 ± 0.01777 | 58.73 |
| paired difference | **−0.06190 ± 0.00646** | |
The per-seed differences are −0.0544, −0.0702 and −0.0610 — negative in all three, with a spread ten
times smaller than the effect. As a paired result at this scale that is about as clean as three seeds
get, and the README is careful to say that "three seeds and one small model scale do not establish
broad scaling behavior."
Two confounds are worth quantifying, and the repository reports exactly the numbers needed to do it.
**Parameters.** LTC adds 345,600 of them, +1.20%. Borrowing the Chinchilla-form sensitivity
$\partial L \approx \alpha\,(A/N^\alpha)\,(\partial N/N)$ with $\alpha = 0.34$, a 1.20% parameter
increase at 28.8M is worth roughly **0.005 nats**. The observed effect is more than ten times that.
The gain is not just parameter count.
**Compute.** This is the one. LTC also runs **12.40% slower** (142,473 vs 162,638 tokens/sec, measured
and reported by the author). The comparison is token-matched, not wall-clock-matched. Spend that same
12.4% on more training tokens for the baseline instead, and the same scaling form
($\beta = 0.28$ on the data term) predicts a gain of roughly **0.061 nats** — which is, to two
decimal places, the entire measured effect.
I want to be precise about what that estimate is and is not. The coefficients come from a scaling law
fitted on a different corpus, tokenizer and budget, so the *absolute* numbers do not transfer; I am
borrowing only the sensitivity, and the error bars on that are wide. The near-exact agreement between
0.061 and 0.062 is a coincidence of a rough calculation, not a measurement. But the direction is
robust: at this scale, a 12% throughput penalty buys enough extra tokens to be the same order as the
observed quality gain. **The missing experiment is a wall-clock-matched run**, and until someone does
it the honest reading is that LTCAttention is better per token and undetermined per second. That the
author reported the throughput cost at all is what makes this check possible — most releases do not.
One further limitation, stated plainly in the repo: the released code is LTC-only, and the baseline
artifacts are "retained only as experiment provenance." So the comparison cannot currently be re-run
from this repository, only re-read.
## What each tradition knows
Setting the implementations aside, the two literatures have complementary blind spots.
**LTC knows about stability and it knows about time.** It has proofs that the time constant and the
state stay bounded under arbitrary input. It treats $\Delta t$ as a real quantity, which means it
handles irregularly sampled sequences natively — a capability the discrete-time literature mostly
gave up without noticing, because tokens arrive on a uniform grid. And it thinks in a unit, seconds,
that forces you to ask how long a memory is supposed to last.
**Gated linear attention knows about scale and it knows about writing.** It has the chunkwise parallel
algorithms that make these recurrences trainable on modern hardware at all, which is the entire reason
the idea reached billion-parameter models. And it has the delta rule — a way to modify one association
without disturbing the others that has no counterpart in the LTC formulation, where the "write" is
just $f \cdot A$ added to a decaying state.
LTCAttention is interesting mostly as evidence that the gap is crossable in either direction: it takes
LTC's bounded input-conditioned $\tau$, GDN's adaptive retention, and applies them to a third
substrate neither paper considered. Whether that particular hybrid pays for its 12% is, on the
evidence available, not yet settled. Whether the two literatures should be reading each other seems
to me much clearer.
---
*Sources: [Liquid Time-constant Networks](https://arxiv.org/abs/2006.04439) (arXiv 2006.04439,
Hasani, Lechner, Amini, Rus, Grosu) for Equation 1, Algorithm 1, and Theorems 1–2, read via ar5iv;
[Gated Delta Networks: Improving Mamba2 with Delta Rule](https://arxiv.org/abs/2412.06464) (arXiv
2412.06464, Yang, Kautz, Hatamizadeh) for Equation 8 and the complementarity argument; and the
[LTCAttention repository](https://github.com/Rikka-Botan/LTCAttention) at its 2026-08-10 state —
`README.md`, `model.py`, `config/`, and `results/fineweb_edu_fullrank_29m_6layer_half_3seeds.json`.
The three figures are LTCAttention's own, flattened onto white. The fused-solver-to-gated-recurrence
derivation, the verification of the query-key factorization, the $\tau_{\min} = T/12$ bound, and both
scaling estimates are mine and are shown in full above so they can be checked. All four interactives
are mine.*
---
# Muse Glimmer: an agentic model designed backwards from a 24 GB budget
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/muse-glimmer
> date: 2026-08-10
> tags: llm, agents, open-weights, on-device, attention, explainer
Most model releases describe an architecture and then mention, near the end, what hardware it runs
on. [**Muse Glimmer**](https://huggingface.co/meta-models/Muse-Glimmer-30B) — released by Meta
Superintelligence Labs on 2026-08-09 under Apache 2.0 — reads the other way round. Nearly every
structural choice in it is answering the same question, which is *how do we fit a competent agent,
its KV cache at 131K tokens, a vision encoder and a speculative-decoding drafter inside 24 GB at
once?*
That is a good question to design against, and the interesting thing is that you can check the
answer. Every number below comes from `config.json`, the safetensors index, and the file sizes in
the four repositories Meta actually shipped. The 24 GB claim is not a vibe; it is arithmetic, and it
closes with 2.4 GB to spare.
## The budget, and what makes it work
Start at the end. The card says 4-bit quantization shrinks the language model to "under 20 GB,"
leaving headroom for the KV cache, the perception encoder and the drafter within a 24 GB envelope.
Here are the measured artifact sizes from the GGUF repository, plus a KV cache computed from the
config rather than quoted:
The KV arithmetic is simple enough to do in one line. With 2 KV heads at head dimension 128 in
bf16, each token costs `2 × 2 × 128 × 2 = 1024` bytes per layer. Thirteen of the fifty-two layers
are global and hold the whole context; the other thirty-nine are capped at a 2048-token window. At
the full 131,072-token context that is **1.83 GB** — and 16.76 + 1.63 + 1.40 + 1.83 = **21.6 GB**.
Now take away the sliding-window pattern and make all 52 layers global. The KV cache becomes
**6.98 GB**, the total becomes 26.8 GB, and it no longer fits on the card the release is named
after. That is the sentence worth keeping: the 3-local-1-global stack isn't an efficiency
refinement applied to a model that already fit, it is the reason the model fits at all. Take away
the 16:1 GQA as well and the KV cache alone is 112 GB.
## The attention stack, verified line by line
The card describes the attention as "[Local, Local, Local, Global] repeating" with "RoPE
(θ = 500,000), local layers only." Both claims are checkable per layer, because `config.json`
carries two 52-element arrays.
`layer_types` gives `L L L G` thirteen times exactly. `layer_rope_theta` is 500,000 on every local
layer and **0 on every global one**, with zero mismatches across all 52. So the layers that see the
entire context run with no positional encoding at all.
That is worth pausing on, because this site has now covered three independent labs converging on it
inside a year: [Kimi K3](/articles/kimi-k3) applies NoPE to its full-attention layers, and
[Maple-Preview](/articles/maple-preview) sets `nope_on_global_attention: true` on the same 3:1
pattern at 24 layers. Muse Glimmer makes it explicit per layer rather than as a flag. The shared
argument is length extrapolation: a layer carrying no notion of absolute distance has nothing to be
surprised by when the context gets longer than anything it saw in training, and the local layers
underneath have already encoded order well enough to reconstruct it.
A few things the config says that the card does not:
- **`final_logit_softcapping: 20.0`** — logits are squashed through a bounded function before the
softmax, a Gemma-style stabilizer that the model card never mentions.
- **`qk_scale_factor: 3.87`** — the attention scale is not the textbook $1/\sqrt{d_k}$. At head
dimension 128 that would be 0.0884; this multiplies it by 3.87.
- **`output_multiplier: 0.19611613513818404`** — which is exactly $16/\sqrt{6656}$, a residual-stream
rescale tied to the hidden size.
- **`post_norm_eps: 1e-08`**, separate from `rms_norm_eps: 1e-05` — implying a post-norm alongside
the pre-norm rather than one or the other.
None of these is exotic on its own. Together they're a reminder that the published table of
hyperparameters is a summary, and the config is the document.
## It is a distillation, and the blog says so
The model card describes what Muse Glimmer is. The blog post describes where it came from, and this
is the part that most changes how you should read the benchmark numbers:
> **Pre-Training.** We trained Muse Glimmer on Muse Spark's outputs using logit distillation,
> leveraging a similar data mix as the teacher.
> **Mid-Training.** We trained the model on longer-context, more agent-heavy data with richer
> reasoning traces, alongside organic data.
> **Post-Training.** We combined supervised fine-tuning with a mix of on-policy distillation and
> reinforcement learning across general, reasoning, coding, and agentic domains.
Distillation from Muse Spark at *both* ends — logit distillation during pretraining, on-policy
distillation during post-training. Muse Glimmer is not a small model trained well; it is a large
model compressed, twice, with RL in between. That framing explains the shape of the results better
than "30B punches above its weight" does: what transferred is the teacher's *behaviour on agentic
trajectories*, which is exactly where the model is strongest.
It also sets up the safety argument later on, which leans on Muse Glimmer being "broadly weaker than
Muse Spark 1.0" — a claim that is much easier to make about a distilled student than about an
independently trained model.
## Speculative decoding, and why the same drafter is worth 3.1× or 1.5×
The second optimization is a companion "drafter" based on [DFlash](https://arxiv.org/abs/2602.06036)
that proposes an entire block of 16 tokens in a single forward pass, which the main model then
verifies in parallel. The shipped drafter is a 5-layer model at the target's full 6656 width —
2.56B parameters, 5.1 GB in bf16, 1.63 GB quantized. Calling it "lightweight" is fair relative to
30B, but it is 8.6% of the model and it has to be resident.
The headline is 3.1× on an RTX 5090 and 1.5× on an M4 Max, and the gap between those two numbers is
the whole mechanism. Single-stream decoding is memory-bandwidth-bound: you read the entire 17 GB of
weights to emit one token. Proposing sixteen and verifying them together amortizes that read across
all sixteen, so the gain depends on how much *spare arithmetic* the device has once the weights are
already moving. A 5090 has an enormous compute-to-bandwidth ratio and converts nearly the whole
block size into speedup; Apple's unified memory narrows that ratio, so verification stops being
nearly free.
The chart carries information the card's table drops. Those error bars span seven prompt categories,
and on the 5090 the DFlash result runs from roughly 132 to 340 tok/s. So the honest statement is
that speculation is worth somewhere between **1.8× and 4.5×** depending on what you ask, and 3.1× is
the midpoint of a wide distribution rather than a number you should expect on your workload.
## The benchmark table Meta loses a third of
Re-tallied against the better of the two rivals in each row:
Muse Glimmer leads **12 of the 22 scored rows**. Qwen3.6-27B takes 8, Gemma4-31B takes 2, and the
losses are not decorative: OSWorld-Verified by 9.7 points and TerminalBench 2.1 by 9.0, both to a
model three billion parameters smaller.
Filter that ledger to *agentic* and the profile becomes legible. The rows Muse Glimmer wins by a
distance — MCP Atlas by 21 points over Gemma, τ³-Banking by 41% relative, DeepSearch QA, Gaia2 — are
the ones measuring tool schemas and multi-turn task completion inside a scaffold. The rows it loses
are computer-use (OSWorld) and long-horizon terminal work (TerminalBench). For a model distilled
specifically on agentic trajectories, that is exactly the shape you would predict, and it is more
informative than a uniform win would have been.
Publishing it in that form is the least common thing about this release. A comparison table where
your own model is beaten in a third of the rows, by a competitor, in your own launch material, is
not the norm.
## The table where losing is the point
The chem/bio section inverts the usual reading of a benchmark, and it is worth explaining because
the presentation is initially confusing. Meta bolds the *most performant* model in each row — and
Muse Glimmer is deliberately not it in four of six:
| | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B | *Kimi K3* |
|---|---|---|---|---|
| MBCT | 41.5% | **50.6%** | 45.9% | *58.9%* |
| HPCT | 52.3% | **54.0%** | 48.7% | *59.6%* |
| VCT | 37.0% | **43.5%** | 33.7% | *48.0%* |
| WMDP (Bio) | **86.5%** | 85.9% | 84.8% | *89.1%* |
| WMDP (Chem) | 75.2% | **80.5%** | 74.8% | *84.2%* |
| Lab Bench (ProtocolQA) | **80.2%** | 75.8% | 69.1% | *81.9%* |
The argument is that Muse Glimmer sits "approximately in line with other models in its size class,
while showing strictly lower capabilities than larger open-weight models, suggesting that it is
unlikely to materially enable new threats upon release." Including [Kimi K3](/articles/kimi-k3) as an
uncontested upper bound on every row is doing real work here: it establishes that whatever Muse
Glimmer can tell you about wet-lab protocols, an already-open model tells you more.
The same inversion runs through the safety rows of the main table, and there Muse Glimmer genuinely
loses. Gemma4-31B has less than half its contextual-integrity violation rate (12.1 vs 26.4) and a
lower prompt-injection attack success rate (25.6 vs 28.4). Muse Glimmer's compensation is utility —
94.2 on AgentDojo against Gemma's 90.8 — which is the familiar helpfulness/safety trade stated in
numbers instead of prose. Whether 26.4% is an acceptable violation rate for a model explicitly
recommended for agents with "deep access to personal context" is a judgment the card leaves to you,
and it does at least give you the number to judge with.
The preparedness section is unusually explicit about its own reasoning: Muse Glimmer "does not fall
under the definition of 'Frontier AI' in Meta's Advanced AI Scaling Framework, since it is generally
less capable than Muse Spark," and its Cyber and Loss-of-Control designations are marked as
**inferred** from that comparison rather than measured directly. Naming an inference as an inference
is good practice. It also means two of the three risk designations rest on the distillation
relationship rather than on evaluations of this model.
## What actually shipped
I checked the "Released Artifacts" table against Hugging Face, because promised artifacts and
present artifacts are frequently different things. All four exist, in three sibling repositories the
card does not link:
| Artifact | Where | Size |
|---|---|---|
| BF16 weights | `Muse-Glimmer-30B` | 59.55 GB, 2 shards |
| 4-bit, 24 GB target | `-GGUF` / `muse-glimmer-30B-kquant-17gb.gguf` | 16.76 GB |
| 4-bit, 32 GB target | `-GGUF` / `muse-glimmer-30B-kquant-dynamic.gguf` | 19.65 GB |
| DFlash drafter | `-GGUF` / `dflash-kquant.gguf`, `-assistant` | 1.63 GB / 5.11 GB bf16 |
| Vision projector | `-GGUF` / `mmproj-kquant.gguf` | 1.40 GB |
| ExecuTorch builds | `-ExecuTorch-PTE` | metal + sm80, text and text-image |
The drafter's own `config.json` confirms the card's spec exactly — 5 layers, `block_size: 16`,
sliding window 2048 on all five, 32 query heads and 8 KV heads. The ExecuTorch repository ships
separate `.pte` files for Apple Metal and NVIDIA sm80, in text-only and text-plus-image variants,
which is how the M4/M5 numbers were produced.
One small correction to the card while I'm counting: it states total parameters as "~29.6B" twice.
The safetensors index says **29,776,626,688** — 29.78B. A 0.6% understatement, and I mention it only
because everything else in that table matched the config to the digit.
## The take
The reason this release is worth reading closely is not the benchmark line, which is good but
contested by a smaller competitor. It is that Muse Glimmer is a clean worked example of designing an
architecture against a deployment constraint and then publishing enough for someone outside the lab
to check the constraint was met.
The three decisions that matter — 16:1 GQA, three sliding layers per global one, and 4-bit
quantization validated at 1.0% degradation — are not independent optimizations. They are one budget,
allocated. Remove any of them and the model stops being the thing the announcement describes. That
is a more useful artifact than a leaderboard position, because the budget is the part that
generalizes: the next person trying to fit an agent on a laptop has a worked example with all its
numbers exposed.
What is missing is the same thing that is always missing on day one. There is no third-party
evaluation of any of these numbers, the methodology report is a link rather than a paper, and the
quantization degradation figure — 1.0% averaged across 15 benchmarks — is exactly the kind of
average that can hide a specific capability falling over. The card says the compression was
validated on agentic tasks; it does not show that table.
---
*Sources: the [Muse Glimmer announcement](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)
(Meta AI Research, 2026-08-10) and the [Muse-Glimmer-30B model card](https://huggingface.co/meta-models/Muse-Glimmer-30B),
plus `config.json`, `model.safetensors.index.json` and the file trees of the `-GGUF`, `-assistant`
and `-ExecuTorch-PTE` repositories, all as of 2026-08-10. Benchmark numbers are Meta's own, with no
third-party replication. The two figures are Meta's, flattened onto white. The KV-cache arithmetic,
the 24 GB budget reconciliation and its counterfactuals, the per-layer verification of the attention
and RoPE arrays, and the re-tally of the comparison table are mine and are computed from the
published config and file sizes. All four interactives are mine.*
---
# The Skaling law: Chinchilla assumes model size and data don't interact, and they do
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/skaling-law
> date: 2026-08-10
> tags: scaling-laws, llm, math, training, explainer
The Chinchilla scaling law is one of the most quoted equations in the field. It says the loss of a
language model decomposes into an irreducible floor plus two independent power-law terms, one for
model size and one for training data:
$$
L(N, D) = E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}
$$
[**Skaling: Chinchilla's Exponents Meet Kaplan's Coupling**](https://arxiv.org/abs/2608.07222)
(arXiv 2608.07222, Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz and Kartik Ahuja, FAIR at
Meta, 7 August 2026) points at the plus sign in the middle and observes that it is a very strong
claim nobody ever tested.
A sum of a function of $N$ and a function of $D$ has a cross-derivative of exactly zero. Not
approximately zero, not small — zero, as an algebraic identity. The additive form *asserts* that how
much a training token is worth does not depend on how big your model is. That assertion was never a
finding. It was a modelling convenience that came along for the ride.
Personal disclosure up front, because it changes how I read this paper: four days ago I published a
[back-of-envelope estimate](/articles/ltc-gated-delta) that leaned on exactly this additive form to
argue that a small model's reported gain might vanish under a compute-matched comparison. The last
section of this piece redoes that calculation with the paper's tools. It moves.
## The saddle
Fit the additive law to a dense grid of trained models and the residuals are not noise. They have
structure.
The paper's description of that first panel is precise: Chinchilla "is accurate in the interior of
the grid but develops large, oppositely-signed errors toward the corners, reaching several percent
where $N$ and $D$ are most imbalanced. This is the saddle-shaped residual expected when the $N$–$D$
interaction is omitted."
A saddle is the signature of a missing product term. If your model of a surface has no $xy$ term and
the true surface has one, the errors you get are positive on one diagonal and negative on the other
— which is exactly what the left panel shows. The paper backs this up with a direct measurement of
the cross-derivative $\partial^2 L/\partial N \partial D$ from local quadratic fits (their Figure 3),
finding it non-zero and structured.
## One exponent
The fix is small enough to state in a line. Kaplan's original 2020 form did couple $N$ and $D$ —
$L(N,D) = [(N_c/N)^{\alpha_N/\alpha_D} + D_c/D]^{\alpha_D}$ — but it tied the inner exponents
together through the ratio $\alpha_N/\alpha_D$, so the per-axis decay rates were no longer
independent. Chinchilla threw out the coupling to get the independence back. Skaling keeps both:
$$
L(N, D) = \left(\frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}\right)^{k} + E
$$
Chinchilla's interpretable base terms and independent inner exponents, raised to a single free outer
exponent $k$. At $k = 1$ it *is* the additive law — Chinchilla is a special case, not a rival. For
any $k \neq 1$ the cross-derivative is non-zero. And because $k > 0$, the loss is still strictly
decreasing in both arguments, so adding capacity or data can never be predicted to hurt.
The fitted value is what makes this more than a formality. On the Farseer grid, $k = 0.41 \pm 0.01$
— nowhere near 1, and tightly determined. Differentiating,
$$
\frac{\partial^2 L}{\partial \ln N\, \partial \ln D} = k(k-1)\,\alpha\beta\left(\frac{A}{N^{\alpha}}\right)\left(\frac{B}{D^{\beta}}\right)R^{\,k-2}, \qquad R = \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}
$$
With $k < 1$ the factor $k(k-1)$ is negative, so the cross-derivative is negative: since
$\partial L/\partial \ln D$ is already negative, making it *more* negative means **bigger models
extract more from the same token**. That is not a surprising claim — it is roughly what everyone
believes — but the additive law is structurally incapable of expressing it.
One honest wrinkle the paper raises itself. The coupled reducible term decays more slowly at large
scale, so it absorbs curvature the additive law can only represent through a larger $E$. Skaling
therefore pushes $E$ down — on Farseer almost to zero (0.03 ± 0.02, against Chinchilla's 0.45 ±
0.01). Since none of the runs reach the scale where loss saturates, the data fix the total loss but
not the split between a decaying term and a constant floor. So $E$ should not be read as a measured
irreducible entropy in either fit. Every other parameter is determined to within a few percent.
## Interpolation quality is not evidence
This is the methodological point I most want people to take away, and it is stated bluntly in the
paper: "High interpolation fit quality is not enough to validate a scaling law."
Chinchilla achieves $R^2 = 0.995$ on the full Farseer grid. By the standard people usually apply,
that is a solved problem. Its extrapolation error is three to four times Skaling's. The failure mode
"is therefore not a poor fit to the interior, but a systematic misprediction of how the loss surface
bends away from the observed region" — which is the only thing anyone ever uses a scaling law for.
Nobody fits a scaling law to predict a run they already did.
The compute argument in the second half of that figure is the one with budget consequences. The
authors pair the coupled form with an **L-shaped sparse grid**: instead of spreading held-out points
across the whole $(N, D)$ plane, restrict training runs to the low-compute edges — a row of small
models across many data budgets, and a column of small data budgets across many model sizes. Skaling
fitted on that L-shape, at roughly a tenth of the FLOPs, extrapolates **better than Chinchilla fitted
on the entire grid** in every held-out regime. On Farseer: 0.89% vs 1.48% on larger models, 1.35% vs
1.98% on more data, 1.51% vs 2.46% far outside both.
It is also worth reading the row that is not Skaling. The nine-parameter Farseer law is *worse* than
three-parameter Chinchilla on several columns and carries huge fold-to-fold variance (±1.93 on far
extrapolation). More parameters bought instability, not accuracy. The Skaling result is a
one-parameter change that improves things, which is a different and much stronger kind of claim.
## The part that changes decisions
Everything above is about fit quality. This is about where the money goes.
The compute-optimal token-to-parameter ratio $D^\star/N^\star$ is the quantity that answers "should
the next dollar buy a bigger model or more data?" The paper recovers it two ways without assuming
any parametric law — a global Gaussian-process surrogate and a local moving-least-squares surrogate,
finding the point where the log-slopes balance — and then compares against what each fitted law
predicts.
The two model-free estimates of the optimum give exponents of **−0.14 and −0.15**. Skaling recovers
**−0.11**. The refitted additive law gives **+0.03** — the opposite sign. Inside the observed data
range all four agree, which is exactly why a good interpolation $R^2$ told you nothing. Outside it
they diverge, and the paper reports that one order of magnitude beyond the data the allocations
differ by more than 10×, with the additive law heading toward hundreds of tokens per parameter while
the empirical fits and Skaling fall to the tens.
Two caveats before anyone reallocates a training budget on this. The empirical exponent is itself
an extrapolation of a fit to a surrogate of a finite grid, and the two surrogates agreeing with each
other is weaker evidence than two independent measurements. And "the additive law has the wrong
sign" is a claim about *these* datasets, at *these* scales, with these architectures. But the sign
disagreement is not subtle, it reproduces across two independently constructed grids, and it lands
on the one number the whole scaling-law enterprise exists to produce.
## What was actually measured
Worth being concrete about the evidence base, because scaling-law papers vary enormously here.
**Farseer** is an existing public grid; the fitting set is 302 configurations totalling
~5.0×10²² FLOPs, with held-out sets for larger models (36 points, 1.5B–6.4B), more data (66 points),
and far extrapolation (7 runs at 2.3B–25B parameters on 126B–453B tokens, beyond both axes).
**SK-Grid** is the authors' own: 134 configurations, 15 model sizes from 134M to 4.9B, 16 data
budgets from 316M to 316B tokens, with far-extrapolation runs at ~10²² FLOPs on 5.8B–10.8B
parameter models. Two more datasets appear in the appendix with the same protocol and the same
ranking.
All laws are fitted identically — Huber loss in log space, L-BFGS-B with 2000 basin-hopping restarts,
analytic gradients — so the comparison is not confounded by one law getting a better optimizer. The
paper also notes that the improvement survives holding the protocol fixed, "confirming the gain
comes from the functional form rather than the protocol."
What is *not* here: no runs at frontier scale, so the far-extrapolation column tops out around 25B
parameters; no test of whether $k$ is stable across architecture families, tokenizers or data
mixtures, which is the obvious next question given that $k$ is now carrying the entire interaction;
and no mixture-of-experts models, where "model size" is ambiguous enough that it is unclear which $N$
even belongs in the formula.
## Redoing my own arithmetic
Four days ago, writing about [liquid time constants and gated delta
rules](/articles/ltc-gated-delta), I used the additive Chinchilla form to estimate a confound. The
setup: a 29M-parameter model, LTCAttention, beat its baseline by 0.062 nats on a token-matched
comparison while running 12.4% slower. I asked what the baseline would have gained if it had spent
that 12.4% on extra tokens instead, got roughly 0.061 nats, and concluded that a compute-matched
comparison might erase the entire result.
I used Hoffmann et al.'s 2022 coefficients. This paper supplies two better options: the same additive
form refitted on Farseer, and the coupled form on the same runs.
The three answers span 2.4×. The estimate I published was the largest of them.
Most of the movement is not the coupling — it is that Hoffmann's coefficients were fitted on a
different corpus and tokenizer, and refitting the *same* functional form on Farseer roughly halves
the answer, from 0.061 to 0.031 nats. The coupled form then trims it further, to 0.026. There is
also a pointed detail: LTCAttention's configuration sits at $D/N = 10.0$, which falls in the band
the paper reports as Chinchilla's **worst** regime — 3.47% MAPE in the optimal-ratio third, where
its pooled number hides the failure.
So the honest revision: a compute-matched baseline would probably have recovered somewhere around
**half** of LTCAttention's measured gain, not all of it. The conclusion I actually drew still stands
— the missing experiment is a wall-clock-matched run, and until someone does it the result is better
per token and undetermined per second — but I stated the confound about twice as strongly as the
evidence supports. That correction is now in the record here rather than only in my own notes.
The broader lesson is the one I would take from this paper even if I had no stake in it. Scaling-law
arithmetic is routinely used the way I used it: pull the canonical coefficients, differentiate, get a
number, cite it as though it were a measurement. It is not. It is a prediction from a functional form
fitted to somebody else's grid, and both the form and the grid are doing real work. When the answer
matters, quote the range.
## The take
The contribution is one exponent, and the reason it is a good paper rather than a small one is that
the exponent is load-bearing. It removes a structural bias that was invisible in interpolation error
and severe at the boundaries, it makes accurate extrapolation possible from a tenth of the compute,
and it flips the sign of the trend in the single number the field uses to allocate training budgets.
What it does not do is settle anything at frontier scale, where nobody has published the grid that
would test it. And there is a mild irony worth naming: the paper's own argument implies that its
$k = 0.41$ is a property of these datasets, and the honest way to use the Skaling law is to refit it
on your own runs rather than to quote 0.41 the way people have been quoting 20 tokens per parameter
for four years.
---
*Sources: [Skaling: Chinchilla's Exponents Meet Kaplan's Coupling](https://arxiv.org/abs/2608.07222)
(arXiv 2608.07222v1, Videau, Youbi-Idrissi, Lopez-Paz, Ahuja, FAIR at Meta, 7 August 2026, CC BY
4.0), read in full via the arXiv HTML rendering. Equation 3, the fitted coefficients in Table 2, the
MAPE figures in Tables 1 and 3, and the compute-optimal exponents in Figure 6 are quoted as
published. Both figures are the paper's own, flattened onto white. The cross-derivative expression,
the recomputation of my earlier LTCAttention estimate, and the sensitivity arithmetic behind the last
interactive are mine, computed from the paper's published coefficients at LTCAttention's reported
N and D. All four interactives are mine.*
---
# BTL-4: reading a model card against its own weights
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/btl-4
> date: 2026-08-06
> tags: llm, open-weights, benchmarks, evaluation, lora, explainer
[BTL-4](https://huggingface.co/badtheorylabs/BTL-4) went up on Hugging Face on 2026-08-05: a 35B
agentic reasoning model from Bad Theory Labs, Apache-2.0, 21 safetensors shards, and a benchmark
table with **78.4% on SWE-bench Verified** in it. There is no technical report. There is no arXiv
paper. There is no third-party evaluation. At the time of writing the repository has 38 likes and,
according to the Hugging Face API, **zero downloads — all-time**. Nobody has run this model.
That combination is common enough now that "how do I read this?" is a real question rather than a
rhetorical one. The answer I want to argue for is that a model card is not the only evidence a
release ships. The artifact itself — `config.json`, the tensor index, the bytes in the shards — is
evidence too, it is machine-checkable, and it is often more informative than the prose. This piece
is that check, run end to end on BTL-4. Some of the card holds up exactly. Some of it does not.
**What this is and isn't.** I did not download 70 GB and run a benchmark; I have not reproduced or
refuted any score here. Everything below comes from public metadata and HTTP range requests
totalling a few megabytes. Where the card is contradicted, it is contradicted by *its own numbers*
or by the artifact it ships, not by a competing measurement of mine. And Bad Theory Labs is not a
drive-by account — it has seven models going back to June 2026, including
[BTL-3](https://huggingface.co/badtheorylabs/BTL-3) and a BTL-4-Compact posted the day after this
one. Read this as an audit of one release, not a verdict on a lab.
## What checks out
Start with the parts that survive, because they are the majority and because a check that only ever
finds problems isn't a check.
**The base-model claim is exactly right.** The card declares `base_model: Ornith-1.0-35B`, which is
an unqualified string rather than a resolvable repository id, so nothing on Hugging Face verifies it
for you. But [`ornith-ai/Ornith-1.0-35B`](https://huggingface.co/ornith-ai/Ornith-1.0-35B) exists —
a real MIT-licensed model with 2.67M downloads — and its `config.json` matches BTL-4's on every
architectural field:
| | Ornith-1.0-35B | BTL-4 |
|---|---|---|
| architecture | `Qwen3_5MoeForConditionalGeneration` | same |
| layers · hidden | 40 · 2048 | same |
| experts · active | 256 · top-8 | same |
| MoE / shared intermediate | 512 · 512 | same |
| vocabulary | 248,320 | same |
| context | 262,144 | same |
| full-attention interval | every 4th layer | same |
| vision tower | 27 layers · 1152 wide | same |
The tensor maps are identical too: **31,666 tensors, same names, in both**. Whatever else is true,
BTL-4 is a derivative of Ornith-1.0-35B and not of something else wearing its name.
**The weights are real and complete.** 21 shards, 70.21 GB, 35.11B parameters in bf16, a coherent
`model.safetensors.index.json`, vision tower included. This is not an empty repository with a good
README.
**The lineage is worth stating**, because it puts BTL-4 next to work already covered here.
Ornith-1.0-35B's architecture is `qwen3_5_moe` — the same family as
[Intern-S2-Mobius](/articles/intern-s2-mobius), which is 40 layers at hidden 2048 with the same
512-wide experts and the same every-fourth-layer full attention. Both descend from Qwen's
`Qwen3.5-35B-A3B`, whose parameter count (35,951,822,704) is also the exact figure I measured for
[Macaron-V1-Tall's](/articles/macaron-v1) base checkpoint. Three unrelated labs, one 35B Qwen
substrate. That is worth noticing on its own.
**And one section of the card is genuinely useful**, which I'll come back to at the end — it isn't
the benchmark table.
## The benchmark table doesn't close
Here is the LiveCodeBench v6 section of the card, quoted in full. Aggregate **66.1%**, and:
| | pass@1 |
|---|---|
| easy | 99.1% |
| medium | 86.7% |
| hard | 60.5% |
followed by: *"The set is 45% hard problems, which is what pulls the aggregate down."*
Those four numbers cannot all be true, and you do not need the benchmark to see it. An aggregate
pass rate is a weighted mean of the per-difficulty rates. Fix the hard share and the aggregate is
pinned inside an interval — lowest when every remaining problem is medium, highest when every
remaining problem is easy.
At 45% hard, the aggregate has to land between **74.9%** and **81.7%**. The card reports 66.1%,
roughly nine points below the floor of what its own difficulty breakdown allows. Going the other
way: to produce a 66.1% aggregate from a 60.5% hard bucket and an 86.7% medium bucket, you would
need **78.6% hard problems and zero easy ones** — which would leave the 99.1% easy row reporting a
score for an empty set.
I want to be careful about what this does and does not establish. It does not tell you the model is
bad, and it does not tell you which number is wrong. Any one of four edits reconciles it: the
aggregate, the hard share, one of the bucket rates, or an unstated detail about how the aggregate
was computed (a different problem set, a different pass@k, a subset that the difficulty table
doesn't describe). What it does establish is that **the table was never checked against itself**,
which is a fact about the release process rather than about the model. Numbers that were run,
recorded, and then arithmetically verified do not do this.
The same section has a smaller tension worth flagging. The card says the runs used "full splits, no
subsetting," and in the next breath specifies "442 problems, 2024-08 → 2025-05." A date window is
LiveCodeBench's intended usage — the whole point of the benchmark is contamination-controlled time
slices — so the window is legitimate. But a date-windowed 442-problem slice is, definitionally, a
subset, and "no subsetting" is the wrong way to describe it. I could not independently confirm
LiveCodeBench v6's true composition for that window, so I can't say which of the card's figures the
real distribution would support.
## What the config says that the card doesn't
`config.json` is written by the training code, not by the person writing the README, which makes it
the more candid of the two documents. BTL-4's contains this:
```json
{
"model_name": "/vol/merged/btl4-pilot",
"transformers_version": "5.13.1",
"unsloth_version": "2026.7.6"
}
```
Three things leak out of five lines. The checkpoint was produced with **Unsloth**, a LoRA and QLoRA
fine-tuning library. It was loaded from a directory called **`merged`**, which is what you call the
output of folding an adapter back into its base. And the run was named **`btl4-pilot`**.
None of that is damning — LoRA is a completely normal way to fine-tune a 35B MoE, and merging is
the normal way to ship one. But the card's training section says only: *"Fine-tuned from
Ornith-1.0-35B on an execution-gated reasoning corpus."* A reader deciding whether a +4.3-point BFCL
gain is likely to generalize would want to know it came from a merged adapter rather than a full
fine-tune, and the card does not say.
The interesting question is whether the weights agree with the config. They do.
## Reading 70 GB without downloading it
A safetensors file opens with 8 bytes giving a header length, followed by that many bytes of JSON
describing every tensor: dtype, shape, and byte offsets into the rest of the file. That means two
small HTTP range requests per shard buy you the complete layout of a 70 GB checkpoint. Once you have
offsets, you can range-request *one specific tensor* out of the middle of a shard and compare it
against the same tensor in another repository, having transferred a few hundred kilobytes.
I ran that against BTL-4 and Ornith-1.0-35B across ten groups of tensors.
The first pass was wrong, and the way it was wrong is the most useful thing in this article. Every
normalization weight came back CHANGED — all forty layers' input and post-attention norms, the final
norm, even norms inside the vision tower. That looked like a substantial finding. It was an
artifact: BTL-4 stores norms as **F32** where Ornith stores them as **BF16**, so I was comparing
8,192 bytes of one format against 4,096 bytes of another and reading the inevitable mismatch as
training.
The check that settles it is arithmetic on file sizes. The two checkpoints differ in total size by
**603,136 bytes**. BTL-4's metadata reports exactly **301,568 parameters stored in F32**; Ornith
reports none. An F32 parameter costs two bytes more than a BF16 one, and 301,568 × 2 = 603,136. The
entire size difference between the two models is the norm upcast and nothing else — which is what
turns "I should exclude those rows" from a hunch into a fact.
## What the change map means
With dtype-mismatched tensors excluded, the pattern is unusually clean:
- **Changed**, in every layer sampled: expert `gate_proj` and `down_proj`, the shared expert's
`up_proj`, the linear-attention `in_proj_qkv`, and full-attention `q_proj`.
- **Unchanged**, in every window sampled: the MoE routers, token embeddings, the output head, the
entire 27-layer vision tower, and the linear-attention `A_log` and `dt_bias`.
That is a LoRA target set, drawn from life. Adapters go on the projection matrices; routers,
embeddings, output heads and frozen encoders are left alone. Combined with `unsloth_version` and
`/vol/merged/`, the artifact is telling a consistent story that the prose omits.
Two of those frozen tensors deserve their own note.
**`A_log` and `dt_bias` are untouched at every layer**, and unlike the big matrices these are small
enough to compare in full — 64 bytes each, byte-for-byte identical. In this architecture family
`A_log` is the per-head base rate of the linear-attention decay gate, the parameter whose exponential
sets how fast a channel forgets. [KDA has a half-life](/articles/kda-half-life) works through what
that number means: it converts directly into a memory horizon measured in tokens. So BTL-4's
forgetting timescales are Ornith's, unmodified. Whatever the fine-tune taught the model about tool
calling, it did not touch the mechanism that decides how long the model can hold something.
**The vision tower is entirely unchanged, and entirely still there.** BTL-4 ships
`processor_config.json`, an `image_token_id`, a `video_token_id`, and 27 untouched vision layers.
The card sets `pipeline_tag: text-generation`, describes a text-only training corpus, and reports no
multimodal evaluation whatsoever. Nothing wrong with that — you inherit a capability you didn't
train and don't claim. But a reader should know that roughly a tenth of what they'd be downloading
is an unexercised, unevaluated image encoder, and that the model's multimodal behaviour is entirely
Ornith's.
## The number with the least behind it
Of the three headline benchmarks, note which one is documented and which is not.
BFCL v4 gets a full protocol sentence: official `ast_checker`, all 1240 cases, run in-house, and —
best practice, this — an explicitly paired comparison against the base with "identical harness,
identical decoding, only the weights differ." That is exactly how a fine-tuning claim should be
stated, and the +4.3 points is the only number on the card that isolates what the training actually
bought. LiveCodeBench gets a protocol sentence too, though the numbers under it don't close.
**SWE-bench Verified 78.4% gets three words: "official harness."** No base comparison, so there is
no way to see what the fine-tune contributed. No statement of who ran it, while the two benchmarks
above it are explicitly labelled in-house. No scaffold named, which for SWE-bench is most of the
result — the agent loop around the model routinely moves that score more than the model does, a
point [the harness effect](/articles/harness-effect) makes at length. No trajectory logs, no
leaderboard submission.
It is also, by a distance, the biggest claim on the page. 78.4% would place a 35B model with roughly
3B active parameters within striking distance of the frontier systems this site has covered — the
Macaron-V1 table has Claude Opus 4.8 at 88.6% on the same benchmark. Extraordinary is the wrong
frame; *unverifiable* is the right one. The claim isn't refuted here. It's simply the one number
with the least behind it, presented with the least detail, on a model that has been downloaded zero
times.
Zero downloads, all-time, is not a snark — it's a structural fact about what any reader can know.
Every number on this card is unreplicated *by construction*, because nobody has yet obtained the
weights to try. Likes are not replication. Until someone runs it, the honest status of all three
benchmarks is "reported, unverified," and the honest status of the LiveCodeBench row specifically is
"reported, internally inconsistent."
## The part of the card that's actually good
None of the above touches the most useful section, and it deserves to be lifted out because it is
the kind of thing most cards leave you to discover in production:
> **Reasoning accumulates across agent turns.** The chat template strips prior reasoning from older
> turns, but this only works if your harness separates it into `reasoning_content`. With vLLM, that
> means `--reasoning-parser qwen3`. Without it, thinking lands in `content`, accumulates every turn,
> and long agent runs degrade.
That is correct, specific, non-obvious, and expensive to learn by yourself. A reasoning model whose
chat template prunes old thinking blocks depends on the serving layer routing them to the right
field; get it wrong and you don't see an error, you see an agent that quietly gets worse over a long
session while your context bill climbs. Whoever wrote that paragraph has actually run this thing in
a loop. The same section is candid that the model is verbose, not a chat model, and token-hungry,
and the generation-settings note — LiveCodeBench moving 60.9% → 66.1% purely by raising the output
budget from 16K to 32K — is a real, useful observation about evaluating reasoning models even if the
endpoint number is the one that doesn't reconcile.
There's a smaller inconsistency in the same neighbourhood: the card claims 262K native context and
the `vllm serve` command it gives sets `--max-model-len 131072`. Both are defensible individually —
you often can't fit the full window on the hardware you have — but nothing explains the gap.
## The take
The reusable part here isn't the verdict on BTL-4, it's the sequence. Four checks, none of which
require downloading a model or running a benchmark, in increasing order of effort:
1. **Does the table close?** Weighted means have to be consistent with their parts. This caught the
LiveCodeBench contradiction in one line of arithmetic, before anything was fetched.
2. **Does the declared base exist, and does the config match it?** Field-by-field comparison
confirmed BTL-4's lineage exactly, which is the strongest positive result in this whole piece.
3. **What does the config leak?** Training code writes provenance the README never mentions —
library versions, working-directory paths, run names.
4. **Do the weights agree with the story?** Safetensors headers plus range requests turn "trust the
training section" into a measurement, for a few hundred kilobytes of traffic.
Applied to BTL-4 the result is mixed rather than damning: a real fine-tune, of the model it says it
is, with complete weights, shipped with one benchmark row that contradicts itself, one number that
carries the most weight and the least evidence, and a training method the artifact discloses more
honestly than the prose does. What I'd want before believing the headline is a base-model row for
SWE-bench and the scaffold used to get it — the same paired-comparison discipline the card already
applies to BFCL, extended to the number people will actually quote.
---
*Sources: the [BTL-4 model card](https://huggingface.co/badtheorylabs/BTL-4) (README, `config.json`,
`model.safetensors.index.json`, safetensors headers), the
[Ornith-1.0-35B](https://huggingface.co/ornith-ai/Ornith-1.0-35B) repository, and the Hugging Face
models API for download, like, and parameter counts, all as of 2026-08-06. The tensor comparison was
performed with HTTP range requests against both repositories' shards; the method, its dtype
correction, and the limits of what a windowed comparison can prove are described in the second tab
of the diff figure above. The LiveCodeBench arithmetic uses only figures printed on the card. No
model was downloaded and no benchmark was re-run. Both interactives are mine; the repository ships
no figures, so there are none to embed.*
---
# Intern-S2-Mobius: a 35B model that separates memory from reasoning
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/intern-s2-mobius
> date: 2026-08-06
> tags: llm, mixture-of-experts, architecture, linear-attention, explainer
[Intern-S2](/articles/intern-s2) was Shanghai AI Lab's case for specialization: a 397B model that
learns straight off the raw page of a scientific paper. **Intern-S2-Mobius** is a different
experiment entirely, and much smaller — 35B parameters, continual-pretrained from Qwen3.5-35B, and
the point isn't science. It's architecture. The model card's claim is that you can pull a
transformer apart into two pieces — a store of learned knowledge and the computation that queries
it — and get a real efficiency win from doing so: reasoning traces up to 5.0× shorter,
average throughput up to 4.6× higher, at matched or better scores than the plain Transformer
it's compared against.
That comparison is the thing to hold onto while reading this. Mobius isn't benchmarked against
GPT-5.5, Gemini, or even other 35B-class open models — every number on its card is Mobius versus
its own base model, Qwen3.5-35B, continual-pretrained the same way. It's an ablation, not a
leaderboard entry. That makes it a cleaner test of what the architecture buys, and a much weaker
basis for "should I use this instead of X."
## What it is
- **35B parameters**, bf16, five safetensors shards on Hugging Face totalling about 73 GB — not
gated, apache-2.0, actually downloadable.
- **`image-text-to-text`**: a vision tower (27 layers, 1152-wide, patch 16 — the same depth and
width as the SigLIP-So400M family of encoders used across a lot of current VLMs) feeds a text
backbone with a 256K-token context window.
- **Continual-pretrained from Qwen3.5-35B**, then SFT and RL, on a new architecture the card calls
**Mobius-v0**, "realized by Xtuner and LMDeploy."
- Deploys single-GPU on LMDeploy, vLLM, or Transformers, with an `mtp` speculative-decoding mode
(`qwen3_5_mtp`) recommended in the quickstart.
The README links InternLM's [ArchSpace](https://github.com/InternLM/archspace) — a public
architecture-experimentation program that turns community proposals into trained, evaluated,
published results — as a related project. It doesn't say Mobius came out of that pipeline, so I'm
not claiming it did; it's worth knowing the program exists, because it's the same lab publicly
running exactly this kind of experiment at scale.
## What "Mobius" names
Not a routing scheme, not a training recipe — an architecture. The card's own framing:
> Instead of binding knowledge storage and reasoning computation layer by layer as in conventional
> Transformer models, Mobius organizes knowledge into a globally shared **Memory** and lets
> multiple **Reasoners** iteratively query and refine hidden states against this shared repository.
Two capabilities follow from that split, per the card: **Backward Residual Connection** (a deep
layer can reach knowledge a shallow layer used, not just what forward propagation handed it), and
**Dynamic Latent Reasoning** (deliberation gets internalized into hidden states instead of written
out as visible chain-of-thought tokens). Both are described in prose. The released code lets you
check what's literally true of the shipped model versus what's evocative marketing language for
the same idea — and it turns out you can, because InternLM shipped the modeling file along with
the weights.
## Forty layers, four memory banks
Here's what `modeling_interns2_mobius.py` actually does. `config.json` sets `num_blocks: 4`. The
model builds exactly four `InternS2MobiusMetaMoeBlock` objects — each one a router plus 2560
routed experts — and holds them in one list, `meta_mlp`. Every one of the 40 decoder layers keeps
its own attention and layernorms, but for its routed-expert lookup it computes
`block_idx = layer_idx % num_blocks` and reads from `meta_mlp[block_idx]`. Layers 0, 4, 8 … 36 all
route into the *same physical weight tensors* — not four separately-trained-but-similar banks, one
set of parameters, referenced by ten different layers.
A standard MoE transformer ties knowledge to depth: layer *k* owns bank *k*, and whatever it
learned lives only there. Mobius reuses the same four banks across the whole stack instead, so a
layer near the input and a layer near the output can draw on the identical knowledge subspace.
That's the concrete mechanism behind "Backward Residual Connection" — not a literal skip connection
running backward through the network, but a shared address space that any depth can query. It's
also a real parameter-efficiency trade: with four banks instead of forty, the routed-expert weight
mass is
$$
\theta_{\text{experts}} \approx N_{\text{blocks}} \times N_{\text{experts}} \times
\big(2\,d_{\text{ffn}}\,d_{\text{model}} + d_{\text{model}}\,d_{\text{ffn}}\big)
= 4 \times 2560 \times 3{,}145{,}728 \approx 32.2\text{B params},
$$
roughly 90% of the model's total, and roughly consistent with the ~36.5B implied by the 73 GB of
bf16 weights on disk. Per token, only one bank is queried per layer and only 8 of its 2560 experts
fire — call it ~28M active FFN parameters per layer (8 routed experts plus the always-on per-layer
shared expert), times 40 layers. That's a back-of-envelope estimate from `config.json`, not a
number the card states; unlike [Intern-S2-Preview-397B's](/articles/intern-s2) plain top-8-of-512
math, Mobius's shared-bank routing makes a clean "active parameters" headline harder to state, and
InternLM doesn't attempt one.
One more thing falls out of matching two arrays in the same config: `layer_types` cycles
linear-attention, linear-attention, linear-attention, full-attention every four layers
(`full_attention_interval: 4`), the same period as the memory-bank assignment. Bank 3 is *always*
the one full-attention layer in its group of four; banks 0–2 are always linear attention — a Gated
DeltaNet variant, the same family covered in [KDA's half-life](/articles/kda-half-life) for Kimi
K3's linear attention. That alignment isn't asserted anywhere in the README. It's just what the two
config arrays do when you line them up.
Whether "iteratively query and refine" is literally true of inference is a fair question to ask of
any of this. The released `InternS2MobiusTextModel.forward()` is a single straight-through pass
over 40 layers — no runtime loop, no repeated pass over the same weights within one layer. What
does repeat, ten times, is the pattern: attend, then query one of four shared memory banks, at
increasing depth. If that reads like an unrolled recurrence rather than free-form iteration, that's
a fair description — it's a coarser, more surgical form of weight sharing than a fully [looped
transformer](/articles/looped-models-done-right), which ties whole layers (attention included)
across depth, or [LOTUS](/articles/lotus-latent-reasoning), which loops the same weights over a
fixed latent region multiple passes per token. Mobius ties only the expert banks, once each, spread
across depth rather than iterated in place.
## The benchmarks
Everything on the card is Mobius vs. Qwen3.5-35B — its own continual-pretraining source, not a
frontier model. On general reasoning:
The average hides a mixed picture. Mobius leads on MMLU Pro (89.05 vs 85.31), IMO Bench (81.25 vs
77.50), HMMT 2026 (85.51 vs 78.50), AIME 2026, GPQA Diamond, AMO, and SimpleQA. It loses on two:
UGD hard (73.02 vs Qwen's 78.02) and HLE (19.11 vs 22.40) — worth stating plainly, since the card's
own bullet points don't mention either.
Scientific tasks show the wider gap, and it's the same shape as the S2-Preview-397B story at a
different scale:
Biology-Instructions carries that average almost alone: 51.40 vs 3.77, a 13.6× gap. Mol-Instructions
(45.73 vs 21.70) and MolecularIQ (59.29 vs 29.13) are more modest but still roughly double. I'd read
this less as "Mobius learned multi-omics" and more as evidence that whatever mix of continual
pretraining and RL Shanghai AI Lab runs across the Intern-S2 family leans hard on scientific data —
consistent with, though far less extreme than, [Intern-S2-Preview-397B's](/articles/intern-s2) own
scientific dominance.
## Shorter traces, faster serving
The headline claim is "nearly 4x speedup reported in the technical report" — a report the model
card references but never links or cites; there's no arXiv listing for Mobius as of this writing. What the card does show directly is Fig. 1: request throughput at batch sizes 16 through
256, averaged across six reasoning benchmarks, with Mobius **2.9× faster at batch 16 and
4.6× faster at batch 256**.
Zoom into the five subplots behind that average and the story isn't uniform. MMLU Pro and GPQA
Diamond show a wide, cleanly growing gap in Mobius's favor — that's most of what drags the average
up. The three math-competition benchmarks look nothing like it. On **AIME 2026 and HMMT 2026 the
lines cross, and the Transformer baseline is the faster of the two at three of the five batch sizes
plotted** — including 2⁷, where AIME's gap is widest in the baseline's favour. IMO Bench does stay
in Mobius's favour at every point, but by a margin closer to 1.1× than to anything in the
headline. The 2.9–4.6× number describes the boxed average panel. It doesn't describe
every benchmark that average is built from, and the chart says so plainly if you look past the box.
Most of the throughput gain traces back to shorter output, not cheaper per-token compute — Fig. 2
gives average trace length directly, and the "Nx shorter" figures on it are exact, not
chart-estimated:
The same pattern repeats: GPQA Diamond and MMLU Pro compress the most and are also where the
throughput gap is widest and cleanest; the math-competition benchmarks compress the least and are
where the throughput lines cross. Shorter traces plus fewer live tokens in the KV cache is a
coherent story for why throughput goes up — it just doesn't go up evenly.
Card gaps worth naming plainly. There's no linked technical report or arXiv paper — "reported in
the technical report" points at a document I could not find. No active-parameter figure is given
(fair, given the shared-bank routing makes one less simple to state than usual). Every benchmark
comparison is against Qwen3.5-35B specifically, not against any external model, so there's no
frontier read and no read against comparably-sized open peers either. And one of the card's own
figures — the reasoning-trace case study — labels its Mobius column **"Intern-Spin-35B"** instead
of Intern-S2-Mobius-35B, an internal-codename leftover that suggests the card was assembled in a
hurry.
## Licence, and whether you can run it
Apache-2.0, same family as [Intern-S2-Preview-397B](/articles/intern-s2). The weights are real:
five bf16 safetensors shards on Hugging Face (`internlm/Intern-S2-Mobius`), about 73 GB total, not
gated, mirrored on ModelScope. That's a workstation-class footprint next to the 397B model's
frontier-hardware requirement — LMDeploy's quickstart serves it on a single GPU (`--tp 1`), MTP
speculative decoding recommended for the throughput numbers above.
## What I make of it
- **The mechanism is real and it's in the code, not just the prose.** `block_idx = layer_idx %
num_blocks` is a two-line change with a genuinely different parameter-sharing shape than a
standard MoE — four memory banks instead of forty, each queried by ten layers spread across
depth. That's checkable, and it checks out.
- **"Dynamic Latent Reasoning" oversells what the inference code shows.** There's no runtime loop —
it's a single forward pass with a repeating depth-wise pattern, which is a more modest and more
precise thing than "iterative refinement" suggests.
- **The efficiency win is real but uneven, and it tracks trace compression.** Where output collapses
— GPQA Diamond 5.0× shorter, MMLU Pro 4.6× — throughput climbs cleanly. Where it barely
moves — HMMT 1.2×, IMO Bench 1.4×, AIME 1.5× — the throughput advantage narrows to
nothing or inverts. That is a coherent mechanism rather than a mystery: the speedup is mostly
fewer tokens, not cheaper tokens. It also means the gain should be expected to shrink on any task
where the model still needs to think at length.
- **This is an ablation, not a leaderboard entry.** Every comparison on the card is Mobius against
its own untouched base model. That's the right comparison for isolating what the architecture
buys. It's the wrong comparison for deciding whether to run Mobius instead of anything else.
---
*Sources: the [Intern-S2-Mobius model card](https://huggingface.co/internlm/Intern-S2-Mobius)
(README, `config.json`, `configuration_interns2_mobius.py`, `modeling_interns2_mobius.py`) and the
[Intern-S2-Preview-397B model card](https://huggingface.co/internlm/Intern-S2-Preview-397B), both
InternLM / Shanghai AI Lab. Benchmark numbers and figures are quoted as reported on the Mobius
model card; no independent technical report or arXiv paper could be located.*
---
# Maple-Preview: a 20B reasoning model where 97% of the weights are −1, 0, or +1
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/maple-preview
> date: 2026-08-06
> tags: llm, quantization, mixture-of-experts, on-device, open-weights, explainer
Most quantization is something you do *to* a model after it is trained. You take bf16 weights, find a
rounding scheme that hurts least, and accept the damage. **[Maple-Preview](https://huggingface.co/deepgrove/maple-preview)**,
released by DeepGrove on 2026-08-04 under MIT, is the other thing: a 20B-A1.49B mixture-of-experts
reasoning model where the weights were *trained* to be ternary, so nearly every parameter in the
network is one of exactly three values — −1, 0, or +1, scaled.
The headline numbers are a 5.31 GB checkpoint and 218 tokens/sec on a Mac mini M4. Both are the kind
of claim worth checking rather than repeating, and this time both check out — with one significant
asterisk about what the released repository actually contains.
## The claim, verified from the bytes
You do not need to download 40 GB to test "is it ternary." A safetensors file begins with a header
giving every tensor's dtype, shape and byte offsets, so two small range requests buy you the layout,
and one more pulls a single row out of the middle of a shard. Count the distinct values in that row.
A normal bf16 weight row of 2048 elements has on the order of two thousand distinct values. A ternary
one has three.
Every row measured came back with exactly three values, perfectly symmetric — `−s`, `0`, `+s` — with
`s` changing from row to row. That is ternary with a **per-output-channel scale**, the BitNet
b1.58 shape. About 38–43% of the weights in each row are exactly zero, which is what absmean
ternarization does to a roughly Gaussian weight distribution: everything inside the rounding
threshold collapses to nothing.
What stays in full precision is as interesting as what doesn't. **96.9% of parameters are ternary**;
the exceptions are the two embedding tables, the norms, and — pointedly — the MoE routers. A router
chooses 8 experts out of 256 based on the *margin* between logits. Crushing that margin to three
levels would scramble which expert fires long before it degraded any individual expert's arithmetic,
so the router is the one 12M-parameter tensor per layer that stays sharp.
## The 5.31 GB claim reconciles, and the leftover is the vocabulary
A ternary weight carries $\log_2 3 \approx 1.585$ bits of information, so that is the floor for any
lossless packing. Working from the actual parameter census — 19.58B ternary, 0.64B full precision —
the arithmetic lands where it should.
5.31 GB implies **1.65 bits per ternary weight**: just above the entropy floor, comfortably below
naive 2-bit, and about where you land packing five trits into a byte ($3^5 = 243$ fits in 256) plus
the per-row scales. The claim is not merely plausible, it is consistent with the measured parameter
split to within a few percent.
The second-order effect is the one I did not expect. Once you have crushed 97% of the model to under
two bits, the **un-quantized embedding tables are about a quarter of the entire file** — 1.27 GB of
5.31 GB, for a 151,936-token vocabulary at 2048 wide, twice over (input and output are untied). At
this compression ratio the interesting problem stops being the weights and starts being the
vocabulary. Anyone chasing the next factor of two on-device has to go after the embeddings.
**What you download is not the 5.31 GB artifact.** The Hugging Face repository ships nine shards
totalling **40.43 GB** — the ternary values stored one-per-bf16, unpacked. The 5.31 GB figure
describes a packed checkpoint that is not in the repository. The README is upfront that the Apple
Silicon result "uses a separate on-device runtime," and I would extend that caveat: the packed
format and the kernels that make 218 tok/s possible are both part of that unreleased runtime. What
is public is the weights and a reference implementation.
## The shipped code does no quantization
This is worth stating plainly because `config.json` looks like it says otherwise. It contains
`"quantize": true` — and nothing in the released code reads it. `MapleConfig.__init__` in
`configuration_maple.py` does not declare a `quantize` parameter, so the flag lands in `**kwargs` and
is stored and ignored. The only occurrence of the word in 1,052 lines of Python is a comment in
`fa3.py`.
The forward pass confirms it. `MapleMLP.forward` is a plain dense matmul on the dequantized bf16
tensors:
```python
def forward(self, x):
gate_weight, up_weight, down_weight = self.gate_proj.weight, self.up_proj.weight, self.down_proj.weight
return torch.nn.functional.linear(
self.act_fn(torch.clamp(torch.nn.functional.linear(x, gate_weight), max=7.0))
* torch.clamp(torch.nn.functional.linear(x, up_weight), min=-7.0, max=7.0),
down_weight,
)
```
There is no packing, no unpacking, no ternary kernel. Run this and you get a correct model that
occupies 40 GB and runs at ordinary dense-MoE speed, with none of the benefit that motivated the
architecture.
The clamps are the tell that quantization-aware training happened somewhere else. `clamp(gate,
max=7.0)` and `clamp(up, min=-7.0, max=7.0)` bound the activations going into the down-projection.
Activation clamping is a standard QAT ingredient — you cannot quantize weights aggressively if the
activations they multiply are free to blow up — and its presence in the inference path is a residue
of the training recipe, kept because removing it would change the model's behaviour.
## The architecture around the quantization
The config describes a design clearly built for a memory-bound device rather than a datacenter.
| | |
|---|---|
| layers | 24 |
| hidden | 2048 · head_dim 128 · 16 heads · 4 KV heads |
| experts | 256, top-8, **no shared expert**, `moe_intermediate_size` 512 |
| attention | 3:1 sliding-window (512) to global |
| position | `partial_rotary_factor` 0.5, `nope_on_global_attention: true` |
| context | 131,072 |
| vocabulary | 151,936 (Qwen tokenizer) |
The `layer_types` array spells the attention pattern out exactly: `s s s G` repeated six times, with
global attention at layers 3, 7, 11, 15, 19 and 23. Only a quarter of the layers hold a full-length
KV cache; the rest are capped at a 512-token window. For a 131K context on a Mac mini that is not a
refinement, it is the difference between fitting and not fitting.
Two details are worth pulling out. **`nope_on_global_attention: true`** means the global layers get
no positional encoding at all — the sliding layers carry position through RoPE (at half the head
dimension, per `partial_rotary_factor: 0.5`) and the global layers are left to infer order from what
the local ones already encoded. The same trick appears in [Kimi K3](/articles/kimi-k3)'s attention
stack, and the argument for it is that removing RoPE from the layers that see the whole sequence is
what lets length extrapolation work.
And **there is no shared expert** — `num_shared_experts: 0`. Most recent MoE designs keep one or two
always-on experts to absorb generic computation. Maple routes everything, which is consistent with
the rest of the design: a shared expert is a dense tensor every token pays for, and this model is
built to minimize exactly that.
## What the benchmarks say, and what the chart leaves out
Maple-Preview averages **78.7** across LiveCodeBench v6, AIME 2026, HMMT 2026 and GPQA-Diamond, at
1.49B active parameters. That beats GPT-OSS 20B (76.3), Qwen3 30B-A3B (76.6), Qwen3.5 9B (76.3),
GLM 4.7 Flash (77.4) and the other ternary entry, Ternary Bonsai 27B (77.1).
It does not beat **Qwen3.5 35B-A3B at 82.9**, and the gap is not evenly distributed. On LiveCodeBench
Maple actually leads (75.1 vs 74.6); on AIME and HMMT it trails by a few points; on GPQA-Diamond it
trails by **10.7 points** (73.5 vs 84.2) and is beaten even by Qwen3.5 9B (81.7). That shape —
competitive on code and competition math, weak on GPQA — is the signature of a model with strong
reasoning and thinner world knowledge, which is exactly what you would predict from a 1.49B active
budget where the knowledge has to survive ternarization.
The frontier chart is the release's strongest visual and its most selective one. Maple sits alone in
the top right, roughly 3.5× the throughput of the nearest model at comparable quality. But notice who
is *not* plotted: **Qwen3.5 35B-A3B and GLM 4.7 Flash, the two models that beat or match Maple on the
score table, do not appear on the speed chart at all.** In fairness that is close to the point — a
35B model in bf16 does not fit on a Mac mini, which is the whole argument for building this way — but
"a new point on the Pareto frontier" is being claimed against a field that excludes the strongest
competitor rather than measuring it. The honest version of the claim is narrower and still
interesting: *among models that fit and run fast on consumer hardware*, nothing else is close.
## Credit where the card gives it
DeepGrove's own limitations section is short and unusually candid for a launch:
> This preview received minimal post-training for agentic tasks and only small-scale general
> reinforcement learning.
and, in the evaluation section, "this preview is focused primarily on raw reasoning and, as such, may
underperform on agentic benchmarks." That is a lab telling you which axis it did not optimize before
anyone can discover it. It also explains the naming — this is `maple-preview`, not `maple`, and the
card says extended training is coming.
## The take
The interesting claim here is not the benchmark row, it is that **quantization-aware training at
1.58 bits now produces a model that competes with bf16 models several times its active size**. That
claim survives inspection: the weights really are ternary, the compression really does land near the
entropy bound, and the resulting artifact really is small enough to matter on a laptop. Ternary
training has been a research thread for a couple of years, mostly at scales small enough to dismiss.
A 20B model scoring 78.7 average is harder to wave away.
What is missing is the half that makes the numbers real. The packed checkpoint, the ternary kernels
and the Apple Silicon runtime are all unreleased, and the reference implementation in the repository
reproduces the model's *outputs* but none of its *economics*. Right now you can verify that
DeepGrove trained what they said they trained. You cannot yet run it the way they ran it.
---
*Sources: the [Maple-Preview model card](https://huggingface.co/deepgrove/maple-preview) — README,
`config.json`, `configuration_maple.py`, `modeling_maple.py`, `model.safetensors.index.json` and the
safetensors headers — as of 2026-08-06. Both figures are DeepGrove's own, downloaded and flattened
onto white. The per-row ternary measurements and the parameter census were taken with HTTP range
requests against the published shards; no checkpoint was downloaded in full and no benchmark was
re-run. Benchmark numbers are DeepGrove's as printed on their table, with no third-party
replication. Both interactives are mine.*
---
# Pokee-Isaac 28B: 10M tokens on one GPU, and an architecture the report never explains
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/pokee-isaac-28b
> date: 2026-08-06
> tags: llm, long-context, agentic, on-device, explainer
Pokee AI's [technical report](https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf) makes two claims about
**Pokee-Isaac 28B**. The specific one: a **28-billion-parameter, non-decoder-only** model that holds retrieval fidelity
across **10 million tokens**, running on a single NVIDIA B200 at up to **137,200 tokens/s prefill** and **335 tokens/s
decode**, scoring **93.3% on RULER** at that length, and priced at **$0.15 / $1.00** per million input/output tokens. The
general one, stated right in the abstract: long-context agentic capability has been cloud-only because of infrastructure
cost, which locks it out of regulated industries, the public sector, and anywhere data can't leave the building — and a
model this small changes that.
The general claim is worth taking seriously. The specific one is where a close read gets uncomfortable: the report
names its architecture "non-decoder-only" twice, in the abstract and the introduction, and then never says what that
means anywhere in nineteen pages. Take both in turn.
**What this is, precisely.** A technical report posted to Pokee AI's own product console
(`console.pokee.ai`) and stamped `arXiv:submit/7908231 [cs.AI]` — a submission-tracking number, not a public arXiv
identifier. At the time of writing it has not been assigned one, has not been peer reviewed, and ships no weights, no
config file, and no independent replication. Every number below is Pokee's own, on Pokee's own infrastructure, unless
marked otherwise. That doesn't make the numbers wrong. It means nobody outside Pokee has checked them yet.
## "Non-decoder-only": the claim the report doesn't explain
Here is the entirety of what the report says about its own architecture. From the abstract: "Pokee-Isaac 28B is a
**non-decoder-only** foundation model that reasons, plans, and uses tools." From the introduction: "we introduce
Pokee-Isaac 28B, a 10M-token context window, **non-decoder-only** model engineered to operate within a compact compute
budget." That's it. Those two sentences are the complete architectural disclosure. There is no architecture section, no
diagram, no equation, no named mechanism, no ablation isolating what the non-decoder component contributes. The word
"decoder" does not appear again anywhere else in the document.
That absence is conspicuous by comparison. [Kimi K3's technical report](/articles/kimi-k3) spends its first several
sections laying out Kimi Delta Attention's exact recurrence, ships the kernel that implements it, and publishes a
`config.json` that pins down which of its 93 layers are which. [KDA's decay mechanism](/articles/kda-half-life) is
something a reader can derive and check independently precisely because Moonshot wrote the state-update equation down.
Pokee's report has the same page budget and spends none of it there — it moves directly from the abstract to
evaluation tables.
One passage nearby is the closest thing to a hint, and it complicates the claim rather than supporting it: "Some
weights in Pokee-Isaac are fine-tuned from **Qwen3.6-27B** (Apache 2.0)." Qwen3.6-27B is, per its own citation in the
same report, a conventional **27B dense** model — ordinary decoder-only transformer. Isaac is a 28B dense model. The
arithmetic sits right there: a ~1B-parameter gap between a decoder-only base and a model described as non-decoder-only.
That is not evidence of anything specific — it could be a new module bolted onto an inherited decoder backbone, a
retrieval or state component, a modified embedding or head, or something else — and the report gives no way to
distinguish between those. I'm flagging the arithmetic because it's the one concrete data point available, not because
it resolves the question. It doesn't.
So: what actually gives Isaac its 10M-token window and makes it "non-decoder-only" is not something this report lets a
reader verify. That is a real gap, not a stylistic one — it's the single most technically interesting claim in the
paper, and it's asserted rather than shown. Treat everything below as evaluation of *what Isaac does*, because that's
what the report actually lets you check; *how* it does it stays a closed question.
## What the retrieval numbers say
Whatever the mechanism, the report does back the context claim with two established long-context benchmarks, run
against five named baselines: **GPT-5.6 Luna**, **Gemini 3.5 Flash Lite**, and **Claude Haiku 4.5** (the cost-optimized
tier of the three big cloud providers), plus **Nemotron 3 Super 120B** and **Qwen 3.5 122B** (open-weight models an
organization can self-host). Frontier flagships — GPT-5.6 Sol, Claude Opus 5, Gemini 3.1 Pro — are explicitly excluded
as costing roughly an order of magnitude more per token and addressing a different deployment envelope. That's a
defensible exclusion, but worth naming: Isaac isn't compared against the actual frontier, only against the cheap tier
of it.
The RULER protocol is NVIDIA's own official pipeline: 10 samples per task configuration, 13 configurations at 256K and
512K, and — since common-words extraction needs a small fixed vocabulary that stops being meaningful past 1M tokens —
12 configurations from 1M onward. Isaac's own scores across the sweep are **96.9 / 96.7 / 95.0 / 95.8 / 96.7 / 93.3**
at 256K/512K/1M/2M/4M/10M. Read that sequence closely and it isn't a smooth decay curve — it dips at 1M, recovers at
4M, then drops again at 10M. At 10 samples per configuration that's within the noise you'd expect, not a story about
Isaac getting worse and then better with more context, but it does mean 93.3% at 10M is a fairly small-sample number,
not a tight measurement.
Every other baseline falls off a cliff. GPT-5.6 Luna and Gemini 3.5 Flash Lite track Isaac closely through 512K, then
hit **context-overflow errors** at 1M and score zero from there — they simply can't be run at that length, which the
report scores as a failure rather than excusing. Claude Haiku 4.5 and Qwen 3.5 122B score zero across the entire
sweep, consistent with their native windows (200K and 262K) being smaller than even the first column tested. Nemotron
3 Super 120B is the one row worth reading carefully: the 256K–512K–1M figures in the table are **NVIDIA's own
self-reported numbers, not Pokee's measurement** — marked with a superscript `s` and dashed in the chart — while the
2M-onward zeros are Pokee's direct measurement. Mixing a vendor's self-reported numbers into one row of your own
comparison table, clearly labeled, is honest; it's still worth noticing when you're reading the row, since it isn't
measured the same way as the rest of the table.
On **MRCR v2** — a harder multi-needle variant that distributes several targets through a long synthetic conversation
rather than one — Isaac leads throughout: **0.607 / 0.743 / 0.500** at 256K/512K/1M (again non-monotonic — it peaks at
512K, not 256K). Gemini 3.5 Flash Lite is the closest competitor and the gap widens with length, from a 0.133 margin
at 256K to 0.295 at 1M. GPT-5.6 Luna collapses to 0.050 at 1M despite scoring 95.0% on RULER at 256K — a reminder that
single-needle retrieval and multi-needle disambiguation measure genuinely different failure modes, and a model can be
strong at one and weak at the other.
### The comparison this site already has an anchor for
The most natural comparison for a 10M-token claim is [Kimi K3](/articles/kimi-k3), the largest context window
documented on this site until now: **1M tokens** on a **2.78-trillion-parameter** open model. Isaac's framing — implicit
in the numbers, not stated by Pokee this directly — is 10× the context at roughly 1% of the parameters. That ratio is
real arithmetic. It is not, however, a like-for-like measurement:
K3's 1M is what Moonshot trained it up to; Isaac's 10M is a RULER score Pokee measured at that length. Those are
different kinds of number — one a training-curriculum endpoint, the other a benchmark result — and neither the report
nor this piece can turn them into a single fair ratio. What the comparison can support is narrower and still
notable: a 28B dense model holding measured retrieval accuracy at a context length ten times past where a 2.8T model's
*training* stopped. Whether Isaac would still say 93% if someone ran RULER on it at 20M or 50M tokens is not
something either report answers.
## Why 137,200 tokens/s and 335 tokens/s are both true
The efficiency section (Table 8, on a single B200-class GPU under the RULER workload) reports **time-to-first-token**
directly and derives **prefill throughput** from it — context length divided by TTFT. Decode throughput is reported
separately and holds close to flat regardless of context length:
The two numbers describe [different bottlenecks](/articles/how-llm-inference-works): prefill processes the whole
prompt as one large matrix-matrix multiply and is compute-bound, so throughput scales with how much parallel work is
available — which is why it actually *rises* with context length, from ~42K tokens/s at 1M to 137K at 10M. Decode
generates one token at a time against an already-populated cache, a matrix-vector operation gated by memory bandwidth
rather than arithmetic, so it doesn't get faster no matter how much context is resident — 335 tokens/s at 1M, 337 at
10M, 322 under four-way concurrency. The report draws out the one number worth remembering: a 10× jump in context
costs about 3× the time-to-first-token (23.6s → 72.9s), not 10× — which is the behavior that makes a 10M window
usable rather than merely addressable. A full 10M-token prefill landing its first output token at 72.9 seconds is a
real number to plan around, not an abstraction.
## The agentic benchmarks: where "matches or exceeds" holds and where it doesn't
The abstract's claim is specific: Isaac "matches or exceeds the strongest cost-optimized cloud systems on **function
calling, multi-turn interactive execution, tool orchestration, and terminal work**." Four categories, four benchmarks.
Worth checking each against the report's own tables, because they don't all say the same thing.
**Function calling — BFCL v4.** The Berkeley Function-Calling Leaderboard, programmatically scored throughout (no LLM
judge), combining five components under fixed weights (0.40 agentic + 0.30 multi-turn + 0.10 live + 0.10 non-live +
0.10 hallucination) over 5,106 scored entries. Isaac leads the panel at **70.94**, just ahead of GPT-5.6 Luna's
**70.61**. The report itself calls this "parity rather than a decisive lead," which is the right read of a 0.33-point
gap — and it's a fair characterization to give credit for. This category holds up.
**Multi-turn interactive execution — τ³-bench.** Sierra's benchmark runs an agent against an LLM-simulated user across
four domains, verified by a five-criteria rubric rather than an LLM judge's opinion of fluency. Isaac leads the
four-domain average at **0.662**, but that average hides two domains where it doesn't win:
| Domain | Isaac | Luna | Gemini | Haiku | Nemotron | Qwen |
|---|---|---|---|---|---|---|
| Retail | **0.789** | 0.623 | 0.719 | 0.667 | 0.614 | 0.693 |
| Airline | **0.760** | 0.720 | 0.700 | 0.500 | 0.688 | 0.660 |
| Telecom | 0.912 | 0.579 | 0.904 | 0.404 | 0.368 | **0.947** |
| Banking | 0.186 | 0.186 | **0.203** | 0.062 | 0.033 | 0.144 |
| Average | **0.662** | 0.527 | 0.631 | 0.408 | 0.426 | 0.611 |
Qwen 3.5 122B edges Isaac on telecom (0.947 vs. 0.912), and Gemini edges it on banking (0.203 vs. 0.186) — the domain
the report itself calls "by a wide margin the hardest," where the policy an agent needs lives across 698 documents
rather than the prompt. Isaac's own banking score, 18.6%, sits below the 25.5% pass@1 the report cites as the
strongest *previously reported* result on that domain. This category holds up on average, not on every domain.
**Tool orchestration — MCP-Atlas.** This is the one built specifically to avoid mock tool surfaces: 500 tasks against
a live 36-server sandbox of real production MCP servers (GitHub, Slack, Google Workspace, Notion, and more), scored by
mean claim coverage under a shared judge. It's also the one category where the abstract's claim doesn't survive
contact with the table:
Isaac places **third of six**, behind both GPT-5.6 Luna and Gemini 3.5 Flash Lite — two of the three cost-optimized
cloud systems named in the abstract's own comparison panel. The report's honest mitigating point is efficiency, not
score: Isaac reaches within 2.1 points of Gemini using 9.10 tool-call turns against Gemini's 14.99, about 60% of the
trajectory length for comparable coverage. That's a genuine and worth-stating efficiency result. It is a different
claim from "matches or exceeds," and on the benchmark built to be hardest to game, the report's own number doesn't
back the headline phrase.
**Terminal work — Terminal-Bench 2.1.** An agent at a bare command line with no enumerated action space, every model
driven by the same harness (Harbor 0.20.0 with Terminus-2), evaluated on the 86 text-compatible tasks of the 89-task
suite:
Second of six, four tasks behind GPT-5.6 Luna, well clear of everyone else including two open-weight models an order
of magnitude larger. The report is direct about this one: "Terminal-Bench is the one benchmark in this report where a
cloud baseline finishes ahead of Isaac, and we report it as measured" — which is honest as far as it goes, but reads
oddly next to MCP-Atlas two sections earlier, where Isaac trails not one but two cloud baselines. Calling Terminal-Bench
"the one" undersells what MCP-Atlas already showed.
So: of the four categories in the headline claim, function calling holds up as genuine parity, multi-turn execution
holds up on average but not on every domain, and tool orchestration and terminal work both show Isaac behind at least
one of the named cost-optimized cloud systems — behind two of them on the benchmark specifically designed to be
hardest to inflate. "Matches or exceeds" is a fair summary of roughly half the evidence and an optimistic gloss on the
rest.
## Security: safest on attacks, third on capability
Isaac is evaluated on **DTAP**, a red-teaming benchmark measuring attack success rate (ASR, lower is safer) and benign
task success rate (BSR, higher is better) across 12 Linux-Docker domains and 6,195 judged tasks. Isaac is the safest
of the six models on both direct and indirect attack rates and their combination — 35.6% combined ASR against a range
of 37.9% to 66.3% for the rest — and shows the tightest balance between direct and indirect attacks (0.8 points),
where models with less refusal training swing 13 to 38 points toward direct attacks specifically. On capability
(BSR), though, Isaac places third: 82.5%, against 85.1% for GPT-5.6 Luna and 83.3% for Gemini 3.5 Flash Lite — thin
margins, but not a win. And there's a footnote worth reading rather than skipping: "the five baselines were run under
the benchmark's stock runner; Isaac was run under the Pokee harness, which is the one condition that still differs
across rows." A safety benchmark is exactly the place where the evaluation harness itself matters, and this one isn't
held constant.
## The deployment argument, taken on its own terms
Strip away the specific benchmark rows and there's a real argument underneath this report, and it's the one I'd give
the most weight to. Long-context agentic capability today is delivered almost entirely from the cloud, because the
infrastructure to serve it any other way has been expensive. That forecloses the option entirely for organizations
that can't send data across a boundary at all — regulated industries, public-sector deployments, on-device
applications — not because of price, but because the data isn't allowed to leave. A model that holds long-context
agentic capability at 28B dense parameters changes what's *possible* to deploy inside that boundary, independent of
whether it's the best model available outside it.
The pricing table backs a narrower, more concrete version of this point better than the headline "$0.15/$1.00 beats
everyone" framing does. Of the five baselines, only two — GPT-5.6 Luna and Gemini 3.5 Flash Lite — can actually be
*bought* at the context lengths this report tests. Claude Haiku 4.5 caps at 200K, Qwen 3.5 122B at 262K, and Nemotron
3 Super 120B's public endpoints all cap at 262K despite a 1M native window — so three of five baselines simply aren't
commercially available at long context, at any price. Against the two that are, Isaac is cheaper on both meters
($0.25 and $0.80 below Luna on input/output; $0.15 and $1.50 below Gemini) while covering an order of magnitude more
context. That's a real and checkable comparison, distinct from the sovereignty argument, and it holds up on its own —
though it's list pricing marked **provisional and subject to confirmation at launch**, so treat the exact numbers as
directional rather than final.
The product page at `console.pokee.ai` fills in what the report doesn't need to say: an OpenAI-compatible endpoint at
`api.pokee.ai/v1/chat/completions`, streaming over SSE with a background mode that survives a disconnect, and three
concrete deployment tiers — a single B200-class GPU for datacenter serving, a single consumer **RTX 4090 or 5090** for
a private workstation, and Qualcomm or Intel Panther Lake NPU silicon for on-device edge inference. None of that page
explains the architecture either — it repeats "purpose-built agentic architecture" without elaborating, which is
consistent with the report rather than a missed opportunity to say more.
Whether this specific model earns that framing is exactly what the benchmark section above complicates. But the
underlying argument — that long-context agentic capability has had no in-boundary path at all, not merely an
expensive one — is real, underserved by the current market, and worth taking seriously as a category even while
staying skeptical of any one vendor's report about their own entry into it.
## Portability, and what's still thin
Beyond the B200 numbers, Pokee reports adaptation to client and edge silicon. On an **Intel Arc Pro B70**, their own
serving stack reaches 1,087–1,500 tokens/s prefill against 305 tokens/s for stock `llama.cpp` on the same hardware — a
3.6–5× gain — and 58.8 tokens/s decode against 25.7, a 2.3× gain (with a note that a 90 tokens/s decode path is still
"in development," meaning the shipped number is below their own internal target). On a 12-core Xe3 **Panther Lake**
SoC, fully on-device with no discrete GPU: 150.7 tokens/s prefill, 22.84 decode. On a **Snapdragon X2 Elite**: 124.95
prefill, 23.54 decode. AMD support is listed as in progress.
One more benchmark result is worth flagging for provenance rather than dismissing: on something called the
"Pinchbench 116-task SuperClaw suite," Isaac scores 0.9567 against 0.929 for a cloud-hosted 744B model and 0.866 for
an 80B local model. Unlike RULER, MRCR, BFCL, τ³-bench, MCP-Atlas, and Terminal-Bench — every one of which is
independently authored and citable — this suite has no citation, no public description, and no other appearance I
could find outside this report. That doesn't make the number false. It means it can't be checked the way the rest of
this report's benchmarks can, and it shouldn't carry the same weight in your own read of the model.
**The report states three limitations itself, plainly, and they're worth repeating rather than summarizing.**
(1) **Text-only** — no image, audio, or video input, which is also why the evaluation runs 86 of Terminal-Bench's 89
tasks and only the text track of τ³-bench. (2) **Coding was not a training priority** for this release, and the
report contains no code-authoring benchmark; Terminal-Bench measures shell-agent execution, not code generation, and
the report explicitly says Isaac's placement there is "neither evidence of coding strength nor evidence of its
absence." (3) **Hardware adaptation is partial** — AMD support is still in progress, and most silicon families aren't
covered yet.
## What I'd want before trusting this further
Take the report's own framing at face value on one thing: everything here is Pokee's measurement, on Pokee's
infrastructure, reported by Pokee, with no third-party replication and no released weights or config to check
independently — a gap the report is upfront about, but a gap all the same. Specific to this piece, five things I
couldn't verify or that need a second source:
- **The architecture.** Nothing in the report or the product page says what makes Isaac "non-decoder-only" or how the
10M window is achieved. This is the biggest open question in the whole document, and it stays open here too — I'd
rather say that plainly than guess at a mechanism the source doesn't support.
- **The MCP-Atlas and Terminal-Bench placements**, where Isaac trails one or two of the exact cost-optimized cloud
systems the abstract claims it matches or exceeds.
- **The DTAP harness difference** — Isaac run under Pokee's own harness against five baselines run under the
benchmark's stock runner, on the one evaluation most directly about safety.
- **The "Pinchbench" result**, which has no independent citation or description to check it against.
- **Pricing and general availability.** Rates are explicitly provisional, and neither the report nor the product page
states a launch date or confirms the model is generally available yet.
## The take
The mechanism claim doesn't survive scrutiny, because there's nothing offered to scrutinize — "non-decoder-only" is
asserted twice and explained zero times, in a report that had the exact same page budget Moonshot used to write down
KDA's recurrence in full. The benchmark claim survives partially: real parity on function calling, a real average
lead on multi-turn tasks that hides two domain losses, and two categories — tool orchestration and terminal work,
including the one benchmark built specifically to resist gaming — where Isaac trails cost-optimized cloud systems by
the report's own numbers, not a critic's.
What does survive, and what I'd actually flag as the interesting part of this report, is the deployment argument
underneath all of it. Long-context agentic capability being cloud-only is a real constraint today, and it genuinely
does foreclose entire categories of deployment — not because of price, but because data can't leave a boundary at
all. A 28B dense model that holds measured retrieval accuracy for even a fraction of a 10M-token claim, running on
hardware small enough to sit in a private workstation, is a meaningfully different option than what existed before —
regardless of whether this particular model, from this particular report, is the one that delivers it best. That
argument deserves to be taken on its merits. This report, on its own, isn't yet the evidence that settles it.
---
*Sources: the [Pokee-Isaac 28B technical report](https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf)
(all benchmark tables, the efficiency and pricing profile, and the two architecture sentences quoted in full above),
and [console.pokee.ai/model](https://console.pokee.ai/model) (API details, deployment tiers, pricing display). Figure
1 here is the report's own Figure 1, cropped from the source PDF and flattened onto white for legibility in both
themes — no relabeling. RULER, MRCR v2, BFCL v4, τ³-bench, MCP-Atlas, Terminal-Bench 2.1, and DTAP are each
independently authored benchmarks cited in the report; none of the scores above are this site's own measurement. The
context-length and prefill/decode interactives are mine, built entirely from numbers in the report's own tables — no
extrapolated or simulated figures appear in either.*
---
# Prime Agent: the interface is a Python REPL, not a tool-call schema
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/prime-agent
> date: 2026-08-06
> tags: agents, coding-agent, harness, open-source, prime-intellect, explainer
Most agent harnesses give the model a menu. You define `grep(pattern, path)`, `read_file(path)`,
`run_tests()`, each with a JSON schema, and the model picks one, the harness parses the call,
runs it, and hands the result back as another message in the transcript. Composition — loop over
these results, retry that one, spawn three of these in parallel — happens in the *conversation*,
one role-tagged message at a time, because the schema has no concept of control flow.
[Prime Agent](https://github.com/PrimeIntellect-ai/prime-agent), Prime Intellect's open-source
coding and research agent, makes a different bet. It gives the model one tool: a persistent Python
interpreter. Composition is not a transcript pattern the harness has to support — it is just
Python. This piece is a tour of that decision, built from reading the actual TypeScript host and
Python runtime in the repo (not just the README), plus the second idea Prime Agent ships alongside
it: a harness that edits its own supplemental state through `/refine`, with one part of itself —
the base system prompt — locked out of the edit path in code, not just in the prompt.
**On the clone.** A first pass at this piece was written against a *shallow* clone, which turns out
to be a much worse instrument than I gave it credit for — it shows one commit and tells you nothing
about a project's real age. Everything below is built from a full clone instead: 4,473 commits of
history, the complete TypeScript host, and the Python runtime. The maturity section near the end is
where that difference bites hardest.
**Read this before the rest.** Prime Agent is open source (MIT) but it is a vendor's product for
its own stack — there is no third-party evaluation of it anywhere. I read the whole README, the
docs under `packages/coding-agent/docs/`, and the source, and there is not one accuracy, pass-rate,
or SWE-bench-style number in any of it. Not a "we're competitive with X," not a chart, nothing. So
treat everything below as an architecture and a set of design decisions — genuinely well-documented
ones — not a measured result. I say more about what maturity signals *do* exist near the end.
This continues two threads already on this site. [Agent harnesses](/articles/agent-harness) argued
the loop wrapped around a model matters as much as the model itself, and [the harness
effect](/articles/harness-effect) showed orchestration — not the model — sets an agent's token
bill. Prime Agent is a concrete instance of both claims taken further: the orchestration layer
here is not a fixed loop around a fixed tool menu, it is a programming environment, and the harness
state that shapes behavior is itself something the agent is allowed to edit.
## One tool, not a menu
Here is the entire tool surface Prime Agent gives the model, from
`packages/coding-agent/src/core/tools/ipython.ts`:
```typescript
const ipythonSchema = Type.Object({
code: Type.String({
description:
"Python scratchpad code or `%%bash` shell cells to execute in the agent kernel. Use the target project's own environment for project imports, tests, scripts, CLIs, and dependency checks instead of direct kernel imports.",
}),
})
```
That is one parameter: `code`, a string. Compare that to a typical schema-based agent, which
carries a dozen or more tool definitions — `read_file`, `write_file`, `bash`, `grep`, `glob`,
maybe a bespoke one per integration — each with its own JSON schema, each repeated in every
request the provider sees. Prime Agent's own `CHANGELOG.md` records the direction of travel: an
early entry reads "Removed the interactive `!` / `!!` bash shortcuts; use IPython for shell
commands." Shell access isn't a separate tool bolted on next to Python. It is a magic cell
(`%%bash`) inside the same interpreter.
The kernel is genuinely persistent. Variables, imports, and open file handles survive across
tool calls — and across context compaction — because they live in the interpreter's process, not
in the token transcript the model rereads every turn:
```python
from pathlib import Path
config_files = list(Path(".").rglob("*.toml"))
large_files = [path for path in config_files if path.stat().st_size > 10_000]
```
`config_files` is still there three turns later. Nothing re-lists the directory, and nothing
re-sends the file list back through the model's context to remind it what it found — the model
holds a *reference* to the data, not a copy of it in its own working memory. That is the
"prompt-as-a-variable" half of the [Recursive Language Model](https://www.primeintellect.ai/blog/rlm)
idea Prime Agent is built on: context becomes something you slice with Python, not a transcript
you re-read.
## Context as variables, made concrete
Take a task like "grep across 60 files for a pattern and summarize the hits." A schema-based
agent issues one `grep` call, gets a list of matches back, and then — if it wants to actually look
at what it found rather than trust the grep output blind — issues one more round trip per file it
wants to inspect, and a final call to write the summary. The loop lives in the conversation: every
iteration is a full model turn, with the tool schemas and message envelope repeated each time.
An RLM turn writes the loop instead of living inside one:
```python
import subprocess
hits = subprocess.run(
["grep", "-rl", "TODO(perf)", "src/"], capture_output=True, text=True
).stdout.splitlines()
summaries = []
for path in hits:
text = Path(path).read_text()
summaries.append(f"{path}: {text.count('TODO(perf)')} occurrences")
print("\n".join(summaries[:10]))
print(f"... {len(summaries)} files total")
```
The `for` loop, the file reads, and the counting all happen inside one `ipython` call. The model
sees one printed summary, not sixty round trips of tool call and tool result. `summaries` stays a
Python list the model can filter, sort, or hand to another cell — it does not have to be re-stated
in the transcript to stay usable.
Drag the slider above. The gap is not a fixed multiplier — it is linear-versus-flat, so it gets
more dramatic exactly where it matters most: large fan-out tasks. The numbers there are a cost
model built to make the *shape* of the tradeoff visible, not a benchmark; Prime Agent doesn't
publish one, so neither do I.
This is also where the honesty has to cut both ways. Working in a persistent kernel does not make
the model's own attention free — if `summaries` is genuinely large, someone (or something) still
has to look at it, and dumping ten thousand lines of `print()` output into the transcript defeats
the entire point. The advantage is that the *decision* about how much of the data to surface is a
line of Python (`summaries[:10]`) instead of a constraint baked into the tool schema. It is a
better failure mode, not an absent one.
## The schema didn't disappear — it moved
Here is the thing I got least precise about the first time, and it is the most interesting
mechanism in the codebase. "One tool" is true of what the *provider* sees. It is not true of what
the model can reach.
When Python in the kernel calls `rlm(...)`, or `goal.get()`, or `agent_message.send(...)`, it is not
doing the work locally. It opens a Jupyter comm target named `host.request` and sends a typed
request back across the ZeroMQ boundary to the TypeScript `AgentSession`, which does the work and
replies. `rlm-runtime.md` is blunt about the division: the Python `rlm` package "is a model-facing
shim; the TypeScript host owns child execution, persistence, usage accounting, and lifecycle," and
"the Python side does not call providers or implement an agent loop."
The dispatch table is built in `_createKernelHostHandlers()` in `agent-session.ts`. I counted
twenty-one entries:
Two separate things fall out of that, and they are worth not conflating.
The cheap one is token economics. A conventional harness pays for its tool surface in every single
request — twenty-one JSON schemas re-serialized into the prompt on every turn, forever. Prime Agent
pays for its surface once, in the kernel bootstrap and the skills' `SKILL.md` files, and the
per-turn cost of the whole bridge is zero. That is the same argument as the turn-count one above,
pointed at a different axis.
The load-bearing one is that most of these handlers are registered **conditionally**. Goals are
wired up only `if (this._includeGoals)`. Compaction only `if (this._includeCompactSkill)`.
Refinement only `if (this._autoRefineAllowedForSession())`. Messaging requires both a controller
*and* that the `agent-message` skill be in the model-visible set. In a session where those flags are
off, the handler is not in the map at all — the model can write the exact right Python and get an
error from the host, because there is nothing on the other end of the comm to answer it.
That reframes the security story in a way the README's "not a security sandbox" warning does not,
and it is a genuine tension inside the design rather than a resolution of it. Arbitrary Python
against the filesystem really is unbounded: the kernel runs with the user's permissions and can do
whatever Python can do. But the *agent-control* surface — spawn a child, open a goal, edit the
harness, message a sibling, read another session's transcript — is not open. It is a typed,
argument-validated, conditionally-registered table in the host, which is exactly the property a tool
schema is supposed to give you. Prime Agent kept the schema and moved it somewhere the model cannot
see or enumerate, then gave the model a general-purpose language for calling into it.
## Subagents are function calls
The same move applies to delegation. `rlm` is preloaded in the kernel as a callable:
```python
handle = await rlm("Review the authentication flow for security issues", name="auth-reviewer")
print(handle.rlm_child_id, handle.name, handle.session_dir, handle.model)
```
`rlm(...)` is admission, not completion — it returns as soon as the TypeScript host has created a
real child `AgentSession` with its own context and session directory, and it never blocks waiting
for the child's answer. That is a real API design decision, not an implementation detail: the
`CHANGELOG.md` for `0.6.0` records changing `rlm(...)` from waiting for the child to finish to
returning a spawn handle at admission, specifically because treating `asyncio.gather()` over
several `rlm()` calls as fan-in was the wrong mental model — spawning three reviewers is three
independent calls, not a scatter-gather:
```python
api_review = await rlm("Review the public API", name="api-reviewer")
test_review = await rlm("Review the test coverage", name="test-reviewer")
integration_audit = await rlm("Run the slow integration audit", name="integration-audit")
```
A child reports back only through an explicit message, never through the `rlm()` return value:
```python
await agent_message.send(message, receiver_role="parent")
```
the same daemon-routed messaging `prime-agent send "..."` uses from the shell. Reach is
deliberately narrow — an agent may message or observe only its parent, siblings, and direct
children (the `0.6.0` changelog calls this "the nuclear family"); reaching a grandchild means
relaying through the intermediate child. That is a real constraint on the "agents can message each
other and orchestrate without routing through the user" claim: it is true, but bounded, not an
open mesh.
## Skills are Python you can call, not prompts you re-paste
Skills follow the standard [Agent Skills](https://agentskills.io/specification) markdown format
(a `SKILL.md` with frontmatter Prime Agent loads lazily), extended with a Python-backed variant: a
skill directory with a `pyproject.toml` gets installed into the kernel's virtualenv and exposed by
import name.
```python
report = await release_audit(repository=".", target_version="0.4.0")
```
That's a real callable, not a re-explained prompt — Prime Agent's built-in `skill-creator` skill
turns a described workflow into exactly this shape: `SKILL.md` plus `src//__init__.py`
plus a documented `run()`. Worth being precise about a distinction the docs themselves flag: an
*installed* Python skill is a package on disk; a continual-harness *skill entry* (below) is a
persisted description of a reusable call. `/refine` can create or update the description after it
sees a repeated pattern, but it never packages the executable capability itself — that stays
`skill-creator`'s job.
## MCP, without adding a tool
The clearest test of whether a "one tool" design is a real commitment or a slogan is what happens
the first time someone wants Linear and Notion in the agent. The default answer everywhere else is
to mount the MCP server's tools into the model's tool list, which is how a clean six-tool agent
becomes a forty-tool agent nobody planned.
Prime Agent's docs refuse the move in the first paragraph: "Consistent with Prime Agent's
single-tool design, MCP integrations are **not** exposed as new agent tools." An integration is a
Python skill whose module subclasses `McpIntegration`, and the MCP connection runs *inside the
kernel* on the official `mcp` Python SDK. The host's only jobs are browser OAuth and keeping a
token fresh in `auth.json`.
```python
import linear
for tool in await linear.list_tools():
print(tool["name"], "-", tool["description"])
help(linear.list_issues) # schema, after list_tools() has run
issues = await linear.list_issues(team="Engineering")
```
Every discovered tool is bound as an async method on the integration object; results come back as
parsed Python rather than JSON to unpack; a tool whose name isn't a valid identifier
(Notion's `notion-search`) falls back to `await notion.call_tool("notion-search", {...})`. Authoring
your own is a `pyproject.toml`, an `mcpServers` entry in settings, and roughly ten lines subclassing
the base.
Two details are more interesting than the API itself.
The first is that **discovery moved into the turn**. In a schema-mounted MCP integration, tool
definitions are resolved when the harness connects and then frozen into the prompt; the model gets
the server's surface whether it needs it or not, and a server that changes its tools mid-session is
a stale-schema bug. Here the docs tell the model to `list_tools()` and `help()` before calling
rather than hardcoding, because "tool names and argument schemas come from the server and can
change." The model pays for the schema only in the turns where it actually looks it up.
The second is a small landmine that says a lot about how the kernel works. The reference
integration's module-level `__getattr__` forwards unknown attributes to the instance, but keeps a
reserved list:
```python
_RESERVED = {"run", "__wrapped__", "__call__"}
```
Forwarding `run` would make the kernel bootstrap, which probes modules for a callable entrypoint,
mistake the whole integration module for a callable skill and break dispatch. That is the flavour of
bug you only get when your tool boundary is Python's attribute protocol instead of a JSON schema —
more expressive, and with sharper edges.
The auth-gating is asymmetric in a way worth knowing before you write one. **Built-in** integrations
(Linear, Notion) ship installed but disabled, are excluded from the prompt, and only get imported
into the kernel once credentials exist. **User-authored** ones are not gated that way at all: drop a
skill into a skills directory and it is visible and imported immediately, failing at call time with
`NotEnabled` until you log in. So your `SKILL.md` has to tell the model how to connect — and tell it
the right way, since `/mcp login` works only for OAuth servers and reports "Unknown MCP integration"
for a bearer-token one.
## The Continual Harness: durable state the agent is allowed to edit
Everything so far is inside one turn. The [Continual Harness](https://arxiv.org/abs/2605.09998)
(arXiv 2605.09998) is about state that outlives the turn — and the session, if you ask for it to.
It has four editable kinds, defined in `refinement.ts`: `prompt` (supplemental behavioral notes),
`memory` (durable facts and decisions), `skill` (a description of a reusable Python call), and
`subagent` (a reusable delegation role). Each entry lives in one of two scopes — `local`, written
to the current session's own `harness/harness_state.json` and gone with the session unless
promoted, or `global`, written to `~/.prime/agent/harness/` and available to every future session.
`/refine` is the mechanism that writes to this state. It reviews the current trajectory and, when
it finds something worth persisting, emits small Create/Update/Delete edits — never a full
rewrite. From the actual system prompt the host sends to the refiner model:
```text
Use the trajectory, current continual harness state, and prior refinement history. Prefer
small evidence-backed edits. If prior refinements caused issues, rollback or replace the
faulty editable entries. Never edit source files directly.
```
## The one thing `/refine` cannot touch
The interesting design choice is not that the harness can improve — plenty of systems do prompt
optimization. It's what's carved out of the edit surface, and how that carve-out is enforced.
`validateEdit()` in `refinement.ts` runs before any edit is applied:
```typescript
if (edit.kind === "prompt" && (edit.id === "base_system_prompt" || computedId === "base_system_prompt")) {
return "base system prompt is not editable";
}
```
That is not a prompt instruction the model could talk itself out of — it's a function that runs on
every proposed edit, in the host, outside the model's control. The base system prompt is
compiled once from the harness's own instructions, and any attempt to create, update, or delete an
entry with that id is rejected before the edit ever lands. Everything the harness learns goes into
one of the four editable kinds instead, injected at the top of the compiled prompt as clearly
subordinate material: "Use these continual harness prompt notes, memories, skills, and subagent
specs when they are relevant. The base system prompt is immutable; prompt entries below are
supplemental notes only."
That matters because it draws a hard line between two very different kinds of self-modification.
The model can accumulate memories, refine delegation roles, and tighten behavioral notes — real,
compounding change to how it behaves — but it can never touch the instructions that define what
counts as a legitimate edit in the first place. Nothing in the four editable kinds can rewrite the
rule that keeps them editable-only. It's the same shape as a constitution that can be amended but
whose amendment procedure is (by design) not itself amendable through the ordinary amendment
process.
Every applied edit is versioned and every refinement pass is appended to
`refinements.jsonl` with before/after entry state, which is what makes rollback possible:
`refineHarness()` accepts a `rollbackId` and, instead of running the LLM proposal pass again,
replays a target refinement's prior state as the new edit. If a `/refine` pass turns out to have
been wrong, the fix is pointing the entry back at an earlier recorded version — not trusting a
second LLM call to undo the first one's mistake correctly.
Read next to [Recursive Harness Self-Improvement](/articles/recursive-harness-self-improvement),
published today, the contrast is worth stating plainly. Sakana and Berkeley's method compares a
harness against its own immediately-previous version and keeps the winner — a research method with
a real information-theoretic argument for why pairwise beats population search, but no product
around it. `/refine` is the shipped, product-side sibling of that same instinct: also self-vs-self
in spirit (evidence from *this* trajectory, checked against *this* harness's own history), but with
no comparison objective, no accept/reject criterion beyond "small and evidence-backed," and a
rollback button instead of a formal proof. One is a method with a Bradley-Terry argument behind it;
the other is a feature with a JSONL log behind it. Neither is a lesser idea for that — they're
answering different questions — but they shouldn't be mistaken for the same rigor.
[MemHarness](/articles/memharness), also published today, is a useful contrast in the other
direction. MemHarness's argument is that retrieved memory should be *reconstructed* — critiqued and
rewritten against the current state — every time it's used, because verbatim replay of a stale
memory can hurt more than having none. The Continual Harness takes the opposite bet on when the
work happens: refinement is a deliberate, evidence-gated event ("prefer small evidence-backed
edits") that happens rarely, and once written, an entry is trusted and injected verbatim into every
future compiled prompt until the next refinement touches it. MemHarness spends compute at *read*
time, on every retrieval; Prime Agent spends it at *write* time, once, on `/refine`. Neither is
obviously right — cheap reads with occasional expensive writes versus expensive reads with cheap
storage — but it's worth knowing you're choosing between them, and Prime Agent has made the choice,
not left it implicit.
## Two memories, and only one of them forgets
A persistent kernel gives an agent two independent memory systems, and I don't think that gets said
plainly enough. The transcript is one: bounded by the context window, and periodically summarized
away. The kernel namespace is the other: bounded by RAM, and *never* summarized. Compaction only
touches the first.
Auto-compaction fires when `contextTokens > contextWindow - reserveTokens` — 16,384 reserved by
default — walks backwards from the newest message accumulating tokens until it has kept
`keepRecentTokens` (20k by default), summarizes everything before that cut into a structured
document, and reloads the session as summary-plus-recent. `long-running-agents.md` states the
kernel's exemption directly: "The IPython kernel persists through compaction, so variables, imports,
helper functions, and task state remain available."
So the same object can be simultaneously forgotten and present. The model may no longer have the
message where it built `summaries`, but `summaries` is still bound in the interpreter. Both halves
of that are useful and both can bite: the good case is that fifteen minutes of expensive analysis
survives a compaction intact; the bad case is that the model retains a variable whose *provenance*
was summarized away, and has to re-derive what it means. The structured summary format is clearly
designed against this — it carries explicit `## Critical Context`, `` and
`` blocks precisely so the pointers outlive the prose.
Three implementation details reveal where the pressure actually is:
- **Tool results are truncated to 2,000 characters** during the serialization that feeds the
summarizer, with a marker recording how much was dropped — because, in the docs' own words, tool
results "especially from `ipython` and optional `bash`, are typically the largest contributors to
context size." The REPL design makes compaction *harder*, and this is the mitigation.
- **Split turns get two summaries.** Normally the cut lands on a turn boundary. When one turn is
itself bigger than `keepRecentTokens` — which is exactly what a long autonomous stretch of kernel
work produces — the cut lands mid-turn on an assistant message, and Prime Agent summarizes the
history and the turn prefix separately, then merges them. Never at a tool result: those must stay
attached to their call.
- **Compaction is explicitly not a stopping condition.** It "does not stop goals, autonomous
continuations, heartbeats, or existing child sessions." A harness that treated a full context as
the end of a task would quietly cap every long-running job at one context window.
Both compaction and the `/tree` branch summarizer accumulate file operations *cumulatively* across
passes, so the record of what was read and modified survives repeated compactions rather than being
re-derived from a summary of a summary.
## What runs when nobody is attached
The last piece, and the one easiest to miss from the README alone: Prime Agent is built for
sessions with no human in front of them. Sessions live in resident daemon worker processes, so closing the terminal detaches a
client rather than stopping the work, and there are four separate mechanisms for producing a prompt
when no user is typing.
- **Heartbeats**, in two flavours. `/heartbeat every 10m ...` is the user's single visible recurring
instruction; `rlm_heartbeat.create(...)` is the agent's own, plural and programmatic. The Python
skill deliberately cannot clear or replace the user-owned one.
- **Schedules** — `prime-agent schedule add worker "0 9 * * 1-5" -- "Review open work"` — persisted
per session, surviving detach. The reliability detail is good: due ticks are claimed before
delivery so a crash cannot replay an uncertain prompt, and missed ticks are coalesced rather than
accumulating into a backlog.
- **Goals**, which store a durable objective plus token usage, elapsed time, continuation count and
an optional budget. Only `await goal.complete()` marks one done. The docs are careful that a goal
is "an explicit user or host action, not something the agent should infer from every task."
- **Autonomous mode**, which is the policy that decides whether to inject another continuation. It
is bounded on four axes at once — continuations, assistant turns, tokens, wall clock — and gated
on shell commands (`--autonomous-gate "npm run check"`) that must pass before the session may
finish, with a failed gate's bounded output returned to the agent for another attempt.
The division there is sharper than most agent products bother with: the goal holds *what* and *how
far along*, autonomous mode decides *whether to continue*. And one line in the gate policy is worth
stealing outright — Prime Agent "avoids rerunning the same failed gate when the workspace has not
changed." An agent that reruns a two-minute test suite against a byte-identical tree is not
verifying anything, it is billing you for a cached failure.
All of it funnels into the same queue. From the session queue onward, a prompt from a heartbeat,
a cron schedule, a goal continuation, autonomous mode, or another agent takes exactly the same path
as one typed by a person, which is why none of these features needed a parallel execution mode.
One consequence of that design shows up in `--mode acp`, added in `0.6.0`, which runs Prime Agent as
an [Agent Client Protocol](https://agentclientprotocol.com) agent so arbitrary clients can drive it.
IPython maps cleanly onto ACP's `execute` tool call. Everything above does not: subagents, autonomous
gate state, compaction, goals, heartbeats and continual-harness refinement have no native ACP
concept, so they ride in a namespaced `ai.primeintellect.prime-agent` `_meta` envelope that vanilla
clients ignore. That is a fair summary of where this design sits relative to the emerging standards —
the single-tool core is portable, and most of what makes it interesting is an extension.
## What programmatic execution costs
The tradeoff the whole design rests on: a persistent interpreter is more capable and much harder
to bound than a fixed tool schema. The README says this plainly, not buried in a docs page:
> Prime Agent executes model-generated Python and project commands with your user permissions.
> Its worker and kernel processes improve lifecycle isolation and recovery; they are **not** a
> security sandbox. Review changes and use trusted repositories, instructions, skills, and
> extensions only.
`rlm.md`'s trust-model section says the same thing about the kernel specifically: it "runs
model-generated Python and project commands with the worker's operating-system permissions. It is
a durable control environment, not a security sandbox." A fixed tool schema at least gives you an
enumerable attack surface — every action the model can take is one of N defined functions, each
individually auditable and individually deniable. A REPL's attack surface is "anything Python (and
`%%bash`) can do," which is a much larger set to reason about, and a much easier one for a
malicious skill or a compromised MCP integration to abuse.
The host bridge splits that claim in two, and the split is the honest version. Against the
*machine* — files, network, processes, credentials on disk — the REPL really is unbounded, and no
amount of typed dispatch changes that. Against the *agent system* — spawning children, opening
goals, editing harness state, steering a sibling, reading another session's transcript — the surface
is exactly as enumerable as a tool schema, because it *is* one: twenty-one named handlers, arguments
validated in TypeScript, most of them absent unless a session flag turned them on. When you read
"not a security sandbox," read it as a statement about the filesystem, not about the agent graph.
Persistence has an operational cost too, separate from the security one: a kernel that gets stuck
stays stuck. The host's own busy-kernel handling spells out the tradeoff directly — interrupting a
runaway cell and it still hasn't stopped, the choices are "wait" (preserve state, keep waiting) or
"kill" (lose every in-memory variable, import, and running task and restart clean). There is no
third option where you get both a responsive kernel and the state back. A schema-based tool call
that hangs just times out; a wedged interpreter is holding real, valuable state hostage to its own
unresponsiveness.
The daemon-backed background sessions and inter-agent messaging compound this rather than replace
it. Sessions run in resident worker processes that survive a detached terminal — genuinely useful
for long tasks, and `long-running-agents.md` is honest that this is a lifecycle property, not a
security one: "Daemon workers are process-isolated for lifecycle and failure containment, not
security-sandboxed. They normally run with the same operating-system permissions as the client."
Add agents that can message and steer each other's active work and the blast radius of one
compromised or badly-instructed agent is no longer just its own kernel — it's whatever its parent,
siblings, and children will act on without a human turn in between. The nuclear-family reach limit
added in `0.6.0` is a real mitigation, but it bounds propagation, it doesn't remove the surface. And
`0.7.0` moved in the other direction on the same axis: agent messages now *always* steer, injecting
into a running turn, with the option to queue politely behind the current work removed from every
API. That is almost certainly the right default for responsiveness, and it does mean an inbound
message from a sibling always interrupts.
None of this makes Prime Agent unusual among agent products with shell and code-execution access —
it makes the tradeoff explicit and names it in the docs instead of marketing around it, which is
more than most.
## Maturity, honestly
This is the section the shallow clone got wrong, and the fix is more interesting than the error.
On that first pass I could not audit the commit history at all — the clone showed one commit — so I
declined to claim anything about the project's age. That was the right call given the instrument,
and the wrong instrument. A full clone answers the question completely, and the answer is not what
either a shallow clone or the changelog suggests.
**4,473 commits, 231 distinct authors, 48 release tags**, first commit 2025-08-09, latest
2026-08-06, with pull request numbers past #660. That is a year-old project with a real
contributor base, not a three-month-old repo — and it is worth saying that my earlier hedge, read
charitably, was still an *underestimate* by an order of magnitude.
But the interesting number is the split. **Mario Zechner authored 3,099 of those commits — 69%** —
and his last one is 2026-05-08. Prime Intellect's first commit lands 2026-05-21. Everything before
the gap is [`pi-mono`](https://github.com/badlogic/pi-mono) under its own name; the `clean up legacy
pi artifacts` commit lands 2026-05-19. In the three months since, **482 commits by 17 authors** have
turned it into Prime Agent.
So "how mature is this?" has two honest answers depending on what you're asking. The *codebase* is
a year old, heavily iterated (December 2025 and January 2026 alone account for 2,096 commits), and
was production software before Prime Intellect touched it. The *product* — the RLM framing, the
continual harness, `/refine`, the daemon-backed agent tree, the MCP-as-skill design — is three
months old and mostly the work of a small team. If you are evaluating engineering quality, use the
first number. If you are evaluating how settled the agent architecture is, use the second, and note
that `0.6.0` and `0.7.0` both shipped breaking API changes inside 24 hours of each other.
The changelog backs that up rather than contradicting it. `0.6.1` and `0.7.0` both landed in the
three days before this was written, and `0.7.0`'s single breaking change is instructive: agent
messages now "always use steering delivery," and the `mode` parameter is gone from the Python, CLI,
RPC, and connection APIs. The three-mode design (`auto`, `steer`, `follow_up`) I would have
described as a feature two days ago has been collapsed into one behaviour. `long-running-agents.md`
still documents all three and still shows `mode="auto"` in its example; the actual `send()` in
`agent-message/src/agent_message/__init__.py` no longer accepts it. Docs lagging source by one
release is a normal cost of moving this fast, and a reason to read the Python, not the markdown.
The rest of what I can check from the repository holds up: version `0.7.0` across all four
TypeScript workspaces (`ai`, `agent`, `tui`, `coding-agent`), a `CHANGELOG.md` per package that
tracks breaking changes deliberately (a house rule bans touching already-released version sections),
GitHub Actions for CI and for building versioned release binaries with SHA-256 checksums, and 414
TypeScript test files plus 4 Python test files under `prime-agent-runtime/test/` across roughly
341,000 lines of TypeScript. The daemon protocol is explicitly versioned
(`DAEMON_PROTOCOL_VERSION`, a schema revision — now at 13 — and compatibility maps for old-client/
new-daemon and new-client/old-daemon pairs) — the kind of care you only add after being burned by
version-skew bugs.
One provenance correction while I'm here. I called Prime Agent "an acknowledged hard fork" of
pi-mono, which is how the docs describe it and is fair as a statement about lineage. The git history
says something more literal: this is not a copy of pi-mono's code, it *is* pi-mono's repository,
history unbroken from Zechner's first commit through to today's. The license reflects it — MIT,
copyright jointly held by Mario Zechner (2025) and Prime Intellect (2026) — and the README credits
pi-mono in its header links. Second: `assets/` in the repo has a brand logo (an SVG butterfly mark)
and nothing else — I checked specifically for architecture or benchmark figures to embed as this
site's house style asks for, and there aren't any. The only other images in the repository are TUI
screenshots under `packages/coding-agent/docs/images/`, and they carry the pre-rename `pi-mono`
branding from before the fork was productized, which would misrepresent the current product if
reproduced here. So this article ships no cover image and no embedded repo figures — the four
interactives above are original, built from reading the code and measuring the repository, not
redrawn from anything Prime Intellect published.
## Where this sits
[Scaling agentic RL](/articles/scaling-agentic-rl) covered Prime Intellect's environments side —
23 agentic tasksets, roughly 365,000 tasks behind one taskset API, each with a graded, reproducible
reward. Prime Agent is the natural agent-side counterpart to that stack: the same company building
the environments an RL loop trains against is also shipping the agent architecture that would run
inside them. I want to be careful about what that observation is and isn't — I did not find any
published result training or evaluating Prime Agent against that taskset catalog, so this is a
structural connection (same company, complementary halves of an agentic-RL stack), not a reported
one. If that pairing produces a number, it belongs in a different article than this one.
What Prime Agent actually is, stripped of both the marketing framing and my own enthusiasm for the
design: an open-source harness that replaces a tool-call schema with a programming environment, and
a harness state that can accumulate evidence-backed edits without ever being allowed to rewrite the
rule that makes those edits legitimate. Both are real, checkable design decisions. Neither comes
with a number attached.
The revision changed my read on one of them. "Replaces a tool-call schema with a programming
environment" is the marketing line and it is half right. What Prime Agent actually did is *demote*
the schema — out of the provider payload, where it is re-billed every turn and every entry competes
for the model's attention, and into a private typed bridge the model reaches through a general
purpose language. Twenty-one operations, argument-validated, conditionally registered. That is a
better idea than abolishing the schema would have been, and it is the part I would steal.
---
*Sources: the [prime-agent repository](https://github.com/PrimeIntellect-ai/prime-agent) at commit
`fix(coding-agent): isolate kernel state tests (#661)`, 2026-08-06 — specifically the docs under
`packages/coding-agent/docs/` (`architecture.md`, `rlm-runtime.md`, `rlm.md`, `compaction.md`,
`mcp-integrations.md`, `long-running-agents.md`, `acp.md`), the TypeScript host
(`agent-session.ts`, `refinement.ts`, `agent-messages.ts`, `tools/ipython.ts`), the Python runtime
under `prime-agent-runtime/src/rlm/` and `packages/coding-agent/skills/`, and the per-package
`CHANGELOG.md` files. All repository statistics — commit counts, author counts, per-month
distribution, tag count, line counts — were measured with `git` against a full clone and are
reproducible from the commands in the source of the commit-history figure. The Continual Harness
paper is [arXiv 2605.09998](https://arxiv.org/abs/2605.09998); the RLM framing is Prime Intellect's
[RLM post](https://www.primeintellect.ai/blog/rlm). The four interactives are mine.*
---
# Recursive Language Models: context as a variable, recursion as a function call
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/recursive-language-models
> date: 2026-08-06
> tags: agents, llm, recursion, context-management, prime-intellect, explainer
The standard agent loop has one data structure at its center: the transcript. The model reads a
growing conversation, emits a tool call that matches a JSON schema, a harness parses it, runs it,
and appends the result back as another message. Everything the model can act on has to live in
that transcript. Everything it does has to fit the schema. [Agent harnesses](/articles/agent-harness)
and [the harness effect](/articles/harness-effect) are both, in different ways, about how much that
loop shape costs — in design effort and in tokens. A **Recursive Language Model** (RLM) doesn't
optimize the loop. It replaces the data structure.
An RLM gives the model a persistent Python REPL instead of a transcript. Two things follow from
that, and they are the spine of this piece:
1. **Prompt-as-a-variable.** The input — a document, a codebase, a 500 MB corpus — becomes a
value bound to a name in the REPL, not text sitting in the context window. The model writes code
to slice, filter, and search it, and only what it chooses to print ever enters its own context.
2. **Programmatic recursion.** Calling another language model is a function call — `rlm(...)` — that
returns a value like any other call. It composes with `for` loops, `if` statements, `map`, and
error handling, because it *is* one of those, not a special harness verb bolted on next to them.
**Idea, implementation, product — three different things, one name.** Alex Zhang (MIT CSAIL) coined
"Recursive Language Model" for a specific inference-time technique: one long prompt held as a REPL
variable, decomposed by recursive sub-LM calls, terminated by an explicit `FINAL()`/`FINAL_VAR()`.
Prime Intellect built two things that carry the same name and the same two mechanisms but are not
the same system as each other or as Zhang's original design: `RLMEnv`, a research-eval
reproduction inside their `verifiers` library, and Prime Agent, a general coding harness that
generalizes both mechanisms — context-as-data, recursion-as-a-call — to its entire working
environment, not just one input string. [Prime Agent](/articles/prime-agent) covers that second one
in full, including its separate Continual Harness feature. This piece is about the idea and how the
paper measures it; read that one for the shipped product.
## Who actually built this
Prime Intellect's own post is straightforward about credit, and I'll be too: "the Recursive
Language Model (RLM), introduced by Alex Zhang in October 2025 as a blog post, and now available as
a full paper," with an acknowledgment thanking him "for his original work on recursive language
models." Zhang's post frames the mechanism plainly — an RLM is "a thin wrapper around a LM that can
spawn (recursive) LM calls for intermediate computation," with an API meant to be a drop-in
replacement for an ordinary completion call: `rlm.completion(messages)` where you'd otherwise write
`gpt5.completion(messages)`. The motivating problem is what he calls **context rot**: model recall
degrades as context grows, independent of whether the context still technically fits the window.
The idea was formalized two months later in [*Recursive Language
Models*](https://arxiv.org/abs/2512.24601) (arXiv 2512.24601, submitted 2025-12-31), authored by
Alex L. Zhang, Tim Kraska, and Omar Khattab, all MIT CSAIL. Their own framing in the abstract: RLMs are "a general
inference paradigm that treats long prompts as part of an external environment and allows the LLM
to programmatically examine, decompose, and recursively call itself over snippets of the prompt."
The paper is explicit about what it's reacting against, and credits the right ancestors rather than
claiming recursion or code-as-tool-use as new:
- **CodeAct** established writing code as the tool-call format, but in a standard coding agent that
code still executes inside the same context-constrained loop — sooner or later the harness has to
compact. RLMs offload the entire prompt as an external variable instead, so the REPL's addressable
state isn't bounded by the model's window at all.
- **MemGPT** manages context explicitly, paging things in and out of a single model's working
memory. RLMs don't build a memory hierarchy; they let the model itself decide what to look at,
programmatically, each time.
- **ReAct**-style sub-calls are verbalized autoregressively — described in natural-language turns
inside one transcript. RLM sub-calls are constructed programmatically and their results are stored
as REPL variables, which is what lets a `for` loop over sub-calls do real accumulated work instead
of restating each result back into the same window.
So the general idea — treat a long input as an external, programmatically addressable environment,
and let recursive delegation happen through control flow instead of prose — has a real, credited
origin, and it is not Prime Intellect. What Prime Intellect has done is build two different things
on top of it. Their research post says plainly: "we at Prime Intellect have implemented our version
of the RLM in [verifiers](https://github.com/PrimeIntellect-ai/verifiers/) so that it is ready to be
used in any environment," landing as the experimental `RLMEnv` — a reasonably faithful reproduction
of Zhang's design, built for running controlled evaluations. Separately, `prime-agent-runtime`'s
`rlm` package — the one this site's [Prime Agent](/articles/prime-agent) piece covers — takes the
same two mechanisms and applies them to an entire general-purpose coding agent: files, shell
commands, skills, and subagents all go through the same persistent kernel, not just one oversized
input prompt. Zhang's design restricts the root model to *metadata* about the prompt (its length, a
prefix) until it explicitly decides to look closer; Prime Agent's root model just has an ordinary
working context plus a kernel, because it's built to be a general agent, not a single-prompt
inference technique. Related, useful, and worth keeping straight — not the same artifact.
## The prompt becomes a variable
In the paper's own algorithm, the root model never receives the prompt as tokens in its context. It
receives metadata — length, a prefix, how to access it — and a REPL where that prompt already sits
as a variable. The loop is: the model writes code, the REPL executes it, truncated stdout comes
back, and this repeats until the model calls `FINAL(answer)` to return a string directly or
`FINAL_VAR(name)` to return whatever a REPL variable currently holds. Nothing about the prompt's
actual content is ever force-fed into the root model's window; the model decides what to look at, a
slice at a time.
Prime Agent's version of this same bet is less specialized but the mechanism is identical in spirit:
a persistent IPython kernel that survives across turns, with `rlm` preloaded in the namespace. Here
is the actual shim that puts it there, from `prime-agent-runtime/src/rlm/__init__.py`:
```python
class _RLMCallable:
async def run(self, prompt: str, **kwargs: Any) -> RLMSpawnHandle:
return await run(prompt, **kwargs)
async def __call__(self, prompt: str, **kwargs: Any) -> RLMSpawnHandle:
return await run(prompt, **kwargs)
rlm = _RLMCallable()
```
`rlm` is not a tool the model selects from a menu. It is a plain Python object with a `__call__`
method, sitting in the kernel's global namespace the same way any import would. Calling it is
calling a function, full stop — which is the whole point: nothing about `await rlm(...)` needs a
harness to specially recognize the string `"rlm"` and route it through a different code path than
any other line of Python.
Here's the same underlying claim made concrete with a task. Say the question is "which of these log
files mentions an out-of-memory kill, and which one is worst." A schema-based harness does this as a
sequence of round trips — each one a full model turn, a parsed tool call, and an appended result:
```text
# illustrative — the general shape of a schema-based harness, not quoted from a specific product
assistant: tool_call grep(pattern="Out of memory", path="logs/")
tool_result: {"matches": ["logs/worker-014.log:88231", "logs/worker-014.log:88245", ...340 more]}
assistant: tool_call read_file(path="logs/worker-014.log", offset=88200, limit=100)
tool_result: "<100 lines of log text>"
# …and one more round trip per file the model wants to actually look inside
```
The RLM version, written in the same idiom as the docs' own `config_files` example
(`packages/coding-agent/docs/rlm.md`):
```python
from pathlib import Path
hits = [p for p in Path("logs").rglob("*.log") if "Out of memory" in p.read_text(errors="ignore")]
worst = max(hits, key=lambda p: p.stat().st_size)
print(f"{len(hits)} files mention OOM; worst by size: {worst}")
```
`hits` is a real Python list, still there next turn if the model wants to `map` something else over
it. The grep, the read, and the size comparison happen inside one cell. The model's context grows by
one printed line, not by one message per file.
That's the mechanism at data-scale, not code-scale: the top bar is what has to happen when the only
way to look at something is to read it into the window — the window caps out at 272K tokens for
GPT-5 regardless of how the corpus is chunked, so a 500 MB corpus needs on the order of
hundreds of read-and-compact rounds just to scan once, and any single round can only ever see a
272K-token slice. The bottom bar is the REPL: the corpus is bounded by machine memory, not context
budget, and only a found, printed slice ever reaches the model. This is the same shape as the
paper's own S-NIAH and OOLONG scaling runs — hold the task fixed and grow the input, and one line
stays flat while the other falls off past the window boundary.
## Recursion is just a call
The second inversion is about delegation. In a schema-based harness, spawning a subagent is a
distinct, specially-recognized action — usually literally called `Task` or `subagent` in the tool
list, with its own parsing path in the harness. In an RLM, `rlm(...)` is not a different *kind* of
call from anything else in the REPL. It's an `async` function that happens to start another agent
instead of, say, reading a file. Here's the actual implementation, trimmed from the same file:
```python
async def run(prompt: str, **kwargs: Any) -> RLMSpawnHandle:
"""Spawn a recursive Prime Agent child and return once its task is admitted."""
if not isinstance(prompt, str):
raise TypeError(f"prompt must be str, got {type(prompt).__name__}")
payload = await host_request("rlm.run", {"prompt": prompt, "kwargs": kwargs})
return _spawn_handle_from_payload(payload)
```
`host_request` opens a Jupyter comm to the TypeScript host, which creates a real child
`AgentSession` and returns as soon as the task is *admitted* — not when it's *done*. That admission-
not-completion design is a deliberate choice recorded in Prime Agent's own changelog (covered in
more depth in [Prime Agent](/articles/prime-agent)), and it's what makes fan-out compose with
ordinary control flow instead of blocking on it:
```python
# real — packages/coding-agent/docs/rlm.md
api_review = await rlm("Review the public API", name="api-reviewer")
test_review = await rlm("Review the test coverage", name="test-reviewer")
integration_audit = await rlm("Run the slow integration audit", name="integration-audit")
```
Three lines, three independent children, one turn. In a schema-based harness the equivalent is three
separate structured messages, each requiring the model to emit a full tool call and the harness to
parse and dispatch it — and, if the harness's subagent tool is synchronous, a wait on each before the
next line can even be written. Here it's `for child in reviewers: await rlm(child)` if you want a
loop, or three independent statements if you don't. Recursion composes with the rest of the language
because it's written in the rest of the language.
The tree above is the concrete version of "a task over 200 files is a `for` loop with 200 model
calls made by the program, not 200 round trips through the model's own context." Each spawned
session's own context holds exactly one task — the child at `RLM_DEPTH=1` never sees its siblings,
never sees the root's other work. What holds the shape of the whole job is the root session's Python
namespace: the list of handles, the loop that produced them, the code that will eventually read their
replies. That's a different place for "the state of the whole task" to live than any single model's
context window, and it's why the unit of work stops being "how much fits in 200K tokens" and starts
being "how many child sessions can the host actually run."
## What's actually measured, and by whom
Two different evaluations exist, and they're worth keeping apart because they measure different
things with different rigor.
**The paper's own numbers** (arXiv 2512.24601) are the more citable evidence, because they're in a
reviewable artifact with a stated protocol: GPT-5 and Qwen3-Coder-480B-A35B-Instruct, compared
against RLM wrappers of themselves, across S-NIAH, OOLONG, OOLONG-Pairs, BrowseComp-Plus, and CodeQA.
The cleanest figures to quote are the ones in the abstract, because they are stated exactly rather
than read off a chart, and because they are relative to the right baselines. Wrapping GPT-5 in an
RLM beats — by a median across the evaluated benchmarks — **26% against compaction, 130% against
CodeAct with sub-calls, and 13% against Claude Code**, "while having comparable cost." That middle
number is the one that matters most for the argument here: CodeAct *also* gives the model code
execution and sub-calls. The gap between it and an RLM isn't code-versus-schema, it's whether the
input lives outside the model's context or inside it. The paper also claims inputs "up to two orders
of magnitude beyond model context windows."
And one contribution the write-ups mostly skip: they didn't only wrap existing models, they
**post-trained one for the paradigm**. RLM-Qwen3-8B beats plain Qwen3-8B by 28.3% on average and,
per the abstract, "approaches the quality of vanilla GPT-5 on three long-context tasks." An 8B model
approaching a frontier model on long-context work by being trained to drive a REPL rather than to
read further is the most interesting claim in the paper, and the one I'd most want replicated.
The headline figure scales input length from roughly 8K to over 1M tokens on S-NIAH, OOLONG, and
OOLONG-Pairs: GPT-5 degrades sharply as input grows, especially past its own 272K-token window where
it structurally cannot see the rest of the input at all, while RLM(GPT-5, depth=1) stays roughly flat
across the same range. In the paper's tables — read from its figures, so treat the exact decimal as
approximate rather than a number I recomputed myself — GPT-5 alone scores around 44% on OOLONG versus
roughly 56–58% for the RLM wrapper at recursion depths 1 and 3; on BrowseComp-Plus (1,000
documents, 6–11M tokens total) GPT-5 alone scores 0% because the input doesn't fit at all,
versus roughly 91–92% for RLM(GPT-5); on CodeQA, GPT-5 scores around 24% versus roughly
62–66% for the RLM wrapper. These are real, single-paper, not-yet-independently-replicated
numbers — but they come with a stated model, a stated task, and a stated context length, which is
more than most of what gets cited as evidence for an agent architecture.
**Prime Intellect's own post** runs a separate, smaller evaluation: GPT-5-mini through their
`RLMEnv` implementation, across four `verifiers` environments — DeepDive (web research), Math-python,
Oolong, and Verbatim-copy — at 50 rollouts each. This is where the honesty has to cut both ways. RLM
helps on DeepDive (with explicit strategy tips pushing sub-LLM calls further), helps on Oolong (the
plain LLM gets close to zero reward on the longest real-data contexts; RLM keeps working out to
roughly 1.5M characters), and helps on Verbatim-copy across most content types. On Math-python it
does not: the post reports RLM performing *worse* than the plain LLM, and ablating the REPL's timeout
up to 600 seconds doesn't close the gap. That's a genuine negative result from the people building
the thing, reported plainly rather than left out — a useful data point on where "run more code" isn't
automatically the right move for a task that's mostly reasoning, not search. Charts, not tables: the
post shows relative comparisons rather than a numeric results table, and says so itself — "this is
not a measurement of any model's absolute performance on any benchmark."
A day before this piece, Prime Intellect published a separate launch post for Prime Agent reporting
Opus 5 scoring 95.5% Best@1 on ARC-AGI-3 (183 test levels), against a self-reported 30.2%
baseline for Opus 5 on its own harness, and a table comparing Prime Agent across models against
Pi-mono, Claude Code, and Codex on OOLONG, OBLIQ-Bench, and LongBenchv2. That's Prime Intellect's own
first-party number, for the product, not in the open-source repository — consistent with what
[Prime Agent](/articles/prime-agent) already found reading the repo itself: no benchmark numbers
ship in the code, only in the marketing post. It belongs in a product-level discussion, not this one;
I mention it only so this piece doesn't read as unaware of it.
## What this costs
None of the above is free, and the paper and the Prime Agent docs are both honest about the price.
**Arbitrary code execution is the primary interface, not a fallback.** A fixed tool schema gives you
an enumerable, individually-auditable set of actions. A REPL's action space is "anything Python (and
a shell cell) can do." Prime Agent's own trust-model documentation says this plainly: the kernel "is
a durable control environment, not a security sandbox." Zhang's design narrows the blast radius
somewhat by keeping the root model restricted to prompt metadata until it asks for more — but the
sub-LM calls and any tool access still execute inside the same interpreter.
**A stateful interpreter can wedge.** Persistent state across turns is the entire point of the
design, but it means a hung cell doesn't just time out cleanly the way a stuck tool call does — Prime
Agent's own busy-kernel handling frames the choice as "wait" (preserve state, keep waiting
indefinitely) or "kill" (lose every in-memory variable and restart clean). There's no option that
gets you both a responsive kernel and the state back.
**Recursion needs an enforced budget, not a polite one.** The recursion tree above has a real cap
behind it: `RLM_MAX_DEPTH` defaults to 1, and the check —
```typescript
// packages/coding-agent/src/core/agent-session.ts
if (this._rlmDepth >= this._rlmMaxDepth) {
throw new Error(
`RLM recursion depth limit reached (RLM_DEPTH=${this._rlmDepth}, RLM_MAX_DEPTH=${this._rlmMaxDepth})`,
);
}
```
— runs in the host before a comm channel even opens, not as a prompt instruction the model could
argue its way past. That matters because the cost of recursion is exponential in depth if fan-out
is uncapped: `fanout^depth` sessions, each billed and each capable of spawning more, versus a `for`
loop's cost growing linearly in the number of iterations. A depth cap enforced in code is the
difference between "a program that does 200 things" and "a program that can, in principle, spawn
without bound."
**Harder to sandbox and audit than a fixed schema.** Every action a schema-based agent can take is
one of N defined functions — individually reviewable, individually denyable. "Any code the model
writes" is a much larger surface for a malicious skill, a compromised MCP integration, or a bad
instruction to abuse, and it's a correspondingly harder surface to review after the fact.
**Debugging shifts from reading a transcript to debugging a program.** A schema-based agent's failure
mode is usually legible from the transcript alone: read the tool calls and results in order. An RLM's
failure mode can be a bug in generated code, a REPL state that's subtly wrong three cells after the
mistake that caused it, or a child session that never replies because nothing in the parent's code
checks for it. That's a different, and for most engineers a more familiar, debugging discipline — but
it is a different one, and treating it like transcript-reading will miss real bugs.
## Where this sits
RLM is an inversion of the model [Lilian Weng's harness framing](/articles/agent-harness) and [the
harness effect](/articles/harness-effect) both describe: a fixed loop around a fixed tool schema,
where the transcript is the only place state can live. RLM doesn't optimize that loop's token
economics — it removes the transcript as the place state lives at all, replacing it with an
interpreter.
It's also a different axis from two other pieces published alongside it. [Recursive Harness
Self-Improvement](/articles/recursive-harness-self-improvement) treats the harness as a single text
prompt and improves it by comparing it against its own immediately-previous version — harness as a
string being optimized. RLM treats the harness as a program the model writes fresh each turn —
harness (or at least the working state) as code being executed. And the Continual Harness
[paper](https://arxiv.org/abs/2605.09998) — *Continual Harness: Online Adaptation for Self-Improving
Foundation Agents*, by Seth Karten, Joel Zhang, Tersoo Upaa Jr, Ruirong Feng, Wenzhe Li, Chengshuai
Shi, Chi Jin and Kiran Vodrahalli — is a third axis again: an online loop that alternates acting with
refining the agent's *own* prompts, skills, memory, and subagent specs *during* a run, without
resetting. (Its Zhang is Joel Zhang, not RLM's Alex L. Zhang — same surname, different author.)
Prime Agent's own Continual Harness feature, covered in [Prime Agent](/articles/prime-agent), draws
on this and composes with its RLM runtime. The connection is more than citational: the paper's lead
author, Seth Karten, is an active committer to the prime-agent repository, with 39 commits in its
history as of 2026-08-06. So the two ideas arrive in the same product from the same people — but they
answer different questions. RLM is about how one task executes; Continual Harness is about how the
harness's own configuration evolves across tasks.
[MemHarness](/articles/memharness) is a useful contrast in the opposite direction from RLM's whole
bet. MemHarness's argument is that a retrieved memory should be reconstructed — critiqued and
rewritten — against the current state every time it's used, because stale verbatim replay can hurt
more than no memory at all. RLM doesn't reconstruct anything: the corpus sits in a variable exactly
as it was written, and the model's job is to write code that finds the right slice of it, not to
have that slice handed to it pre-digested. Reconstruction spends compute making memory trustworthy
before use; RLM spends compute letting the model decide what's worth looking at, each time, from an
unmodified source. Different failure modes follow from each: a MemHarness-style system can
misreconstruct; an RLM-style system can simply fail to look at the part that mattered.
The idea itself is Zhang, Kraska, and Khattab's, credited honestly by the company building on it.
What Prime Intellect has actually built is two separate systems that take the same two mechanisms —
context as a variable, recursion as a call — and apply them at different scopes: one faithful to the
original single-prompt design, one generalized into an entire agent. Both are real, checkable
architecture decisions. The paper's numbers are the most rigorous evidence either has; Prime
Intellect's own eval is smaller, honestly mixed, and shown as charts rather than a table; and the
product-level ARC-AGI-3 number belongs to a different piece than this one.
---
# TencentDB Agent Memory: the four-tier pyramid is really a cache-stability hierarchy
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/tencentdb-agent-memory
> date: 2026-08-06
> tags: agents, memory, context-management, open-source, retrieval, explainer
[TencentDB Agent Memory](https://github.com/TencentCloud/TencentDB-Agent-Memory) is Tencent Cloud's
open-source memory layer for coding agents — MIT, TypeScript, about 148,000 lines across three
services, with adapters for OpenClaw, Hermes, Claude Code and CodeBuddy. The pitch is the one every
agent-memory product makes: stop re-explaining your project to every new session.
The pitch is not the interesting part. The interesting part is a decision buried in a source comment,
which reframes the whole design and is worth stealing whether or not you ever run this software.
## The pyramid everyone builds
Start with what the README shows you. Conversations are captured raw and refined by an async
pipeline into four levels:
So far this is the standard picture, and on its own it does not tell you much — "summarize old things
into shorter things" describes most memory systems ever built. The obvious reading is that the
pyramid is about *abstraction*: L1 is more general than L0, L3 more general than L2.
Read the code and a different axis appears.
## Sorted by volatility, not abstraction
`MemoryProxy/src/injection/injectors/tdai-profile-memory-injector.ts` opens with a comment
explaining how each tier reaches the model, and the three answers are all different:
- **L3 persona** — injected in full. "稳定且通常较短" — stable and usually short.
- **L2 scenarios** — **only the scene-navigation index** goes in: a list of paths plus a one-line
summary each. The stated reason is that L2 full text "经常上千 chars × N 个" — often thousands of
characters times N blocks. The agent reads the full scene through a tool once it has decided the
scene matters.
- **L0 and L1** — not injected at all on this path. The comment is blunt: "不再自动召回" — no longer
auto-recalled. They are exposed as read-only search tools.
The proxy README gives the reason in one clause, and it is the whole thesis of the system:
> injects Skills, Knowledge and Memory L2/L3 into the system prompt on demand; L0/L1 are exposed as
> read-only tools for the model to query proactively, **avoiding upstream KV-cache invalidation**.
That is a token-economics argument, not a knowledge-management one. Anything you put in the system
prompt that differs from last turn breaks the provider's prompt cache and makes you re-pay for the
entire prefix — so the question "which tier goes where" is really "which tier is stable enough to
sit in a cached prefix." Sorted that way, the pyramid falls out for a completely different reason
than the abstraction story suggests. L3 is at the top not because it is the most abstract but
because it is the most *stable*. L0 is at the bottom because it grows every single turn.
The discipline shows up again in the deployment guide, which is where you can tell someone has run
this in production:
> For multi-node deployments you must use `storage.backend=cos` and explicitly set
> `injection.externalGatewayUrl`, otherwise each instance caches independently and causes upstream
> KV-cache misses.
Treating a prompt-cache miss as a documented operational failure mode — in an *installation* doc — is
not something most agent-memory projects think to do. It rhymes with what [the harness
effect](/articles/harness-effect) argued from the other direction: the orchestration layer, not the
model, sets the bill.
The same repository makes the opposite choice on its other integration path. `MemoryCore`'s
`auto-recall` hook — used by the OpenClaw plugin — *does* automatically retrieve L1 and inject it
into context before the agent runs. So the product ships two integration surfaces with two different
injection policies, and only the proxy one is cache-preserving. Neither doc mentions the divergence.
If you are evaluating this, which path you take changes the token behaviour substantially.
## The retrieval underneath is real
It would be easy to ship "hybrid search" as a marketing phrase. This is not that. `auto-recall.ts`
runs FTS5 BM25 for the keyword side and cosine similarity over a vector store for the dense side,
then merges the two rank lists with reciprocal rank fusion, with the constant spelled out and
attributed:
```typescript
// RRF merge: k=60 is a standard constant from the RRF paper
const RRF_K = 60;
```
Scoring a record at rank *r* as $1/(k + r)$ and summing across lists is the standard formulation, and
k = 60 is the value from Cormack et al. If the backing store can do dense-plus-sparse-plus-RRF
server-side it short-circuits to one API call; on the SQLite path it runs both sides in parallel and
fuses client-side. There is a graceful degradation if FTS5 is unavailable — the keyword list comes
back empty and RRF operates on the dense side alone.
This is the same fusion idea this site's own search uses, and it is the right default: BM25 finds the
document that says the exact identifier you typed, embeddings find the one that means what you meant,
and RRF combines them without needing calibrated scores from either.
## What actually bounds an injection
The README says results are "further capped by item count, character budget, and timeout limits to
prevent memory from overwhelming the context window." Checking the shipped defaults in
`MemoryCore/src/config.ts`, that is two-thirds true.
`maxResults` defaults to 5, `scoreThreshold` to 0.3, `timeoutMs` to 5000 — all binding. But
`maxCharsPerMemory` and `maxTotalRecallChars` both default to **0**, and the budgeting function
short-circuits when they are:
```typescript
if (!maxCharsPerMemory && !maxTotalRecallChars) {
return lines;
}
```
So the character budget exists, is properly implemented with truncation markers and drop counts, and
ships turned off. Five results still bounds things, but five results of unbounded length is a
different guarantee than the sentence implies — and a single sprawling L1 memory is exactly the case
a character budget exists to catch. It is a one-line config fix, not a design flaw, but you have to
know to make it.
A smaller drift in the same file: `l1IdleTimeoutSeconds` is documented in its own doc comment as
"default: 30" and initialized to `600`. Twenty times the documented value.
## The guide it injects into your agent
One more thing the code shows that no doc mentions. Alongside the memories, MemoryCore injects a
usage guide telling the model how to retrieve more — and it is hardcoded **in Chinese**, in a
repository whose README, install guide and contributing guide are all bilingual:
```text
### ⚠️ 调用次数限制
每轮对话中,tdai_memory_search 和 tdai_conversation_search 合计最多调用 3 次。
```
*"Per conversation turn, `tdai_memory_search` and `tdai_conversation_search` may be called at most
3 times combined."* The guide goes on to instruct the model that if three searches turn up nothing,
the information is not in memory and it should answer from what it has rather than keep searching.
Two observations. The **3-call ceiling is a good idea** — an agent that can search its own memory
without limit will, and each miss costs a round trip. Naming the budget in the prompt and telling the
model what to do when it is exhausted is more thoughtful than most retrieval integrations manage.
And the **language is a real deployment consideration**: a fixed Chinese-language instruction block
enters the context of every agent this wraps, including English ones. Models handle it, but it
consumes tokens in a tokenizer that is not optimized for it and it sets the instruction language for
that portion of the prompt.
## Permissions, checked against the code
The visibility model is the part I expected to be thinnest and it is the most carefully built. The
README promises `private` means private "not even team admins," and `permission-checker.ts` backs it
with a dated comment explaining the choice:
```typescript
case "private":
// 私密语义(2026-07 变更):严格私密,只有 owner_user_id 能访问。
// 团队 admin 也不放行 —— 因为第 2 步 owner 判定已优先返回 ALLOW,
// 走到这里说明当前 user 不是 owner,即使是 admin 也一律拒绝。
return { allowed: false, reason: "visibility_restricted" };
```
A July 2026 semantics change, the reasoning preserved in the source, and the consequences enumerated
underneath it — including that admin `list-accessible` calls must not return other people's private
assets. That is a team that had the "should admins see everything?" argument and wrote down how it
ended.
One gap worth naming, because the README's framing does not survive it. `restricted` is described as
"precise access via User / Role / Agent ACLs," and for ordinary members that is exactly what the code
does — an explicit ACL match is the only way in. But the check is gated on `membership.role !==
"admin"`, so **team admins skip the ACL entirely** and fall through to role defaults. Defensible —
someone has to administer the thing — but "strict ACL whitelist" is true for members and not for
admins, and the docs do not say so.
## The architecture, briefly
Three services. **MemoryCore** owns storage and the L0→L3 pipeline. **MemoryKnowledge** builds the
Wiki and CodeGraph assets. **MemoryProxy** is the clever piece: a transparent LLM proxy that forwards
OpenAI `/v1/chat/completions` and Anthropic `/v1/messages` verbatim, doing session setup, injection
and write-back on the way past. Point your coding agent's base URL at it and you get team memory
"without changing a single line of code."
That is a genuine integration strategy rather than a shortcut. It also means the proxy sits in the
path of every request and every response, holding your model credentials, which is a trust decision
worth making deliberately rather than by following a quickstart.
The unification is the real product claim: Chat Memory, Skills, Wiki and CodeGraph are all registered
as **Memory Assets** with owner, version, status, visibility and agent bindings, retrieved through
one permission-scoped surface. The README's comparison table puts it well — RAG answers "what can be
found?", and this also answers "who can use it, which version is valid, and which agent should
receive it." Whether that ontology is worth its complexity depends entirely on whether you have a
team; for one person with one agent it is overhead.
## The number
There is exactly one benchmark in the repository, and it is in the README:
| Benchmark | Without | With | Relative |
|---|---|---|---|
| PersonaMem | 48% | 76% | +59% |
That is the entire evaluation. No harness, no model named, no agent configuration, no seed count, no
link to a run. I searched the repository for any other mention of PersonaMem and found two — the same
table in the Chinese README. So there is no reproduction script here, and the claim is first-party
and unreplicated.
To be fair on two counts: a memory layer improving a *memory* benchmark is not a surprising result,
and the repository's own Notes section is refreshingly frank about what is unfinished — CodeGraph
"currently prioritizes public HTTPS repositories," the Hub supports manual binding while "fully
automated memory routing is still under iteration," and Team Memory is labelled Beta. A project that
tells you which parts are not done yet has earned some patience about the parts it has not measured.
The provenance is also handled properly. The acknowledgements credit
[CodeGraph](https://github.com/colbymchenry/codegraph) for code the CodeGraph module "uses," Nous
Research's Hermes Agent for part of the Skill management code, and Karpathy's LLM-wiki gist for the
Wiki design — specific about what was borrowed rather than a generic thank-you list.
## The take
Most agent-memory projects are a retrieval index with an ontology bolted on, and the ontology is
where the marketing lives. This one has a real idea underneath it, and the idea is not the pyramid.
It is that **memory has to be sorted by how often it changes, because the cost of memory is not
storage, it is the prompt prefix you invalidate by updating it.** Once you see the four tiers as a
cache-stability ordering rather than an abstraction ordering, the delivery mechanism for each one
stops being arbitrary: stable things get injected, semi-stable things get injected as an index,
volatile things become tools with a call budget.
That principle is portable to any agent you are building, with or without this software. What comes
with the software is a competent hybrid retriever, a genuinely careful permission model, a
transparent proxy that is a real integration story and a real trust decision, one unreplicated
benchmark number, two integration paths that disagree about injection policy, and a character budget
you should turn on before you rely on it.
---
*Sources: the [TencentDB-Agent-Memory repository](https://github.com/TencentCloud/TencentDB-Agent-Memory)
at its 2026-08-06 state — `README.md`, `INSTALL.md`, `MemoryProxy/README.md`, and the TypeScript in
`MemoryCore/src/config.ts`, `MemoryCore/src/core/hooks/auto-recall.ts`,
`MemoryCore/src/metadata/service/permission-checker.ts` and
`MemoryProxy/src/injection/injectors/`. Both figures are the project's own, flattened onto white;
the pyramid is its English-language variant. Chinese source comments are quoted verbatim with my
translations. The PersonaMem figure is the project's own and is not independently replicated. Both
interactives are mine.*
---
# ABot-World-0: a 5B world model that wins on efficiency, not the leaderboard
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/abot-world
> date: 2026-08-03
> tags: world-models, video-generation, diffusion, open-weights
Most interactive-world demos are a video model you watch. ABot-World-0, from Alibaba's AMAP CV Lab, is one you steer: upload a starting image, hold WASD, and a 5B-parameter model streams a continuously-controllable 720p world back at up to 16 frames per second, from a single RTX 5090, with 1.2 seconds between an action and its first frame on screen. Code, weights (Apache-2.0), and a 500-hour action-annotated dataset are all released. It does not top the benchmark it's measured on. That's the more useful fact to lead with, not bury.
## Bidirectional teacher, causal student, then fix the drift
The model is a distilled video-diffusion world model, trained in three stages on top of the `Wan2.2-TI2V-5B` backbone. First, a high-quality **bidirectional** action-conditioned teacher learns good dynamics over full temporal context -- accurate, but it has to see the whole clip at once, so it can't stream. Second, teacher forcing plus causal ODE distillation compress that into a **causal**, few-step student that approximates the teacher's denoising trajectory turn by turn, which is what makes low-latency interaction possible at all. Third, and this is the paper's actual technical contribution: **LongForcing**.
Plain causal distillation only ever supervises short, clean trajectories -- it never sees what its own errors compound into over a minute of rollout. LongForcing closes that gap directly: let the causal student generate a long self-rollout, then correct the *distribution* of that self-rollout against an extended-horizon teacher, rather than only imitating clean short clips. It is, in effect, training the model against its own failure mode instead of only against ground truth. Measured over a 60-second rollout against a plain causal-forcing baseline, the effect shows up as less accumulated visual damage, not a benchmark score:
The baseline doesn't fail suddenly -- it drifts. Causal Forcing looks comparable to LongForcing for the first 15-20 seconds on all four curves, then peels away: color saturation creeps up, the image blurs, patches start repeating. That's exactly the accumulated-error problem autoregressive video generation is known for, and LongForcing's fix is to train against long rollouts directly rather than assume short-horizon quality generalizes.
## Real-time is five separate wins, not one
"Few-step generation does not automatically translate into real-time interaction" is the paper's own line, and Table 2 backs it up in a way that's genuinely counterintuitive: adding a faster attention kernel by itself does nothing, because the model doesn't fit in memory to begin with.
Every one of those five changes is load-bearing. Skip the VAE swap and the faster attention kernel just gets you a faster out-of-memory error. That's a more honest way to read "single desktop GPU" than treating it as one clever optimization -- it's a full-stack co-design where the first fix is the one that makes the rest of the stack possible to even measure.
## WorldRoamBench: a real third-party number, and it doesn't sweep
WorldRoamBench is not Alibaba's benchmark -- that independence is worth stating plainly, because it means ABot-World-0's score wasn't set by the people reporting it. Against Genie 3, HappyOyster, LingBot-World (14B), and HY-World 1.5 (8.3B):
ABot-World-0 sits second, not first, on Strict Accuracy -- and that pattern holds across the rest of the benchmark's sub-metrics too:
Two honest qualifications on top of what the chart above already shows. First, neither Genie 3 nor HappyOyster has a disclosed parameter count, so the efficiency claim is only verifiable against the two comparators whose sizes are public -- LingBot-World and HY-World 1.5 -- not against the benchmark's actual leader. Second, ABot-World-0 running on a single consumer GPU is not a property this benchmark measures at all; WorldRoamBench scores output quality and controllability, not deployment cost. The efficiency story and the benchmark score are two separate claims, and only one of them is what WorldRoamBench actually tested.
Reproducibility here is unusually complete for this space: code, weights, and a 500-hour action-annotated dataset (`ABot-World-Explorer-500h`) are all released under Apache-2.0. The release README documents a staged rollout from 2026-07-09 through 2026-08-03 -- today, by this piece's own dateline. That's a meaningfully higher bar than a paper with numbers and no artifacts.
## What's missing
The paper's qualitative claims -- physically plausible responses despite no explicit physics training, coherent hour- and day-scale rollouts, generalization to out-of-domain controls -- are demonstrated with cherry-picked keyframe strips, not a systematic user study. That's standard for this genre of paper, not a special flaw of this one, but it means "plausible physical responses" is an illustration, not a measured claim the way WorldRoamBench's numbers are. The data-collection system behind all of this, WorldExplorer, is also worth a sentence on its own: it's closed-loop and distribution-aware, meaning it uses the current model's own failure modes to decide where to collect more data next, rather than collecting blind -- a genuinely different approach from scraping video and hoping coverage works out, though the paper's evidence for how well that targeting works is qualitative too.
## The take
ABot-World-0 is not the best model on WorldRoamBench. HappyOyster beats it on six of seven reported sub-metrics, and the benchmark's own leader has no disclosed size to compare against. What ABot-World-0 actually demonstrates is that a 5B model, with the right three-stage distillation and a genuinely load-bearing systems stack, beats two larger open rivals on every metric measured while being the only one of the group that runs interactively on one desktop GPU. That's a real, checkable claim, and it's a more interesting one than a clean sweep would have been -- a paper that only won everywhere would have less to say about where the wins actually come from.
---
*Built on Alibaba AMAP CV Lab's [ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU](https://arxiv.org/abs/2607.19191) (Jiang et al., 2026) and the [amap-cvlab/ABot-World](https://github.com/amap-cvlab/ABot-World) repository (Apache-2.0). Figures 1, 3, and 10 are reproduced from the paper for commentary, flattened onto white; the systems-ablation and WorldRoamBench explorers are original visualizations of the paper's Table 2 and Table 3 data, not measured traces. Benchmark numbers are as reported in the paper and on WorldRoamBench.*
---
# AngelSpec: specialize the drafter, share the verification budget
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/angelspec
> date: 2026-08-03
> tags: inference, speculative-decoding, llm, systems, explainer
Speculative decoding's basic trade is well covered on this site already: a cheap draft model
proposes tokens, the target model verifies them all in one pass, and the target only ever emits
what it would have sampled anyway. [EAGLE-3](/articles/eagle-3-speculative-decoding) fixed the
draft model's own scaling law; [DSpark](/articles/deepseek-dspark) paired a semi-autoregressive
drafter with a load-aware verifier. [AngelSpec](https://arxiv.org/abs/2607.25852) (Liu, Cen, Shi,
et al., Tencent) is a different kind of paper: it is not one new trick, it is what a production
team building on top of that whole lineage — [multi-token prediction](/articles/multi-token-prediction),
DFlash, DFlare, DSpark, EAGLE-3's Training-Time Test — actually ships for
[Hunyuan Hy3](/articles/hunyuan-hy3). Two ideas carry the paper. First: don't train one universal
drafter, train a different architecture for each workload's entropy. Second: don't give every
request a fixed verification budget, treat verification depth as a resource the whole batch shares.
I am reading this as **engineering**, not a new algorithm. DFly extends DFlash's target-conditioning
and DFlare's layer fusion; the Training-Time Test principle is EAGLE-3's. AngelSpec's own
contribution is combining them into one training framework, specializing them per workload, and
running the combination on real serving traffic. That last part is the rare thing here — most
speculative-decoding papers stop at static benchmarks.
## One drafter doesn't fit both workloads
Chat is high-entropy and open-ended; the next few tokens are genuinely hard to guess. Code and math
are the opposite — once you're two lines into a `for` loop or three steps into an algebraic
simplification, a lot of what comes next is close to deterministic. A single drafter trained on a
uniform mixture of both has to compromise. AngelSpec's answer is to stop compromising: train an
autoregressive multi-token-prediction (MTP) drafter on conversation-heavy data for chat, and a
block-parallel diffusion drafter — **DFly** — strengthened with code and math data, for everything
else.
## MTP: fix position 2 and 3, not position 1
The MTP drafter reuses one physical Transformer block recurrently at increasing logical depth,
each depth predicting one more token ahead. The problem EAGLE-3 already diagnosed for feature
prediction shows up again here for direct token prediction: depth 1 is trained against a clean
prefix, but at inference depth 2 has to condition on depth 1's own (possibly wrong) output — a
distribution it never saw in training. AngelSpec's fix is the same principle EAGLE-3 calls
Training-Time Test: unroll the drafter over its own predictions during training, so what it
practices on matches what it sees at inference.
The loss stack that gets it there is a genuine progression, not a single choice: hard-label
cross-entropy, then forward-KL, then an adaptively-blended KL/total-variation objective ("LK
loss"), finally switching to an end-to-end objective that directly optimizes expected accepted
length —
$$
\mathcal{L}_{e2e} = 1 - \frac{1}{|I| D} \sum_{i \in I} \sum_{m=0}^{D-1} \prod_{k=0}^{m} \alpha_{i,k}
$$
— a product, not a sum, because one rejection anywhere in the prefix invalidates every position
after it. Training directly on raw total variation from a cold start is worse than plain KL (the TV
gradient is too weak far from alignment); the paper's own ablation shows the cold-start-then-switch
recipe is necessary, not decorative.
The payoff shows up exactly where the theory predicts: the first drafted token barely moves,
positions two and three — the ones with no defense against drift before TTT — improve the most.
## DFly: a hybrid backbone, then a cheap causal patch
DFly starts from DFlash's move (feed every draft layer the same shared cross-layer-projected
target context) and adds DFlare's move (a layer-specific weighted fusion of target features),
combined rather than chosen between:
$$
g^{(i)}_t = \text{RMSNorm}\big(c_t + f^{(i)}_t\big)
$$
$c_t$ is DFlash's shared basis, $f^{(i)}_t$ is DFlare's depth-dependent refinement — the hybrid adds
only $D \times T$ scalar fusion weights over DFlash alone, precomputable once training finishes.
Block-parallel diffusion drafts $B$ tokens in one shot, which is where the latency amortization
comes from — but a one-shot draft has no mechanism to make token $t{+}2$ aware that token $t{+}1$
was just chosen. DFly's **hidden-correction head** patches that in afterward, cheaply: a small
SwiGLU pass folds the previous position's embedding into each hidden state before the LM head runs,
turning independent marginals into a causal chain
$$
q\big(X_{t+1:t+B} \mid x_{\le t}\big) = \prod_{i} q_i\big(x_{t+i} \mid x_{\le t}, x_{t+1:t+i-1}\big)
$$
while the expensive backbone stays fully parallel — only this small head runs sequentially. Tested
against a Markov-style low-rank correction (DSpark's approach), hidden-correction wins on both
accepted length and, notably, on the break-even latency it needs to beat MTP — the more accurate
head is also the cheaper one to run.
## The numbers
On Hy3-A21B, cumulative ablation (backbone → AR head → domain data) takes mean accepted length from
3.77 to 4.75; against the other drafters on the same target:
That's +59.7% over MTP and +29.8% over DFlash on this target — the paper is upfront that DSpark
isn't in this row (it's only benchmarked against Hy3 as MTP/DFlash/DFly; DSpark's own comparison
runs on Qwen3-8B, where it still wins MT-Bench, consistent with AngelSpec's own framing that DFly
targets code and math, not chat). Production throughput on Hy3-295B-A21B, 8×TP, tells the
concurrency story:
DFly wins the average speedup at every tested concurrency, 4 through 64. The more interesting
detail is what happens at the high end: at concurrency 64, DFlash's own speedup actually **drops**
below MTP-3's (1.89× vs 2.08×), while DFly stays ahead at 2.11×. DFly isn't just faster — it's the
one that degrades least gracefully into the regime where the GPU is already saturated with
verification work.
## D-cut: verification depth is a shared resource, not a per-request setting
Here's the fact that makes D-cut make sense: median target-model verification (`execute_model`)
latency runs 19.77–64.16ms; drafting and sampling (`sample_tokens`) runs 0.89–4.49ms. Verification
dominates decode-step cost by roughly an order of magnitude. So the thing worth optimizing at serving
time isn't the drafter — it's how much of that expensive verification you spend, and where.
The mechanism is a genuine reallocation, not a threshold. Per request $i$, expected progress from
keeping $n_i$ drafted positions is estimated from the drafter's own prefix-confidence product,
$\hat A_i(n_i) = \sum_{k=0}^{n_i} s_{i,k}$. D-cut doesn't pick $n_i$ per request — it flattens every
position across the **whole batch**, ranks by that same confidence score, and takes a global top-K:
$$
K_\rho(B) = \max\big(B,\ \lceil \rho\, B (D{+}1) \rceil\big)
$$
restricted to four ratios, $\rho \in \{0.25, 0.5, 0.75, 1.0\}$, chosen each step by a pre-profiled
runtime latency table that picks whichever $\rho$ maximizes projected throughput, not just kept
length. It only ever discards drafts — verification stays exact, so the target distribution is
untouched.
## Live traffic: the validation that actually matters
Static benchmarks are where DFly's story ends for most papers in this space. AngelSpec adds one
more figure, replaying real Hunyuan production traffic on 8×H20 at concurrency 2 through 64 — and
this is the evidence that made me want to write the piece up.
DFly alone saturates past concurrency 48 (~848–860 tok/s, flat). D-cut keeps rising: +3.0% at c48,
+9.2% at c56, **+15.7% at c64**. At matched per-user decode speed (~15.3 tok/s), D-cut sustains 981
tok/s at c64 versus DFly's 858 tok/s at c56 — 14% more aggregate throughput at the same latency.
Against plain autoregressive decoding, DFly's own speedup peaks at 1.33× (c24–c40) and then falls
back to 1.25× at c64; D-cut keeps climbing to 1.45× at c56 and 1.44× at c64. All of that for a
pruning cost of just **1.5%** average reduction in accepted length (2.50 → 2.46), rising to only
**2.8%** even at the most contended concurrency tested (2.50 → 2.43).
Read the caption on the paper's own Figure 4 carefully — it says the comparison **understates
D-cut**: "DFly uses full-and-piecewise CUDA graph capture and D-cut piecewise capture only." D-cut is
running with a documented implementation disadvantage relative to DFly and still wins. That is an
unusually candid thing for a paper to put in its own headline figure's caption.
## What's honest here, and what isn't new
Two disclosures matter more than most papers in this space bother to make. All production numbers
— throughput Tables 7–8, the live-traffic Figure 4 — run on **NVIDIA H20**, the export-compliant
part, not a flagship H100 or B200. The paper doesn't claim these numbers generalize to other
accelerators; it just tells you what it actually ran on. And the "this comparison understates
D-cut" line above is the second: a paper flagging that its own reported advantage is a conservative
lower bound is rarer than papers that quietly let an asymmetric comparison flatter them.
What isn't new: DFly is DFlash plus DFlare plus a hidden-correction head borrowed from TreeFlash's
idea; MTP's training recipe leans on EAGLE-3's Training-Time Test and an external LK-loss paper;
D-cut's "verification as a shared resource" framing groups itself explicitly with DSpark as the
other method doing this, rather than claiming to invent the idea. None of that is a knock — production
systems earn their keep by combining existing pieces well, not by mandating a new algorithm — but
it means the honest read of AngelSpec is "well-executed systems integration with real production
validation," not "a new speculative-decoding algorithm."
A few more gaps worth knowing before you cite this: DSpark is compared on Qwen3-8B, not on the
paper's own Hy3-A21B target — so the strongest same-class competitor is missing from the main Hy3
table. The released DFly is mode-specific (Table 6: no-think and high-think drafters don't transfer
across each other), which roughly doubles the drafters you maintain for a model family serving both
modes. And every reported baseline number — including DFlash and DSpark's — was retrained and
measured by the authors inside their own stack; there's no independent third party re-running any of
it.
## The take
The interesting move in AngelSpec isn't a new speculative-decoding trick — it's refusing to ship one
universal answer. Chat and code/math have different entropy profiles, so they get different
drafters. Verification cost and drafter confidence vary across requests and load, so verification
depth becomes a batch-level resource instead of a fixed setting. Neither idea is exotic on its own;
what makes the paper worth reading is that both survive contact with real Hunyuan traffic on
honestly-disclosed hardware, with the one place it could have inflated its own result — the CUDA-graph
asymmetry — disclosed instead of hidden.
---
*Source: [AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding](https://arxiv.org/abs/2607.25852)
(Hong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao,
Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, Jianchen Zhu — Tencent), arXiv:2607.25852. Figures 2,
3, and 4 are reproduced from the paper for commentary; the interactives are mine, built on the
paper's own reported numbers and formulas.*
---
# AutoCompact: teaching an agent to decide when to forget
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/autocompact
> date: 2026-08-03
> tags: agents, llm, context-management, explainer
Get the caveat out of the way first, because it changes how to read everything below: [AutoCompact](https://autocompact.github.io/) has **no paper, no arXiv listing, and no released code**. I checked the arXiv API by title and by author — zero hits. The project page is the only artifact, and the two charts that carry its headline claims (Figures 4–6) are unlabeled-axis line plots of "pass rate vs. inference-cost budget," not a results table. The one hard number in the whole post is "+10.6% RL gain on average on SWE-bench Verified," stated in prose, not shown as a table row you could check against a baseline.
So this is a research blog post, and I'm treating it as one. I still think the idea is worth explaining, because the mechanism — training a model to make a decision it doesn't naturally make, using an LLM judge to correct its trajectory rather than hand-labeling it — is a genuinely interesting instance of a pattern that shows up across post-training right now. Just don't read the rest of this as a verified result.
## The problem: compaction is usually a clock, not a decision
Long-horizon coding agents run out of context. The standard fix — what the post says OpenAI uses for ChatGPT and Codex — is a **fixed-threshold** compaction: once the trajectory crosses some token count, everything before it gets replaced by a summary. It doesn't matter whether the agent is mid-hypothesis or about to write the fix; the clock doesn't know the difference.
AutoCompact's bet is that *when* to compact is a decision the agent should make about its own task state, not a number a harness enforces on it. The model gets a `compact()` tool call it can invoke at any point. When it does, the trajectory so far gets replaced by a generated summary — objective, the localized issue, files touched, what's verified, what's next — while the original task description and the most recent turns stay verbatim. Everything else (dead-end hypotheses, reverted edits, redundant file reads) gets dropped.
## Training it: a judge corrects the trajectory, not the label
The harder problem is that models don't call `compact()` well without training — they don't reliably notice when a phase has ended. AutoCompact's answer is judge-guided correction rather than hand-authored demonstrations. At each step of a rollout, a judge (GPT-5.5-Codex) sees only the history visible so far, plus the fact that `compact()` exists, and makes one of three calls: leave the model's proposed action alone, **replace** it with a `compact()` call if this is a good moment to summarize, or **repair** a summary/continuation that's missing state or drifting off-track. That turns "teach the model to self-manage its own context" into "step-level correction under an annotation protocol" — something you can run at scale with a well-prompted LLM instead of a small army of human raters.
From 379 SWE-rebench tasks, after filtering malformed and off-track examples, this produces **1,052 SFT examples** — a genuinely small cold-start set, split roughly 24% teaching *when* to trigger, 53% teaching *what to preserve*, 23% teaching *how to continue* after compaction. From that SFT checkpoint, GRPO reinforcement learning on SWE-Gym with a binary pass/fail reward pushes further: active compaction rate rises from **44.3%** of tasks (SFT) to **58.5%** (SFT+RL) — the RL stage doesn't just improve quality, it makes the model reach for `compact()` more often, because doing so is apparently what correlates with solving the task. Two more self-reported quality numbers from the post: generated summaries retain relevant state **99.8%** of the time and specify a concrete next action **97.8%** of the time.
None of the pass-rate comparisons on the project page come with an exact number. All three head-to-head evaluations — no forced compaction vs. adaptive compaction, SFT vs. SFT+RL, and a forced 16k-token regime testing whether adaptive timing still beats a fixed threshold at the *same* limit — are qualitative line charts ("AutoCompact solves more tasks at every budget shown"), not tables. The "+10.6%" figure is the only exact number stated for RL over SFT, and even that is a TL;DR-level claim without a breakdown by task or budget.
This is also a context-management idea, which puts it in conversation with [how a harness manages context more generally](/articles/agent-harness) — Lilian Weng's point that durable state belongs on disk, not in an ever-growing prompt. AutoCompact is one layer higher: it's not asking *where* state should live, it's asking *when* the model itself should decide to shed it. Both are betting that context is the scarce resource and the harness (or the model, here) needs an explicit policy for spending it.
## The take
The training method — judge-corrected trajectories teaching a behavior the base model doesn't do on its own — is a pattern worth knowing regardless of whether AutoCompact's specific numbers hold up: it's a cheap way to get supervision for a decision (when to compact, when to stop, when to ask) that's hard to hand-label at scale but easy for a strong model to critique step by step. What I can't tell you is how good the resulting agent actually is, because there's nothing outside one team's own charts to check it against. If code or a paper ships later, the interesting question is whether the qualitative "wins at every budget" story survives being reduced to a table.
---
*Source: the [AutoCompact project page](https://autocompact.github.io/) (Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong; July 30, 2026). No code or paper release exists at time of writing; all figures on the source page are JS-rendered widgets, not downloadable images — the timeline above is my own illustration of the mechanism, using an invented trajectory and thresholds, not a reproduction of anything on the page.*
---
# A.X-K2: a sparse-attention upgrade that costs nothing, trained natively in FP8
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/ax-k2
> date: 2026-08-03
> tags: llm, mixture-of-experts, attention, sparse-attention, quantization, fp8, long-context
SK Telecom's **A.X-K2** is a 688B-parameter, 33B-active Mixture-of-Experts model, trained from scratch as
part of Korea's Sovereign AI foundation-model project and shipped under Apache 2.0. It is the successor to
A.X-K1, and the tech report ([SKT, 2026](https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf))
is unusually specific about what changed: a full architecture table, a full FP8 training recipe, full RL
hyperparameters, and — the detail worth building an article around — an ablation that reports the *quality
cost* of its own efficiency trick, and the cost is close to zero.
Two things carry this piece. First, **Sparse Gated Attention (SGA)**: a lightweight top-k indexer bolted
onto gated Multi-head Latent Attention, where SKT states the LongBench score before and after adding
sparsity — 62.80 to 62.99 — rather than only the speedup. Second, A.X-K2 is **trained natively in FP8**,
forward and backward, from the first optimizer step. Not quantized after the fact. Those two facts turn
out to be connected: the same architectural choices that make SGA cheap (GatedNorm suppressing outlier
activations) are the choices that make FP8 training survivable at this scale.
## What the weights say
Total tokens: about 8.5T (8.2T pre-training, the remainder post-training) — fewer than A.X-K1's roughly
10T, and SKT is direct about why that's the headline, not the parameter count: *"despite training on
fewer tokens than A.X K1 (∼10T), A.X K2 shows substantial improvements across the board — over 30
percentage points on some benchmarks — reflecting substantial gains in token efficiency."* The
architecture, read straight off `config.json` and Table 1 of the report:
| | |
|---|---|
| Total / active parameters | **688B / 33B** |
| Layers | **61** (1 dense, 60 MoE) |
| Hidden size | 7,168 · 64 attention heads (Q = KV) |
| Routed / shared experts | **256 / 1**, 8 active + 1 shared per token |
| Expert routing | sigmoid score, `noaux_tc`, 8 groups, group top-k 4, routing scale 2.5 |
| Attention | MLA + head-specific output gate + QK-norm, plus a top-*k* sparse indexer (k = 2,048) |
| KV-lora / Q-lora rank | 512 / 1,536 |
| Context | 128K native (ABF), 256K via YaRN (factor 2.0), zero-shot to 512K (factor 4.0) |
| Vocabulary | 163,840 (unchanged from A.X-K1), 5 languages |
| Training precision | native **FP8** (MXFP8, E4M3, block 32) forward + backward |
| Checkpoint | block-scaled FP8 (E4M3, 128×128), ~646 GB on disk |
Relative to A.X-K1 (519B-A33B), the entire parameter growth is in expert count: 192 to 256, chosen because
it's a power of two and a multiple of 128 for sharding, targeting a compute-derived total-parameter budget
under a fixed 70-day, 512-GPU schedule. Active parameters didn't move. That's a scale-up in *capacity*, not
in per-token compute — worth holding onto before the next section.
## Sparse Gated Attention
Start from what A.X-K1 already had: Multi-head Latent Attention (MLA) with a **head-specific output
gate** — a learned, input-dependent gate applied to the attention output before the `Wo` projection,
present in every layer throughout pretraining, not something bolted on for long context. The report's
framing for why the gate matters: it "introduces non-linearity into the attention output, mitigates
attention sinks, and improves loss convergence." Attention sinks are the few uninformative token positions
(often the first token) that vanilla softmax attention dumps disproportionate probability mass onto — a
gate that suppresses that mass gives every downstream consumer of the attention output a cleaner signal.
**SGA** is the second half of the name: a lightweight indexer, adopted from DeepSeek-AI's sparse-attention
design, that scores every cached key and keeps only the **top 2,048 tokens per query** — selected at
individual-token granularity, not in fixed blocks. (That's a genuine design fork from
[MiniMax Sparse Attention](/articles/minimax-sparse-attention), which scores and selects in 128-token
blocks specifically so memory access stays contiguous. A.X-K2 trades that contiguity for finer-grained
selection.) MLA then runs exactly over the selected set, `KV[I_topk]`, instead of the full cache. Here is
the paper's own architecture figure for the resulting block:
Read it as a data path: GatedNorm output splits into the Indexer→Selector pair (which decides *what* MLA
gets to read) and the Output Gate (which decides how much of what MLA computes gets through). The
mechanism I built to walk through the dynamics — how the fixed 2,048-token budget shrinks as a fraction of
a growing context, and what changes when the Selector is switched off entirely:
The report calls the gate and the indexer **mutually reinforcing**, and the causal story runs one
direction: because the output gate already suppresses attention-sink mass throughout pretraining, the
attention distribution the indexer is trained to imitate is better-calibrated before the indexer ever sees
it — so its top-*k* budget goes to genuinely relevant positions instead of partly being spent re-discovering
which tokens are sinks. The indexer itself is trained with a KL-divergence loss against the (already
gated) attention distribution, introduced in a dedicated Stage 3C after the model is natively trained to
128K context.
The number that makes this worth an article: on LongBench, A.X-K2 scores **62.80 before** the sparse
adaptation and **62.99 after** — sparsity made it very slightly *better*, not worse. A lab publishing the
comparison that shows its own efficiency trick is nearly free — rather than only the speedup — is worth
crediting on its own. And the adaptation recipe has a second, smaller honest claim attached: unlike
DeepSeek-V3.2 and GLM-5, which warm the indexer up against the *dense* (full) attention distribution
before switching on sparse selection, SKT trains the indexer against the **sparse** top-k selection from
the outset — a "sparse warmup" — because it's cheaper (sparse attention is less compute per step than
dense) and, in their experiments, cost no measurable downstream quality relative to the dense-warmup
alternative used elsewhere.
The efficiency payoff shows up exactly where you'd expect — long-context serving. Reading the report's own
inference sweep at 120K input tokens (concurrency 32, dp8/ep8): total-token throughput goes from roughly
9,100 tok/s for A.X-K1 to roughly 12,200 for A.X-K2 with a bf16 KV cache, and roughly 14,600 with an FP8 KV
cache — with per-token latency and time-to-first-token moving the same direction, each a few tens of
percent lower for K2 than K1 at that length. (Approximate, read off the report's chart, not a published
table — but the direction and rough magnitude are unambiguous.)
One more piece worth naming here because it recurs in the next section: **GatedNorm** replaces A.X-K1's
dual-normalization scheme entirely — a single input-dependent gate applied right after RMSNorm, instead of
stacking normalization layers around attention and the MLP the way A.X-K1 (and Gemma-style designs) did.
SKT ran the ablation at 20B-A3B scale and found GatedNorm alone matches the loss curve of the full
dual-norm design; stacking a second post-MLP norm on top added nothing. The reason GatedNorm matters
beyond training stability: it suppresses **massive activations** — the small number of hidden units that
run orders of magnitude larger than the rest and persist across layers — which is exactly the failure mode
that wrecks narrow low-precision formats. Which is where the second half of this piece starts.
## The scale, next to what else is disclosed
A.X-K2 sits in the middle of this range by total parameters and at the small end by active parameters.
[Kimi K3](/articles/kimi-k3) — 2.8T total, 104B active — took the opposite bet on the same axis: where A.X-K2
grew total capacity 519B → 688B while holding active compute flat at 33B (a pure expert-count expansion,
192 → 256), K3 tripled its active parameters alongside its total, spending its extra headroom on a bigger
per-token forward pass rather than more parked capacity. Both are legitimate ways to spend a training
budget; they're just different bets about where the marginal FLOP is worth spending.
## Trained natively in FP8
Everything above assumes a working low-precision model. A.X-K2 gets there by training natively in FP8 from
the start rather than quantizing a full-precision model afterward — MXFP8, E4M3, block size 32, forward
*and* backward pass, with FP32 master weights and BF16 optimizer state as the only higher-precision parts
of the recipe. I made the same argument at a different precision two weeks ago in
[Neutrino-1](/articles/neutrino-1): quantization is a decision you make before training starts, not a knob
you turn on a finished checkpoint. Neutrino-1 showed the cliff that decision avoids — ternary weights
rounded post-hoc land at 24.2–24.7 on 5-shot MMLU, against a 25.0 chance line, while the same ternary
format trained in from scratch reaches 72.1. A.X-K2 is the same principle, replayed at FP8 instead of
ternary, at 688B instead of 8B.
The practical consequence: A.X-K2 has no BF16 form to compare itself against, because none was ever
trained. Its FP8 checkpoint *is* the master weights, not a rounded-down copy of something else. Serving it
in NVFP4 — a further post-hoc step, applied only to expert weights (W4A4) — is a much smaller step down
than Neutrino-1's ternary rounding, because the base it's stepping down from was already trained
natively in a narrow format:
The report's own robustness table backs this up on eleven benchmarks: NVFP4 tracks FP8 within about a
point on most of them — CLIcK 84.21 → 84.06, MMLU 82.27 → 82.00, KoBEST-BoolQ 96.72 → 96.01 — with GSM8K
(−2.50) and MATH (−2.42) as the honest outliers, and HumanEval and KoBEST-COPA actually improving slightly.
Compare that spread to Neutrino-1's cliff and the shape of the difference is the whole argument: rounding
*into* a format a model never trained in collapses to chance; stepping *further down* from a format it was
already native in costs a couple of points at most.
This same commitment shows up again, more sharply, inside RL post-training — and it's the cleanest evidence
in the whole report that "native FP8" is an infrastructure discipline, not just a training-time flag. RL
needs the trainer (Transformer Engine, on Blackwell) and the rollout engine (vLLM) to agree numerically. But
Blackwell defaults to MXFP8 while vLLM's mature MoE FP8 path targets the older *blockwise* FP8 recipe built
for Hopper — so if you leave each side on its native default, they diverge. Figure 7 above shows what that
divergence does: a trainer running MXFP8 against a blockwise-FP8 rollout looks fine early, then the reward
curve stalls and drifts down. SKT's own diagnosis is the sentence worth keeping: *"Applying TIS does not
prevent this collapse, indicating that token-level intervention alone cannot remove the underlying
trainer–rollout precision mismatch."* Truncated Importance Sampling is a standard token-level correction for
exactly this kind of train/inference distribution drift, and it doesn't work here — the fix has to be
architectural (a patched Transformer Engine branch that forces blockwise FP8 on Blackwell, matching vLLM's
format end to end), not a loss-side patch. That's a small, honest, specific admission: a common trick from
the RL toolbox failed, and they said so instead of quietly switching methods without comment.
One more low-precision data point, smaller but concrete: on Rebellions' ATOM-Max NPU, A.X-K2 reports
**107% performance-per-watt** relative to a comparable NVIDIA L40S GPU — a real deployment-hardware number,
not a simulation.
## The benchmarks
A.X-K2 leads this five-model, five-benchmark slice outright, and the gap on **Apex** is the most striking:
45.8 against a next-best of 28.1 (DeepSeek-V4 Flash) — more than double the third-place score. Beyond
what's in that chart, SKT reports two non-benchmark math results worth noting because they aren't
self-scored evals: **35/42 on IMO 2025** (the gold-medal threshold is 35, with a perfect 7/7 on each of the
first five problems), and correct proofs for all eight KMO26 second-round problems, using an iterative
proof-refinement method borrowed from DeepSeekMath-V2's approach.
Long-context quality holds up on RULER, staying above 92 out to 128K and only easing to 86.6 at the full
256K:
| Context | 4K | 8K | 16K | 32K | 64K | 128K | 256K | Overall |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| RULER | 97.5 | 97.2 | 97.5 | 96.5 | 94.3 | 92.7 | 86.6 | **94.6** |
Needle-in-a-haystack retrieval is a clean 100 at every position tested at both 256K (YaRN factor 2) and a
zero-shot 512K (YaRN factor 4) — including after NVFP4 quantization, which is the same "the base format
survives further compression" story as the precision ladder above, applied to retrieval instead of MMLU.
## Where it's honest about losing
Every number above is self-reported by SK Telecom, on their own harness, with no independent reproduction
I could find. The eval protocol is disciplined — all baseline open-weight models run through OpenRouter
fixed to the model's own publisher, `xhigh` reasoning effort, pass@1 averaged over multiple generations
(8 for math) rather than single-shot — but it's still one lab grading a comparison it designed. The clearest
weak spot, and SKT names the reason itself:
9.3 is worst of all seven models with a reported score — GLM-5.1 leads at 29.1, and even the next-worst
(Nemotron 3 Ultra at 13.4) beats A.X-K2 by 44%. The model card's own explanation: *"Agentic performance is
moderate — A.X K2 trails the strongest compared models on BrowseComp — reflecting limited agentic RL during
post-training."* That's the right way to publish a weak number — attribute it to a specific, checkable
cause (the RL data mixture allocates only 18% of SFT tokens and a modest RL slice to agentic tool use,
against much heavier agentic investment in a model like [Kimi K3](/articles/kimi-k3)) rather than burying
it. One methodology caveat the report itself surfaces: A.X-K2's only tool on this benchmark was Brave
Search's API, capped at ≤10 searches per problem — a real constraint, not full open browsing — and the
report doesn't state whether every compared model ran under the same cap. That could shift the absolute
number somewhat; it's very unlikely to explain a 3× gap to the next-worst model.
BrowseComp isn't the only place A.X-K2 comes second. On **GPQA Diamond** it's mid-pack, and the surprise is
which model beats it on a Korean-language benchmark:
DeepSeek-V4 Flash — not a Korean-focused lab — beats SK Telecom's own Korean-sovereign model on a Korean
benchmark, by 2.3 points. A.X-K2 still wins the *other* two Korean benchmarks in the comparison (KMMLU-Pro,
CLIcK), so this is one loss inside a category it otherwise leads, not a category-wide miss — but it's
exactly the kind of specific, checkable number a self-reported table should surface rather than smooth
over. Rounding out the mid-pack results: LiveCodeBench v6 at 84.0 (DeepSeek-V4 Flash leads at 89.4), SciCode
at 41.0 (near the bottom of the field; Kimi-K2.6 leads at 53.5), and IFBench at 75.9 (DeepSeek-V4 Flash
leads at 81.2). None of these are collapses — they're a model that wins decisively on math and most of
Korean, and trails on strict-instruction-following, code-execution benchmarks, and — sharply — on
open-web agentic search.
Two more limitations the model card states plainly, worth repeating because they're easy to omit: A.X-K2
is text-only (no native multimodality, listed as future work), and SKT explicitly did not run a dedicated
quantitative bias or fairness evaluation.
## The take
Two disclosures make A.X-K2 worth writing about on their own, independent of where it lands on any single
leaderboard. It's one of the only sparse-attention releases I've seen that reports the ablation showing its
sparsity is nearly free (62.80 → 62.99 on LongBench) instead of only the speedup — crediting the reader
with the question "what did this cost?" instead of hoping nobody asks. And it's trained natively in FP8
end to end, with the RL infrastructure section going out of its way to show a standard fix (TIS) failing
against a real precision mismatch rather than quietly working around it off-page. Set against
[Neutrino-1](/articles/neutrino-1)'s ternary cliff and [MiniMax Sparse Attention](/articles/minimax-sparse-attention)'s
block-granularity bet, A.X-K2 reads as the same 2026 pattern — quantization and sparsity are training-time
commitments now, not deployment-time knobs — applied at a scale and with a level of self-disclosure that
makes the whole argument checkable, including the parts (BrowseComp, KoBALT) where the honest answer is
that it lost.
---
*Sources: the [A.X K2 Technical Report](https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf)
(SK Telecom, dated 2026-07-28 — architecture, training recipe, RL infrastructure, evaluation tables) and the
[model card and config](https://huggingface.co/skt/A.X-K2). Figures 1 and 3 here are the report's Figures 2
and 7, reproduced for commentary; the benchmark comparison figure is the report's Figure 1. All benchmark
numbers are SK Telecom's own, on their own harness, with no independent reproduction found. The
inference-efficiency numbers in the Sparse Gated Attention section are approximate, read off the report's
chart rather than a published table. Interactive diagrams are mine. Related: [Neutrino-1](/articles/neutrino-1)
on training-native vs. post-hoc quantization, [Kimi K3](/articles/kimi-k3) on the other end of the MoE
sparsity-ratio spectrum, and [MiniMax Sparse Attention](/articles/minimax-sparse-attention) on block- vs.
token-granularity top-k selection.*
---
# Chimera: unbundling RoPE into a diffusion Transformer that extrapolates 6× on video
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/chimera-diffusion
> date: 2026-08-03
> tags: diffusion, linear-attention, video-generation, scaling-laws, positional-encoding, explainer
Visual generation is hitting the same wall language models hit a few years back: the tokens keep multiplying.
A high-resolution image is thousands of tokens, a video clip is tens of thousands, and once you want text,
image, and video sharing one context, full attention's quadratic cost stops being a rounding error and starts
being the budget. Language models solved their version of this with linear and hybrid attention. The catch,
as [Chimera](https://arxiv.org/abs/2607.28611) — Adobe Research's new hybrid visual diffusion Transformer —
points out, is that those solutions don't transfer directly: a diffusion backbone has to preserve spatiotemporal
locality and support genuinely bidirectional interaction across modalities, neither of which a causal language
model has to worry about.
Chimera's answer is a single-stream backbone that processes text, image, and video tokens together, mixing them
with **Kimi Delta Attention (KDA)** for cheap $O(N)$ state tracking, periodic **Multi-head Latent Attention
(MLA)** for exact global interaction, and **modality-aware short convolutions** for local structure — with no
positional embeddings anywhere in the stack. The paper backs this with **HeteroP**, a hyperparameter-transfer
scheme built for a backbone that is not one uniform shape, and fits genuine Chinchilla-style scaling laws on top
of it. The headline results: a real zero-shot **6× video-length extrapolation** (5-second training clips
generalizing to 30 seconds with only 6.5% FID degradation, versus 50%+ for two full-attention baselines), and a
compute-efficiency claim over a matched Wan2.1 baseline that the abstract prints as **7.3×** but the paper's own
arithmetic, two pages later, computes as **6.8×**. Both numbers are worth seeing, and the mechanism behind the
extrapolation result is the more interesting story.
## One stream, three mechanisms, no positions
Text tokens (from a frozen T5-style encoder) and visual tokens (from a frozen Wan2.1 VAE, patchified) are
concatenated into a single sequence and pushed through the same stack. Visual tokens are flattened in
**temporal-major raster order** — position $(i,j,k)$ in a $(T, H, W)$ grid maps to sequence index
$m = iHW + jW + k$, with a single image just a one-frame video — so there is one token order for every modality,
not a per-modality scheme bolted on afterward.
Inside a block, the attention sublayer is either KDA or MLA on a fixed **3:1 KDA-to-MLA schedule**: three linear
layers, then one global layer, repeating. The first and last blocks use a dense SwiGLU feed-forward; every other
block routes through **sparse MoE** — 56 experts, top-8 active, no shared experts, balanced by an auxiliary-loss-free
bias added only at top-K selection (not to the mixture weights). With that bias, batch-level `MaxVio` — the metric
tracking how far the busiest expert's load sits above the average — settles near 0.5; strip the bias out and it
blows past 5, close to the theoretical collapse bound of 6. Every sublayer — attention and FFN alike — is wrapped in
**Identity Hyper-Connections (iHC)**: the residual stream is duplicated into $M{=}4$ parallel copies with
token-dependent read/write gates, a simplified version of hyper-connections that fixes the residual-mixing matrix
to the identity instead of learning a doubly-stochastic one via Sinkhorn iterations — cheaper, at the cost of losing
learned cross-stream mixing.
None of KDA, MLA, or NoPE is new by itself — they're the same three ideas [Kimi K3](/articles/kimi-k3) uses to run
a 1M-token language model, adopted here essentially as-is and pointed at diffusion instead of next-token
prediction. **KDA** keeps a fixed-size recurrent state per head instead of a growing KV cache:
$$
S_t = \big(I - \beta_t k_t k_t^{\top}\big)\,\mathrm{Diag}(\alpha_t)\,S_{t-1} + \beta_t k_t v_t^{\top}
$$
with $\alpha_t \in (0,1)^{d_k}$ a **per-channel** forget gate and $\beta_t$ a scalar write strength — the exact
recurrence [KDA has a half-life](/articles/kda-half-life) walks through: each channel forgets a fixed fraction of
its state per step, so it has a half-life $n_{1/2} = \ln(0.5)/\ln(\alpha)$ measured in tokens. That article works
the math out for language tokens; Chimera is the same forget-gate law now doing memory management for pixels and
frames instead of words. **MLA** restores exact bidirectional interaction on top, compressing keys and values
through a low-rank projection, with one twist: the direct key path is left **unrotated** — no RoPE at all, on
either mechanism. That's the part worth slowing down on.
## What RoPE was actually doing
Every visual diffusion Transformer before this one, and most language models, lean on RoPE to inject position.
Chimera's authors run a mechanistic audit of what that rotation is actually buying, on Qwen3-4B first and then on
the visual diffusion models FLUX.2 and Wan2.2, and decompose each attention logit into its per-frequency-pair
cosine contributions (the pieces sum to the real logit to within $3 \times 10^{-13}$, so this isn't an
approximation of the mechanism — it's an exact accounting of it).
The clearest case is layer 0, head 1 in Qwen3-4B, a strong "previous-token" head: 98.5% of its queries attend most
strongly to the immediately preceding token. Trace *why*, and it's one channel pair — the fastest-rotating one,
turning roughly one radian per token — whose cosine happens to peak exactly at offset 1. The slow, high-amplitude
pairs sitting alongside it contribute a flat, content-driven background that doesn't move the peak at all. Search
across every head for the best "attend exactly $n$ tokens back" detector and the same pattern holds: the strongest
head hits a 0.56 fraction at $n{=}2$, but only 0.07 at $n{=}10$ — RoPE-based position selection is a short-range
tool, not a general one, in both language and visual models alike.
From that audit the paper pulls out three things RoPE is doing at once, inside the same rotated channels:
1. **Position selection** — a canonical previous-token/induction-head trick, and (per the numbers above)
reliable only at short offsets.
2. **Recency decay** — an *implicit* average property of the summed rotations, not an explicit mechanism, and one
a trained head can learn to bypass.
3. **Layout encoding** — the channel partition across positional axes (time, height, width, text index) is set by
hand at design time; every additional axis divides the available channels further and never adapts.
Chimera's move is to stop asking one set of rotated channels to do all three jobs and give each its own dedicated
module instead:
Token order for KDA comes for free from the recurrence itself — a scan is inherently order-aware, no rotation
required. Position selection moves to the **modality-aware short convolution**: a depthwise kernel mixing tokens at
explicit index offsets, which is a cheap, parameter-light way to do exactly the short-range job the audit found
RoPE was actually good at. Recency decay moves to **KDA's own forget gate** $\alpha_t$ — explicit and
content-adaptive, rather than an emergent side effect. And layout encoding falls out of the convolution's native
shape: **causal 1D** along the token index for text, **causal-in-time 3D** over the $(T,H,W)$ grid for video — each
modality's structure is encoded by the operator itself, with no channels spent partitioning anything. Text and
visual tokens each get their own kernel here, implemented as one fused Triton pass that measures 2.2–2.3× faster
forward, 1.5–1.8× faster forward-plus-backward, and up to 4× less peak activation memory than the naive
gather-convolve-scatter version.
With all three biases reassigned, MLA is left to do only content matching — the paper's framing is that MLA
without RoPE is the limiting case where every channel has zero rotary frequency, so its logits depend purely on
content. KDA's queries and keys carry no positional phase either. Nothing in the stack is tied to how long the
training sequences were, which is the actual, mechanistic reason extrapolation works — not a property tacked on
after the fact, but the direct consequence of where each inductive bias now lives. It's the same bet
[Kimi K3](/articles/kimi-k3) makes for a 1M-token language model: no RoPE means nothing to rescale when the
context grows past training length. Chimera is that bet, replayed in a diffusion Transformer over image and video
tokens instead of a causal LM over text.
## HeteroP: a scaling ratio per tensor, not per model
Fitting a Chinchilla-style law needs a family of models at different sizes, each trained with hyperparameters as
good as they'd be at full scale — otherwise you're not comparing model sizes, you're comparing tuning quality.
The standard fix is µP-style hyperparameter transfer: tune a small proxy model, then derive the large model's
learning rate, init, and weight decay from a single width ratio. Chimera's backbone breaks that assumption,
because it isn't one uniform width. Widening the model changes the KDA head width, the MLA compression rank, the
MoE expert width, the router width, and the timestep-conditioning MLP width all differently — a single global
ratio, tuned for the backbone, is the wrong ratio for the rest.
**HeteroP**'s fix is to stop pretending there's one ratio. For each parameter group $W$, it computes its own width
ratio from that group's own **functional fan-in**, plus one shared depth ratio from the block-count ratio:
$$
(m_W, m_L) = \left(\frac{\mathrm{fan\text{-}in}(W)}{\mathrm{fan\text{-}in}(W^{(0)})},\ \frac{n_{blk}}{n_{blk}^{(0)}}\right)
$$
Concretely: hidden weights get init variance and learning rate scaled by $m_W^{-1}$ and weight decay scaled by
$m_W$ (keeping the LR-times-decay product invariant); attention and FFN residual branches get an additional
$m_L^{-1}$ output scaling, a depth correction borrowed from CompleteP; input adapters, norms, and the readout keep
their base LR and init (standard µP convention), with the readout's forward pass separately rescaled by $m_W^{-1}$.
The proxy model is width 512, depth 4 (22M activated, 59M total parameters); the largest is width 2048, depth 32.
The validation is direct: under HeteroP, the optimal base learning rate sits at about $10^{-3}$ across a 56× range
in activated parameters (20M to 1.12B) and an 8× range in depth (4 to 32 layers). Under standard
parameterization — one global ratio for everything — the optimum drifts sixfold, from $10^{-4}$ to $6\times
10^{-4}$, and several of the high-learning-rate runs at large scale diverge outright. That drift isn't just an
inconvenience for the scaling-law fit, it actively biases it: trained without HeteroP, the same image model
family gives a fitted exponent of $N_{opt}\propto C^{0.588}$ (envelope) or $C^{0.581}$ (isoFLOP) — inflated by
0.08–0.10 over HeteroP's 0.505/0.481 — which the paper's own extrapolation shows prescribes a compute-optimal
model roughly **2× oversized (and correspondingly undertrained)** three orders of magnitude of compute out from
where it was fit. Get the transfer wrong, and the law tells you to build the wrong-shaped model at scale.
## What the scaling law says
With HeteroP holding hyperparameter quality constant across scale, the paper fits $\hat L(N,D) = E + AN^{-a} +
BD^{-b}$ — activated parameters $N$, visual-latent-token count $D$ — using three independent estimators (a
training-loss envelope, an isoFLOP profile, and a direct parametric fit) for image and video pretraining
separately. All three estimators agree, and they disagree with each other by modality:
| Modality | $N_{opt}$ exponent (across 3 estimators) | Split |
|---|---|---|
| Image (256²) | 0.48–0.52 | Nearly balanced between model size and data |
| Video (180p) | 0.53–0.56 | Modestly favors model size at higher budgets |
The parametric fits — $\hat L_{image} = 0.126 + 5.28N^{-0.315} + 33.8D^{-0.336}$ ($R^2{=}0.993$) and
$\hat L_{video} = 0.124 + 8.07N^{-0.330} + 145.2D^{-0.394}$ — land on nearly identical irreducible-loss terms
(0.126 vs 0.124), consistent with both modalities sharing the same denoiser and VAE. The paper adds an axis prior
scaling-law work doesn't have: the compute-optimal **image-to-video data ratio**. It drifts from roughly 4:1 to 3:1
as compute grows from $10^{18}$ to $10^{19}$ FLOPs (image gets relatively cheaper to learn from per token as
budget grows), while the video-loss-optimal ratio stays pinned at 1:1 — the video-heaviest mixture the authors
actually tested, so that half of the result is a boundary effect, not a discovered optimum, and the paper is
upfront that it didn't search past it.
## The number that doesn't quite add up
Guided by those laws, the paper trains an 11B-total / 2B-activated Chimera and compares it against matched 2B
full-attention baselines — Wan2.1 and Z-Image — all four models trained in-house to the same $5\times10^{20}$-FLOP
budget on identical data. At a shared training loss of 0.149, Wan2.1 needs $4.29\times10^{20}$ FLOPs and Z-Image
needs $3.75\times10^{20}$; Chimera-dense (no MoE, no iHC, no HeteroP) reaches it in $2.55\times10^{20}$ — a clean
**1.7×**. The complete configuration (MoE + iHC + HeteroP) reaches it in $6.27\times10^{19}$ FLOPs.
Do the division on the two numbers printed a page earlier and you get $4.29\times10^{20} / 6.27\times10^{19} =
6.84$ — which is exactly what Section 5.5's own sentence says: "a 6.8× compute-efficiency gain over Wan." The
abstract, the introduction, and the plotted label in Figure 12 above all instead say **7.3×**, from the same pair
of FLOPs figures.
I'm not accusing anyone of anything here — I re-fetched the paper's own HTML and confirmed both numbers appear
verbatim: "6.8 × compute-efficiency gain over Wan" in the Section 5.5 prose, right next to the $4.29\times10^{20}$
and $6.27\times10^{19}$ FLOPs figures it's computed from, and "7.3 ×" in the abstract, the introduction, and baked
into Figure 12's own plotted label. $4.29 / 0.627 = 6.84$, not $7.3$, using the numbers exactly as printed. It's
possible the true, unrounded internal FLOPs values reconcile to 7.3× and the 3-significant-figure numbers printed
in the text are what drifted — the paper doesn't say either way, and I found no footnote reconciling the two. What
I can say: **using the checkable numbers, the arithmetic supports 6.8×, not 7.3×.** If you're going to cite
Chimera's headline efficiency gain, cite 6.8×, or go verify the underlying FLOPs yourself.
Worth separating, too: the component ablation (dense → +MoE → +iHC → +HeteroP) reports a **4.1×** cumulative gain
at loss 0.149 — MoE alone gets to 1.5×, +iHC to 1.7×, +HeteroP the rest of the way. That 4.1× is measured *relative
to Chimera-dense*, not to Wan2.1, so it isn't a third candidate for the headline number — it's answering a
different question (how much of the complete system's win comes from which piece), and it's consistent with
either the 6.8× or 7.3× reading of the Wan-relative number, since $1.7 \times 4.1 \approx 7.0$, which lands between
the two and settles nothing on its own.
## Zero-shot length extrapolation: NoPE's actual payoff
This is where the RoPE audit cashes out. Chimera is trained only on 5-second, 81-frame clips, then asked — with
**no length-specific fine-tuning at all** — to generate 30 seconds, 6× its training length. Every metric below is
computed only on the final 5 seconds of each generated clip, isolating the extrapolated region, over 512 generated
vs. 512 reference videos at matched prompts, seeds, resolution, and fps:
The numbers: Chimera's FID goes from 77.1 at 5 seconds to 82.1 at 30 — a **6.5%** degradation. FVD moves from
685.8 to 829.5, up 20.9%. Wan2.1-T2V-1.3B's FID degrades **50.5%** over the same stretch, HunyuanVideo-1.5's
**53.6%** — both well past the point where a video model's later seconds are visibly falling apart. Chimera also
posts the lowest *absolute* FID and FVD of the three at 30 seconds, not merely the smallest percentage move — it
isn't winning by having started worse and degrading less, it's ahead the whole way.
This is the sibling result to [SANA-Video 2.0](/articles/sana-video2), the other linear-attention video approach
covered here, and the two make an interesting contrast. SANA-Video 2.0 keeps the same 3:1 linear-to-global
attention idea and Block Attention Residuals for cross-depth flow, but keeps RoPE and optimizes for raw
single-GPU latency at a fixed, modest clip length. Chimera keeps RoPE out entirely and stakes the design on
exactly the axis SANA-Video 2.0 doesn't test: generalizing far past the lengths it was trained on. Different bets,
same underlying conviction that softmax attention over every token pair was never the part of video generation
worth paying full price for.
## What $O(N)$ buys in memory and latency
A softmax KV cache costs $O(N \cdot H \cdot d_h)$ — grows with sequence length. KDA's recurrent state costs
$O(H \cdot d_h^2)$ — fixed, independent of $N$. Measured directly: a matched KDA/MLA and MHA/MLA backbone, both
around 2B activated parameters at the same 3:1 ratio, batch size 1, BF16, 512 text tokens plus 18×28 visual
tokens per frame, on one NVIDIA A100-SXM4-80GB —
— the linear variant supports 1.68× longer sequences before it runs out of memory, and runs 2.14× faster at the
255k-token point both backbones can reach. The paper is careful about what this comparison actually shows:
FlashAttention removes the quadratic attention *workspace*, but not the quadratic *arithmetic* — so this isn't a
straw-man comparison against an un-optimized baseline, it's the honest gap that remains after the standard fix.
## Benchmarks, and what 600 H100-days buys
Trained for only about 600 H100-days, Chimera is competitive on text-to-image quality with models that cost far
more to build:
On GenEval, Chimera lands at 0.82 overall — tied with Z-Image-Turbo, matching FLUX.1-dev, beaten only by
Seedream 3.0's 0.84 — and it beats both FLUX.1-dev and Z-Image-Turbo on DPG-Bench specifically. The paper also
quotes Z-Image-Turbo's *own* reported training budget, about 12.4K H100-days, as roughly 20× Chimera's — worth
reading as a cross-lab comparison rather than a controlled one: different codebases, different clusters, and a
number each lab measured on its own infrastructure, not a shared benchmark. The GenEval/DPG-Bench baseline rows
themselves are the field's normal practice, too — each competitor's own published number, not re-run under
Chimera's exact sampling protocol. Worth knowing before quoting either comparison as settled.
## Honest limits
The paper is unusually direct about what it hasn't shown yet:
- **MoE underperforms its LM-scaling expectation.** Sparsity buys only about **1.5×** compute efficiency here,
well short of the commonly cited $\sqrt{\text{sparsity}}$ heuristic — the authors attribute this to weak expert
specialization (routing stays close to uniform across tokens and timesteps) and call it out as an architectural
ceiling, not a training bug. A negative result reported plainly rather than smoothed over.
- **Muon underperforms AdamW throughout** their tests. They hypothesize Adam's implicit low-rank bias matters for
diffusion training specifically, but say so as a hypothesis, backed by a brief spectral-analysis follow-up, not
a systematic sweep.
- Several structural ratios are **held fixed across the entire scaling study** and never made scale-dependent:
iHC's stream count, MoE's expert count and top-K, and MLA's KV-compression ratio. The paper flags directly that
whether the optimal compression ratio is itself scale-dependent is left to future work.
- Only **text-to-image and text-to-video generation** are evaluated — no multimodal *understanding* task, despite
the single-stream design being a natural fit for one (the paper name-drops this as a target direction, not a
result).
- The timestep-conditioning MLP's width is scale-sensitive enough to destabilize training if mismatched to the
backbone — a 4096-dim MLP paired with a 1024-wide backbone went unstable — and it's patched ad hoc rather than
covered by the main HeteroP table.
Set against that, the scaling-law and compute-efficiency measurements themselves are unusually rigorous for the
genre: Wan2.1 and Z-Image aren't cited from their own papers here, they're re-implemented and trained in-house at
matched 2B scale on identical data, specifically so the efficiency comparison isn't citing someone else's number
under someone else's conditions.
## The take
The RoPE audit is the part of this paper worth remembering past the benchmark tables. It's not "we removed
positional embeddings and it worked" — it's a demonstration, with an exact per-frequency accounting, that RoPE was
quietly doing three separate jobs through the same rotated channels, that one of those jobs (position selection)
only works at short range anyway, and that giving each job its own dedicated, non-attention mechanism is what
actually buys extrapolation — not a side effect of going linear, a direct consequence of where position now lives
in the model. HeteroP is the less flashy but equally load-bearing half: none of the scaling-law numbers mean
anything if the hyperparameters drift as you change scale, and a heterogeneous backbone needs a heterogeneous
transfer scheme to keep them from drifting.
The 6.8-versus-7.3 gap doesn't undercut any of that — it's a rounding-sized discrepancy in one headline multiplier,
not in the mechanism. But it's exactly the kind of thing worth checking yourself before repeating a number,
which is the whole reason to read the arithmetic instead of just the abstract.
---
*Source: [Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers](https://arxiv.org/abs/2607.28611)
(Ge, Jiang, Wang et al., Adobe Research, 2026), read from the arXiv HTML render. Figures 1–3 here are the paper's
Figures 2, 12, and 16b, reproduced for commentary; all benchmark and scaling-law numbers are the paper's own,
self-measured against in-house-trained baselines except where marked as cited. The RoPE re-assignment, HeteroP
drift, and length-extrapolation diagrams are mine — the first is a schematic of the paper's own finding, the second
reproduces the real numbers from its Figure 7 ablation on an illustrative loss curve, the third traces the actual
tested points from its Figure 16b. Related: [Kimi K3](/articles/kimi-k3) for where KDA, MLA, and NoPE come from;
[KDA has a half-life](/articles/kda-half-life) for the forget-gate math this piece leans on; and
[SANA-Video 2.0](/articles/sana-video2) for the other linear-attention video architecture on this site.*
---
# DCFormer: an ICML oral that let attention heads borrow each other's circuits, and quietly shipped anyway
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/dcformer
> date: 2026-08-03
> tags: attention, transformers, architecture, scaling, llm
**Dynamically Composable Multi-Head Attention** (DCMHA) is a two-year-old idea: May 2024,
an ICML oral, a real mechanism, a real headline number — DCPythia-6.9B beats open
Pythia-12B on Pile validation perplexity (5.95 vs 6.01) with close to half the
parameters. By any normal measure that should have traveled. Semantic Scholar currently
shows it at roughly a dozen citations, one flagged "influential." That's a modest
footprint for an ICML oral two years out. What it does have is a real deployment: Caiyun
Technology, the authors' own industry affiliation, put it into a production language
model and an AI-RPG platform, and kept extending the training code through early 2025.
The honest frame is not "field-changing" — it's a technically solid piece of work that
quietly shipped without the citation graph to match, and is worth understanding on its
own terms.
There is a **second, unrelated** paper that reuses the same name: "DCFormer: Efficient 3D
Vision-Language Modeling with Decomposed Convolutions" (arXiv 2502.05091, 2025) is about
decomposed convolutions for 3D vision-language models — nothing to do with attention
heads. Everything below is arXiv 2405.08553, Xiao, Meng, Li, and Yuan, "Improving
Transformers with Dynamically Composable Multi-Head Attention." Check the ID if you go
looking for it.
## Two problems with heads that never talk to each other
Standard multi-head attention runs $h$ heads in parallel, each on its own learned
$Q$/$K$/$V$ projection, and concatenates the results. The heads never see each other's
work. That independence is also the design's two weaknesses. First, the **low-rank
bottleneck**: each head's attention score matrix is a rank-limited function of a
head-dimension-sized projection, and Bhojanapalli et al. (2020) showed that widening the
per-head QK dimension relieves it — but widening every head's projection is expensive.
Second, **head redundancy**: with nothing coupling them, heads are free to learn
overlapping, partially duplicated functions instead of covering the space of useful
attention patterns efficiently.
DCMHA's answer isn't to make heads bigger. It's to let them **compose** — combine their
scores and weights across heads, per token — which the paper shows buys the same kind of
expressivity gain that a wider QK projection would, without actually widening anything.
## What Compose does
Two calls to a function named `Compose` are inserted into ordinary multi-head attention:
one right after the scores are computed (pre-softmax), one right after softmax turns
them into weights (post-softmax).
For a fixed query/key pair, stack the $H$ heads' scores (or weights) into one vector
$A_{:ij} \in \mathbb{R}^H$ — "the attention vector." `Compose` turns that into a new
vector $A'_{:ij}$ by summing five branches:
- **B1**, a static base projection (in practice, DCMHA drops this in favor of a plain
skip connection, with no measurable loss);
- **B2/B3**, a query-wise dynamic low-rank projection and a query-wise dynamic gate,
both generated from $Q_i$ by a small FFN;
- **B4/B5**, the same pair generated from $K_j$ instead.
Every one of B2 through B5 is **input-dependent** — the weights that decide how much of
head $h'$ leaks into head $h$'s new value are computed fresh from the actual query or key
vector at that position, not fixed at training time.
## Why it has to be dynamic, not just wider
The paper proves something specific about the *static* version of this idea first.
Compose one head's score with a fixed matrix $C \in \mathbb{R}^{H \times H}$, and that is
provably identical to concatenating an $H$-fold expanded QK projection (Theorem 2.1); do
the same to the post-softmax weights, and it's identical to an expanded V/O projection
(Theorem 2.2). In other words: a **static** composition matrix buys you exactly what a
wider head dimension buys you — the fix for the low-rank bottleneck — and nothing more.
**Talking-Heads Attention** (Shazeer et al., 2020) is that static case: it already
composes both scores and weights, just with one fixed matrix reused for every token,
every input, forever. DCMHA's own ablation measures the gap this leaves on the table —
adding the static projection alone gets Pile validation perplexity from 11.68 down to
11.17, but the full dynamic Compose reaches 10.79. The static version is doing real work;
the dynamic version is doing about 60% more of it, by this measure. Query-wise and
key-wise branches contribute nearly as well on their own as together, and post-compose
(on the weights) alone beats pre-compose (on the scores) alone, 11.05 vs 11.54 — the low-
rank projection branches (B2/B4) matter more than the gates (B3/B5).
## Why it's cheap
The reason DCMHA doesn't cost what a full head-to-head transform would is a
decomposition, not a shortcut. Conceptually, composing every head with every other head
needs an $H \times H$ transform per query/key pair — quadratic in the number of heads.
DCMHA factors that tensor into a query-wise term plus a key-wise term (row + column), and
factors each of *those* into a rank-$R$ product plus a diagonal gate (low-rank + diagonal
decomposition). The cost drops from $H^2$ to $2HR + H$, and — a nice side effect — the
key-wise half can be computed once and cached alongside K/V, which is exactly what a
serving stack needs.
At the paper's own 6.9B-scale example ($D_h = 128$, $R = 2$): roughly 1.3% extra
parameters and 1.9–3.3% extra FLOPs, depending on sequence length. Rank $R = 2$ turns out
to be close to a sweet spot in the ablation ($R{=}1$: 10.87 ppl, $R{=}2$: 10.83,
$R{=}4$: 10.89 — non-monotonic, and not worth pushing higher).
## The headline number
Trained on The Pile, matched Chinchilla-style token budgets, three model families: the
scaling curves show **DCFormer-834M matches a plain Transformer trained with roughly
1.87 times the compute**, and DCFormer++ (RoPE + SwiGLU added to both sides) matches its
own baseline at roughly 1.67 times. That gap doesn't shrink with scale — DCMHA's relative
improvement decays more slowly than the RoPE+SwiGLU improvement does, which is the
favorable direction.
The result that carries the abstract is the 300B-token run against the actual Pythia
suite:
DCPythia-6.9B's 5.95 beats Pythia-12B's 6.01 — a model with close to half the
parameters, ahead on the metric that matters for pretraining. It also edges out on
average 0-shot downstream accuracy (56.7 vs 56.5) and 5-shot (57.7 vs 57.2). The gap over
its own size class is not close: Pythia-6.9B sits at 6.29, meaningfully behind.
| Model | Pile ppl | Flan ppl | Avg 0-shot acc |
|---|---|---|---|
| Pythia-2.8B | 6.63 | 8.16 | 53.1 |
| DCPythia-2.8B | 6.36 | 7.68 | 54.5 |
| Pythia-6.9B | 6.29 | 7.85 | 55.1 |
| **DCPythia-6.9B** | **5.95** | **7.13** | **56.7** |
| Pythia-12B | 6.01 | — | 56.5 |
The gap is largest on the Flan Collection (instruction-following/few-shot/CoT data) and
grows with scale, which the authors read as DCMHA disproportionately helping the harder,
more compositional end of the task distribution — a reading the paper backs up with a
purpose-built test.
## A synthetic test built to need composition
The authors built a 74-task, 888-example diagnostic where getting the right answer
requires *simultaneously* attending to the right source token and applying the right
output transformation (e.g., mapping an object to its superclass) — precisely the
combination a head that only ever reads its own fixed QK/OV circuit should struggle with:
Perplexity on this set drops from 10.05 to 7.36 alongside the accuracy jump — a much
bigger swing than on Pile or Flan, and the paper's own explanation is the one you'd
expect: this task rewards recombining an existing head's QK circuit with a different
head's OV circuit on the fly, which is the one thing static heads structurally cannot do.
The head-diversity analysis (captured variance of concatenated QK and OV projection
matrices, lower meaning more diverse heads) backs this qualitatively too — DCPythia shows
markedly more QK-circuit diversity than Pythia, and moderately more OV-circuit diversity.
## The honest costs
None of this is free, and the paper says so plainly. Composition is I/O-bound, not
compute-bound, and the reference implementation has no fused kernel — plain JAX for
training, plain PyTorch for inference:
| Size | Training throughput (DCFM++ / TFM++) | Inference throughput (DCFM++ / TFM++) |
|---|---|---|
| 2.8B | 74.5% | 81–88% |
| 6.9B | 83.1% | 89–95% |
| 13B | 84.4% | 90–95% |
| 33B | 89.2% | 90–95% |
The overhead shrinks as models scale up, and the authors are explicit that a fused
kernel — FlashAttention-style — is headroom they haven't taken. A separate lever recovers
most of it directly: raising the local-to-global attention ratio and composing only
query-wise (dropping the key-wise branches) pushes DCFormer++-6.9B's training throughput
from 83.1% back up to 92.5% of baseline, at a small, still-net-favorable cost in
perplexity.
Two more honest limits, stated in the paper's own words. First: **DCMHA doesn't
transplant onto a pretrained model.** Continual-pretraining a 1.4B LLaMA-style checkpoint
into a DCFormer for a tenth of its original training steps produced no real improvement —
the composition that matters most happens in early layers, and early-layer gradients are
too small during fine-tuning to move already-settled MHA weights. DCFormer has to be
trained from scratch. Second: the paper is explicit that **matching SOTA was never the
goal** — DCPythia deliberately keeps every other Pythia hyperparameter fixed, to isolate
what DCMHA alone contributes, rather than stacking it with every other efficiency trick
to chase a leaderboard number.
The mechanism transfers outside language too: on ImageNet-1K, DCViT-S/16 at 1.03× the
baseline's parameters (68.0 top-1 at epoch 90) matches ViT-M/16 at 1.72× the parameters
(67.1 top-1) — the same roughly 1.7× parameter-efficiency story, in a different domain,
on one held-out test.
## A different axis from the memory-side attention variants
If you've read the [field guide to attention mechanisms](/articles/attention-mechanisms)
on this site, it's worth being precise about where DCMHA sits relative to that map. MQA,
GQA, and MLA all operate on what that piece calls the **memory axis** — they *share* or
*compress* the K/V heads to shrink the KV cache, trading some quality for less memory
bandwidth at decode time. DCMHA doesn't touch the cache at all: the number of physical
K/V heads is unchanged, nothing shrinks. It operates on an orthogonal axis entirely — not
how many heads you cache, but what each head is allowed to compute, by letting it borrow
another head's QK or OV circuit, per token. A model could in principle combine GQA's
cache savings with DCMHA's composition; the paper doesn't test that combination, so
read it as plausible, not demonstrated. (For the general design question of specializing
attention below the layer level, [HydraHead](/articles/hydrahead) is the other piece on
this site working that seam, from a different angle.) The broader landscape of
architecture choices — attention, position encoding, MoE, diffusion — is mapped at
[/architectures](/architectures).
## How much of this actually caught on
Being fair to the number in the abstract requires two caveats the paper itself doesn't
hide. The Pythia baselines are **re-run by the authors** under matched settings, not
copied from the original paper — a genuinely controlled comparison, and the authors say
so directly ("our aim is not to obtain SoTA results, but to clearly quantify the gain").
And the compute-equivalence multipliers (1.87×, 1.67×, 1.85×, 1.97×) come from fitting
scaling-law lines to **three data points per curve** — reasonable given the cost of
training each point, but a thinner fit than, say, Chinchilla's own study.
Two years after an ICML oral, roughly a dozen citations and one flagged "influential" is
a modest academic footprint — the kind of number that would normally suggest an idea
that didn't pan out. What actually happened looks different: Caiyun Technology, the
paper's own industry co-author's employer, shipped DCFormer into a production language
model and upgraded an AI-RPG platform to run on it, and the GitHub repository was still
being extended — DeepSpeed ZeRO support, Hugging Face Trainer integration — well into
2025. No successor paper has benchmarked against DCFormer as a state-of-the-art baseline
to beat; the adoption signal here is industrial, not academic. That's a real but
different kind of validation than a citation count measures, and it's the honest way to
read this one: not a paper that changed the field's direction, but a working piece of
architecture that one production system actually adopted, sitting quietly under-cited.
## The take
Fixed, independent attention heads leave two things on the table: a low-rank bottleneck
that a wider head dimension fixes at a real cost, and redundancy nothing forces heads to
avoid. DCMHA's Compose function fixes both by recombining heads' scores and weights
per token, through a decomposition cheap enough that a 6.9B model pays about 1.3% more
parameters for it. The result — DCPythia-6.9B beating Pythia-12B on perplexity at
roughly half the parameters — is real, reproduced by the authors under controlled
settings, and backed by a synthetic test built specifically to need what static heads
can't do. It just hasn't been the paper everyone cites. Production adoption at one
company and a quiet GitHub repository are what two years actually bought it — which is a
fine outcome for a piece of architecture, even if it isn't the one the citation count
would lead you to expect.
---
*Built on [Improving Transformers with Dynamically Composable Multi-Head Attention](https://arxiv.org/abs/2405.08553)
(Da Xiao, Qingye Meng, Shengping Li, Xingyuan Yuan; Beijing University of Posts and
Telecommunications / Caiyun AI; ICML 2024, Oral) and its
[code release](https://github.com/Caiyun-AI/DCFormer). Figures are the paper's own
Figures 2 and 4, reproduced for commentary. Tables and numbers are the authors' except
where marked as this site's own illustrative simplification (the Compose bar demo, the
head-cost demo); interactive diagrams are mine.*
---
# Dream-Cubed: diffusion directly on Minecraft's block IDs, where inpainting comes free
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/dream-cubed
> date: 2026-08-03
> tags: diffusion, generative-models, minecraft, world-models, 3d
Most generative work on 3D game worlds goes through a pixel renderer or a learned latent
space before it touches anything the game itself understands. **Dream-Cubed** skips both.
It trains a diffusion model directly on Minecraft's own vocabulary — integer block IDs,
the same representation the game engine uses — and gets a specific, useful property for
free as a result: place a block by hand anywhere in a chunk, and the model will build
around it *exactly*, with no special training and no extra inference-time machinery. The
paper is a three-month-old preprint from a team at NYU and Sakana AI, posted with no peer
review and, as of this writing, no citations. It also lands in a niche that's suddenly
crowded — the paper names two concurrent Minecraft/voxel generators of its own accord —
and its own limitations section is unusually candid about how little its evaluation
actually proves. All of that is worth knowing before the mechanism, which is genuinely
worth understanding.
## Blocks as tokens, not pixels
Each training sample is a $32 \times 32 \times 32$ tensor of integer block IDs — dirt,
water, stone, whatever the game placed there — over a vocabulary of 117 block types in
the core procedurally-generated set, extended to 177 once six professionally human-
authored maps are folded in. The dataset totals 1,667,781 procedural chunks plus 358,762
human-authored ones: **2,026,543 chunks**, tens of billions of tokens at $32^3 = 32{,}768$
voxels per chunk, spanning fifteen biomes from ocean to village to cave.
One backbone serves both diffusion families the paper compares: a 280M-parameter 3D
Diffusion Transformer, 25 blocks, hidden dimension 768, 8 attention heads. A 3D
convolution patchifies each chunk into non-overlapping voxel patches; fixed 3D sine-cosine
position embeddings and a biome-label embedding condition every block through AdaLN
modulation, the same conditioning mechanism DiT uses for timestep and class in image
diffusion.
## Two diffusion families, one backbone
**Masked discrete diffusion (MD4).** Add a `[MASK]` token to the block vocabulary. The
forward process independently masks each voxel with probability
$p_{\text{mask}}(t) = \sin(\pi t / 2)$ for $t \sim \mathrm{Uniform}(0,1)$ — at $t{=}0$
nothing is masked, at $t{=}1$ everything is. The network sees the corrupted chunk and
predicts the original block ID at every masked position, trained with cross-entropy on
masked positions only. Sampling starts from an all-`[MASK]` chunk and iteratively unmasks
positions over a fixed number of steps.
**Continuous diffusion (DDPM), in an embedding space.** Every block name — "dirt",
"sand" — is embedded once via OpenAI's `text-embedding-3-small`, giving a frozen 16-
dimensional lookup table with a semantic prior baked in for free. A standard cosine noise
schedule runs $x_t = a_t x_0 + b_t \varepsilon$ over 1000 steps, trained with v-
prediction. At the end of sampling, the continuous output is decoded back to discrete
block IDs by nearest-neighbor lookup against the embedding table.
Both are trained on the identical backbone, the identical data, the identical compute
budget — the paper's stated goal is a controlled, apples-to-apples comparison of the two
diffusion formulations, not a fight either one is rigged to win.
## Why inpainting is free under one formulation and not the other
This is the mechanism worth sitting with. MD4's forward process decides, **independently,
per voxel**, whether that voxel gets masked. A voxel the user has placed by hand simply
isn't a candidate for that decision — it's excluded from the start, at every step, all
the way through sampling. There is no moment where the model has to reconcile "what I was
trained to expect here" with "what's actually here," because the fixed voxel was never
part of the corruption process to begin with.
A continuous DDPM cannot get the same thing for the same reason. Its forward process adds
Gaussian noise to *every* position, at every step, following one global schedule — clean
voxels aren't a special case the network was ever trained to see. If you clamp a user-
placed block to its clean embedding partway through sampling — the natural thing to try —
the model is still conditioned on that position carrying noise amplitude $b(t)$ at
timestep $t$, and it sees zero instead. That's a real mismatch between what training
taught the network to expect and what inference is handing it, not a cosmetic one.
The paper is upfront that it doesn't solve this: closing that gap needs extra machinery —
RePaint-style repeated re-noising and re-sampling — which it flags in an appendix as an
unresolved comparison point, not a capability it demonstrates for the DDPM side. Exact
conditioning is what falls out of the masked formulation for nothing; it's what the
continuous one would have to be re-engineered to approximate.
If you've read the piece on [iLLaDA](/articles/illada-diffusion-language-model) on this
site, the mechanism will look familiar: masked discrete diffusion over text tokens is the
same "absorbing-state" idea MD4 applies here to voxels — a masking probability schedule,
a network trained to fill in exactly the masked positions, bidirectional context by
construction. iLLaDA's own masking ratio is closer to a straight linear schedule
($t$ itself, roughly); Dream-Cubed's MD4 uses the $\sin(\pi t/2)$ reparameterization from
Shi et al.'s original MD4 paper — a detail, not a different mechanism. What changes here
is the alphabet the diffusion runs over: block IDs instead of vocabulary tokens, arranged
on a 3D grid instead of a 1D sequence. Same masking idea, different token space — see also
the [masked-diffusion-lm entry](/architectures) in the architecture map for where this
sits relative to the wider non-autoregressive-LM family.
## Outpainting is the same trick, tiled
Generating a world larger than one $32^3$ chunk uses a sliding window: partition the
larger canvas into overlapping cells, generate them in sequence, and for every cell after
the first, treat the already-generated overlap with its neighbors as more fixed context —
recursively the same "these voxels are excluded from masking" trick, just applied at
world scale instead of one seeded pattern.
The cost of this is real and disclosed: a single 5×5 outpainted world takes over an hour
of H100 inference time, generated cell by cell, sequentially — the paper calls inference
speed "a practical barrier to all envisioned applications," not a solved problem.
## What the numbers actually say
**MD4 and DDPM land in a statistical tie on the paper's own metric.** Adjusted FID
(generated minus a reference FID from held-out chunks) averages 59.26 for MD4 at patch
size 2 versus 59.29 for DDPM at the same patch size — indistinguishable overall, with MD4
winning 9 of 15 biomes and DDPM winning 6.
**Patch size is where the two formulations actually separate.** MD4 holds up at patch
sizes 2 (4,096 tokens per chunk) and 4 (512 tokens), with visible artifacts only at patch
8; DDPM works at patch 2 but **fails outright at patch 4** under the identical
configuration. That's the one place in the paper where discrete and continuous diffusion
give clearly different answers, and it favors the discrete side.
**Naive frequency matching doesn't work for rare, structured content.** Three data
mixtures were compared: a balanced split, natural biome frequency, and a village-boosted
split. Natural frequency wins on average FID — but ocean chunks, over-represented 5.3×
relative to balanced, and forest, at 1.7×, improve, while village and cave, both rare and
structurally complex, get worse. Boosting village samples specifically recovers the
village-biome losses. The honest reading: matching real-world frequency is not
automatically the right training mixture once some categories are both rare and hard.
The human study is small and its own authors say so: 19 Minecraft-experienced
participants (all from the authors' own institution), roughly 1,000 two-alternative
forced-choice trials, free pan/zoom/rotate:
Both MD4 configurations beat real chunks at statistical significance (patch 2: p less
than 0.001, patch 4: p = 0.042) — a result the authors attribute candidly to classifier-free
guidance pushing generated samples toward a "prototypical" idealized biome, more uniform
than messy real terrain, rather than claiming their model has somehow out-built reality.
MD4 (patch 2) and DDPM (patch 2) tie head-to-head at 49.4%, consistent with the FID
result above.
| Min FID gap between two models | Agreement with human preference | n | p |
|---|---|---|---|
| 0 | 54.3% | 512 | 0.029 |
| 5 | 61.7% | 227 | less than 0.001 |
| 10 | 62.9% | 159 | less than 0.001 |
| 15 | 66.1% | 109 | less than 0.001 |
Agreement between FID and human raters rises with the size of the FID gap being
compared, but tops out at 66.1% even at the largest gap tested — a coin flip with a thumb
on the scale, offered by the authors themselves as evidence that FID is only weakly
informative here, and only at large gaps.
The evaluation the whole paper hangs its numbers on has real, self-acknowledged holes.
FID is a render-based metric: it cannot see building interiors, cave systems, or whether
a door is actually reachable — none of the things that make a Minecraft structure
*functional* rather than merely picturesque. It's computed from only 1,500 rendered
images per model, and costs roughly 60 GPU-hours to run — expensive enough that the
authors say it can't be used for model selection during training, only for a final
after-the-fact score. And the human study, small as it is, evaluated only biome-
conditioned generation. It never tested the inpainting or outpainting capability the
paper actually leads with — the figures above are demonstrations, not measured results.
The dataset itself is drawn from Minecraft version 1.12.2 (2017), for tooling
compatibility, so newer blocks and biomes aren't represented at all.
Training cost is disclosed cleanly: 4×H100 GPUs, classifier-free guidance with 20%
label dropout during training and a guidance scale of 4.0 at inference, patch-2 models
run 20 epochs and patch-4 models 160 (matched for equal token exposure), roughly 192
GPU-hours total across every model in the paper. Inference is the bottleneck end of the
system: about 2.5 minutes per chunk at patch 2, 25 seconds at patch 4.
## A crowded moment, honestly disclosed
Dream-Cubed cites its own competition directly rather than presenting itself as
singular: **Scaffold Diffusion** (a NeurIPS 2025 workshop paper that conditions on an
input occupancy scaffold instead of generating from nothing) and **PERSIST** (arXiv
2603.03482, roughly a month earlier, which uses a 3D DiT with rectified flow matching but
as one component inside a video-generation system, not a standalone generator) are both
named as concurrent work on the same general problem, in the same few months of 2026.
**Solaris** (arXiv 2602.22208) is cited as another concurrent voxel/world-modeling effort
in the same window. None of WorldGAN, Scaffold Diffusion, or XCube — the paper's
narrative comparison points — are actually benchmarked against Dream-Cubed's FID or
human-preference numbers on shared ground; the positioning against prior work is
qualitative throughout, and there is no table anywhere in the paper showing Dream-Cubed
beating a previously published Minecraft or voxel generator on a metric both were scored
on. For a reader wanting a settled state-of-the-art claim, that table doesn't exist yet.
Code, data, and all pretrained models are released
([github.com/SakanaAI/DreamCubed](https://github.com/SakanaAI/DreamCubed)), which is
worth crediting on its own — a preprint this new, this openly scored against its own
limitations, and this fully released, is a reasonable way to publish work you don't yet
have citations to back up.
## The take
The technical point is narrow and real: masked discrete diffusion turns user-block-
conditioning, inpainting, and outpainting into a structural guarantee — unmasked voxels
were never part of the corruption process, so they can't drift — while the equivalent
constraint on a continuous DDPM has to be bolted on after the fact, and the paper is
explicit that it doesn't fully solve that side. Working directly in block-ID space,
skipping pixels and learned latents entirely, is what makes that guarantee possible in
the first place. Everything past that point is evidence you should discount
appropriately: FID and DDPM come out statistically tied on the paper's own numbers, the
human study that exists didn't test the paper's headline capability, and the field
around this exact problem got crowded within the same few months this was written. Read
Dream-Cubed for the mechanism and the pictures it produces — both hold up on inspection —
and treat the quantitative claims as a first data point from one preprint, not a result
that's been through the wringer yet.
---
*Built on [Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on
Billions of Cubes](https://arxiv.org/abs/2604.22847) (Tim Merino, Sam Earle, Ryunosuke
Iwai, Julian Togelius, Edoardo Cetin; NYU / Sakana AI, preprint, April 2026) and its
[code and data release](https://github.com/SakanaAI/DreamCubed). Figures are the paper's
own Figures 6 and 7, reproduced for commentary. Tables and numbers are the authors'
except where marked as this site's own illustrative simplification (the block-grid and
schedule-mismatch demos use hand-picked, not trained, values); interactive diagrams are
mine.*
---
# Explorative Modeling: factor the training loop, not the generation loop
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/explorative-modeling
> date: 2026-08-03
> tags: generative-models, diffusion, training-methods, robotics, explainer
Every generative model has to solve the same problem: a prompt like "generate a dog" has billions of
valid answers, not one. [Diffusion](/articles/set-diffusion) handles that by denoising in dozens of
steps; autoregression handles it by predicting one token at a time. Both are the same trick wearing
different clothes — break *generation* into enough small steps that no single prediction has to average
across the billions of valid dogs. **Explorative Modeling** (Gladstone, Ji, and Du — UIUC and Harvard)
asks a different question: what if you left generation alone and factored the *training loop* instead?
Generate $K$ candidate outputs per example, score them all against the target, and only backpropagate
through the winner. They call this **exploration**, argue it is a third axis for scaling generative
models — next to parameters and data — and show it can *substitute* for step-factored generation
entirely, which is what lets them train models that generate in a single forward pass and still match
diffusion.
## Why generation gets factored in the first place
Squared-error regression is maximum likelihood — under a fixed-variance Gaussian. That sounds like a
technicality, but it has teeth: the maximum-likelihood-optimal output under a multimodal target is the
*mean* of every valid answer, and the mean of "every plausible dog photo" is not a photo of a dog. It is
a brown blur. Same failure in language: force a next-token model to average over "the cat sat on the
_____" and you get a smear of probability mass spread across every plausible word, not a committed
answer. The paper gives this capacity a name,
**generative expressivity**: the number of distinct modes a training objective's loss minimizer is
*allowed* to capture. A direct regressor has expressivity $E=1$ — one output, always the average — and
no amount of extra parameters or data raises it, because expressivity is a property of the *objective*,
not the model.
That is the reason diffusion and autoregression look the way they do. Autoregression conditions each
token on everything already generated, so by the time it predicts token $i$, most of the multimodality
is already resolved by the tokens before it — each individual prediction is closer to unimodal.
Diffusion does the analogous thing over noise levels: each denoising step only has to move a slightly
noisy sample a little cleaner, not solve the whole distribution at once. **Factoring generation is a
device for keeping generative expressivity high**, one small step at a time. It is also why training and
inference stop matching: a diffusion model trained on isolated denoising steps gets unrolled over
hundreds of them at test time, and even single-step distillations still anchor their *training* targets
to the multi-step trajectory. Sampling and training are never the same procedure, so exposure bias never
fully goes away.
## Forward XM: buying expressivity with candidates, not steps
If factoring generation is one way to raise expressivity, exploring more candidates is another. Fix a
data target $x$, draw $K$ generations $\hat y_1, \dots, \hat y_K \sim G_\theta$ from the model, and train
only on the closest one:
$$
\mathcal{L}_{\text{Forward}}(\theta) \;=\; \min_{i \in \{1,\dots,K\}} J(\hat y_i, x) \tag{1}
$$
This is the entire mechanism — no new architecture, no new loss family, just a `for` loop around
generation and a `min` before `.backward()`. Explore $K$ candidates with a plain regressor and its
expressivity rises to at least $K$: with enough candidates, one of them lands near enough to any given
mode that the model can commit to it instead of averaging.
Play with $K$ above and the mechanics of equation (1) are exactly what's on screen: every candidate is
scored against the same target, the closest one gets the gradient, and the rest are discarded for this
step. Pulling $K$ up can only tighten the best-of-$K$ error — never loosen it, since you are taking a
minimum over a strictly larger set — but each extra candidate is a full extra generation, so the
compute for a training step scales linearly with $K$. That is the whole cost of exploration, and it is
paid once, during training.
The same de-blurring shows up in the model's actual outputs, not just the training loss:
XM-1 is not a badly-trained model — it is the theoretical best a direct regressor *can* do, the blurred
mean equation (1) predicts. XM-50 is the same architecture, same data, same loss, with one difference:
50 candidates scored per step instead of one.
## Forward and Reverse XM, and what they actually optimize
Forward XM (fix the data, explore the model's generations) is one direction. **Reverse XM** flips it:
fix one generated sample $\hat y \sim G_\theta$ and explore over $K$ data targets $x_1, \dots, x_K \sim
\mathcal D$, training toward whichever is closest:
$$
\mathcal{L}_{\text{Reverse}}(\theta) \;=\; \min_{i \in \{1,\dots,K\}} J(\hat y, x_i) \tag{2}
$$
These are not cosmetically different. The paper works out what each one minimizes in the limit. Write
$p_\theta$ for the model's output distribution blurred by the reconstruction kernel (Gaussian, for
squared error) and $p^*_\sigma$ for the data blurred the same way. Then, in their smooth relaxations:
$$
\underbrace{\mathrm{KL}(p^* \,\|\, p_\theta) + H(p^*)}_{\text{Forward XM}}
\qquad\text{and}\qquad
\underbrace{\mathrm{KL}(g_\theta \,\|\, p^*_\sigma) + H(g_\theta)}_{\text{Reverse XM}}
$$
Forward XM's entropy term, $H(p^*)$, belongs to the *data* — a constant the model can't touch — so
Forward XM is just maximum likelihood over its $K$-candidate mixture, for every $K$. Mass-covering,
never collapsing, but the recall comes at the price of running $K$ full generations per step, so it
struggles to scale to very high-multimodality targets. Reverse XM's entropy term, $H(g_\theta)$, belongs
to the *model* — something it can shrink by narrowing its own spread — so Reverse XM drifts toward
collapse on its own and needs an explicit entropy bonus to stay at the true reverse-KL optimum instead.
The paper is candid that Reverse XM's fix is "largely left for future work"; Forward XM is what every
downstream result in the paper actually runs.
## Substitutable, not just additive
Here is the move that turns this from "a training trick" into "a scaling axis." Factoring generation
exists only to supply expressivity. Exploration supplies the same quantity a different way. If that's
right, the two should be interchangeable — you should be able to trade generation steps for exploration
and land in the same place. The paper tests this directly with **Jumpy** models, a family that
interpolates between direct regression (one jump) and full continuous-time flow (infinite jumps) by
varying the number of steps. Take two Jumpy models, one with fewer jumps and one with more, and add
exploration to both: the model with *fewer* jumps — the more end-to-end one — gains more from
exploration than the one that already had step-factorization doing the work. That is the substitution
effect, measured rather than asserted: the less a model already leans on factored generation, the more
it has to gain from factoring training instead.
Push that trade all the way and you get **XM's other headline**: a model that samples exactly the way
it trained, in one forward pass, with no separate multi-step inference procedure to keep in sync. The
paper calls a model "end-to-end" when it never faces inputs at inference it wasn't trained on — no
denoising schedule to unroll, no exposure bias from a mismatched sampling procedure. This is the same
argument that ended hand-designed feature pipelines after AlexNet, aimed now at the one corner of deep
learning that never fully got the memo: [diffusion language models](/articles/illada-diffusion-language-model)
and [autoregressive decoding](/articles/how-llm-inference-works) both still train on one procedure and
sample with another; exploration is what lets a model close that gap without giving up quality.
The trade is exactly compute, moved to a different place in the pipeline:
[MrFlow](/articles/mrflow-diffusion-acceleration) and [Set Diffusion](/articles/set-diffusion) both
attack the *inference* side of this same step-factorization: reshuffle where a fixed step budget gets
spent, or change which tokens get decoded together, but the sample is still built from many forward
passes. Explorative Modeling is a different lever entirely — it doesn't make the multi-step generator
cheaper, it removes the requirement to be multi-step in the first place, by paying for expressivity up
front instead of on every draw.
## Results: three modalities, two robot benchmarks
**Image generation.** Added to RAE, a near-state-of-the-art ImageNet latent-diffusion recipe, exploration
(XRAE, using XM-2) reaches a near-SOTA **1.43 FID** without classifier-free guidance:
| Method | FID (no CFG) ↓ |
|---|---|
| DiT | 9.62 |
| SiT | 8.61 |
| VA-VAE | 2.17 |
| REPA-E | 1.70 |
| Latent Diffusion + RAE | 1.55 |
| **XRAE (RAE + XM-2)** | **1.43** |
That 47% parameter-efficiency figure is unrelated to the next number, which happens to share a digit:
RAE itself converges 47x faster than SiT (a separate, prior result the paper is building on), and
stacking XRAE's 6.2x sample efficiency on top of *that* puts the whole recipe at roughly **300x faster
to converge than plain SiT** — the paper's arithmetic, not an independent measurement. One negative
result worth keeping: minibatch optimal-transport coupling, an alternative de-blurring trick, made FID
*worse* (46.3 → 54.5 at the Small scale) — exploration wins here specifically, not "adding any anti-blur
trick" generically.
**Scale doesn't dilute the gain — it grows it.** Going from XM-5 to no exploration, the improvement
climbs from 13% to 23% as model size scales up, and from 7% to 36% as data scales up. That is the
opposite of what you'd expect from a scaling axis that's about to run out of room — the paper's reading
is that generative expressivity becomes a *larger* bottleneck as the other two axes get pushed harder,
because parameters and data stop being the limiting factor first.
**Video (Something-Something V2).** FID/FVD improve monotonically with more explored modes. The more
interesting number is generalization, not fit: best achievable FVD is **30.0 with exploration versus
37.5 without** — less overfitting on a fixed dataset, which the paper frames as a compute-generalization
tradeoff: extra training compute spent on exploration buys generalization the way more data usually
does.
**Robot policies (Behavior Cloning, Robomimic).** This is where "single forward pass, matches diffusion"
gets tested against a real baseline:
Explorative Policy matches Diffusion Policy on Lift and Can (both 100%), and beats it on Square (96%
vs. 94%), Transport (74% vs. 72%), and ties on Tool Hang (86%) — at **1 forward pass instead of 100**.
**Goal-conditioned world models (Maze2D), vs. Diffuser:**
Average score edges up too (130.0 vs. 127.2), at 16-256x fewer denoising steps depending on the maze
size (4 vs. 64 on U-Maze, 1 vs. 256 on Medium). This pairing — matching or slightly beating a strong
multi-step baseline, at two orders of magnitude less inference compute per sample — is the article's
headline for a reason: it is the plainest demonstration that the compute the paper claims you save at
inference is compute it actually spent, once, at training.
## What I make of it
- **The conceptual reframe is the real contribution.** "Factor the training loop instead of the
generation loop" is a genuinely different axis, not a repackaging of an existing trick — best-of-$K$
training has appeared before, but treating it as *substitutable* for diffusion/AR step-factorization,
and confirming that substitution empirically with the Jumpy-model ablation, is new.
- **"Third scaling axis" is the authors' framing, argued from one paper's worth of experiments** — real
equations, a real KL derivation, and empirical scaling curves that trend the right way, but not yet a
claim anyone outside this group has stress-tested. Treat it as a strong hypothesis with supporting
evidence, not an established fact.
- **The evaluation is real but narrow at the edges that matter most.** The robotics results — the ones
carrying the "matches diffusion at 100-256x less inference compute" headline — are Robomimic
proficient-human/state-observation behavior cloning and Maze2D planning: small, well-studied
benchmarks, and the paper says outright the control experiments got "barely any tuning." That's stated
as a limitation working in XM's favor (untuned and already competitive), but it also means these are
not yet frontier-scale robot-learning results, and there's no evidence here about vision-based control,
long-horizon manipulation, or real hardware. Gains on autoregressive language models are the weakest
reported of any modality — the paper's own explanation is that next-token prediction is already close
to unimodal, so there's less blur left for exploration to fix.
- **Code is Apache-2.0 and real, which counts in its favor** — `--xm_best_of_k K` is a genuine flag in a
runnable repo, not a promise. But as of this writing the repository explicitly marks the code behind
the headline results as not yet released: the RAE image-generation runs use a separate codebase "to be
released separately," and masked-diffusion-language-model and control-task (robot policy / world model)
code are both marked "coming soon." What's public today is the general XM training scaffolding, not a
drop-in reproduction of the paper's own numbers.
- **No third-party replication yet** — the paper is a July 2026 preprint. The authors are transparent
about a related limit: Diffusion Policy's numbers had to be reproduced under their own setup because
they used a newer Robomimic version than the original paper, so the baseline is a good-faith
re-run, not a quoted number — a small but real point of honesty worth crediting.
The clean way to hold all of this: exploration is a training-time payment for a capability generation
factorization normally buys at inference-time, over and over. That's a real trade, mechanically well
argued, and it works on real benchmarks. Whether it holds at frontier model scale, on harder control
tasks, or once other labs have run the numbers, is still open.
---
*Built on [Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation](https://arxiv.org/abs/2607.27372)
(Gladstone, Ji & Du — UIUC and Harvard, 2026). Code: [github.com/alexiglad/XM](https://github.com/alexiglad/XM)
(Apache-2.0). Figures 1-3 are reproduced from the paper; all numbers are from its Tables 1-3 and Sections 4.1-4.2.*
---
# Fara 1.5: an open, vision-only browser agent at 4B, 9B, and 27B
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/fara-1-5
> date: 2026-08-03
> tags: agents, computer-use, open-weights, vision-language-models
[Qwen-CUA](/articles/qwen-cua), published earlier today, controls a full desktop from screenshots alone with a 397B-A17B mixture-of-experts model. Fara1.5, from Microsoft Research, does the narrower version of the same job -- browser only, no desktop -- at 4B, 9B, and 27B parameters, and ships all three sizes under the MIT license, weights included. That is two orders of magnitude smaller than Qwen-CUA's backbone, scoped to one surface instead of an entire OS, and open in a way Qwen-CUA specifically isn't: Qwen-CUA's repository is Apache-2.0, but its own README states plainly that "model weights are not included" -- paper and demo only. Fara1.5 publishes the checkpoints. The interesting question isn't whether the biggest Fara model is good. It's what the 4B-to-27B ladder says about how much capability a vision-only agent actually needs.
## The same narrow interface, a much smaller model
Fara1.5 is a multimodal decoder-only model built on Qwen3.5, at three sizes, all trained the same way. Given a goal, the current screenshot, and the last three steps of history, it reasons in text and then emits exactly one action from a fixed vocabulary: click, type, scroll, drag, `visit_url`, `web_search`, `go_back`, plus meta-actions for longer horizons -- `memorize` to persist a fact past the three-screenshot window, `ask_user` to pause on a critical point, `finish` to stop. Coordinates are predicted directly from pixels, the same design choice Qwen-CUA makes for an entire desktop: skip the DOM and accessibility tree, and bet that whatever a human can operate with a screen and two input devices is a general enough interface. Fara1.5 just makes that bet at a fraction of the parameter count, and only for the browser.
The safety mechanism is worth naming precisely because it's a real constraint, not a suggestion: the model is trained to trigger `ask_user` at eight defined "critical point" types -- across three dimensions (permission granted or not, task fully specified or not, action reversible or not) -- covering things like entering personal information, submitting payment or shipping details, or sending a message on the user's behalf. Microsoft's recommended deployment wrapper, MagenticLite, is sandboxed and pausable specifically so a human is in the loop at those points. `ask_user` and `memorize` are context-management tools in exactly Lilian Weng's sense from [Agent harnesses](/articles/agent-harness) -- deciding what to carry forward and when to hand control back -- just built into the action vocabulary itself rather than the surrounding scaffold.
## Trained on trajectories it generated for itself
Nearly all of Fara1.5's training data comes from FaraGen1.5, its own synthetic-data pipeline. A solver -- GPT-5.4, paired with a user simulator that withholds task details the way a real user would -- attempts tasks in two kinds of environments: the live, open web, and six sandboxed synthetic apps (Mail, Calendar, Stream, ML, Stay, Scheduler) whose functional code was itself generated by a coding agent (GitHub Copilot CLI) rather than scraped or mocked. Every resulting trajectory then has to clear three independent verifiers before it counts as training data.
That data becomes roughly 2 million training samples, 60% still ordinary open-web trajectories, the rest split across synthetic environments, deliberately ambiguous form-filling, grounding, and a small slice of VQA and drag gestures. It's a real answer to the standard complaint about computer-use training data -- human demonstrations are slow and expensive to collect -- but it's worth being precise about what "generated" means here: the solver, the user simulator, and the verifiers are themselves LLM judgments, not ground truth. A verifier checking "did this trajectory ask before an irreversible action" is exactly as reliable as the model doing the checking.
## What the ladder buys
Fara1.5-27B reaches 72.3% on Online-Mind2Web, ahead of Gemini 2.5 Computer Use (57.3%), OpenAI Operator (58.3%), and Yutori Navigator n1 (64.7%) -- three proprietary systems, all evaluated on an independently maintained academic benchmark, not one Microsoft built. That's a genuine result: an open, MIT-licensed family beating closed competitors on a benchmark none of them control. But it only holds at the top of the ladder.
WebTailBench v1.5 -- Microsoft's own 609-task eval set, worth flagging as self-authored rather than independent -- shows the same monotonic climb, no crossover to check it against:
Read against the predecessor, Fara-7B, the jump looks even sharper: Fara1.5-9B improves +29.3 points on Online-Mind2Web, +13.1 on WebVoyager, +8.3 on WebTailBench, +18.1 on ScreenSpot-Pro grounding, +8.9 on OSWorld-G Refined. That comparison is real but not clean -- it conflates a parameter increase (7B to 9B) with a full training-pipeline change (FaraGen1.5 replacing whatever generated the original Fara's data). Stated as a training-pipeline improvement, it overclaims; stated as "the current generation beats the last one," it's exactly as strong as it sounds and no stronger.
The Online-Mind2Web and WebVoyager comparisons against Operator, Gemini 2.5 CU, and Navigator n1 are self-reported by Microsoft on benchmarks those three systems don't control -- a meaningfully better setup than grading your own exam, but still not an independently run leaderboard. No third-party replication of these specific numbers was found for this piece.
## Where the vision-only bet costs something
The model card is direct about the downsides of skipping the DOM: English-only, vulnerable to visual deception and prompt injection embedded in page content, error accumulation over long multi-step trajectories, and explicitly **not suitable** for legal, health, or financial use. None of that is unique to Fara1.5 -- Qwen-CUA's paper documents the same shape of limitation for the same underlying reason -- but a 4B vision-only model has less capacity to notice something is wrong mid-trajectory than a 397B-A17B one, and the model card doesn't pretend otherwise.
## The take
Two orders of magnitude smaller than Qwen-CUA, scoped to a browser instead of a desktop, and shipping actual weights under MIT where Qwen-CUA ships code and a paper but withholds the checkpoints: Fara1.5 is a genuinely different point in the design space, not a smaller copy of the same idea. The headline -- 27B beats three proprietary computer-use agents on a benchmark none of them own -- is real. The more useful reading of the paper is the ladder underneath it: at 4B, Fara1.5 merely ties the weakest of those three baselines; the win only fully arrives at 27B. Vision-only browser control is not a capability a small model gets for free. It's bought, roughly a third of it per step up the ladder, exactly as parameter-scaling laws would predict.
---
*Built on Microsoft Research's [Fara1.5: Scalable Learning Environments for Computer Use Agents](https://arxiv.org/abs/2606.20785) (Awadallah et al., 2026) and the [microsoft/fara](https://github.com/microsoft/fara) repository (MIT license). Figures 5 and 7 are reproduced from the paper for commentary, flattened onto white; the FaraGen1.5 pipeline diagram and scaling-vs-baseline chart are original illustrations of the paper's Figure 2 and Table 3 / Figure 7 data, not measured traces. Benchmark numbers are as reported in the paper.*
---
# Language models are injective, so a KV-cache is not a summary — it's the prompt, in another basis
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/injective-language-models
> date: 2026-08-03
> tags: llm, interpretability, privacy, theory, kv-cache
Every individual piece of a transformer is lossy. LayerNorm throws away a per-token scale and
shift. Softmax attention collapses many key-value pairs into one weighted average. Low-rank
projections shrink dimensions on purpose. It would be reasonable to conclude that the
hidden state a transformer produces for a prompt is a compressed, irreversible summary of that
prompt — the way a hash or a JPEG is.
[Nikolaou, Mencattini, Crisostomi, Santilli, Panagakis, and Rodolà](https://arxiv.org/abs/2510.15511)
prove that conclusion is wrong. Decoder-only transformer language models, as a whole, are
**almost surely injective**: two different prompts essentially never produce the same
last-token hidden state. Not "usually don't" — provably, with probability one, for any model
trained by gradient descent for any finite number of steps. And they don't stop at the proof.
**SipIt** is an algorithm that exploits injectivity to reconstruct a prompt *exactly* from its
hidden states, in time linear in the prompt's length, with 100% accuracy in their tests. The
paper is ten months old, already accepted at ICLR 2026, and has picked up roughly 30 citations
(7 flagged "influential") in that time — fast uptake for a theory paper.
The result has an immediate consequence worth sitting with: a hidden state, or a cache built
from one, is not a lossy fingerprint of what a user typed. It is that text, in a different
basis, and it can be read back out.
## Non-injective parts, injective whole
The paper's move is to stop looking at individual components and instead look at the whole
map. Write $f: \mathcal{V}^{\le K} \times \mathbb{R}^p \to \mathbb{R}^d$ for the model: a
vocabulary $\mathcal{V}$, a context bound $K$, parameters $\theta \in \mathbb{R}^p$, and
$r(s;\theta)$ for the last-token hidden state of prompt $s$. The claim is that for any two
distinct prompts $s \ne s'$,
$$
\Pr_\theta\big[\, r(s;\theta) = r(s';\theta) \,\big] = 0.
$$
The argument runs through **real-analyticity**. Embedding lookups, affine projections,
softmax, LayerNorm with $\varepsilon > 0$, and every real-analytic activation in common use
(GELU, SiLU, SwiGLU, GeGLU) are all real-analytic functions of their inputs and parameters.
Real-analytic functions are closed under addition, multiplication, division away from poles,
and composition — so the entire network, prompt fixed, is a real-analytic function of
$\theta$.
That matters because of a classical dichotomy. Fix two prompts $s \ne s'$ and define
$h(\theta) = \lVert r(s;\theta) - r(s';\theta) \rVert^2$. Because $h$ is real-analytic, exactly
one of two things is true: either $h \equiv 0$ everywhere, or the zero set
$\{\theta : h(\theta) = 0\}$ has Lebesgue measure zero — a thin, lower-dimensional slice of
parameter space, not a region with any volume. The paper rules out $h \equiv 0$ by hand: it
exhibits one concrete $\theta$ where $s$ and $s'$ provably map to different states (for
instance, freeze the network down to embeddings plus positions and point at two distinct
rows). So the collision set for that pair is measure zero.
That picture is the whole argument. Any initializer with a continuous density — Gaussian,
uniform, Xavier — places exactly zero probability mass on a measure-zero set, so at
initialization the odds of landing on the collision curve are zero. Training doesn't change
that: a gradient step $\varphi(\theta) = \theta - \eta \nabla L(\theta)$ is itself
real-analytic, so its Jacobian determinant $\det D\varphi(\theta)$ is real-analytic and not
identically zero, which makes $\{\theta : \det D\varphi = 0\}$ measure zero too. Away from that
set, the inverse function theorem says $\varphi$ is a local diffeomorphism, and a
diffeomorphism cannot squash a positive-volume region down onto a lower-dimensional set. Push
an absolutely-continuous parameter distribution through enough of these steps — full-batch,
mini-batch, or even adversarially chosen batches — and it stays absolutely continuous. Induct
over any finite training horizon $T$ and injectivity holds with probability one after training,
not only at init. The same argument extends to any finite set of prompts being pairwise
distinct simultaneously, not just one pair at a time.
The theorem is explicit about how a collision *would* have to happen: two vocabulary items
given exactly identical embedding rows, or two positional encodings set exactly equal by hand
while everything else is tuned to suppress positional signal. Both are measure-zero,
hand-engineered pathologies — never reached by continuous initialization plus gradient
descent, but not physically impossible if someone builds them on purpose. That is exactly what
"hand-engineered" does in the picture above.
## What "almost surely" does and doesn't buy you
This is worth being precise about, because it is the obvious objection. "Probability one"
is a statement about a *distribution* over parameters — the set of continuous initializers,
pushed through finitely many gradient steps. It is not a certificate stamped on any one
specific, already-trained, already-quantized checkpoint sitting on a GPU. A model built by
deliberately engineering a collision (identical embedding rows, say) would sit exactly on the
measure-zero set and would not be covered by the "almost surely" — the theorem says that
model is vanishingly unlikely to arise by chance, not that no such model can exist.
The other half of the objection is floating point. The proof is a statement about real numbers;
a GPU computes in fp16, bf16, or fp32 with rounding at every step. The paper's own practical
collision test uses `torch.allclose` with `rtol=1e-5, atol=1e-8` — a floating-point *tolerance*
check, not exact real-number equality. So what gets verified empirically is "no near-collisions
above this threshold," which is a weaker, computational claim standing in for the idealized
real-valued one. The paper doesn't claim its proof covers the discretized forward pass directly
— only the empirical tests do, and only up to that tolerance.
The empirical side is where finite precision gets tested directly. Table 2 measures the
minimum pairwise distance at the final layer under FP4, INT8, and FP32 for three models — for
Llama-3.1-8B: **2.281 (FP4) · 6.597 (INT8) · 1.274 (FP32)**. Quantization didn't shrink the
separation margin in these tests; if anything the coarser formats measured *larger* minimum
distances. That's reassuring, but it's a different computational object than the real-analytic
map the theorem is stated over, checked empirically rather than derived from the proof.
## Zero collisions, at a scale that matters
A separate run — 100,000 prompts sampled from Wikipedia, C4, The Pile, and GitHub Python,
roughly 5 billion pairwise comparisons — measured the minimum pairwise distance across four
different models at three depths (layer 1, the middle layer, the last layer), against a
collision threshold of $10^{-6}$:
| Model | Layer 1 | Layer L/2 | Layer L |
|---|---|---|---|
| Llama-3.1-8B | 0.001 | 0.129 | 0.620 |
| Mistral-7B-v0.1 | 0.002 | 0.187 | 1.274 |
| Phi-4-mini-instruct | 0.014 | 1.336 | 9.020 |
| TinyStories-33M | 0.029 | 1.434 | 2.793 |
Zero collisions, and the separation grows by roughly two to three orders of magnitude from the
first layer to the last — consistent with the boxplot above. The closest pairs the authors
found anywhere, on manual inspection, were near-duplicate code and documentation snippets
differing only by a trailing newline token — still far above the threshold.
Then the authors went looking on purpose. They took the ten closest prompts in their sample and
appended *every* vocabulary token as a one-token continuation to each, an exhaustive
collision hunt rather than a random sample: **over 343 billion prompt pairs per model**. Zero
collisions, on both GPT-2 Small and Gemma3-1B. That is the number in the abstract, and it is
worth being precise about its scope: it is the most exhaustive test in the paper, and it ran
on two of the *smaller* models tested. The bigger models — Phi-4 (14B) and Llama-3.1-70B —
were checked under the sampled protocol above (Table 3), not the exhaustive one; injectivity
at 70B+ scale rests on the same theorem plus a smaller, sampled empirical check, not the
343-billion-pair stress test.
Two more things the tests deliberately probed, both reported candidly rather than
cherry-picked: separation does not shrink as prompts get longer (above), and — a genuinely
counter-intuitive result the authors report without softening it — inverting **random,
out-of-distribution token sequences is faster than inverting natural language** (146s vs.
107s mean, GPT-2, 100-token prompts). Their read: natural-language hidden states sit on a more
structured, clustered manifold, which is flatter and worse-conditioned for the gradient-guided
search described next; OOD states are more dispersed, giving sharper gradients to follow.
One distinction worth making explicit, because it's easy to blur: injectivity is a claim about
the **hidden state**, not about what a model eventually says. Two prompts can produce the
identical next-token answer — "the sum is 12", a translation landing on the same word, a
completion ending in "dog" — while their hidden states remain measurably distinct underneath.
The paper stress-tests exactly this: translation pairs, arithmetic pairs, and ten thousand
different Wikipedia prefixes all forced to the same fixed suffix and the same output token
all still show real, measurable separation at the hidden-state level, growing with depth just
like everything else. Collapsing to the same output is common and expected; collapsing to the
same internal state is what the theorem rules out.
## SipIt: turning the proof into an algorithm
Injectivity is a static fact about the map. **SipIt** (Sequential Inverse Prompt via ITerative
updates) is what you get when you notice the map is also *causal*: the hidden state at
position $t$ depends only on the prefix already fixed and the token at $t$. That means the
one-step map $v_j \mapsto h_t(\pi \oplus v_j)$, for a fixed correct prefix $\pi$ and candidate
$v_j$ ranging over the vocabulary, is itself almost-surely injective — the same argument, run
one position at a time. So the algorithm doesn't need to solve the whole sequence at once:
```
for t = 1..T:
for each candidate v_j (policy: gradient-guided, or random):
if the candidate's predicted hidden state matches h_t within tolerance ε:
append v_j to the reconstructed prefix; move to position t+1
```
**Correctness (Theorem 3.1):** this recovers the true sequence with probability 1, in at most
$T \cdot |\mathcal{V}|$ candidate checks in the worst case — linear in the prompt length $T$
for a vocabulary of fixed size, which is the "linear time" the abstract promises. **Robustness
(Theorem 3.2):** it still recovers the exact sequence under bounded perturbation of the
observed state, as long as the perturbation stays under half the minimum pairwise distance
among candidate continuations at that step — which is exactly the separation margin measured
above, and exactly why that margin *growing* with depth matters practically, not just
theoretically.
In practice SipIt doesn't try candidates in vocabulary order. It uses a **gradient-guided
policy** — clip the gradient norm to 1, periodically re-project the running estimate back to
the nearest true token embedding every 50 proposals — rather than the brute-force random order
its own ablation uses as a baseline. And it is explicit about its threat model: it assumes an
attacker who already holds the **full per-position hidden-state sequence at some layer**
$\ell$ — the paper's own examples are "a leaked KV-cache, a shared-inference pipeline, or an
API exposing intermediate activations." Recovering a prompt from *only* the final embedding is
asserted to be theoretically possible under the same theorem, but no efficient algorithm for
it is demonstrated — that's left as future work. Everything below is about the case SipIt
actually solves: someone already has the hidden states.
Also worth naming: Thomas et al. (2025), the paper's own "most closely related" citation,
recovers prompts from hidden states with a similar sequential structure but
without an injectivity guarantee behind it — so it has to score close to the entire vocabulary
at each step before committing to a token. SipIt's early exit, explored in under a quarter of
one percent of the vocabulary in these experiments, is what the guarantee buys on top of the
same basic idea.
## How fast, and how much of the vocabulary
On 100 prompts (90% real sentences, 10% random tokens), 20 tokens each, GPT-2 Small:
| Method | Mean time (s) | Accuracy |
|---|---|---|
| HardPrompts (gradient prompt search) | 6132.59 ± 104.61 | 0% |
| BruteForce (SipIt, random-order ablation) | 3889.61 ± 691.17 | 100% |
| **SipIt** (gradient-guided) | **28.01 ± 35.87** | **100%** |
HardPrompts — the standard gradient-based *approximate* prompt-search baseline, adapted by the
authors from its original vision-language objective to a text-only one — never lands on the
exact sequence: it optimizes toward *a* prompt, not *the* prompt. Brute-force random search
gets there eventually, at over two minutes an average token. SipIt matches brute force's
accuracy at roughly **1/140th the time**, purely by trying candidates in a smarter order.
That gap holds up against real vocabularies, not just GPT-2's ~50K tokens. Under FP4
quantization, 50 prompts, 10 tokens each:
| Model | Vocab size | Accuracy | Time (s) | Vocabulary explored |
|---|---|---|---|---|
| Mistral-7B-v0.1 | 32,000 | 100% | 111.78 ± 46.50 | 0.19 ± 0.08% |
| Llama-3.1-8B | 128,255 | 100% | 549.48 ± 265.75 | 0.21 ± 0.10% |
The unquantized appendix ablation lands within noise of the same numbers (0.21% and 0.22%
explored respectively) — quantizing the model barely moves how much of the vocabulary SipIt
has to touch. Every measurement here is a single NVIDIA A100-SXM (64GB), no custom kernels
for SipIt itself — the ~28-second figure is an unoptimized, single-GPU number, not a
lower bound on how fast this can go.
## The consequence: a KV-cache is the prompt, in another basis
Put the two halves together. The hidden state is provably (almost surely) a lossless
encoding of the prompt that produced it, and there is a linear-time algorithm that decodes it
back exactly, needing only a sliver of the vocabulary and no training data of its own. That
means the sentence "the model doesn't store your prompt, it just computes with it" is not
true in the way people mean it. The computation *is* the storage. A hidden state is not a
fingerprint or a hash of the input — it's the input, run through an invertible function.
This is precisely what a [KV-cache](/articles/how-llm-inference-works) is built from. Every
serving stack that skips recomputing attention for tokens it has already seen — prefix caching
in vLLM and SGLang, multi-tenant inference sharing a cache across requests, cache offload from
GPU to CPU DRAM or disk when memory is tight, activation logging for debugging or evals — is,
under this paper's result, holding recoverable prompt text, not an opaque compressed artifact.
One nuance worth being exact about: SipIt's proven target is the residual-stream hidden state
$r(s;\theta)$ itself, not the $K$ and $V$ tensors a serving stack actually caches — those are
per-head linear projections of that hidden state, generically lower-dimensional per head. But
stacked across every layer and every head, a full KV-cache is a far *higher*-dimensional
linear image of exactly the same per-position hidden-state sequence the paper's own threat
model names as its motivating example: "a leaked KV-cache, a shared-inference pipeline, or an
API exposing intermediate activations." The paper doesn't run SipIt against raw $K$/$V$
tensors — it inverts hidden states directly — so read "the cache is invertible" as the natural
extension the authors themselves point at, not a number they measured.
A concrete, current example: [Kimi K3](/articles/kimi-k3)'s reinforcement-learning
infrastructure writes idle KV prefixes out to an external CPU DRAM pool between rollouts, so
paused sandboxes stay cheap. Under this paper's result, that pool isn't holding compressed
activations — it's holding recoverable prompt text, sitting outside the GPU's usual trust
boundary, on a different piece of hardware entirely. That's not a criticism specific to K3 —
it's the same design every prefix-cache and cache-offload system makes for the same
performance reasons — it's just a live instance to point at.
And compression doesn't obviously buy you out of this: quantizing the cache, as
[TurboQuant](/articles/turboquant-kv-cache) does for entirely different (memory and
throughput) reasons, doesn't collapse distinct prompts into a shared entry either — Table 2
above shows minimum pairwise distances at FP4/INT8 holding up or growing relative to FP32.
Shrinking the cache for efficiency and erasing what's recoverable from it are different
problems, and solving the first doesn't solve the second.
There's a regulatory angle here too, which the paper raises directly. The Hamburg Data
Protection Commissioner argued in 2024 that a model's trained *parameters* aren't personal
data, because training folds the data into abstract, non-retrievable representations. The
authors' point is narrower and, on their result, correct as far as it goes: that argument is
about parameters at rest, not about **inference-time hidden states**, which this paper shows
are lossless, recoverable encodings of whatever a specific user typed, right now. Any system
that stores or transmits those states — including as a cache — is storing something closer to
the original text than "abstract representation" suggests.
(If you've read the [Jacobian lens](/articles/jacobian-lens) piece: that method also reads
information out of a hidden state, but by linearizing the model around a corpus average — an
approximation. SipIt's guarantee is exact, because it has an injectivity theorem underneath it
instead of a linear approximation.)
## How new is this, and what to weigh
The paper is genuinely recent — submitted October 2025 — but it isn't a fringe preprint sitting
uncited. It's accepted at **ICLR 2026**, and by the time of writing has around **30 citations**,
**7** of them flagged "influential" by Semantic Scholar, which is a fast citation trajectory for
a ten-month-old theory paper. Code (`SIPIT`) is public.
A few things to weigh before taking every number at face value. **Baselines are re-implemented,
not re-run verbatim**: HardPrompts is the authors' own adaptation of a gradient prompt-search
method originally built for vision-language models, ported to a text-only $\ell_2$ objective —
its 0% accuracy reflects that adaptation, not a hostile misreading of someone else's code.
**All numbers are self-reported**; no independent third-party reproduction exists yet at ten
months old. **Every timing number is single-GPU, no custom kernels** — treat "28 seconds" as
this implementation's number, not a hardware-independent constant. And the theorem's scope is
decoder-only transformers with real-analytic activations — the paper surveys 18 widely-used
LLMs and finds all 18 use real-analytic FFN activations (SwiGLU, SiLU, GeGLU, GELU), but a
classic ReLU network sits outside the theorem's direct coverage, since ReLU isn't
real-analytic.
None of that undermines the core claims — billions of comparisons across multiple independent
experimental setups, a working algorithm with two proven theorems behind it, and honest
reporting of the results that don't flatter the paper (OOD prompts inverting faster than
natural language; the biggest models tested under the less exhaustive protocol). It's the
right amount of scrutiny for a result this consequential, not a reason to discount it.
## The take
The intuition that hidden states are lossy comes from staring at individual layers — LayerNorm,
softmax, low-rank projections — each of which really is lossy on its own. The paper's point is
that lossiness doesn't compose the way that intuition assumes: the full map from prompt to
last-token state is, almost surely, injective, and an algorithm exists that inverts it exactly,
in linear time, using a sliver of the vocabulary. The finite-precision and hand-engineered
caveats are real, and worth stating precisely rather than waving away — but they don't touch the
core result, which is that a hidden state is not a summary of a prompt. It's the prompt.
Anything built to store, cache, offload, or ship hidden states around — for speed, for
multi-tenancy, for debugging — is, whether it says so or not, in the business of storing exact
user text.
---
*Built on [Language Models are Injective and Hence Invertible](https://arxiv.org/abs/2510.15511)
(Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis,
Emanuele Rodolà; EPFL / Sapienza University of Rome / University of Athens / Archimedes, Athena
RC; accepted ICLR 2026), and its [SipIt code release](https://github.com/giorgosnikolaou/SIPIT).
Figures are the paper's own Figures 3 and 9, reproduced for commentary. Tables and numbers are
the authors' except where marked as this site's own illustrative simplification (the SipIt
walker's compressed vocabulary); interactive diagrams are mine.*
---
# Inkling-Small: what on-policy distillation actually buys a reasoning model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/inkling-small
> date: 2026-08-03
> tags: llm, mixture-of-experts, distillation, reinforcement-learning, explainer
Thinking Machines' [Inkling](/articles/inkling) shipped with an unusually candid pitch: *"not the
strongest overall model,"* a broad multimodal base meant for fine-tuning. **Inkling-Small** is the
smaller sibling promised in that release, and its own pitch is just as specific: *"an efficient
open-weights model that achieves comparable performance to Inkling at a quarter of its size."* Same
architecture family, same 256-expert MoE backbone, same 1M-token context — but a materially
different post-training story. Inkling-Small was **post-trained from an earlier checkpoint using
on-policy distillation with Inkling as the teacher**, then pushed through **two weeks of scaled
agentic-coding RL**. The result, per Thinking Machines: Inkling-Small now **surpasses Inkling** on
reasoning and agentic coding benchmarks, while Inkling **keeps the edge on knowledge and
factuality**.
That is the interesting part of this release — not the size, the recipe. On-policy distillation is
a genuinely different training signal from ordinary distillation, and it is worth being precise
about why. Along the way there is also a smaller, checkable finding: the parameter counts both
Thinking Machines channels quote for Inkling-Small and Inkling do not match what the released
weights actually contain.
## Same backbone, one size down
Inkling-Small shares its architecture with Inkling almost feature-for-feature — the
[full mixture-of-experts and hybrid-attention design is covered in the Inkling piece](/articles/inkling),
so here is just the shape, read from each model's `config.json`:
| | Inkling-Small | Inkling |
|---|---|---|
| Hidden size | 4096 | 6144 |
| Layers | 42 | 66 |
| Attention heads / KV heads | 32 / 8 | 64 / 8 |
| Sliding-window heads / KV heads | 32 / 8 | 64 / 16 |
| Sliding-window size | 512 | 512 |
| Routed experts | 256 | 256 |
| Active experts / shared experts | 6 / 2 | 6 / 2 |
| MoE expert intermediate size | 2048 | 3072 |
| Dense-layer intermediate size | 16384 | 24576 |
| Context length | 1,048,576 | 1,048,576 |
| Multi-token-prediction heads | 8 | 8 |
Same routing scheme (256 routed experts, 6 active, 2 always-on shared experts, sigmoid router with
post-top-k norm), same [relative position bias](/articles/how-llm-inference-works) and short-conv
mixing, same encoder-free image/audio path, same million-token context. Inkling-Small is a
narrower, shallower cut of the identical design — fewer layers, a smaller residual stream, and a
tighter expert width. What changes is everything downstream of pretraining.
## Off-policy vs on-policy: who generates, who scores
Ordinary distillation — call it off-policy — has the teacher generate the training data. The
teacher produces a sequence of tokens (an answer, a reasoning trace, a full trajectory), and the
student is trained by cross-entropy to reproduce those exact tokens. It is imitation: match the
teacher's output distribution on the teacher's own text.
That works fine for short, single-step outputs. It runs into a specific problem for anything
autoregressive and long — like a chain of reasoning. At inference time the student has no teacher
transcript to fall back on; it has to sample its own next token from its own distribution, condition
on that, sample the next one, and so on. The moment a sampled token differs even slightly from what
the teacher would have produced at that step, the student is in a state its training never covered —
and every token after that is generated conditioned on an increasingly unfamiliar prefix. This is
**exposure bias**: errors compound because the training signal only ever showed the model
teacher-generated prefixes, never its own.
**On-policy distillation** removes the reference trajectory entirely. The *student* generates its
own rollout, token by token, under its own policy — and the teacher's only job is to score the
tokens the student actually produced (as a per-token reward or a log-probability target, depending
on the recipe). There is no teacher transcript to drift away from, because training never showed the
student one. Whatever state the student's own sampling puts it in, that is exactly the state it gets
graded and corrected in. Drag through the two modes below and scrub the rollout step to see the
difference concretely:
This is precisely why on-policy distillation matters more for a reasoning model than for a plain
chat model: a reasoning trace is long, autoregressive, and self-referential — later steps depend
directly on the model's own earlier steps. A student trained only to imitate a teacher's specific
path is fragile exactly where it counts, the moment its own sampling wanders off that path. A
student whose own rollouts are the only thing ever scored has no such cliff to fall off.
Thinking Machines describes the Inkling-Small recipe directly: *"we post-trained an earlier
checkpoint, Inkling-Small (preview), in part using on-policy distillation with Inkling as the
teacher. Starting from that checkpoint, we continued scaling agentic coding RL for two weeks."*
Inkling — the larger, already-trained sibling — is the sole teacher; the smaller model generates,
Inkling grades.
This site has covered two other takes on the same idea, and the contrast is worth naming. [Kimi
K3's post-training](/articles/kimi-k3#post-training-nine-experts-then-one) trains **nine** separate
RL specialists (three domains times three effort levels) and then uses **Multi-Teacher On-Policy
Distillation** to collapse all nine back into one shipped checkpoint — many teachers, all of them
versions of the model itself. [Agents-A1](/articles/agents-a1) does something similar with six
domain specialists, routing each training trajectory to the one teacher that owns its domain.
Inkling-Small's version is the simplest point in that space: **one** teacher, and it is not a
specialist expert of the student — it is a wholly separate, larger, already-shipped model. Same
underlying mechanism (student generates, teacher scores the student's own tokens), different
teacher cardinality and a different relationship between student and teacher.
## Two weeks of RL — read the disclosure level honestly
After the on-policy distillation stage, Thinking Machines says it *"continued scaling agentic
coding RL for two weeks."* Read that number for what it actually specifies and what it does not.
It tells you the wall-clock duration of one training phase. It tells you nothing about cluster
size, GPU count, rollout throughput, number of environments, or total compute — so "two weeks" from
a 64-GPU pod and "two weeks" from a full GB300 NVL72 cluster are the same sentence describing
wildly different amounts of work. [Scaling agentic RL](/articles/scaling-agentic-rl) is mostly an
environments-and-infrastructure problem — verified, reproducible task environments at scale is
usually the actual bottleneck, not algorithm novelty — and none of that infrastructure detail is
disclosed here either: no environment count, no rollout count, no reward model description beyond
"agentic coding." Compare that to Kimi K3's post-training write-up, which at least names concrete
infrastructure numbers (sandbox counts, checkpoint latencies) for its agentic RL stage. Thinking
Machines' own blog names the training hardware for the base models (NVIDIA GB300 NVL72) but not
specifically for this RL phase. Two weeks is a real number and a real signal that the recipe kept
running rather than stopping early — it is just not, by itself, a compute disclosure.
## The parameter count: stated vs measured
Both Thinking Machines channels — the announcement blog and the Hugging Face model card — quote the
same rounded parameter counts for both models:
> "Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active,
> trained on NVIDIA GB300 NVL72 systems." — Thinking Machines blog
> Params (B) (activated/total): Inkling-Small "12/276", Inkling "41/975" — HF model card,
> evaluations table
Fetching each repository's `safetensors` metadata directly from the Hugging Face API
(`api/models/thinkingmachines/{Inkling-Small,Inkling}`, checked 2026-08-03) gives a different
number — the literal count of parameters in the released weight files:
| | Stated (blog + HF card) | Measured (HF `safetensors.total`) | Difference |
|---|---|---|---|
| Inkling-Small | 276B total | **265,956,439,090** (≈265.96B) | +10.04B, ≈3.8% above measured |
| Inkling | 975B total | **952,377,623,626** (≈952.38B) | +22.62B, ≈2.4% above measured |
No accusation implied here — both numbers come from official Thinking Machines channels, and this
is simply what the weight files measure against what both channels quote. It is consistent across
both models and both channels, so it reads as a rounding-and-carry-forward convention rather than a
one-off typo. The active-parameter figures (12B / 41B) cannot be checked the same way — they
describe how many parameters fire per token, which depends on live MoE routing at inference and
cannot be read off static weight metadata. Take those as self-reported.
The **"a quarter of its size"** framing is worth checking on its own terms too. A literal quarter
means Inkling should be 4x Inkling-Small. On the stated numbers, 975 ÷ 276 ≈ 3.53x; on the measured
numbers, 952.38 ÷ 265.96 ≈ 3.58x. Either way, Inkling-Small is closer to 28% of Inkling's size than
25% — "a quarter" is a round-down of a real but smaller ratio, not a precise figure. One more data
point that tracks the same rough ratio: Tinker's stated output pricing is $1.20 per million tokens
for Inkling-Small against $4.05 for Inkling — Inkling-Small at about 30% of Inkling's price, in the
same neighborhood as the ≈28% size ratio.
## Benchmarks — where it wins, and where it doesn't
Thinking Machines' own evaluation suite backs the headline claim: Inkling-Small beats its own
larger sibling on most reasoning, coding, and agentic benchmarks.
That pattern — Small ahead of its own larger sibling, and ahead of every similarly sized open peer
Thinking Machines tested — holds cleanly on SciCode, GPQA Diamond, ARC-AGI-1/2, CritPt, and
Toolathlon Verified. It is not universal: on **SWE-Bench Pro** Inkling-Small (55.9%) sits in a
three-way near-tie, edged out slightly by MiMo V2.5 (56.1%) and Minimax M2.7 (56.2%) even as it
still beats its own sibling Inkling (54.3%). And it does not hold at all on knowledge-recall tasks.
Thinking Machines states that exception directly: *"Inkling maintains an advantage on knowledge
coverage and factuality."* SimpleQA Verified is the sharpest case:
Inkling-Small loses to its larger sibling by more than 23 points here, and a similar gap shows up
on **AA Omniscience** (Inkling-Small −9.0 vs. Inkling +2.1). Both are pure knowledge-recall
benchmarks, not reasoning or agentic ones — exactly where the blog's caveat says to expect the
loss, and the evaluation table backs it up cleanly. The next-largest gaps are **Tau³ Banking**
(15.5% vs. 23.7% — an 8-point, roughly one-third relative deficit) and **FORTRESS adversarial**
safety (71.6% vs. 78.0%); smaller, single-digit-point trails also show up on **AIME 2026**,
**Global-MMLU-Lite**, and the multimodal audio/voice suite. None of these are reasoning or agentic
benchmarks either — they cluster around knowledge, safety-adversarial robustness, and multimodal
recall, consistent with "knowledge and factuality" being the one axis the bigger model still owns.
**Scope the evaluation methodology, not just the scores.** All of this is Thinking Machines
grading its own model against a provider-selected comparison set (Qwen3.5, MiMo V2.5, Minimax
M2.7, DeepSeek V4 Flash, Nemotron 3 Ultra, plus closed models Claude 4.5 Haiku, Gemini 3.5
Flash-Lite, and GPT 5.6 Luna) — a self-report, not a neutral third-party leaderboard. The card
itself discloses several protocol caveats worth carrying forward: SWE-Bench Verified and
Terminal-Bench 2.1 use an **internal, bash-only harness**, and external models' numbers on those
two are self-reported by their own vendors, not run in-house. Terminal-Bench 2.1 also zeroed out
"a small number of solutions... found to be contaminated from web search." HLE-with-tools numbers
for MiniMax M2.7, Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna were run in-house by
Thinking Machines, not vendor-reported. None of this invalidates the results, but it means the
comparison set and the harness are both chosen by the same lab whose model is winning most of the
charts.
## The take
Inkling-Small is a useful data point for a specific question: what does distillation from a bigger
sibling actually buy you, mechanically? The answer here is not "compress the teacher's knowledge
into a smaller container" — Inkling-Small is clearly *worse* at raw factual recall than Inkling,
which is exactly what you would expect if the distillation target was never "know what the teacher
knows." The target was "generate reasoning and agentic trajectories the teacher scores well" — and
on-policy distillation is the mechanism that makes that the actual training signal, because it
grades the student's own rollouts instead of teaching it to imitate someone else's. Layer two weeks
of agentic-coding RL on top of that checkpoint and the result tracks: gains concentrate exactly in
reasoning and agentic coding, and the one place the recipe doesn't touch — static factual
knowledge — is the one place the bigger sibling keeps its lead.
The parameter-count gap is a smaller story, but it is the kind of thing worth checking rather than
repeating: two official channels, one consistent 2.4-3.8% overstatement, verifiable in about two
API calls. None of it changes what Inkling-Small actually is — an Apache-2.0, genuinely open-weights
model that beats its own much larger sibling on most reasoning and coding benchmarks. It is just a
reminder that "check the primary source" is worth doing even when the primary source is the model
card itself.
---
*Sources: the [Inkling-Small announcement](https://thinkingmachines.ai/news/inkling-small/) and the
[Hugging Face model card](https://huggingface.co/thinkingmachines/Inkling-Small) (architecture,
training recipe, evaluations, pricing), cross-checked against the
[Inkling flagship card](https://huggingface.co/thinkingmachines/Inkling) and this site's
[Inkling piece](/articles/inkling). Parameter counts were independently verified via the Hugging
Face API's `safetensors.total` field for both repositories on 2026-08-03, not taken from either
card. All benchmark numbers are Thinking Machines' own, on their own evaluation suite, with the
harness caveats noted inline. Related reading: [Kimi K3's Multi-Teacher On-Policy
Distillation](/articles/kimi-k3#post-training-nine-experts-then-one),
[Agents-A1's domain-routed on-policy distillation](/articles/agents-a1),
[mixture-of-experts from scratch](/articles/mixture-of-experts-from-scratch), and [scaling agentic
RL](/articles/scaling-agentic-rl). Neither Thinking Machines source publishes a static architecture
or benchmark figure for this release — the blog's charts are rendered client-side from inline data,
not static images — so the diagrams here are original illustrations of the mechanism, not
reproductions of a paper figure.*
---
# Instella-MoE: a 16B MoE that never touches an NVIDIA GPU
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/instella-moe
> date: 2026-08-03
> tags: llm, mixture-of-experts, amd, rocm, explainer
Every big open MoE release from the last two years shares one unstated assumption: it was trained
on NVIDIA GPUs. [AMD's Instella-MoE](https://github.com/AMD-AGI/Instella-MoE) breaks that
assumption on purpose. It is a **16B-total, 2.8B-active** Mixture-of-Experts model, and every
stage of it — pretraining, mid-training, long-context extension, SFT, DPO, RL — ran on **AMD
Instinct MI300X and MI325X** GPUs under ROCm, with nothing borrowed from a CUDA cluster. AMD
[released six checkpoints](https://huggingface.co/collections/amd/instella-moe), one per stage,
plus the training, inference, and RL codebases, under the tagline "fully open" — a word this piece
is going to hold them to.
Two things make this worth a full read rather than a spec-sheet skim. First, the architecture has a
real new idea in it: **Gated MLA**, a per-token gate on the attention output, which turns out to be
the same move [Kimi K3](/articles/kimi-k3) landed on independently. Second, the systems story —
**FarSkip-Collective** — is a genuinely interesting trick: it makes MoE training and serving faster
by *deliberately* feeding the model stale, outdated activations. That sounds like a bug. It is the
entire point.
## Six checkpoints, all the way down
Most "open" model releases mean one thing: a final `safetensors` file and a model card. Instella-MoE
releases the **entire pipeline** — a checkpoint at every stage, not just the one you'd chat with:
`Instella-MoE-16B-A3B-Pretrain` → `-Midtrain` → `-Base` → `-SFT` → `-DPO` → `-Think`. Every one of
those six ships with its own weights, training config, and a named, token-counted data recipe. That
is a materially different claim from "we released the weights." Pick a stage below and see exactly
what shipped for it:
AMD draws its own line here, and it is a useful one: in the blog post the fully-open bucket is
**OLMo-3, SmolLM3, and OLMoE** — full data, full code, checkpoints along the way — while
**Moonlight-16B-A3B, Qwen3.5, and Gemma-4** get filed as "open-weight": you get the final weights and
usually a report, not the recipe. Instella-MoE places itself in the first group, and the checkpoint
list above is the receipt.
What is genuinely missing, and it matters: **no training compute figure, no cluster size, no
wall-clock duration** anywhere in the repo or the blog. "Fully open" usually implies you could, in
principle, reproduce the run — and the one number every reproduction attempt needs first is exactly
the one AMD didn't publish.
## The shape of the model
Strip away the training story and here is what actually runs at inference time:
| | |
|---|---|
| Total / active parameters | **16B / 2.8B** |
| Layers | 27 decoder layers |
| Hidden dimension | 2048 |
| Attention | Gated Multi-head Latent Attention (Gated MLA) |
| MoE routing | 2 shared experts (always on) + 6 of 64 routed experts (top-6) |
| Pretraining objective | next-token + Multi-Token Prediction |
| Tokenizer / vocabulary | DeepSeek-V3 tokenizer · 128,896 tokens |
| Context | 4K pretrained → 64K via YaRN + document masking |
| Training frameworks | Primus (Megatron-LM based) · Miles (RL, SGLang + Slime) |
| Inference | SGLang v0.5.9 with FarSkip-Collective overlays |
| Hardware | AMD Instinct MI300X + MI325X, ROCm |
The MoE layer is a shared-plus-routed design: 2 experts run on every token no matter what (the
model's general-purpose knowledge), and a router picks 6 more out of 64 candidates per token — the
same [top-*k* routing](/articles/mixture-of-experts-from-scratch) and shared-expert idea DeepSeek-MoE
popularized, applied at a 16B/2.8B ratio. The pretraining objective adds
[Multi-Token Prediction](/articles/multi-token-prediction) on top of ordinary next-token loss, DeepSeek-V3
style — training the model to predict a short run of future tokens, not just the next one, which both
improves the base model and hands you a natural draft head for speculative decoding later.
### Gated MLA: a per-token filter on attention
Standard Multi-head Latent Attention compresses the KV cache into a small latent vector, then
reconstructs keys and values from it — cheap to store, same attention math otherwise. Gated MLA adds
one more piece: after attention produces its output, a **dedicated linear projection reads the input
token and produces a gate**, one value per channel, and that gate multiplies the attention output
**before** it goes through the final output projection.
Concretely: attention answers "what did this token look up." The gate answers a second, separate
question — "how much of what it found is actually worth keeping" — and answers it per channel, per
token, learned from data. AMD's own framing is direct: the gate lets the model "selectively
attenuate low-utility attention responses for each token." Attention decides what to look at; the
gate decides how much of the answer to trust.
The reason this is worth pausing on: **the same idea shows up independently in [Kimi K3](/articles/kimi-k3)**,
which also augments MLA with an input-dependent output gate — Moonshot's version is a heavier,
full-rank gate; AMD's is a single lightweight linear projection. Two labs, thousands of miles apart,
training on different hardware stacks, converged on "put a learned gate after MLA's output" as a
cheap way to buy expressivity. When two independent teams reach for the same fix, that is usually a
sign the fix is addressing something real in the base mechanism, not a one-off trick.
## FarSkip-Collective: paying with staleness to buy overlap
Here is the problem FarSkip-Collective solves. In expert-parallel MoE training, each MoE layer's
routing decision depends on that layer's own, freshly-computed attention output. Once the router
picks experts, the tokens have to physically move across GPUs to wherever their chosen experts live
— an **all-to-all** collective. That communication cannot start until the fresh activation exists,
and the expert compute that follows cannot start until the communication finishes. Compute and
communication are chained, not parallel, and on a large expert-parallel cluster that chain is
expensive: the GPUs sit idle every time the network is busy, and vice versa.
FarSkip-Collective's fix is to break the dependency that causes the chain. Instead of routing on the
fresh activation, it deliberately routes the MoE (and attention) sub-blocks on an **outdated,
partial activation** — a slightly stale copy of the signal that was already available earlier. Stale
data has one property fresh data doesn't: it's already sitting there, so the communication that
depends on it doesn't have to wait for this layer's compute to finish. It can start **alongside**
that compute instead of after it.
That is the whole trick, and it generalizes past this one model: the separate FarSkip-Collective
paper (Dukler et al., MLSys 2026) reports **97.3%** prefill communication-computation overlap and
**88.9%** training all-to-all overlap, validated on models from 16B up to 109B parameters — including
converting Llama 4 Scout to the FarSkip architecture via self-distillation and landing within **1%**
of the original's accuracy. For Instella-MoE specifically, AMD reports the trade paid off exactly as
advertised:
**+12.7%** pretraining throughput from overlapping expert-parallel communication, and **up to 39.2%
lower Time to First Token** when serving with expert parallelism — a systems win that costs nothing
in serial correctness, because the model is trained end-to-end to expect stale inputs at those
points rather than having staleness bolted on after the fact at serving time.
## Where it lands
AMD ran its evaluations through **OLMES**, Allen AI's open evaluation harness — a real third-party
framework, even though AMD is the one running it. On standard benchmarks, the base checkpoint lands
second among six comparably-sized models, ahead of every "fully open" peer:
Instella-MoE-Base runs at **2.8B active parameters** — less than every model above it except
Moonlight, and well under OLMo-3-7B's 7B dense parameters. It also leads all six on
`WinoGrande` at **86.5**, and posts `HumanEval+` **65.7**, a solid coding number for a base
checkpoint that hasn't seen SFT yet.
Long context is where the honest counterexample lives. At 64K tokens on HELMET and RULER, the
**dense** 7B OLMo-3 actually wins:
| Model | HELMET avg | RULER avg |
|---|---|---|
| OLMo-3-7B (dense) | **43.1** | **80.2** |
| Instella-MoE-Base | 41.5 | 79.4 |
| SmolLM3-3B-Base | 37.6 | 78.6 |
AMD reports this without burying it. A sparse 2.8B-active model narrowly losing to a dense 7B on
long-range retrieval is a plausible, checkable result, not a suspicious one — and it's a useful
reminder that "active parameters" isn't the only variable that determines long-context strength.
After SFT, the post-training funnel adds up:
SFT → DPO → Think is a steady climb, not a single post-training jump, and the RL stage's gain is
concentrated exactly where you'd expect from an instruction-following-focused reward: `IFEval` moves
from **77.08** after DPO to **83.70** after RL, the single largest jump on the sheet. The RL recipe
itself is worth a note for anyone who followed [Rollout Routing Replay](/articles/rollout-routing-replay):
Instella-MoE's IF-RL stage uses R3 alongside GRPO/DAPO-style tricks (zero-gradient filtering, active
sampling, token-level loss, no KL term) — the same fix for MoE-RL's rollout/training routing mismatch,
here as one ingredient in a larger recipe rather than the whole story. A second RL stage,
Multi-Teacher On-Policy Distillation, then anchors the model back to both an IF-specialized teacher
and the DPO checkpoint, so the instruction-following gain doesn't come at the cost of the math and
code ability DPO already had.
## What's still closed
"Fully open" is doing real work as a claim here, so it deserves the same scrutiny as any benchmark
number.
**What's released, precisely:** all six stage checkpoints on Hugging Face; the full training
codebase (`Primus`-based, Megatron-LM lineage) under **MIT**; the inference overlays and RL codebase
(`Miles`, built on SGLang); per-stage YAML configs; and named, token-counted data mixtures for every
stage, down to the individual sub-corpus (`Nemotron-CC-Math-v1`, `Dolma3 Dolmino`, `cranecode`,
`instella-gsm8k-synthetic`, and dozens more).
**What isn't released:** the model **weights** carry a **Research RAIL license** — academic and
research use only, not the permissive MIT the code ships under. So the license itself is split: the
code is as open as it gets, the weights are open-but-restricted. Beyond licensing, three things are
just missing: training compute and cluster size (undisclosed anywhere), wall-clock training
duration, and a peer-reviewed technical report — the citation AMD gives today is for a *different,
earlier* dense 3B Instella model plus the separate FarSkip-Collective systems paper, not an
Instella-MoE-specific report. The blog itself says the report is "coming soon."
There's a methodology gap worth naming too: it isn't stated whether the comparison models' scores
(Qwen3.5, Gemma-4, Moonlight, SmolLM3) were re-run by AMD under the same OLMES harness, or taken from
those models' own published numbers. Either is defensible, but the blog doesn't say which, so treat
cross-model comparisons as directionally trustworthy rather than exactly apples-to-apples.
AMD is candid about the rest: the models are released "for research purposes only," explicitly not
intended for "safety-critical applications" or "health and medical applications," shipped "without
any safety promises," and multilingual ability "has not been tested." That's an unusually direct
limitations section for a benchmark-forward launch blog, and it's worth taking at face value rather
than reading past it.
## What training off NVIDIA proves, and what it doesn't
The part of this release that will get repeated the most is also the simplest to state: a
competitive, 16B-parameter MoE, trained through every post-training stage including RL, ran entirely
on AMD Instinct hardware. That is a real data point. It says the ROCm software stack — Primus for
pretraining, Miles and SGLang for RL and serving, FarSkip-Collective's overlays making expert
parallelism efficient on this hardware specifically — can carry a full modern LLM pipeline, not just
a pretraining demo.
It does not say AMD hardware is cheaper, faster, or as mature to develop against as the CUDA
ecosystem for this workload — none of the numbers that would let you compute a cost-per-FLOP or a
wall-clock comparison are published. It does not say the result would look the same at 10× the
scale. And it's one vendor's own benchmark of its own model on its own hardware, run through a
credible third-party harness but not independently reproduced elsewhere yet. What it *does* rule out
is the null hypothesis that this can't be done at all outside NVIDIA — six checkpoints and a working
RL pipeline are hard to argue with on that specific point, even while the cost question stays open.
## The take
Two real ideas, evaluated honestly, and a rare complete-pipeline release: Gated MLA is a cheap,
convergent fix (the same one Kimi K3 found independently) for getting more out of MLA's compressed
attention; FarSkip-Collective is the more interesting systems idea, because it's not "make the
network faster" — it's "make the network's timing not matter" by feeding the model activations that
are already a step behind, on purpose. Together with a genuinely complete checkpoint trail —
pretrain through RL, not just a final drop — Instella-MoE beats every other "fully open" peer AMD
names and trails only a larger, open-weight-only Qwen3.5-4B. The unresolved part is exactly the part
AMD chose not to publish: what the whole run cost, on how many GPUs, for how long. Until the
technical report lands, that number stays the reader's to estimate, not AMD's to claim.
---
*Sources: the [Instella-MoE GitHub repository](https://github.com/AMD-AGI/Instella-MoE) (architecture,
training stages, license, data preparation), the
[ROCm technical blog](https://rocm.blogs.amd.com/artificial-intelligence/instella-moe/README.html)
(benchmarks, FarSkip-Collective and Gated MLA framing, figures), the
[Hugging Face model collection](https://huggingface.co/collections/amd/instella-moe) (six checkpoints),
and the [FarSkip-Collective paper](https://arxiv.org/abs/2511.11505) (Dukler et al., MLSys 2026 — overlap
percentages, cross-scale validation, Llama 4 Scout conversion). Figures reproduced here are the blog's
own Figures 1–3. Benchmark numbers are AMD's, via OLMES; the training-cost and cluster-size figures this
piece flags as missing are missing because AMD has not published them, not because they were left out
here. Interactive diagrams are mine; the FarSkip timeline and gate values are illustrative, not measured
traces.*
---
# JOSIE-2: a 4M-token fine-tune, and a labeling bug in its own benchmark table
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/josie-2
> date: 2026-08-03
> tags: open-weights, llm, fine-tuning, explainer
[JOSIE-2](https://huggingface.co/collections/Goekdeniz-Guelmez/josie-2) is Gökdeniz Gülmez's third-generation personality-tuned model family: three sizes — 2B, 4B, 9B — each fine-tuned from the matching Qwen3.5 base, MIT-licensed, trained entirely on Apple Silicon. The release note makes a specific, checkable claim: *"JOSIE-2-2B-OSS consistently outperforms its 4B base model. JOSIE-2-4B-OSS consistently outperforms its 9B base model."* A 2B model beating a 4B, and a 4B beating a 9B, would be a genuinely notable result. I went and checked it against the model cards and `config.json` directly, because that's the kind of claim worth verifying before repeating.
## What the cards actually say
Each `JOSIE-2--OSS` repo carries a `base_model` field in its frontmatter and a `model_name` field in `config.json`. Both agree, on all three repos: **2B is fine-tuned from Qwen/Qwen3.5-2B, 4B from Qwen/Qwen3.5-4B, 9B from Qwen/Qwen3.5-9B.** Same size, every time. There is no size-up training anywhere in the released weights — the release note's "outperforms its 4B/9B base" framing describes a comparison that, per the models' own configs, never happened.
That mismatched "4B" badge on a row that says "Qwen3.5-2B" is the same bug in miniature. The bigger version of it lives in the part of each card that Hugging Face's benchmark widget actually reads — the `model-index` YAML. There, **every one of the three cards labels its baseline row `Qwen/Qwen3.5-4B (base)`, including the 2B and 9B cards**, even though the numbers next to that label differ card to card (82.5/49.1 on the 2B card, 83.4/48.9 on the 4B card, 92.6/69.5 on the 9B card) in a way that only makes sense if each card's numbers really are its own base model's, mislabeled:
Toggle between what's published and what `config.json` verifies, and the numbers don't move — only the caption does. That's what makes this read as a copy-paste templating bug rather than a fabricated result: the underlying scores look real and internally consistent per-card, but the machine-readable label attached to them is wrong on two of three cards. The 9B card has a second, unrelated gap: its own reasoning-mode benchmark row is simply blank, published as "comming soon" in the card's own chart — the one number that would most directly support "the 9B model in its best mode," and it doesn't exist yet.
I want to be fair to the release here: this reads as a labeling artifact, not a fabricated claim. The scores are plausible and self-consistent within each card. The problem is narrower and more mundane — a shared table template where the label field wasn't updated per model, which happens to be exactly the kind of error you'd only catch by checking `config.json` against what the benchmark table says, which almost nobody does.
## What's actually worth taking seriously
None of this makes the underlying work uninteresting. The whole JOSIE-2 family — three sizes — was fine-tuned on **the same dataset**: roughly 3,500 samples, about 4 million tokens total, generated by a pipeline that itself leaned on larger models (Gemma 4 31B, Qwen3.5-9B-Base, GPT-5.4, and an unreleased JOSIE-2-35B-A3B-RTG model) to synthesize training data far more capable than the 4M-token set alone would suggest. That's a striking ratio: a few thousand curated examples, reused across three model sizes, apparently doing real work — the ARC-C and TruthfulQA gains over each model's *actual, same-size* base are consistent and positive across all three sizes, even once you correct the label.
The more genuinely interesting finding in the cards is emergent, not benchmarked: JOSIE-2's reasoning traces sometimes "roast" or insult the user mid-thought, and the cards are specific that *"no reward was introduced to enforce a uniformly polite, corporate, or sanitized internal monologue"* — the behavior wasn't trained in, it showed up during RL because nothing trained it out. It's confined to the hidden reasoning trace, not the user-facing reply, and the cards are candid that they can't yet separate "reasoning-first supervision improves the policy" from "the training data's honesty framing drives the gain" — an open question, stated as one.
## The take
Check the claim, not just the headline: a 2B model plausibly outperforming its own 2B base and a 9B model plausibly outperforming its own 9B base is still a real, useful result from a tiny dataset — it just isn't the size-up story the release note tells, and the cards' own tables currently say something they don't mean to say. That's worth fixing on Gülmez's end, and worth checking on everyone else's before the "2B beats a 4B" framing gets repeated as fact.
---
*Sources: the [JOSIE-2 collection](https://huggingface.co/collections/Goekdeniz-Guelmez/josie-2) and individual model cards (`Goekdeniz-Guelmez/JOSIE-2-2B-OSS`, `-4B-OSS`, `-9B-OSS`), their `config.json` files, and each repo's `benchmarks.png`, cross-checked directly for this piece. The figure is the 2B card's own chart, unedited aside from flattening onto a white background; the interactive reproduces the model-index labels and config-verified bases as published.*
---
# KDA has a half-life: linear attention forgets like a radioactive isotope
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/kda-half-life
> date: 2026-08-03
> tags: linear-attention, kimi, attention, explainer, math
Here is a small thing I keep turning over. The forgetting mechanism inside **Kimi Delta Attention** — the linear
attention in [Kimi K3](/articles/kimi-k3) — is *the same mathematics as radioactive decay*. Not "reminiscent of",
not "a useful analogy". The same two-line derivation, with tokens where a physicist writes seconds.
## Two laws that are one law
A radioactive sample loses a fixed *fraction* of its remaining atoms per unit time. That gives the exponential
law everyone meets in school:
$$
N(t) = N_0 e^{-\lambda t}
$$
A KDA channel loses a fixed *fraction* of its remaining state per token. Take the recurrence and strip it to the
decay term — set the write strength to zero and watch a stored value with no new input arriving:
$$
S_n = \alpha^{\,n} S_0
$$
Those are the same function. Since $\alpha = e^{\ln \alpha}$,
$$
\alpha^{\,n} = e^{n \ln \alpha} = e^{-\lambda n}, \qquad \lambda = -\ln \alpha
$$
The retention factor $\alpha$ and the decay constant $\lambda$ are two spellings of one number. A channel with
$\alpha$ close to 1 is a long-lived isotope; a channel with small $\alpha$ is one that barely outlives its own
creation.
## So it has a half-life
Once you accept that, the half-life comes for free. Ask for the $n$ where half the signal is gone:
$$
\alpha^{\,n_{1/2}} = \tfrac{1}{2}
\quad\Longrightarrow\quad
n_{1/2} \, \ln \alpha = \ln \tfrac{1}{2}
\quad\Longrightarrow\quad
n_{1/2} = \frac{\ln 0.5}{\ln \alpha}
$$
That is the whole result, and it is worth internalizing because it converts an opaque hyperparameter into a
number with units you can reason about. **α = 0.99 gives a half-life of about 69 tokens.** Not "some decay" —
sixty-nine tokens, roughly a long sentence. Drag it:
The lever is brutally nonlinear near 1, which is the part worth feeling rather than reading. Going from
α = 0.99 to α = 0.999 does not extend memory by a tenth of a percent; it multiplies the half-life by ten, from
about 69 tokens to about 693. Each additional nine buys another factor of ten. That is why linear-attention
gates are usually parameterized in log space — the useful resolution all lives in the last few decimal places,
and a linear parameterization would spend nearly all its range on channels that forget immediately.
A useful sanity check: half-life is a property of the *ratio*, not the magnitude. A channel at α = 0.99 has lost
half its signal after 69 tokens, three quarters after 138, and about a thousandth of it survives to 690 — ten
half-lives, the same "ten half-lives and it's gone" rule of thumb used for isotopes.
## The interesting part: α is per channel
If KDA had one global α this would be a cute observation and nothing more. It doesn't. In K3, α is a
**channel-wise** vector — the report writes the state update as
$$
S_t = \left(I - \beta_t k_t k_t^{\top}\right) \mathrm{Diag}(\alpha_t)\, S_{t-1} + \beta_t k_t v_t^{\top}
$$
where $\alpha_t \in (0,1)^{d_k}$ is a **per-channel** one-step retention factor and $\beta_t$ is the delta-rule
write strength. `Diag(αₜ)` is the load-bearing notation: every one of the $d_k$ channels gets its own decay
constant, so a single head carries a whole spectrum of half-lives simultaneously.
This is what makes a fixed-size state genuinely useful rather than merely cheap. The head is not choosing
between "remember recent things sharply" and "remember old things vaguely" — it runs both at once, on different
channels. The fast channels behave like a local window: they hold the current clause and dump it. The slow
channels are closer to a running summary that survives the entire context. Attention over a KV cache gets its
long-range recall by *storing everything*; KDA gets a version of it by storing a small number of things at
deliberately different rates.
It also reframes what "training the gate" means. The model is not learning *whether* to forget. It is learning a
distribution of timescales — effectively allocating channels across memory horizons, the way a filter bank
allocates across frequencies.
## What K3's config actually pins down
The released weights make a couple of things concrete. K3 runs **69 KDA layers out of 93**, three of every four,
with a Gated MLA layer as the fourth — so most of the model's sequence mixing is this decay process, and the
full-attention layers are the periodic exact-recall anchor. Head dimension is 128, and the gate is full-rank
(`use_full_rank_gate: true`) rather than a low-rank approximation — though see the update below for exactly how
the per-channel variation is produced.
The config also carries `gate_lower_bound: -5.0`. Read as a floor on log-α, that bounds the fastest a channel is
allowed to forget: $\alpha \ge e^{-5} \approx 0.0067$, which is a half-life of about **0.14 tokens** — a channel
that has essentially dumped its state by the very next step. The ceiling is the interesting end and it is open:
as α approaches 1 the half-life grows without bound. To keep half your signal across a full 1M-token context you
need α ≈ 0.99999931. That number has seven leading nines, which is exactly why the bound is expressed in log
space.
**Update, 2026-08-03: the inference above is confirmed.** I originally flagged the log-α reading of
`gate_lower_bound: -5.0` as an assumption — the config does not state the functional form, and the report does not
either. [kimi-k3-in-c](/articles/kimi-k3-in-c), an independent C99 reimplementation, computes the gate exactly
that way:
```c
const float a = expf(A_log[h]); /* per HEAD */
const float u = a * (z[i] + dt_bias[i]);
const float gi = lb * sigmoidf_(u); /* in (lb, 0] -> this is log alpha */
alpha[i] = expf(gi); /* in (e^lb, 1] */
```
With `lb = -5.0`, α is bounded to $(e^{-5}, 1] \approx (0.0067, 1]$ — the 0.14-token floor holds.
One refinement the code makes that the config alone did not: `A_log` is stored **per head**, and the per-channel
variation comes from the `z + dt_bias` term inside the sigmoid. So a channel's decay is a per-head base rate
modulated per channel, rather than a fully independent per-channel parameter. The implementation carries a pointed
warning about this — the checkpoint stores `head_dim` floats but only the first `H` are nonzero, so indexing
`A_log` per channel is *"a silent, fatal error"*. The `Diag(αₜ)` structure and the resulting spread of timescales
are unaffected.
## Why this is more than a nice analogy
Two things fall out of it that are practically useful.
**It gives you a unit.** "The gate decays the state" is unfalsifiable prose. "This channel has a half-life of 69
tokens" is a claim you can check against a model's behaviour — and it tells you immediately that a channel with a
7-token half-life cannot be the thing carrying a fact across a document, no matter what the attribution heatmap
suggests.
**It explains the parameterization.** Every design choice around these gates — log-space parameterization,
bounded gates, careful initialization near 1 — follows from the shape of $n_{1/2} = \ln 0.5 / \ln \alpha$. The
function is nearly flat for most of $(0,1)$ and then explodes in the last sliver. Any scheme that samples α
uniformly wastes almost all of its capacity on channels that forget within a few tokens.
The same algebra runs through every gated linear-attention variant, not just KDA — Mamba's $\bar{A}$, the decay
in RetNet and RWKV, the forget gate of an LSTM. They differ in how α is produced and whether it depends on the
input. They agree on the underlying law, which has been sitting in physics textbooks the whole time.
---
*Sources: the [Kimi K3 technical report](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf) for
the KDA recurrence and the hybrid layer composition, and the released
[Kimi K3 `config.json`](https://huggingface.co/moonshotai/Kimi-K3) for the layer split, head dimension,
`use_full_rank_gate` and `gate_lower_bound`. The half-life framing and the derivation are mine; the channel
α values in the spectrum widget are illustrative, chosen to span the range, while every half-life shown is
computed exactly from them.*
---
# kimi-k3-in-c: 2.78 trillion parameters, one CPU, 8 GB of RAM
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/kimi-k3-in-c
> date: 2026-08-03
> tags: llm, inference, quantization, kimi, systems, c, explainer
```console
$ ./bin/k3 ~/k3model --trunk ~/k3trunk --preset laptop \
--tok ~/k3model --prompt "The capital of France is" --gen 8 --incremental
--- generated text ---
Paris.",
+ "The Eiffel
----------------------
8 tokens in 261.5 s, 32.69 s/token average
PEAK RSS for the whole run: 8.24 GB
```
That is [Kimi K3](/articles/kimi-k3) — 2.78 trillion parameters, 1.56 terabytes on disk —
answering a question correctly from a laptop-sized memory budget, on one CPU, with zero GPUs. The
engine that did it is [kimi-k3-in-c](https://github.com/FareedKhan-dev/kimi-k3-in-c): 4,779 lines of
portable C99, a single author, Apache-2.0. It is slow — 32.69 seconds for that one token — and it is a
base model, so `" Paris."` is a continuation of a sentence, not a chat reply. Neither of those facts
changes what the run demonstrates: a frontier-scale mixture-of-experts model, read and multiplied
straight off an NVMe drive, fits in memory most people already own.
I went through the source rather than just the README — `src/core/k3_ops.c`, `src/cache/k3_cache.c`,
`src/io/k3_st.c` — because the interesting claims in a project like this live in code comments, not
marketing copy, and this codebase's comments are unusually good. This piece is about what they say.
## The four reductions
Every parameter at bfloat16 is **5,560 GB**. That number is where the project starts and it is not
close to fitting anywhere. Four decisions, each one shipped in the checkpoint or written into this
engine, bring it down to a measured **8.24 GB** — a 675× reduction, with no weight dropped and no
approximation:
1. The routed experts already ship at **half a byte** — MXFP4 — multiplied straight out of their packed
form, never expanded to floats first.
2. **KDA** gives 69 of the 93 layers a recurrent state that does not grow with context.
3. **MLA** caches one 576-wide latent per position instead of ninety-six heads of key and value.
4. The dense **trunk streams** a layer at a time instead of sitting resident, which turns the last
floor into a dial.
The first three are architectural decisions Moonshot made when training Kimi K3; I covered why they
exist — the routing, the gating math, the attention-residual stack — in
[the K3 architecture piece](/articles/kimi-k3). What is new here is the fourth one, which belongs
entirely to this engine, and the fact that all four survive being reimplemented from scratch in C and
checked against the released weights.
## Reduction one: the experts already ship at half a byte
The routed experts are not quantised by this engine — they arrive from Moonshot already in **MXFP4**,
a microscaling 4-bit float. Every weight is a 4-bit code indexing a 16-entry table, and every 32
consecutive weights share one 8-bit exponent. One routed expert is exactly 33,030,144 parameters. At
half a byte plus a shared scale, that is **17,547,264 bytes** — 17.55 MB. Dequantised to fp32 first, the
same expert is 132 MB.
A token touches 16 experts in each of 92 MoE layers — 1,472 experts. Multiply that out and the
difference stops being an abstraction:
- Dequantise everything first: **194 GB** of pure format conversion, per token, before one multiply
happens.
- Read the nibbles directly: **25.83 GB**.
The comment above the kernel that does this is the best sentence in the codebase: *"This is not an
optimisation; it is what makes streaming experts possible at all."* And the mechanism is worth sitting
with, because it inverts an intuition every ML engineer carries: quantisation is supposed to trade
compute time for memory. Here it does the opposite. A matrix-vector product is memory bound — the
arithmetic is cheap, the wait is for bytes to arrive — so reading 7.5× fewer bytes makes the
packed kernel **faster** than dequantise-then-multiply, not slower.
```c
/* y[rows] = W[rows][in] . x[in], with W read straight out of packed MXFP4 and never
* materialised as floats. This is not an optimisation; it is what makes streaming
* experts possible at all. */
void k3_matmul_mxfp4(float *y, const float *x, const unsigned char *packed,
const unsigned char *scales, int in, int rows, int group)
{
...
for (int r = 0; r < rows; r++) {
for (int g = 0; g < ngrp; g++) {
const unsigned char sb = sr[g];
if (sb == 255) continue; /* NaN scale: contribute nothing */
/* expand the group of 32 to floats, dot product, then apply ONE scale */
...
acc += sub * (double)K3_E8M0[sb];
}
y[r] = (float)acc;
}
}
```
The reason the inner loop is fast is a small lookup table: `K3_E2M1_PAIR[256][2]` maps a whole byte to
both of its decoded values, so the loop does one 8-byte load instead of masking and shifting each
nibble out separately. Groups of 32 elements are exactly 16 packed bytes, which is why the group size
was chosen there — and the scale factors out of the inner sum entirely, applied once per group instead
of once per weight.
## A floating-point contract, not a convention
Kimi K3's headline claim is stronger than "it runs small": *"the same model runs in 8 GB and in 224 GB
and produces byte-identical output at every budget between."* Not close. Identical. That does not
happen by accident in floating point, because addition is not associative — `(a + b) + c` and
`a + (b + c)` round differently once you are past a handful of terms, and this engine sums thousands of
them per output, across a scalar path, an OpenMP path, and an AVX2 path, at any thread count.
`k3_matmul` fixes the order rather than trusting the compiler with it:
```c
void k3_matmul(float *y, const float *x, const float *W, int in, int out)
{
for (int o = 0; o < out; o++) {
const float *row = W + (size_t)o * in;
double a0 = 0.0, a1 = 0.0, a2 = 0.0, a3 = 0.0;
int i = 0;
for (; i + 3 < in; i += 4) {
a0 += (double)row[i ] * (double)x[i ];
a1 += (double)row[i + 1] * (double)x[i + 1];
a2 += (double)row[i + 2] * (double)x[i + 2];
a3 += (double)row[i + 3] * (double)x[i + 3];
}
double acc = (a0 + a1) + (a2 + a3);
for (; i < in; i++) acc += (double)row[i] * (double)x[i];
y[o] = (float)acc;
}
}
```
Four accumulators, partitioned by `i % 4`, reduced as `(a0 + a1) + (a2 + a3)` — written out by hand
rather than left for the compiler to vectorise however it likes, because *that specific split is the
summation order the AVX2 path must reproduce exactly*. A `__m256d` register holds four doubles; loading
four elements per iteration places element `i` in lane `i % 4`, the same partition as the scalar
accumulators, and reducing with the same parenthesisation gets the same bits back. Two more details
carry the guarantee: the accumulators are **double**, because a float accumulator loses precision the
comparisons can see at hidden size 7168; and the AVX2 code uses a separate multiply and add, never a
fused multiply-add, because `-ffp-contract=off` means the scalar path rounds twice and an FMA rounds
once — a hardware capability that would otherwise quietly change the output. `k3_matmul_bf16` mirrors
the same layout for the bf16 trunk, so all three paths — scalar, OpenMP, AVX2 — agree to the bit. And
because output rows never depend on each other, threading the outer loop changes nothing about the
arithmetic either: every row is still summed by exactly one thread in exactly this order, so the result
is identical at any thread count.
That determinism is also honest about its one exception. `k3_matmul_mxfp4` is *not* bit-identical to
dequantise-then-multiply, and the comment says so without hedging: it sums each group of 32 under its
own accumulator and applies that group's scale before combining groups, while a plain matmul sums the
whole row under one accumulator — a different order. But the bound on that difference is derived, not
asserted. Every individual product inside a group is **exact** in double: an E2M1 value carries 3
mantissa bits, `x` carries 24, the product needs 27 of the 53 bits double has to spend, so only the
additions round at all. Reassociating exact terms moves the result by about one unit in the last place
of a double — roughly `1e-16` relative. The test gate requires agreement to `1e-6`. The margin between
what the reordering actually costs and what the test demands is nine orders of magnitude. A codebase
that states plainly where it is *not* exact, and then bounds how far off, is doing something most
numerical code does not bother to do.
## KDA in the code
I wrote about [KDA's decay as a half-life](/articles/kda-half-life) from the technical report alone,
and had to *infer* that the forget gate is parameterised in log-alpha space — the report gives the
mechanism but not the exact functional form, so I flagged the parameterisation as unverified. This C
source confirms it outright. `k3_kda_decay` computes the gate per head, then folds it per channel:
```c
void k3_kda_decay(float *g, float *alpha, const float *z, const float *A_log,
const float *dt_bias, int H, int D, float lb)
{
for (int h = 0; h < H; h++) {
/* PER HEAD. The checkpoint stores head_dim floats but only the first H are
* nonzero. Indexing this per channel is a silent, fatal error. */
const float a = expf(A_log[h]);
for (int d = 0; d < D; d++) {
const int i = h * D + d;
const float u = a * (z[i] + dt_bias[i]);
const float gi = lb * sigmoidf_(u); /* in (lb, 0] */
g[i] = gi;
alpha[i] = expf(gi); /* in (e^lb, 1] */
}
}
}
```
`gi = lb * sigmoid(u)` is log-alpha directly, and `alpha[i] = expf(gi)` is exactly the exponential I
had to guess at from the outside. With the checkpoint's `gate_lower_bound` of −5, alpha lands in
`(e^-5, 1]`, about `(0.0067, 1]` — a per-channel retention factor, with a per-*head* base rate (`A_log`
is indexed by `h`, not by the channel index `i`) modulating it. An independent reimplementation
confirming an inference from the outside is a satisfying result on its own, and it is the reason these
two pieces belong read together.
The comment on `A_log` is worth pausing on for a second reason: it is one of five *invariants* the
codebase states up front as places a plausible-looking implementation silently produces the wrong
model — no crash, no NaN, just a different function that still writes fluent English. `A_log` being
per-head rather than per-channel is invariant one.
The recurrence itself is the delta rule, in four stages that the comments number:
```c
void k3_kda_step(float *S, float *o, const float *q, const float *k,
const float *v, const float *alpha, float beta, int dk, int dv)
{
/* 1. channel-wise decay: scale ROW i of S by alpha[i] */
for (int i = 0; i < dk; i++) { ... }
/* 2. read the state along k: u = S^T k */
...
/* 3. rank-one delta write. (v - u) is the prediction error: this is what makes
* it a DELTA rule rather than plain accumulation. */
for (int i = 0; i < dk; i++) {
const float ki = k[i];
float *row = S + (size_t)i * dv;
for (int j = 0; j < dv; j++) row[j] += ki * beta * (v[j] - u[j]);
}
/* 4. output from the ALREADY UPDATED state: o = S^T q */
...
}
```
Decay the state, read from it along the key, write back the *error* between the value and what the
state already predicted — not the value itself — then read the output from the state that write just
produced. That third stage is what turns a running sum into a rule that corrects itself: writing `v`
directly would just accumulate; writing `v - u` writes only what the state did not already know. Step 4
reading from the post-write state, not the pre-write one, is the second place a plausible-looking bug
hides with no visible symptom.
## The cache the project exists for
The dense trunk is 108.81 GB and every layer of it runs on every token — nothing to skip there, so it
streams from a packed file with a pinned prefix and one rotating ring slot. The routed experts are the
opposite kind of problem: 1.45 TB of the 1.56 TB checkpoint, and only 1,472 of the 82,432 experts fire
per token. The header comment on the cache that handles them does not undersell its importance: *"This
is the part the project exists for."*
Left uncached, one decode step reads 25.83 GB of experts. At the roughly 1.2 GB/s a commodity NVMe
device sustains on cold random reads of that size, that alone is about **21 seconds per token** from
storage. The cache holds those experts in the same MXFP4 bytes the matmul consumes directly — caching
dequantised floats would cut the number of experts that fit by 7.5× for nothing, since nothing
downstream ever wants the expanded form.
The replacement policy is LRU with pinning, and the victim search is a plain linear scan, on purpose:
```c
/* Least recently used unpinned slot. Linear, deliberately: a few hundred comparisons
* against a 17.55 MB read is not where the time goes. */
static int pick_victim(K3Cache *c) { ... }
```
A few hundred integer comparisons next to a 17.55 MB disk read is not a place worth a heap. And the
cache keeps a request histogram — 82,432 counters, 330 KB — purely so a hot set can be identified and
pinned, because, as the comment puts it, *"which experts are hot is not knowable in advance"*: without
measuring it, pinning is guesswork.
### The bug the code confesses to
The slot table has three states, not two, and the comment explains why with a candour I have rarely
seen in a repository:
```c
/* >= 0 holds that key
* K3_SLOT_EMPTY holds nothing, free to take
* K3_SLOT_INFLIGHT reserved by a batch prefetch whose read has not finished
*
* The third state exists because of a real bug. The batch prefetch marks a slot empty
* before reading into it ... But the empty test below is a FAST PATH that returns
* immediately, ahead of the pinned check and the LRU scan -- so the next expert in the
* same batch was handed the SAME slot, several parallel reads wrote into one buffer, and
* the MoE multiplied garbage. It cost one wrong token (65 instead of 2494) on the real
* model and nothing at all in the fixtures, because no fixture exercises the streaming
* cache. */
```
Read that last clause again. The bug was invisible to the entire test suite, because the fixtures test
kernels and the fault lived in the cache. It surfaced as **exactly one token** — `65` where the model
should have emitted `2494` — in a run that otherwise produced fluent, plausible text. That is the
failure mode that should worry anyone building inference infrastructure: not a crash, not a NaN, but one
silently wrong token inside an output that reads perfectly well.
That component reuses the measured, steady-state numbers, and they are less flattering to the cache
than a quick simulation suggested. K3's training process uses a technique called Quantile Balancing
specifically to keep expert usage flat across the pool — good for training, and exactly what defeats an
LRU cache, which needs a hot subset to be worth anything. Below about 36 GB of cache arena, the bytes
read per token do not move at all, a fact the engine's own measurements caught and reported rather than
smoothing over: a full-recompute trace predicted a 36% hit rate at 8 GB; steady-state incremental
decode measured 0%.
## Why the trunk stays at 16 bits
There is an obvious asymmetry in all of the above. The experts are 4-bit. The trunk — 108.81 GB of it —
is bfloat16, and the engine has no bit-width knob for it at all. If quantisation is what made the
experts streamable, why not quantise the part that has to stay resident?
Because they measured it. A sensitivity study over 31 real attention tensors, quantised symmetrically
per row, gives **about 1% mean relative weight error at int8 and about 17% at int4** — a ratio of
roughly 18 that holds across every tensor type. The worst individual rows at int4 reach 45%, 56% and
**65%**. So the trunk's precision is not an oversight or a to-do; it is a decision with a number behind
it, and the absent knob is the decision being enforced rather than left to a flag.
The study is honest about its own limit: it measures **weight error**, not output quality. No downstream
logit or token comparison was run at int4, so the cost is bounded rather than observed. That is a
narrower claim than "int4 would break the model", and the repo makes the narrower one.
A second measurement worth stealing: on their hardware, `O_DIRECT` cold reads run at 3.2 GB/s and are
*faster* than buffered ones. The repo flags this as **"the opposite of the usual expectation, and it is
why the engine opens the trunk `O_DIRECT`"**. When you are streaming a terabyte past a model, the page
cache is not helping you — it is another copy.
## The memory dial
Put the streaming trunk and the streaming cache together and memory stops being a wall and becomes a
knob. The engine ran the identical prompt through twelve cgroup-enforced budgets, from 8 GB to 224 GB,
with `MemorySwapMax=0` so an over-budget rung fails outright instead of quietly swapping:
Every one of those twelve runs produced the same token ids. Not similar — identical, at a budget span
of 28×. Going from 8 GB to 224 GB buys 1.70× the speed, and the paper trail behind that
number is worth respecting: three back-to-back runs of one identical configuration on a quiet machine
spanned 33% just from device timing noise, so the engine's own docs treat anything under that as
unproven. The 28×-memory-for-1.70×-speed result clears the noise floor with room to spare;
a great many of the smaller steps in between do not, and the source says so rather than reporting every
row as significant.
One more result from the same measurement campaign is worth stating because it runs against instinct:
at a fixed 128 GB total, giving memory to the trunk before the expert cache is 1.69× faster, even
though the winning split reads *79% more* expert bytes from disk than the losing one. Optimising the
number that looks obviously important — cache hit rate — actively picks the slower configuration,
because the trunk is re-read in full on every single token while the experts are only ever sampled.
## What this is not
This is a hobby project, version 0.1.0, one author, Linux x86-64 only. 32.69 seconds per token is not
usable for anything interactive, and the project does not claim otherwise — the README's own words are
"slow, and answering correctly." Every measurement in this piece is the author's own, taken on one
workstation; nothing here is independently replicated the way a benchmark suite would be. What sets it
apart from most projects making similar claims is that it ships the receipts: raw TSVs and JSON traces
under `docs/data/`, a replicated-noise-floor study that argues against several of its own smaller
results, and a fixture ladder that gates every kernel against a PyTorch reference before the released
checkpoint is ever touched.
None of that makes this a serving solution — nobody should run a chatbot on it. What it demonstrates is
narrower and, I think, more interesting: that a 2.78 trillion parameter model can be read, multiplied,
and audited on hardware someone already owns, by one person, in under 5,000 lines of a language
older than most of the engineers who trained the model. That is a pedagogical and archival result, not
a production one, and it is worth having regardless.
---
*This engine is a C implementation of [Kimi K3](/articles/kimi-k3) — see that piece for why K3 uses
KDA, MLA, and Stable LatentMoE in the first place. Its MXFP4 kernel is the concrete, memory-bound case
behind the general argument in [how LLM inference actually works](/articles/how-llm-inference-works):
decode is bound by bytes moved, not by arithmetic, which is exactly why reading fewer bytes wins even
when it means reading them in an awkward packed format. And its confirmation of KDA's log-alpha gate
closes a loop from [the half-life piece](/articles/kda-half-life) I published the same day. If you're
comparing this to training-time low-precision work like [Neutrino-1](/articles/neutrino-1), the
distinction is the direction: that piece is about training a model to tolerate ternary weights from
the first gradient step. This one is about multiplying weights a much larger model already shipped in
4-bit form, without ever training anything.*
---
# LFM2.5-Encoder: classification in one forward pass, zero completion tokens
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/lfm2-5-encoders
> date: 2026-08-03
> tags: explainer, encoders, classification, nlp, open-weights
The default way to classify text with an LLM today is to prompt a decoder and parse whatever
comes back. Write instructions, describe the labels, ask for JSON, decode the answer one token
at a time, then hand the string to a parser and hope it's valid. It works, and it's also the
slow way to do something that has a much cheaper shape: a fixed-size answer picked from a fixed
set of labels doesn't need to be *generated* at all.
That's the pitch behind Liquid AI's **LFM2.5-Encoder** models, released a week ago alongside a
fine-tuning tutorial in their [cookbook](https://github.com/Liquid4All/cookbook) repo: **one
forward pass, zero completion tokens.** I read the code behind that line, not just the slide it's
on, and it holds up exactly as stated.
## Two new encoders, built from a decoder
**LFM2.5-Encoder-230M** and **LFM2.5-Encoder-350M** are bidirectional encoders — the BERT
shape, not the chat-model shape. Their parameter counts, read straight from the safetensors
metadata rather than the rounded name: **229,693,184** and **354,483,968**. Both use the LFM2
hybrid backbone (interleaved gated short convolutions and grouped-query attention), hidden size
**1024**, vocabulary **65,536**, context length **8,192 tokens**, and cover 15 languages. They ship
under the LFM Open License v1.0, and both were created on Hugging Face on **2026-07-27** — about
a week old as I write this.
The interesting part is how they were built: each size starts from the corresponding causal
**LFM2.5 decoder** checkpoint and is converted into an encoder with three changes.
1. **Bidirectional attention** replaces the causal mask, so every position can see the whole
sequence instead of only what came before it.
2. The short-convolution layers switch from causal padding to **symmetric center padding**, so a
kernel mixes information from both neighbors instead of only the left one.
3. Pretraining continues with **masked-language-modeling at 30% token masking** — twice BERT's
original 15%, which is a real methodological choice, not a rounding difference, in how hard the
denoising task is made.
This is the same bidirectional-vs-causal distinction the [architectures gallery](/architectures)
draws for BERT: full attention lets every position condition on the entire input, which is what
you want when the job is producing *a representation* rather than continuing a sequence.
Generation needs causal masking so the model can't peek at the future it's about to write;
classification has no future to hide — the whole document is already there, and every token
should get to see all of it before you summarize what the document means.
## The mechanism: mean-pool, one Linear, sigmoid
Liquid's fine-tuning tutorial (`examples/lfm-encoder-classification/` in the cookbook) is a
complete, runnable project — `train.py`, `predict.py`, a YAML config, sample data — and the model
class inside it, `DocumentClassifier`, is short enough to read in full. It loads the pretrained
encoder, throws away the masked-token prediction head it was pretrained with, and keeps only the
backbone:
```python
outputs = self.backbone(input_ids=input_ids, attention_mask=attention_mask)
hidden = outputs.last_hidden_state # [batch, seq_len, 1024]
mask = attention_mask.unsqueeze(-1).to(hidden.dtype)
pooled = (hidden * mask).sum(dim=1) / mask.sum(dim=1).clamp_min(1.0)
logits = self.classifier(self.dropout(pooled)) # one nn.Linear(1024, num_labels)
loss = nn.functional.binary_cross_entropy_with_logits(logits, labels.float())
```
That's the entire model on top of the backbone: an attention-masked mean over the last hidden
state,
$$
\text{pooled} = \frac{\sum_i m_i \, h_i}{\sum_i m_i}
$$
(padding tokens carry $m_i = 0$ so they don't dilute the average), then one `nn.Linear` sized
`[hidden, num_labels]`, trained with `binary_cross_entropy_with_logits` — the loss for **multi-label**
classification, where a document can carry zero, one, or several labels, as opposed to
cross-entropy's "exactly one correct class." At inference, `predict.py` does the other half in
three lines:
```python
with torch.inference_mode():
probabilities = torch.sigmoid(model(**inputs).logits)[0].cpu().tolist()
predicted = [label for label, p, t in zip(labels, probabilities, thresholds) if p >= t]
```
No `generate()`, no sampling, no stop tokens, no text to hand to a parser. One forward pass in,
a fixed-length array of probabilities out, compared against per-label thresholds tuned on a
validation split. That is what "zero completion tokens" means in the code, not just the slide.
## Sigmoid, not softmax — and why it has to be
The reason this needs its own head, rather than reusing whatever classifier head ships with a
decoder fine-tune, is the difference between picking one thing and scoring several independently.
**Softmax** turns a set of logits into a probability distribution that sums to 1 — it is built to
choose exactly one winner, which is correct for single-label problems (a document *is* sports,
politics, or tech, never more than one). **Sigmoid** applied per label makes each label its own
independent yes/no question: $\sigma(z_i) = 1/(1+e^{-z_i})$, with no normalization across labels,
so two labels can both clear the threshold, or none can. A support ticket about a double charge is
legitimately both a billing issue and a technical one; softmax would be forced to pick a single
"real" answer and quietly discard the other.
## Where it lands: 4th of 14, honestly
Liquid's own eval — a 17-task suite spanning GLUE, SuperGLUE, and five multilingual tasks,
averaged over 5 seeds with standard deviations reported — puts LFM2.5-Encoder-350M **4th of 14**
models at a mean score of **81.02**, and LFM2.5-Encoder-230M **6th** at **79.29**.
That ranking is worth sitting with rather than rounding up. LFM2.5-Encoder-350M sits a point below
ModernBERT-large and two points below XLM-R XL — a model **ten times its size**. It is not the
best encoder on this benchmark, and Liquid doesn't present it as one. It's the fourth-best, at a
fraction of the parameters of the model above it, next to a considerably larger one. The 230M
model separately beats ModernBERT-base despite being the larger of the two by parameter count — a
real, checkable comparison the raw numbers support either way you read them.
The chart also settles a question Liquid answers candidly in the same blog post: why build a new
general-purpose encoder instead of reusing their existing retrieval models? Because those
retrieval-tuned siblings — **LFM2.5-ColBERT-350M** and **LFM2.5-Embedding-350M** — score **76.18**
and **75.68** on this same suite, both below the general-purpose LFM2.5-Encoder-350M's 81.02. In
their own words: *"Because retrieval is only a subset of what encoders enable, we chose to build a
general-purpose encoder rather than adapt the existing retrievers."* A model tuned to make
embeddings cluster well for search is not automatically a good classification backbone, and
Liquid's own numbers show the gap rather than hiding it.
On raw speed, Liquid also reports the 350M encoder running about **3.3× faster than
ModernBERT-base at 8,192 tokens on CPU**; a separate blog claim puts the 230M model specifically
at roughly **3.7×** faster than ModernBERT-base at the same length (about 28 seconds versus over a
minute and a half). Those are two distinct comparisons, not one number restated — worth keeping
straight if you quote either.
## The tutorial's own result
The cookbook ships a second, harder example beyond the 4-label sample data: fine-tuning
LFM2.5-Encoder-350M on **ECtHR-A** (`coastalcph/lex_glue`), European Court of Human Rights cases
labeled by which of 10 Convention articles they violate — real long documents, real multi-label
targets, 9,000 / 1,000 / 1,000 train/validation/test examples, CC BY 4.0. Trained at the full
8,192-token context, one seed, 3 epochs:
| Split | Metric | Score |
|---|---|---|
| Validation | micro-F1 (after per-label threshold tuning) | 0.8060 |
| Test | micro-F1 | 0.7913 |
| Test | macro-F1 | 0.7062 |
| Test | micro average precision | 0.8400 |
The README says outright that these numbers are from one seed — no variance reported, unlike the
17-task pretraining eval above. Read it as "this recipe works on a real long-document benchmark,"
not as a tuned, reproducible leaderboard number. Thresholds are tuned only on the validation split
and never touch test until a separate, explicit `--evaluate-test` flag is passed — a small
detail, but the right one for anyone checking the tutorial's methodology.
## What you give up
An encoder with a classification head cannot do several things a decoder can, and it's worth
naming them plainly rather than only listing what it's good at:
It cannot generate. There's no explanation, no rationale, no free-text answer — only probabilities
over a label set that's fixed at training time. Adding a new label means retraining the head (a
small, cheap step, but a step), not writing a new prompt. And it needs supervised examples per
task: this is a fine-tuning recipe, not a zero-shot classifier out of the box, even though Liquid's
own HF Spaces (prompt routing, PII detection, policy linting) show the same base encoder
fine-tuned across several different classification tasks.
The honesty gaps worth naming too: these models are about a week old, with limited independent
adoption to point to yet. The 17-task eval is Liquid's own compilation — methodologically solid
(5 seeds, reported std, a full 14-model field rather than a curated subset), but not yet
replicated by anyone outside Liquid. And the cookbook repo carries no LICENSE file at its root as
of this writing, which matters if you plan to reuse the tutorial code itself, distinct from the
separately-licensed model weights.
## The take
The argument here isn't that encoders are back or that decoders are wrong for classification —
it's narrower and more useful than that: match the tool to the shape of the answer. If you already
know the output is a choice from a fixed set of labels, a bidirectional encoder can produce that
choice as a probability vector in one forward pass, with nothing to decode and nothing to parse.
If you don't know the shape of the answer in advance — you need explanation, planning, or free
text — that's what the decode loop and [its own cost structure](/articles/how-llm-inference-works)
are for. LFM2.5-Encoder is a clean, current example of the first case done right: a small model,
an honestly-reported 4th-of-14 ranking against models many times its size, and a fine-tuning recipe
short enough to read start to finish in one sitting.
---
*Built on Liquid AI's [LFM2.5-Encoders blog post](https://www.liquid.ai/blog/lfm2-5-encoders), the
[LiquidAI/LFM2.5-Encoder-230M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-230M) and
[LiquidAI/LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) model cards,
and the [`examples/lfm-encoder-classification`](https://github.com/Liquid4All/cookbook/tree/main/examples/lfm-encoder-classification)
tutorial in Liquid4All/cookbook. Parameter counts are read from HF safetensors metadata, not the
rounded model names. The 17-task benchmark and both embedded figures are Liquid AI's own,
reproduced for commentary; the two interactive diagrams are original illustrations of the
mechanism using illustrative example data, not measured traces.*
---
# Towards looped models done right: what actually separates Ouro from Huginn
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/looped-models-done-right
> date: 2026-08-03
> tags: llm, looped-transformers, recurrent-depth, ablation-study, architecture, mixture-of-experts, explainer
Ask why Huginn-style looped transformers tend to beat Ouro-style ones and most people reach for the
same answer: random state initialization. It is the assumption inherited wholesale from deep
equilibrium models — randomize the recurrent state so the loop cannot just memorize a fixed point,
and the model is forced to learn a genuinely path-independent computation. It sounds right. It is also,
according to a new controlled-ablation report from the Institute of Foundation Models (IFM), mostly
wrong.
**Towards Looped Models Done Right** does not propose a new looped architecture. It audits the two
existing lineages this site has already covered — the padded-latent loop in
[LOTUS](/articles/lotus-latent-reasoning) and the shipped two-pass loop in
[Nanbeige4.2-3B](/articles/nanbeige-4-2-3b) both descend from this same family of
[looped / recurrent-depth transformers](/architectures) — and asks a narrower, more useful question:
Ouro and Huginn differ along *three* design axes at once, so which one is actually doing the work? The
answer, walked through below, is not the one folk wisdom would guess.
**Read this as a first look, not a finished paper.** This report lives only as a Notion "living blog,"
explicitly billed as Part I of a series and continuously updated — there is no arXiv listing. Code is
marked "Release Soon," meaning nobody outside IFM can rerun these numbers yet, and there has been no
third-party replication. Every result below comes from ablations run by one group, at 730M dense /
8B-resident MoE scale — a real, carefully controlled experiment, but not yet an independently checked
one, and not yet evidence about frontier scale.
## One formalism, two lineages, three tangled axes
Both Ouro and Huginn are instances of the same tied-iterative model. Take token embeddings, run them
through a prelude, loop a shared recurrent body $T$ times, then run a coda:
$$
\mathbf e = P_\theta(\mathbf x_0),\quad
\mathbf z_0 = \phi_\theta(\mathbf e, \boldsymbol\xi),\quad
\tilde{\mathbf z}_t = W_\theta(\mathbf z_t, \mathbf e),\quad
\mathbf z_{t+1} = R_\theta(\tilde{\mathbf z}_t),\quad
\mathbf h = C_\theta(\mathbf z_T)
$$
$P_\theta$ is the prelude, $R_\theta$ the tied recurrent body (the only part that repeats), $C_\theta$
the coda, $\phi_\theta$ the state initializer (optionally seeded with noise $\boldsymbol\xi$), and
$W_\theta$ the per-step write that decides how much of $\mathbf e$ gets re-injected at each pass.
Set every one of $P_\theta$, $W_\theta$, $\phi_\theta$ to the identity and you get **Ouro-style**: the
entire network is the recurrent body, looped over the full sequence, with $\mathbf z_0 = \mathbf e$
directly. Untie a real prelude and coda from the loop, persistently re-inject $\mathbf e$ into the core
at every pass, and optionally randomize the initial state, and you get **Huginn-style**. Three
independent knobs — iteration envelope, input interface, latent-state design — get flipped between the
two architectures simultaneously in the existing literature. Nobody had isolated which flip mattered.
## The controls that make this an ablation, not a vibe check
A myth-busting result is only as good as what it holds constant. IFM's models are matched on logical
depth — Ouro-style loops a 28-block stack four times, Huginn-style runs an 8-block prelude, a 12-block
core eight times, and an 8-block coda; both total 112 block executions — and on parameters (730M
stored / 2.9B unrolled-equivalent dense; 8B-resident / 0.8B-active, 32B/3.2B unrolled-equivalent for the
MoE runs), tokens (TxT360, swept across 58/115/230/460 tokens-per-parameter), and a ten-benchmark suite
spanning knowledge (ARC-C, HellaSwag, MMLU, TriviaQA), context reasoning (BBH-CoT, DROP), math (GSM8K,
MATH500), and code (HumanEval+, MBPP+). Whichever topology wins a given comparison, it isn't winning on
a hidden depth or size advantage:
With the controls fixed, IFM walks the transformation path from Ouro to Huginn one axis at a time —
sandwich envelope, then input injection, then random state init — and measures what each addition
actually buys.
## Q1: does untying the prelude and coda matter?
Yes, and only for a specific kind of task. Adding a sandwich envelope — untying a prelude and coda from
the loop so only the middle core repeats — lifts **MATH500 by 12.00 points** and **DROP by 2.61 points**
at the 460 tokens-per-parameter budget, and the gain persists across all four token budgets tested. But
knowledge-heavy benchmarks and strict-output-contract tasks like code show no consistent gain, sometimes
a small decline. Read plainly: the envelope helps *instance-conditioned, multi-step reasoning*. It does
not help stored-knowledge recall or spec-compliant code generation, and the report is upfront that it
shouldn't be expected to.
## Q2: does persistent input injection matter?
Also yes, and it's the widest-reaching single change in the whole report. Writing the prelude's
representation $\mathbf e$ into the core at every pass — not just once at the start — uses a learned,
per-channel gate:
$$
D(\mathbf z_t, \mathbf v) = \boldsymbol\alpha \odot \mathbf z_t + \boldsymbol\delta \odot \mathbf W_{\text{in}}\mathbf v,
\qquad \boldsymbol\delta = \operatorname{softplus}(\mathbf b_\delta),
\qquad \boldsymbol\alpha = \exp\{-\boldsymbol\delta \odot \exp(\mathbf a)\}
$$
On the middle-loop (sandwich) topology, adding this write lifts **MMLU +2.53, BBH-CoT +6.63, DROP +1.39,
HumanEval+ +5.49, MBPP+ +4.23** — five benchmarks, all up, some by a lot. Bolt the same write onto the
full-stack Ouro topology (injecting raw token embeddings instead of a prelude-encoded $\mathbf e$) and
the same-direction gains show up smaller (MMLU +1.79, BBH-CoT +1.80, DROP +0.91, HumanEval+ +2.44,
MBPP+ +4.50) — evidence the effect is real and not an artifact of one topology.
It is not a free lunch. The same write **hurts** quantitative reasoning: middle-loop MATH500 drops
**3.60 points**, GSM8K drops **2.51** (the full-stack Ouro version is less damaged: MATH500 −1.60,
GSM8K +0.53). Persistently re-showing the model its own input, it turns out, competes with letting the
loop's state evolve freely enough to carry a multi-step derivation.
Put the envelope and the injection together and the combined model beats full-stack Ouro on **8 of the
10 benchmarks** (losing only ARC-Challenge and HellaSwag). Step through both changes yourself — the
diagram below is my own redrawing of the same construction path, with each stage's measured delta
attached so you can see exactly which wire produced which number, before the third, more surprising
change gets added:
## Q3: the myth — random state init and shared H/L hierarchies
This is the report's contrarian core. Swap the direct initial state $\mathbf z_0 = \mathbf e$ for a
randomly sampled one, $\mathbf z_0 \sim \mathcal N(0, I/d)$ — the equilibrium-model-inherited move
everyone assumes is load-bearing — and two benchmarks improve (**ARC-C +3.34, GSM8K +1.22**) while four
get worse by more than a point (**MMLU, MATH500, HumanEval+, MBPP+**). Net: **direct init wins 6 of 10
benchmarks**, and it's cheaper, since it skips sampling noise at every forward pass. The report's own
words: "random initialization is not a necessary ingredient for loop language models... [it] should
instead be viewed as a task- and objective-dependent inductive bias, rather than as a universally
beneficial design choice."
A second candidate "obviously helps" ingredient fares no better. HRM/TRM-style hierarchies split the
loop into a slow high-level state and a fast low-level state cycling underneath it; IFM tests a version
that **shares one recurrent body** across both states (isolating the state-hierarchy idea from the
separate-modules idea) and finds gains over a point on three benchmarks, losses over a point on three
more — MATH500 hit hardest — and roughly flat on the rest. Their conclusion: "a shared-module H/L
hierarchy provides no consistent benefit." (They flag, honestly, that this doesn't rule out a
separately-parameterized HRM-Text-style version — that variant is "still under evaluation.")
Here is the believed-important story against what got measured, for both:
## The net effect: two wires did almost all the work
Chain every change together — Ouro, plus envelope, plus injection, plus random init — and you land on
full Huginn, which does beat Ouro on all ten dense benchmarks at the 730M/336B-token setting. But laid
out per-benchmark across the whole construction path, the shape of the win is obvious: most of the
climb happens in the first two steps, and the last step (random init) barely moves several benchmarks
and actively costs a few.
## Does the story change at MoE scale?
The two levers that mattered — envelope and injection — hold up when the recurrent body becomes a
mixture-of-experts. At 8B-resident / 793.9M-active parameters (500B tokens, top-2 routing, 25 experts),
Huginn-MoE beats Ouro-MoE on 8 of 10 benchmarks, with the largest gains on **GSM8K (+4.70)** and
**MATH500 (+3.60)**. It also routes more evenly — a normalized load-balancing loss where lower means more
balanced:
A causal check backs up that the routing difference is meaningful, not noise: force loop iterations 2
through 8 to reuse iteration 1's expert *identities* (keeping the iteration-specific mixture weights)
and accuracy drops on all six evaluated tasks. Whatever the loop is learning to route each pass, it
matters.
Against a **112-layer feedforward MoE** reference (32B resident parameters), the feedforward model still
wins overall — 7 of 10 benchmarks — but Huginn-MoE beats it on DROP and GSM8K and matches it on MATH500,
while using **75% fewer resident parameters**. The mean gap to the feedforward reference shrinks from
4.96 points in the dense setting to 1.71 points in the MoE setting. Looping doesn't close the gap to a
much bigger feedforward model outright, but MoE narrows it substantially — the same shape of result as
Nanbeige's own parameter-efficiency argument, below.
## Set against a shipped model: Nanbeige4.2-3B
The most useful check on any ablation study is an independent result that wasn't trying to test the
same hypothesis. [Nanbeige4.2-3B's technical report](/articles/nanbeige-4-2-3b) is exactly that: a
production model that made its own looping decisions under deployment pressure, not a controlled
academic sweep.
Nanbeige's architecture is closer to Ouro-style — a homogeneous loop over the full stack, run twice —
and its report reached three conclusions of its own: **two passes is the sweet spot** (more loop count
bought little and made training less stable), **training the looped architecture from scratch beats
upcycling** a pretrained feedforward model into one, and **sharing the KV cache across passes
underperformed**, so they paid full attention cost at every pass rather than take the cheaper shortcut.
None of Nanbeige's three findings directly tests IFM's three axes — Nanbeige never tried an untied
prelude/coda, persistent injection, or random state init — so this isn't a replication in either
direction. But the two reports rhyme in an interesting way: every place Nanbeige tested a cheap shortcut
inside the loop (share the KV cache, upcycle instead of retraining, add more passes without changing the
topology), the shortcut lost. Every place IFM tested a richer per-pass mechanism (untie the envelope,
inject persistently), it won — and the one change that added complexity without adding real per-pass
information (random init) was the one that didn't clearly help. Read together, the two reports point at
the same underlying rule: what a looped model does *each pass* — how much fresh computation and fresh
input it gets — matters more than how many times it loops or how its state gets seeded. Where they don't
overlap at all is loop count itself: Nanbeige's "two is enough" is a statement about a plain full-stack
loop; IFM's Ouro-style baseline already runs four passes and Huginn-style runs eight core iterations
inside a smaller envelope, so the two reports are sweeping different variables and shouldn't be read as
agreeing or disagreeing on "how many loops."
There's a second, more mechanical echo. [LOTUS](/articles/lotus-latent-reasoning) — the site's other
looped-transformer piece — already does something IFM's ablation independently flags as one of the two
levers that matter: every LOTUS iteration recomputes $\mathbf e + \mathbf h^{(t-1)}$, persistently
feeding the fixed input embeddings back into the loop rather than only conditioning on them once. That's
a specific instance of the same principle behind IFM's write operator $D(\mathbf z_t, \mathbf e)$: keep
re-showing the loop its input. LOTUS applies that idea inside a frozen-backbone latent-reasoning setup
at inference time rather than IFM's from-scratch pretraining setup, so the two aren't directly
comparable — but it's a second, independent place persistent input injection shows up as doing real
work.
## What to trust, and what to hold loosely
**The scope, precisely.** Every result above is at 730M dense / 8B-resident MoE scale — there is no
evidence yet these findings hold at frontier (100B+) scale, and the report says so. The random-init
result is explicitly narrower than "random init never helps": IFM's own caveat is that these evaluations
"do not directly measure multi-start path independence or extrapolation to recurrent depths beyond those
used during training" — the property random init is classically supposed to buy. The H/L null result is
scoped to a *shared-module* hierarchy only; a closer HRM-Text replica with separate modules is still
under evaluation and not included here. And this whole report has no arXiv listing, lives on a
continuously-updated Notion page billed as "Part I" of a series, and ships with no code yet — treat every
number as provisional until an arXiv version, released code, or a third-party rerun shows up.
## The take
The intuitive story about looped transformers has always centered on the recurrent state itself —
randomize it, and the model is forced to learn something more general. IFM's controlled ablations say
that story is backwards for at least these two designs at this scale: the state's initialization is
close to a wash, sometimes a net loss, and a fashionable H/L hierarchy adds nothing consistent once you
control for everything else. What actually separates a strong looped model from a weak one is much less
exotic — untie a prelude and coda from the loop so the recurrent core can specialize, and keep showing
that core its input at every pass instead of just once. Neither idea needs a random number generator.
That's a genuinely useful result for anyone building a looped model today, and it lines up with what
Nanbeige found the hard way in production: cheap shortcuts inside the loop (shared KV cache, more passes
without restructuring, upcycling instead of retraining) tend to cost you, while spending real compute
and real information on each pass tends to pay. The caveat that matters most is the one IFM states
themselves — this is one group's ablation, at sub-billion-to-8B scale, with no code out yet and no arXiv
paper. "Part I" means there's more coming. Whether the random-init myth holds at 100B+ parameters, and
whether a properly separate-module HRM-Text hierarchy fares better than the shared one tested here, are
open questions the authors name as future work, not settled ones.
---
*Source: [Towards Looped Models Done Right — Part I: Topology, Input Injection, Recurrent-State
Design](https://ifm-research.notion.site/Towards-Looped-Models-Done-Right-3ade511912ec8128987dfeb7a5580043)
(Huang, Shi, Chen, Wen, Liu, Xing, Ma; Institute of Foundation Models, 2026). All benchmark deltas,
equations, and figures are quoted or reproduced from the report; the matched-depth, construction-path,
and believed-vs-measured diagrams are my own illustrations of the same numbers.*
---
# MemHarness: agent memory should be reconstructed, not replayed
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/memharness
> date: 2026-08-03
> tags: agents, memory, reinforcement-learning, llm, explainer
Most memory-augmented agents treat a retrieved experience the way a tape recorder treats a
cassette: press play, get back exactly what was stored. **MemHarness**, from a
Zhejiang University / Shanghai AI Lab team, argues that's the wrong model of memory
entirely — and points at cognitive science to say so. Human recall isn't playback; it's
reconstruction, rebuilt each time from fragments and reshaped to fit the moment. Their
agent does the same: before acting on a retrieved memory, it first critiques that memory
against what it's looking at *right now*, and rewrites it if the two don't match.
Single paper, one lab, no third-party replication found. Everything below — every
percentage, every table — is MemHarness's own reported numbers on two benchmarks
(ALFWorld, WebShop), one 7B backbone (Qwen2.5-7B-Instruct). The paper reports no
hardware, no wall-clock latency, and no variance across seeds for any of its evaluation
numbers — see **Honesty check** near the end before you take any number as a settled
fact. The interactive diagrams below are clearly labeled: two use the paper's own
measured table values, one is an illustrative toy walkthrough of the mechanism.
## The failure mode: negative transfer from a memory that no longer fits
Retrieval-augmented agents work like this: finish a task, distill what happened into a
short natural-language "experience," store it, and next time a similar task comes up,
pull the closest matches back into context. The problem is *closest* is doing a lot of
work. An experience learned from one kitchen layout, one inventory state, one shopping
page, gets pasted into a context where the fridge is already full or the shelf has
something else on it — and the agent, having no reason to doubt its own memory, follows
the stale instruction anyway. MemHarness calls this the "replay" paradigm, and its
central claim is that replay's failures are systematic, not occasional: the retrieved
experience is abstract and general by construction, while the state at decision time is
concrete and constantly changing, and nothing in a replay pipeline reconciles the two.
This isn't a new observation for this site — [Agent harnesses: engineering the loop
around the model](/articles/agent-harness) already named "context and memory lifecycle"
as one of the open problems in agent engineering, and pointed out that a file-backed
harness effectively turns the file system into the agent's long-term memory. MemHarness
is a concrete answer to a sharper version of that problem: it's not enough to *store*
memory durably, the harness also has to decide, at read time, whether a piece of stored
memory still applies — and if not, what to do about it. MemHarness's answer is to make
that decision itself a trained, learned skill rather than a fixed retrieval-and-paste
rule.
## Three stages: retrieve, reconstruct, act
Formally, the agent holds a memory bank $\mathcal{B} = \{m_i\}_{i=1}^N$ of entries
$m_i = (e_i, o_i^{src})$ — an abstracted experience $e_i$ paired with the **source
observation** $o_i^{src}$ it was distilled from. At each step $t$, retrieval returns the
top-$k$ closest entries, $\mathcal{E}_t = \mathcal{R}(q_t, \mathcal{B})$. A pure replay
policy conditions the next action directly on whatever comes back:
$$
a_t \sim \pi_\theta(\cdot \mid \mathcal{T}, h_t, \mathcal{E}_t)
$$
MemHarness inserts one step in between. The same policy first produces guidance $g_t$ by
critiquing the retrieved experiences against the current history $h_t$, then maps that
into final guidance $\tilde{g}_t$, and only then generates the action:
$$
g_t \sim \pi_\theta(\cdot \mid \mathcal{T}, h_t, \mathcal{E}_t), \qquad
\tilde{g}_t = f(g_t), \qquad
a_t \sim \pi_\theta(\cdot \mid \mathcal{T}, h_t, \tilde{g}_t)
$$
The reconstruction input concatenates the task, the recent history, and every retrieved
(experience, source-state) pair:
$$
x_\text{recon} = \mathcal{T} \oplus h_t \oplus \bigcup_{i=1}^{k} (e_{t,i},\, o_{t,i}^{src})
$$
and $f$ is a simple conditional: if the policy decides nothing retrieved applies, it
emits the literal token ``, and $\tilde{g}_t$ falls back to a fixed self-reasoning
prompt $p_\text{self}$ instead of forcing a bad match into the action context.
The mechanics of that middle stage — comparing a retrieved memory's source state against
the live one, and deciding pass-through / adapt / reject — are easiest to see with a toy
example. Toggle through the three cases below:
Note what each branch does differently from plain replay. When the state genuinely
**matches** the memory's source, reconstruction is a no-op and replay would have been
fine anyway — the interesting cases are the other two. When the state has **drifted**,
reconstruction rewrites the *target*, not just the wording, while replay keeps repeating
an instruction the environment has already invalidated. And when retrieval turns up
nothing usable, MemHarness can say so explicitly and fall back to the agent's own
reasoning — a replay pipeline has no equivalent move; it either injects a weak match or
injects nothing silently.
## Training: one policy, three roles, GRPO end to end
The same weights play all three parts — retriever-decider, reconstructor, actor — and
the whole thing is trained with **GRPO** on a sparse outcome reward plus a small format
bonus:
$$
R(\tau_i) = R_\text{outcome} + 0.1 \cdot R_\text{format}
$$
$R_\text{outcome}$ is 10 for a successful episode and 0 otherwise; $R_\text{format}$
checks that every step emits exactly one valid `` block, one valid ``
block, that memory is retrieved through valid `` blocks (one to five
times per episode), and that everything is in English. Rewards are group-normalized —
sample $G=8$ rollouts per prompt and standardize against the group:
$$
A_i = \frac{R(\tau_i) - \text{mean}(\{R(\tau_k)\}_{k=1}^{G})}{\text{std}(\{R(\tau_k)\}_{k=1}^{G})}
$$
and the policy update is the standard clipped-surrogate-plus-KL objective, with clip
range $\varepsilon = 0.2$ and KL coefficient $\beta = 0.01$:
$$
\mathcal{J}(\theta) = \mathbb{E}\!\left[\frac{1}{\sum_i |\tau_i|}\sum_{i=1}^{G}\sum_{j=1}^{|\tau_i|}
\Big(\mathcal{L}^{\text{CLIP}}_{i,j}(\theta) - \beta\, \mathbb{D}_{\text{KL}}[\pi_\theta \| \pi_\text{ref}]\Big)\right]
$$
None of this is a new RL recipe — it's the same token-level, group-relative machinery
covered in [Token-level RL is a first-order approximation to the reward you actually
want](/articles/first-order-rl). What's specific to MemHarness is that the *reconstruction
step itself* is inside the RL loop and gets credit for the same sparse outcome reward as
the final action, rather than being a fixed prompt template bolted on the side. That's
also why an ablation later in this piece — replacing the trained reconstruction with a
generic, untrained LLM doing the same rewriting job — measurably underperforms: rewriting
text is not the same skill as rewriting text so that it wins the episode.
Before RL, there's a short cold-start SFT stage — 200 trajectories with GPT-5.1-generated
retrieval and reconstruction turns, plus 200 trajectory-to-memory summarization examples
per benchmark — whose only job is to teach the interaction protocol (when to emit
``, how to format guidance). The paper is explicit that this stage is
about "protocol and format alignment rather than task-skill acquisition," and the numbers
back that up: the cold-start model alone scores a *worse* 7.6% on ALFWorld than the
untrained base model's 14.5%, because it has learned to follow a longer protocol without
yet having learned to solve the task.
The memory bank itself lives in **Milvus**, embedded with **BGE-M3**, retrieved by
cosine similarity at $k=3$. It isn't hand-curated — during training, the policy
distills roughly half of its own generated trajectories (balanced between successes and
failures where possible) into new memory entries, so the bank grows out of the same
policy that reads from it. Write-time deduplication skips a new entry if it's
too similar (cosine $> 0.85$) to something already stored — enabled for WebShop,
disabled for ALFWorld — and retrieval-time deduplication thins a larger candidate pool
before truncating to the top-$k$.
## Does it beat the baselines
On the headline numbers: MemHarness reaches **85.2%** average success on ALFWorld's six
task categories and **75.6%** on WebShop, ahead of every baseline the paper reports —
including foundation models an order of magnitude larger:
A few things worth being precise about here. The 16-row full table (not all shown
above) mixes closed-source frontier models (GPT-4o, Gemini-2.5-Pro), prompt-only
memory agents (ReAct, Reflexion, Mem0, ExpeL, MemP, SimpleMem), and RL-trained agents
(RLOO, GRPO, MemRL, EvolveR, and two "+GRPO" memory hybrids) — and with one exception
(**EvolveR**, explicitly marked "reproduced"), the paper doesn't say whether the other
baseline numbers are copied from those methods' original papers or re-run by the authors
under this setup. Given every RL-based and prompt-based baseline shares the same
Qwen2.5-7B-Instruct backbone as MemHarness, it reads as an in-house re-implementation for
a controlled, like-for-like comparison — which is the right thing to do for fairness, but
it also means there's no independent number to check any of them against. **Mem0** in
particular scores worse than the untrained base model on WebShop (2.0% vs. 7.8%) — plausible
for a general-purpose memory library not tuned to this task, but a reminder that "memory
system" is not automatically an improvement.
## Why raw memory can hurt: the ablation
The paper's most useful table isn't the leaderboard, it's the ablation, because it
isolates *why* MemHarness wins rather than just *that* it wins. Same policy, same GRPO
recipe throughout — only the memory wiring changes:
Two results here are worth sitting with. First, **RL + Raw Memory** — verbatim replay,
grafted onto the same trained policy — actually *loses* to having no memory at all on
ALFWorld (70.1% vs. 76.4%), which is the paper's sharpest evidence that unreconstructed
memory is not a free win; it can be actively confusing. Second, **w/o memory** — the
fully-trained MemHarness policy with retrieval switched off at test time — still beats
the no-memory-ever RL baseline on both benchmarks (83.0% vs. 76.4% on ALFWorld, 73.6% vs.
66.1% on WebShop). The paper reads this as evidence that training the policy to
reconstruct memories also sharpens its general reasoning, independent of whether memory
is available at inference — the reconstruction objective works partly as a training-time
signal, not only a run-time lookup.
The training curves back this up with a second, independent kind of evidence — not an
end-of-training snapshot, but what happens over the run:
Trajectories where the policy accepted a reconstructed memory track the overall
success-rate curve closely; trajectories where it rejected one lag behind and stay
noisier throughout training. That's a consistency check on the whole framework: if
"accept vs. reject" were an arbitrary or miscalibrated signal, there'd be no reason for
it to correlate with which trajectories actually succeed.
MemHarness also holds up when the environment itself is unfamiliar. On ALFWorld's
out-of-distribution split — unseen room layouts and object placements — it scores
**85.9%**, while stripping reconstruction back out (raw memory injected, same OOD
environments) drops to **76.3%**, and disabling reconstruction only at test time (same
trained policy) lands at **82.4%**. The direction of every result here matches the
in-distribution ablation: verbatim replay is the worst way to use memory precisely when
the environment has changed most.
## The mechanism, under a microscope
Everything so far shows *that* reconstruction helps. The paper also runs two controlled
probes asking a narrower question: does the policy's reconstruction step actually compare
the current state against the memory's recorded source state, or is it just producing
plausible-sounding rewrites without really checking anything?
The **source-state ablation** answers this directly: strip $o_i^{src}$ out of the
reconstruction prompt entirely, and rejection rate barely moves — but success rate drops,
because the policy now accepts guidance it has no way to judge as stale. Swap in a
*random* memory's source state instead — a state that's guaranteed not to match — and
rejection rate jumps sharply (8.7%→13.3% on ALFWorld, 56.0%→63.3% on WebShop). That
asymmetry is the tell: removing the comparison signal doesn't change behavior much
because the policy simply can't tell anymore, while corrupting it with a wrong-but-present
signal actively triggers more rejections. The **counterfactual probe** — asking a strong
LLM to make a minimal edit to 1,000 real states so a previously-applicable memory should
no longer apply, then scoring only the reconstruction output — shows the same pattern
from the other direction: minimal edits shift outputs measurably away from "unchanged"
and toward "adapted" or "rejected" on both benchmarks, with WebShop rejecting far more
often than ALFWorld in both the matched and edited conditions (72–79% vs. 0–6%), which
the paper attributes to WebShop's longer, more heterogeneous page observations making a
fuzzy accept riskier than a clean reject.
## Honesty check
- **Self-reported, single lab, no replication.** Every number above is from this one
paper. I found no independent reproduction, and the community-discussion page on
alphaXiv had nothing beyond the paper's own abstract and tables at the time of writing.
- **Baselines are the paper's own reruns, not cited published numbers**, as far as the
text discloses — with the single exception of EvolveR, marked "(reproduced)" in the
table. That's a reasonable design for a fair, same-backbone comparison, but it also
means none of the sixteen rows in Table 1 have an outside number to be checked against.
- **No hardware, no latency, no cost.** The paper never states what GPUs it trained or
evaluated on, never reports wall-clock time, tokens/sec, or dollars, and never measures
the added inference cost of the reconstruction step itself. Reconstruction is a second
full decode pass through the same 7B policy on every step where memory is retrieved
(retrieval decision, then reconstruction, then action) — at minimum one extra
generation versus a direct-replay or no-memory baseline — and that overhead is not
quantified anywhere in the paper.
- **No variance, no seeds.** Every success-rate and rejection-rate number is reported as
a single figure with no standard deviation, confidence interval, or multi-seed spread
disclosed for the evaluation runs.
- **Two benchmarks, one model scale.** ALFWorld and WebShop are both well-worn,
relatively short-horizon (15–50 step) simulated environments; the paper does not test
a larger backbone, a real-world tool-using agent, or a benchmark with a genuinely
different observation modality. The conclusion names this directly: "future work will
explore scaling to larger models and open-ended environments" — which is the authors'
own way of saying this hasn't been tried yet.
- **No explicit Limitations section.** The paper has no dedicated limitations
discussion; what's above is reconstructed from the ablations, the conclusion, and what
the method section does and doesn't measure.
- **What is solid:** the mechanism probes (source-state ablation, counterfactual editing)
are a genuine attempt to falsify the "it's just fluent rewriting" explanation, and they
point the same direction from two independent angles. That's better methodological care
than a bare leaderboard table, even without outside replication.
## The take
The idea underneath MemHarness is simple enough to state in a sentence — compare the
memory's source state to the current one before you trust it — and the paper's real
contribution is making that comparison a *trained* skill inside the same policy, credited
by the same sparse outcome reward as the action itself, rather than a hand-written
heuristic bolted onto retrieval. The ablations back the framing better than the
leaderboard does: raw memory replay measurably *loses* to no memory at all on one
benchmark, and the reconstruction-trained policy keeps a chunk of its advantage even with
memory switched off entirely, which says the training signal is doing more than teaching
better lookups.
What it hasn't shown yet is whether any of this survives outside two small, well-studied
simulators at 7B scale, and whether the extra reconstruction pass is worth its
unmeasured latency cost in a setting where that matters. "Reconstruct, don't replay" is a
good design principle for any agent harness that reads back its own memory. Whether this
specific recipe for teaching it — GRPO, a `` escape hatch, a Milvus bank refreshed
by the policy's own trajectories — is the way to get there past ALFWorld and WebShop is
still an open question the paper itself doesn't claim to answer.
---
*Built on [MemHarness: Memory Is Reconstructed, Not Replayed](https://arxiv.org/abs/2607.28272)
(Wu et al., 2026; arXiv:2607.28272). Figures are reproduced from the paper for commentary.
The interactive diagrams use the paper's measured table values except where marked
illustrative; see the Honesty check above for what is and isn't independently verified.*
---
# MiniMax H3: open weights, four excluded countries, zero benchmarks
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/minimax-h3
> date: 2026-08-03
> tags: licensing, open-source, multimodal, video-generation, explainer
MiniMax [announced H3](https://www.minimax.io/blog/minimax-h3) on July 31, 2026 and put weights on Hugging Face two days later — an omni-modal system that takes text, image, video, or audio in and produces video with native stereo audio out, up to 2K resolution, 15 seconds, 24 FPS. The architecture underneath is a real, disclosed piece of engineering: an encoder built on the **full pretrained weights of Qwen3-VL-32B**, sampled from its 50th layer, feeding a **33B-parameter dense Omni-Transformer** with no modality-specific attention or feed-forward blocks — only the input/output layers and a set of AdaLN branches (about 13B of the 33B) are modality-specific.
I'm not writing about the capability, though. I'm writing about what shipped alongside it, because the license and the evidence base are the actual story here.
## The license
MiniMax H3's weights carry the **MiniMax H3 Community License Agreement**, and its territorial scope is precise enough to quote directly. The license grants use across the "Applicable Territory," defined as worldwide **excluding** the "Excluded Territories" — and the Excluded Territories are named explicitly: **the European Union, the United Kingdom, the Republic of Korea, and the United States of America.**
That is a different thing from a normal open-weight release. Apache-2.0 and MIT — the licenses this site's other open-weight coverage almost always carries — don't have a geography clause at all. This one draws the line at specific jurisdictions, and the four it picks are, not coincidentally, four of the jurisdictions with the most developed AI-regulatory frameworks in the world. There's a second condition stacked on top for everywhere else: commercial deployments need "separate, prior written authorization" once they clear **$20 million/year** in revenue, plus a requirement to "prominently display 'MiniMax H3'" on the interface of anything built with it.
That figure matters for the licensing question too: even inside the "open" release, two of the three modules the diagram shows aren't open at all. H3-Context-IR — which the README calls "critical to the quality of the final output" — and H3-Regenerate-2K, the 2K upsampling stage, are both hosted services you call MiniMax's API for. What's actually downloadable is the middle box, in two task-specific checkpoints (text/first-last-frame→video and reference→video), both CFG-distilled BF16. Native sparse attention, used in the final training stage, is withheld from this release as well.
## The benchmark that isn't there
I looked for a number to weigh the license against and didn't find one. There is no VBench score, no Elo comparison, no named baseline model anywhere in the blog post or the Hugging Face card. The one performance claim in the entire release is pricing, and even that has no dollar figure attached: *"At 2K, H3's per-second price is less than a third of mainstream models, and at 768p, it's less than half the price of mainstream models' 720p."* Less than a third of what? Which mainstream models? The post doesn't say.
To be precise about what's disclosed and what isn't: the architecture (encoder choice, layer sampled, transformer size, VAE compression ratios) is specific and checkable. The training data is not — "built entirely from real, natural data" is the only description given, with no token counts or dataset composition. The capability claims are entirely qualitative.
## The take
Put the two things next to each other: a model you may be legally barred from using depending on which of four major jurisdictions you're in, released with no numbers that would let you decide whether it's worth working around that restriction if you could. Neither fact is hidden — the license text is public and precise, and the absence of benchmarks is just an absence, not a false claim. But "open-source," used as freely as MiniMax uses it in the blog copy, is doing less work here than it usually does. A license with a four-jurisdiction carve-out and a revenue-gated authorization clause is a commercial license with a wide default grant, not a permissive one. Whether that's a reasonable posture for a company shipping an expensive-to-train omni-modal model is a separate question from whether it should be called "open" without the qualifier.
---
*Sources: the [MiniMax H3 blog post](https://www.minimax.io/blog/minimax-h3) and [Hugging Face model card](https://huggingface.co/MiniMaxAI/MiniMax-H3) (MiniMax, July–August 2026), including the license file's Excluded Territories clause. The figure is MiniMax's own system-overview diagram; the territory checker is my own illustration of the license's geographic scope, not a legal opinion.*
---
# pdf-inspector: classifying PDFs without a single model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/pdf-inspector
> date: 2026-08-03
> tags: open-source, rust, pdf, heuristics, explainer
Most PDF-to-text pipelines start with a coin flip: run OCR on everything and pay for it, or guess which files need it and get burned when you guess wrong. [pdf-inspector](https://github.com/firecrawl/pdf-inspector) (Firecrawl, MIT, 6.3k stars) skips the guess. It classifies a PDF as text-based, scanned, or mixed — page by page — and converts the text-based pages to Markdown, in under 200ms, with **one dependency** (`lopdf`) and **no ML models, no OCR, no external services**. Firecrawl's own number for why this matters: roughly **54% of PDFs** don't need OCR at all, and this is how you find out which 54% without running an OCR engine to check.
That's the whole pitch: a bounded, well-understood problem — is this page's text selectable? — solved with parsing and arithmetic instead of a model.
## The heuristics, not the pitch
The interesting part of pdf-inspector isn't that it's fast. It's *what it checks*. Reading `src/detector.rs`, three heuristics stand out, each with an inline rationale in the source:
**Path-op density.** A page can have plenty of drawing operators and still not have real text — some PDFs render glyphs as outlined vector paths rather than as selectable characters. The detector flags this specifically: massive path-drawing volume, almost no text-showing operators, almost no distinct characters actually rendered. Try the three conditions live:
**Font decodability, with a fallback chain.** Not every font that's *present* in a PDF is used — the detector only inspects fonts actually invoked via a text-show operator. If those fonts are Type0/Identity-H without a `ToUnicode` CMap, decoding produces garbage characters. Rather than flag that immediately, it tries two fallbacks first — CID values that happen to look like passthrough Unicode, then an embedded TrueType `cmap` table lookup — before finally giving up and marking the page `suspected_garbled_text`.
**A newspaper-layout detector.** Even a page classified `TextBased` can still need OCR: dense multi-column prose with a low font-change-to-text-op ratio reads badly as extracted text even when every character decodes correctly. The thresholds for this one are, per the source comments, calibrated against a named 50-page *Wall Street Journal* test PDF plus DPA/contract PDFs and SEC filings — a heuristic tuned against specific real documents, not an abstract rule.
None of this is a neural network. It's operator-stream scanning over the page's content stream, with page sampling (8 evenly-spread pages by default, not "bail on the first bad page" — the source comment explains why: an image-only cover page followed by dense text, like most annual reports, would trip an early-exit strategy into over-flagging OCR).
## Where it lands on a benchmark
Firecrawl's own July 31, 2026 benchmark, run on an Apple M4 Pro against the 200-PDF `opendataloader-bench` corpus, with OCR disabled and only non-ML local engines in the comparison:
pdf-inspector edges liteparse by 0.002 overall, but wins tables decisively (TEDS 0.814 vs 0.693) and is roughly **1.6× faster** (0.470s vs 0.750s for the full corpus) — while losing on headings (MHS 0.788 vs 0.811). pymupdf4llm and markitdown aren't close on tables or speed. This is a genuinely tight three-way race at the top, not a rout.
The comparison set is explicitly scoped: "only local engines without model-based PDF parsing are shown; OCR was disabled." This is not a claim of beating Docling, LlamaParse, or other vision-model-based extractors — it's the fastest option in the non-ML, non-OCR lane, and the README says so directly.
## Honest gaps
This is Firecrawl's own benchmark, on Firecrawl's own hardware, published in Firecrawl's own README — there's no independent re-run I could find, though the corpus and evaluator are public and a reproducible-results branch is published, so a third party *could* check it (I did not). It's single-machine (one Apple M4 Pro), so there's no cross-platform or server-CPU number to point to. And the version story is a little tangled: the README's benchmark says it tested "pdf-inspector 0.2.6," which matches none of the three independently-versioned language bindings cleanly (Rust crate 0.1.7, npm package 1.11.2) — normal for a multi-target Rust project, but worth knowing if you go looking for "the" version number.
## The take
There's no dramatic headline number here and no diagram to embed — the README's own "figure" is a Markdown table. What's worth taking from pdf-inspector is smaller and more useful: three specific, well-reasoned heuristics (path-op density, a font-decodability fallback chain, a newspaper-layout detector tuned against named documents) that collectively do a job people increasingly reach for an ML model to do, at a fraction of the cost, on the specific slice of the problem where a heuristic is the right tool. Not every classification problem needs a model. This one apparently doesn't.
---
*Source: the [pdf-inspector README and benchmark](https://github.com/firecrawl/pdf-inspector) (Firecrawl, MIT, refreshed 2026-07-31) and `src/detector.rs`. The heuristic thresholds and benchmark numbers are the project's own; the interactive is my reconstruction of the vector-text condition for explanation, not a copy of the crate's code.*
---
# Qwen-CUA: a computer-use agent that only ever sees pixels
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/qwen-cua
> date: 2026-08-03
> tags: agents, computer-use, reinforcement-learning, moe, systems
Qwen Team and XLang Lab published [Qwen-CUA](https://github.com/xlang-ai/Qwen-CUA) on 2026-08-02: a computer-use agent that never sees anything but a screenshot and never acts through anything but keyboard and mouse events. No DOM tree, no accessibility metadata, no task-specific API. The backbone is a 397B-A17B Qwen mixture-of-experts model, and a scaled variant, Qwen-CUA-Max, pushes past one trillion total parameters. The headline number is 86.2 on OSWorld-Verified. That is a real result, but it is one of eight benchmarks, and it is not even the most interesting fact in the paper.
The part worth taking apart is the context-management scheme that makes long-horizon screenshot-only control workable at all: fold the visual history in blocks of 10, not one screenshot at a time. It sounds like a minor implementation detail. It is actually the difference between a rollout fleet that reuses its KV-cache and one that recomputes a fresh prompt prefix on every single turn.
## A narrow interface on purpose
Most production computer-use systems cheat a little. They read the DOM, they call an accessibility API, they get coordinates for free. Qwen-CUA's interface is deliberately narrower than that — the model observes a screenshot and emits one action from a fixed keyboard-and-mouse vocabulary (Appendix A):
| Category | Actions |
|---|---|
| Keyboard | `key`, `key down` / `key up`, `type` |
| Mouse | `move`, `left`/`right`/`middle click`, `double`/`triple click`, `left click drag`, `left mouse down`/`up`, `scroll`/`hscroll` |
| Control | `screenshot`, `wait`, `terminate` (success or failure), `call user` |
That last one, `call user`, is the tell that this is meant to run unattended: when the task cannot proceed autonomously — a login wall, a genuinely ambiguous instruction — the agent is allowed to stop and ask, rather than guess and keep going.
This is close to the opposite design point from a coding harness. [Agent harnesses](/articles/agent-harness) walks Lilian Weng's tool taxonomy for coding agents — `bash`, `edit`, `grep`, `git_status` — and her argument for why: the tools are "deliberately simple and generic" because the model has already seen a million shell sessions in training. Qwen-CUA leans on the same instinct pointed at a different substrate. It does not give the model `bash` or a DOM query; it gives it exactly what a person gets — a screen and two input devices — on the bet that native computer use is "a sufficiently general interface for interacting with almost any software accessible to a person." The tradeoff is real: a shell command can rename 400 files in one call, while Qwen-CUA has to click, drag, and type its way through the same job one primitive at a time. The payoff is that the interface never goes stale — it works on software that has no API at all, which is most software.
## The mechanism: folding the visual prefix in blocks of 10
Screenshot-only control has an obvious cost: every turn adds an image to the context, and a long task can run for a hundred turns. Two bad options present themselves. Keep every screenshot, and the context blows past any practical budget. Use a sliding window and drop the oldest ones, and the agent forgets what it did five minutes ago — the exact state that explains why the screen looks the way it does now.
Qwen-CUA's answer is in Figure 3 of the paper: scale the *active* visual history to 20 screenshots (up from 1 in Qwen2.5, 5 in Qwen3, 10 in Qwen3.5 — successive Qwen generations have simply been raising this number), and once the active window would exceed 20, fold the oldest 10 screenshots at once into a fixed textual placeholder. The reasoning and actions tied to those folded screenshots stay in the conversation; only the pixels get replaced.
The "at once" is the whole design. Fold one screenshot per turn — the obvious way to enforce a 20-image budget — and the folded-prefix boundary moves on every turn, so the text before the newest screenshot is different from what it was a moment ago. Fold 10 at a time instead, and the boundary only moves every 10 turns: steps 21 through 30 all extend the exact same prefix. Step through it below.
Training uses the identical operator. Reinforcement-learning episodes are sliced into context-bounded chunks by advancing the same fold boundary, each slice inherits the full terminal reward, and only the model's own generated tokens count toward the loss. Train and inference see the same folding rule, which is the detail that keeps this from being an inference-time hack layered on top of training that never saw it.
## Why prefix stability is a rollout-economics problem
Here is the part the paper is explicit about and worth spelling out: a stable prefix is not a memory nicety, it is a **KV-cache reuse story**. An inference server that serves the same prompt prefix repeatedly can cache the attention keys and values for that prefix once and reuse them for every subsequent request that shares it — skipping the prefill compute for everything except the new tokens at the end. A prefix that changes on every turn gets none of that: every request looks new to the cache, so every request pays full prefill cost. The paper names this directly, citing Anthropic's cache-aware batched-pruning guidance for computer use as the precedent for the design.
Multiply that by scale. Qwen-CUA's training infrastructure is a cloud rollout fleet with close to 100,000 vCPUs and tens of thousands of concurrent environments, generating roughly 40,000 verifiable tasks' worth of trajectories. At that volume, the difference between "the prefix changes every turn" and "the prefix is stable for 9 turns out of 10" is not a rounding error in the compute bill — it is close to an order-of-magnitude difference in how much of the prefill work has to be redone per rollout step. Folding 10 at a time instead of 1 at a time is, in effect, a decision about how much of a 100,000-vCPU cluster's time goes to recomputing text it has already computed.
It is the same underlying instinct as the file-backed context strategy in [Agent harnesses](/articles/agent-harness) — treat context as a bounded, managed resource instead of an ever-growing transcript — aimed at a different bottleneck. Weng's harness spills durable state to files so the model's context stays flat. Qwen-CUA can't spill screenshots to a filesystem the model can `grep`; there is no text index over pixels. So it does the analogous thing structurally: collapse old state into a fixed, cheap textual stand-in, and do it in a way that happens to also keep the serving engine's cache warm. Same principle — bound what has to be reprocessed — solved with the tool available to a vision-and-text model instead of a coding agent.
## Training: verifiable rewards at rollout-fleet scale
The RL recipe is RLVR — reinforcement learning with verifiable rewards — using **Soft Adaptive Policy Optimization (SAPO)**, a smooth, temperature-gated alternative to PPO-style hard clipping (Gao et al., 2025, not original to this paper). The gate temperature is asymmetric: `τ_pos = 1.0`, `τ_neg = 1.05`, so tokens on non-positive-advantage trajectories decay faster than tokens on positive ones — called out as important specifically for long multimodal trajectories on an MoE backbone. Task-pool calibration runs 8 trial rollouts per candidate task and keeps only the ones with a mix of successes and failures, discarding tasks that are already saturated or unreachable.
| Config | Value |
|---|---|
| Group size (valid trajectories/task) | 16 |
| Oversampling before filtering | 20 candidates |
| Outer batch size | 128 prompts (up to 2,048 valid trajectories/update) |
| Optimizer | AdamW, LR `1e-6` constant, no warmup |
| SAPO `τ_pos` / `τ_neg` | 1.0 / 1.05 |
| Total updates | 1,000 |
| Max turns/episode | 100 |
| Max context (after slicing) | 144K tokens |
| Slice interval | every 10 turn-pairs |
The distributed setup is 512 H200 GPUs across 64 nodes, split disaggregated-style (32 training, 32 rollout, `verl`-style), with SGLang serving the rollout side. A full 1,000-update run takes about 5 days, roughly 61,440 H200 GPU-hours, holding upwards of 2,000 environments active concurrently at better than 75% average utilization. Across the training curve, the cross-domain validation score climbs from about 0.734 before RL to a peak of 0.770 at checkpoint 40 — the checkpoint the paper actually ships — before drifting slightly to 0.762 by the final checkpoint 50. That's a real, disclosed detail: the best model on the training curve is not the last one.
Data comes from three sources layered together: environment-interaction tasks built off a feature taxonomy, user-interactive tasks with a simulated user holding back task-specific knowledge (the OSWorld 2.0 setting), long-horizon tasks chained through verifiable phase states, and personalized workflows collected from human trajectories in everyday and professional software — CAD tools and Blender included — with reasoning reconstructed via model-assisted chain-of-thought from the raw (task, screenshot, action, resulting state) tuples.
## Where it actually lands: eight benchmarks, not one
The 86.2 on OSWorld-Verified is real, and it is the best score in the set on that particular benchmark. It is also the exception. Across the other seven benchmarks the picture is more mixed — Qwen-CUA leads outright on two of eight, is close behind on several, and loses outright on the rest, most notably safety. Pick a benchmark:
Scaling the same recipe to Qwen-CUA-Max (over 1 trillion total parameters) moves OSWorld-Verified from 86.2 to 87.6, and helps more on partial-credit long-horizon completion:
1T)", value: 87.6, highlight: true },
]}
/>
On the safety benchmark, RedTeamCUA, Qwen-CUA is a clear improvement over its own predecessor and a clear loss against Claude Opus 4.8. RedTeamCUA runs indirect prompt injection through ownCloud, Rocket.Chat, and Reddit environments and jointly reports benign task success and attack success rate (ASR — how often the injected instruction actually hijacks the agent):
A 20.2-point reduction in attack success versus the previous Qwen generation is a genuine gain. It is also more than 20 times Opus 4.8's ASR. The paper states its own limits plainly here: "RedTeamCUA therefore shows improved resistance to indirect prompt injection, not a deployment-safety guarantee." Worth repeating rather than softening.
## Efficiency: the gain is not longer reasoning
One honest, checkable claim in the paper: Qwen-CUA's OSWorld-Verified score does not come from generating more tokens per task. It reaches 86.2 at 3,605.8 output tokens per task; Claude Opus 4.8 needs a similar budget to reach 80.0 and roughly 21,800 tokens to reach 83.3.
The second panel is where the paper pre-empts its own obvious gotcha. On OSWorld 2.0, Qwen-CUA averages 218.9 turns per task against 83.5 for GPT-5.5 and 105.7 for Opus 4.8 — a turn count that looks far worse. But GPT-5.5 and Opus 4.8 can batch several actions into one turn; Qwen-CUA emits exactly one native action per turn by construction. The turn-count gap is mostly an artifact of how each interface packages low-level actions, not evidence that Qwen-CUA needs more attempts to do the same work. The paper says as much itself rather than leaving a reader to work it out. A related experiment adds a Bash tool alongside native computer use on MyPCBench: trajectories get shorter for every model tested, but task completion drops too, for Qwen-CUA and Qwen-3.7 specifically — the paper frames this as an unresolved "capability-efficiency frontier," not a win.
## Grading your own exam
Here is the fact that belongs next to the 86.2, not three pages after it: **XLang Lab built OSWorld and OSWorld-Verified, and XLang Lab co-authored this paper.** The lab that defines what counts as a passing score on the headline benchmark is also a lab reporting how well its own model does on that benchmark. The paper does not flag this anywhere as a conflict of interest — it is simply true of the author list and the benchmark's provenance, stated here as a fact about who is grading whom, not as an accusation of anything specific.
The eval protocol has a second, quieter honesty issue: baselines are not run under matched inference budgets. Per the paper's own settings, Qwen-3.7 is evaluated in non-thinking mode, GPT-5.5 runs with `xhigh` reasoning effort, and Claude Opus 4.8 runs at its max inference setting. "Most scores for comparison models are taken from official reports released by the corresponding benchmark or model providers" — for the ones the authors reproduced themselves, the settings differ by model, and the paper does not report what a matched-budget comparison would look like. The Gym-Anything table is the one place a second Opus 4.8 setting (medium) appears alongside max, and the two settings score 43.7 versus 47.3 — a 3.6-point swing from inference budget alone, which gives some sense of how much slack "differing settings" can hide.
There is a third thing worth naming that I could not find explained anywhere in the paper. Figure 1's legend lists six systems, not the four in Table 1 and everywhere else in the text — it adds **Muse-Spark-1.1**, scoring 80.8 on OSWorld-Verified and 47.3 on Gym-Anything. Searching the full extracted paper text, that name appears exactly once: in the Figure 1 legend. It is not in Table 1, not in the eval-settings section, not in the references, not identified anywhere else in 24 pages. I don't know what it is or why it only appears in one chart.
Two more disclosed-but-real caveats round this out. MacAgentBench's "clock" domain scored 0.0% for every model across all 12 tasks; the paper reports manually inspecting the trajectories, finding they looked like correct completions, and keeping the official 0.0% score anyway rather than quietly correcting it — which means the reported 69.2 aggregate is very likely a slight undercount, in Qwen-CUA's favor by omission, and the paper says so. And Gym-Anything's headline 46.3 runs on 97 of 197 possible environments; the other 100 were excluded because their Windows, Android, or Linux setups didn't work, not because they were held out for any principled reason. Both caveats are in the paper. Neither is in the abstract.
Finally: several contributors are marked in the author list as having departed the Qwen Team by the time this was published, including researchers who worked on the original OSWorld and OpenCUA lines. The paper doesn't explain the departures, and neither can I — it's listed here because it's a real, checkable detail about who built this and who was still there to see it ship.
## The take
Native computer use is not new — UI-TARS, OpenCUA, Aguvis, and AutoGLM already established that a single model can ground pixels to actions without a separate grounding stage. What Qwen-CUA adds is mostly an engineering answer to what happens when you actually try to run that idea at rollout-fleet scale: fold visual history in blocks, not one screenshot at a time, so a 100,000-vCPU cluster spends its time on new work instead of recomputing prefixes it already has. The paper's own framing for where this goes next is worth keeping: "we view native computer use not as the only action interface, but as the universal grounding and fallback layer of a hybrid agent" — paired eventually with something more like the coding-harness tool table in [Agent harnesses](/articles/agent-harness), not replacing it. The honest scorecard is two benchmark wins out of eight, a real safety improvement that still trails the safest competitor by more than 20x on attack success, and a headline number graded in part by the lab that wrote the exam. All three of those things can be true about a genuinely useful piece of systems engineering at the same time.
---
*Built on Qwen Team & XLang Lab's [Qwen-CUA: Native Computer Use for (almost) Everything](https://github.com/xlang-ai/Qwen-CUA) (2026-08-02). Figures 1, 3, and 6 are reproduced from the paper for commentary, flattened onto white and cropped from the original PDF; the interactive fold timeline and benchmark explorer are my own illustrations of the mechanism and Table 1's data, not measured traces. Benchmark numbers are as reported in the paper.*
---
# Qwen3.8-Max: 16 days, 265 commits, zero humans in the loop
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/qwen3-8-max
> date: 2026-08-03
> tags: qwen, agents, llm, benchmarks, moe
Alibaba announced [Qwen3.8-Max](https://qwen.ai/blog?id=qwen3.8) today: **2.4 trillion parameters, 95B active**, built on the architectural foundation of Qwen 3.5. It is also, by Alibaba's own framing, the first Qwen-Max-class model getting open weights at all — those weights are announced, not released; they ship "next week." Right now the only way to use Qwen3.8-Max is the API, through [QwenCloud](https://www.qwencloud.com/).
A 2.4T/95B split puts activation sparsity at about 25× (2,400 / 95). That is close to [Kimi K3](/articles/kimi-k3)'s roughly 27× (2.78T total, 104.2B active) — two labs, released weeks apart, converging on almost the same ratio of total-to-active parameters at the very top of the open-weight-adjacent scale. Where K3 backs that ratio with a 47-page technical report anyone can audit against a released `config.json`, Qwen3.8-Max's architecture claims are, for now, a paragraph in a blog post. The open weights next week will be the point where the second half of that comparison becomes checkable.
But the benchmark grid is not the interesting part of this release. The interesting part is what Alibaba says the model did **completely unsupervised**, for days at a time.
## The case studies are the real headline
Every frontier lab now publishes agentic benchmark numbers. Fewer publish concrete, checkable claims about what their model actually built when nobody was watching it. Qwen3.8-Max's release includes five of those, spanning a 24-hour coding contest to a 365-simulated-day economy, and they are, collectively, some of the most specific long-horizon autonomy claims I have seen from any lab this year — specific enough that at least one of them (the coding harness) has a public commit history you can go read yourself.
### 16 days, no one watching: oh-my-cli
Alibaba tasked Qwen3.8-Max with building `oh-my-cli` — a CLI tool — from an empty repository, and kept it running. The loop it built for itself: an issue state machine moves work through `ready → leased → active`; an agent claims a task, implements it, and triggers Build, Unit Test, E2E, and Desktop Lifecycle validation; failures route back to the originating issue for another pass; passing PRs merge. Community feedback and the model's own test results both feed back in as new issues, so the harness is quite literally evolving its own capabilities (`/goal`, `/resume`, Dynamic Workflow, Session Replay, Desktop) as it runs.
As of July 30, 2026 — about 16 days in — the repository held **265 commits, 127 PRs, and 151 issues**, all without a human merging, reviewing, or filing anything. What makes this claim unusually checkable is that the trace is public: [github.com/qwen-code-dev-bot/oh-my-cli](https://github.com/qwen-code-dev-bot/oh-my-cli). Most "our agent ran autonomously for weeks" claims ask you to take the vendor's word for it. This one, you can go read commit-by-commit.
### Reproduce a paper, then beat it
Handed only a citation — [arXiv 2605.22389](https://arxiv.org/abs/2605.22389), "Unified Data Selection for LLM Reasoning" — and a set of GPUs, Qwen3.8-Max had to write the entire pipeline from nothing: no starter code, no scaffold. The paper's claim is that when you have more training data than compute to use it on, the examples worth keeping are the ones full of "hard decision points" — places in a worked solution where the model was genuinely torn between next steps.
Over roughly 125 hours (about five days) of continuous, unattended work, Qwen3.8-Max wrote about **7,600 lines of code**, took **over 1,100 actions**, and ran **33 rounds of GPU training**. The first ~37 hours went into rebuilding the paper's pipeline from zero and reproducing all six of its findings — including the headline result, that the paper's selection method beats random selection by +7.7% on AIME24 after fine-tuning Qwen3-8B on the selected data.
Then it kept going. The next ~88 hours ran a self-improving loop — form a hypothesis, write the code, run it on GPUs, analyze the result, try again — across four rounds and 18 self-generated ideas, each round's diagnosis shaping the next round's hypothesis:
| Round | Best idea that round | AIME24 | Gain vs. baseline |
|---|---|---|---|
| — | Paper's method, reproduced (baseline) | 49.58% | — |
| 1 | Split the data by difficulty before selecting | 50.42% | +0.84 |
| 2 | Weight examples by an entropy–score gap | 51.67% | +2.09 |
| 3 | Tune the selection width | 51.25% | +1.67 |
| 4 | Count the hard decision points ("nhighgate") | **52.29%** | **+2.71** |
The final method — round 4's "nhighgate" idea — is not a tweak Alibaba fed it. It is something the model proposed, tested, and kept because it worked better than the paper it was asked to reproduce.
### Beat 526 human teams in 24 hours
Alibaba entered Qwen3.8-Max into a live competition on its own Tianchi platform — the [WWW2025 Multimodal Dialogue Intent Recognition Challenge](https://tianchi.aliyun.com/competition/entrance/532277) — against **526 human teams**, under a strict 24-hour clock. The task: read customer-service conversations, text and screenshots both, and classify what the customer actually wants.
Qwen3.8-Max read the rules and built a solution in code: fine-tuned and ensembled BERT, MacBERT, and RoBERTa for the text side; fine-tuned Qwen2.5-VL-7B, backed by Chinese-CLIP for uncertain cases, for the screenshots; then fused everything through a weighted-voting system, calibrated by cross-validation and re-weighted after every submission. Across **45 submissions** in 24 hours, accuracy climbed from **0.60 to 0.853**, finishing ahead of **458 of the 526 human teams — 87% of the field**.
## Autonomous chip design: a staircase that keeps stepping down
The most technically specific of the five case studies is silicon. Qwen3.8-Max was given a stub RTL workspace, an empty-module task description, and an evaluation script — no golden reference design, no human in the loop — and told to design a GCD/RSA cryptographic hardware accelerator that stays bit-exact across 4-, 6-, 8-, and 16-bit configurations under a randomized `cocotb` verification framework, while minimizing synthesized gate count.
Over one continuous run of **about 500 turns and 71 evaluations across 13 milestones**, working entirely in a sandbox wired to Iverilog (simulation), Yosys (synthesis), and OpenROAD (physical layout), the model's first functionally correct design measured 8,298 gates. It drove that down to **678 gates** — a 91.8% reduction — and the milestones along the way show real architectural rewrites happening deep into the run, not just early low-hanging fruit:
The single biggest step, by far, was recognizing that the 16-bit hardware modulo divider inside `modular_multiplier` could be replaced with an iterative shift-subtract architecture — one change, at turn 22, cutting 6,288 gates in a single move, over 80% of the entire reduction. Everything after that is smaller, later, and arguably harder: register and FSM pruning at turns 60–113, module fusion at turns 170–252, and gate-level refinement all the way out to turn 500. A model that only found the big early win and then plateaued would be a much less interesting story than one that kept finding real (if progressively smaller) structural improvements for 400 more turns.
Alibaba then re-ran the final RTL through a physical place-and-route flow (OpenROAD, Nangate45) to check whether the front-end gate-count win actually routes. It does: the die shrank from 106×106 to 46×46 µm² (−81%), wirelength dropped from 33,369 to 4,187 µm, and the design closed timing at 500 MHz with **positive** slack (+0.66 ns), up from a failing −4.46 ns at the start. Optimizing gate count without checking place-and-route is a common way to produce a design that looks good on paper and doesn't actually work in silicon; Alibaba closed that loop.
## 365 simulated days of running a business
The last of the five case studies is not a coding task at all. **E-Commerce Bench** simulates a full year of operating online stores against desensitized real Taobao/Tmall transaction data — 12 store types, 60 product categories, nearly 600 suppliers, 7,000 products — starting from ¥100,000 in capital. The model has to choose products, negotiate with suppliers, manage inventory, price dynamically, and handle returns, all while surviving seasonal demand swings, sudden supply shocks (typhoons, material shortages), and a settlement system with real cash-flow pressure. Buried in the supplier matrix: **152 fraudulent merchants**, running classic scams — membership-fee traps, low-price bait, goods not as described.
Two things stand out in how Qwen3.8-Max played this. First, supplier negotiation is modeled with distinct personalities and concession strategies per supplier, and Qwen3.8-Max's negotiation efficiency measurably *improved* over the year — the same products from the same suppliers got progressively cheaper, and that experience generalized to similar products, where Alibaba says other models' negotiation efficiency plateaued mid-year. Second, it front-loaded capital early to establish position rather than playing conservatively, then converted the resulting inventory and operating gains back to cash before the simulation ended — timing that matters because unconverted assets left on the books at year-end hurt the final score.
The result: a final balance of **¥416,252 — a 4.16× return** — 38% ahead of second-place GLM 5.2 and 152% ahead of Qwen3.8-Max's own predecessor, Qwen3.7-Max. Alibaba frames this as evidence of "adaptive learning from transactional feedback" across more than 2,000 rounds of interaction, rather than a model that locked in a strategy early and rode it out. That framing is plausible given the negotiation-efficiency detail above, but it is worth remembering this is Alibaba's own simulation, built on Alibaba's own marketplace data, scored by Alibaba.
## Scaling real-world work
Underneath all five case studies is an infrastructure bet Alibaba is explicit about: jointly scaling RL environments and compute lifts "general working competence" across several harnesses at once (QwenWork, Claude Code, Codex, OpenClaw, Hermes), and doing that required three things to scale together rather than one at a time — environments along independent axes (task, workspace, harness) so growth compounds instead of requiring bespoke integration per new environment; a **universal reward system** unifying execution-based checks, rubric-conditioned judging over text and rendered visual output, and agentic inspection, so there is one reward mechanism instead of a pile of task-specific verifiers; and an **online data balancer** that keeps every training batch balanced across task, difficulty, workspace, and harness, which is what keeps gradient variance from blowing up RL training at scale.
That chart is worth reading carefully rather than just squinting at the upward trend: the curve peaks at 4,000 environments (0.725) and is already *down* to 0.689 by 5,000 — the shipped checkpoint is not the last one on the curve, which is the kind of disclosed detail that makes the rest of the curve more credible, not less.
That second chart is the practical payoff of training against a harness-agnostic reward system: Qwen3.8-Max does not have one harness it happens to be tuned for. Point it at QwenWork, Claude Code, Codex, OpenClaw, or Hermes and the CoWorkBench score moves in a band of about 73–76; Fable5 and Opus4.8, each shown in only their own native harness, land in a similar range without ever being tested for harness portability the same way. This is directly the concern [The harness effect](/articles/harness-effect) raises from the other side — that orchestration, not the model, is what actually determines an agent's cost and reliability on a task — and Qwen3.8-Max's answer is to train the reward system to not care which harness is wrapped around it, rather than picking one harness and optimizing hard for it.
The same Dynamic Workflows capability that lets it self-orchestrate shows up in a quant-research vignette Alibaba includes alongside the five headline case studies: given a one-line task description, Qwen3.8-Max built a complete ETF-rotation strategy over several hours, pruning overfit factors when it noticed design-period and validation-period metrics diverging, and separately parallelized factor mining from six short descriptions into 50 research directions each, dispatching roughly 330 sub-agents through about 6,000 backtests to find factors with excess Sharpe ratios of 0.64–1.48. Whether that generalizes past a demo is unverifiable from a blog post, but the mechanism described — noticing an overfitting signal and automatically triggering pruning, mid-run, without being told to — is the same "acting on evidence instead of a fixed script" pattern that shows up in the chip-design and paper-reproduction case studies above.
## Multimodal and hybrid agents
Qwen3.8-Max's visual pipeline gets a similar "watch itself work" framing: while executing a task, the model inspects its own intermediate results — page layout, object orientation, spatial relationships, animation quality — and revises when something looks wrong (a television facing backward, a misaligned interface). Alibaba's phrase for this is a "native feedback loop across planning, execution, verification, and iteration," which is a reasonable description if the examples given hold up, though none of them are independently reproducible from the blog post alone.
The concrete new benchmark here is **RecreationBench**: the model observes a real running application as a black box — no source code, no network access — across five platforms (Ubuntu, macOS, Windows, Android, web), and has to rebuild the whole thing from scratch through interaction alone. Alibaba frames Qwen3.8-Max's showing here as "frontier-level Hybrid Agent capability" — the pairing of writing code (does the heavy lifting) with operating a GUI directly (reaches whatever a human can see and click, and reports back what a live system actually does).
That second half — driving a computer through screenshots and input events alone — is exactly [Qwen-CUA](/articles/qwen-cua)'s whole premise, published by a different Qwen team one day earlier. It is worth putting the two numbers next to each other: Qwen3.8-Max reports **86.1** on OSWorld-Verified; Qwen-CUA, a dedicated 397B-A17B computer-use specialist trained specifically for this, reports **86.2**. A general-purpose 2.4T model and a purpose-built computer-use agent land within a tenth of a point of each other on the benchmark that agent was built for — which either means Qwen3.8-Max's general agentic training has genuinely absorbed computer-use skill, or that OSWorld-Verified has a ceiling both are bumping into. Both readings are consistent with the data; the blog post doesn't say which.
## The benchmark tables — and where they don't hold up
Here is the full picture, reproduced from Alibaba's own release. The pattern is not "Qwen3.8-Max wins everything" — it wins some things outright, loses some things clearly, and several of its best numbers come with an asterisk worth reading before you trust them.
### Coding Agent
| Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 84.6 | 84.6 | 88.8 | 74.5 | 86.6 |
| SWE-bench Pro | 69.2 | 80.0 | 64.6 | 60.6 | 67.7 |
| DeepSWE 1.1 | 59.0 | 70.0 | 73.0 | 21.6 | 56.6 |
| NL2Repo-Bench | 69.4 | -- | -- | 47.2 | 55.9 |
| FrontierSWE | 70.0 | 88.8 | -- | 40.7 | 73.5 |
| MLS-Bench-Lite | 42.8 | 49.9 | 46.2 | 31.7 | 41.0 |
| PaperBench | 80.3 | 88.8 | 90.5 | 64.8 | **93.0** |
| AndroidBench | 69.8 | 84.5 | 74.0 | 56.5 | 75.1 |
| QwenSWEBench | 84.0 | 86.3 | 73.5 | 63.4 | 80.7 |
| QwenQoderBench | 62.7 | 63.1 | 53.8 | 36.8 | 58.4 |
| QwenReactBench | 1694 | 1770 | 1564 | 1538 | 1724 |
| QwenSVGBench | 1648 | 1690 | 1758 | 1499 | 1713 |
### General Agent
| Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max |
|---|---|---|---|---|---|
| CoWorkBench | 72.3 | 75.9 | 71.5 | 64.6 | 74.8 |
| WorkSpaceBench | 66.8 | 68.7 | 65.6 | 61.4 | 67.7 |
| JobBench | 48.4 | 57.4 | 45.4 | 31.3 | 53.4 |
| SkillsBench | 65.1 | 70.9 | 73.5 | 61.2 | 70.2 |
| Agents' Last Exam (Pass / Score) | 27.0 / 45.1 | -- / -- | 30.6 / 53.6 | 11.8 / 31.1 | 27.0 / 52.4 |
| Automation-Bench (Pass@1) | 27.2 | 29.1 | 29.7 | 14.2 | 27.3 |
| Toolathlon Verified (Pass@1) | 76.2 | 77.9 | 74.9 | 49.7 | 72.5 |
| WideSearch | 72.9 | 81.2 | -- | 75.2 | **81.9** |
| HLE w/ tools | 57.9 | 64.5 | 58.0 | 53.5 | 56.2 |
### General Capabilities
| Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max |
|---|---|---|---|---|---|
| GPQA Diamond | 92.0 | 92.6 | 94.1 | 92.4 | 92.6 |
| HLE | 45.7 | 53.3 | 47.2 | 41.4 | 43.6 |
| IFBench | 62.2 | 63.5 | 72.7 | 79.1 | **82.8** |
| $OneMillion-Bench (expert score) | 41.8 | 55.9 | 53.8 | 44.4 | 52.5 |
| HealthBench | 52.4 | -- | 55.3 | 54.5 | **60.2** |
| PLawBench | 69.6 | 70.2 | 72.3 | 58.9 | **73.2** |
| PRBench-Legal | 52.7 | 57.6 | 57.6 | 48.5 | 57.6 |
| PRBench-Finance | 51.9 | 55.8 | 55.5 | 46.8 | **58.3** |
| MRCR v2 256K (8-needle) | 83.2 | -- | 93.8 | 86.7 | 92.9 |
| LongBench v2 | 69.1 | -- | 67.1 | 65.3 | 66.3 |
**Read the footnotes before you trust the wins.** Alibaba's own notes on this table say, plainly:
- **Terminal Bench 2.1**: Qwen3.8-Max is evaluated with Claude Code at avg@10 (5h timeout, 131,072 max tokens). Every other model is scored at "the best published score across harnesses" — Opus4.8/Fable5 via Artificial Analysis, GPT5.6 Sol via OpenAI's own post. Best-of-published vs. one model's avg@10 is not the same measurement.
- **SkillsBench**: a different harness per model — Opus4.8 and Fable5 on Claude Code, GPT5.6 Sol on Codex, the entire Qwen series on OpenCode.
- **DeepSWE 1.1**: Qwen3.8-Max is scored on whichever of Claude Code / mini-SWE-agent is higher, and Alibaba notes it does best specifically on Claude Code — the harness closest to what it trains against.
- **PaperBench**'s 93.0 — the highest score in the table — is judged by **Claude Opus 4.6**, a competitor model, not an automated or human grader.
- **$OneMillion-Bench** and **PLawBench** are both judged by **gemini-3.1-pro-preview**.
- Footnote 1, verbatim: "Fable5 results may involve fallbacks."
- QwenSWEBench, QwenQoderBench, QwenReactBench, QwenSVGBench, CoWorkBench, and WorkSpaceBench are **Alibaba's own in-house benchmarks**, evaluated in-house.
None of this means the wins are fake. It means "Qwen3.8-Max leads Terminal Bench 2.1" and "PaperBench's judge is a Claude model" are both true at once, and a benchmark table alone won't tell you that — you have to read footnote 2 and footnote 8.
Explore the same numbers benchmark-by-benchmark, with the eval-setup note attached to whichever row has one:
### Multimodal (selected rows)
The full multimodal table runs to roughly fifty rows across six categories; here are the ones that matter most for the agentic and visual-agent story above, using the table's own column set (Gemini3.1-Pro and Qwen3.7-Plus replace GPT5.6 Sol/Qwen3.7-Max from the tables above — Alibaba compares against a different baseline set for multimodal).
| Benchmark | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max |
|---|---|---|---|---|---|---|
| MMMU-Pro | 75.6 | 81.2 | 80.5 | 83.0 | 79.0 | 82.3 |
| LogicVista | 76.7 | 85.7 | 82.6 | 89.7 | 84.3 | **91.9** |
| HiPhO | 69.3 | 78.6 | 85.4 | 86.8 | 84.1 | **90.0** |
| OSWorld-Verified | 83.4 | 85.0 | 76.2 | 83.2 | 73.3 | **86.1** |
| OSWorld 2.0 (binary / partial) | 20.6 / 54.8 | -- / 66.1 | 7.8 / 30.6 | -- / 62.6 | 2.8 / 21.5 | 19.4 / 46.7 |
| WebArena-Verified | 67.9 | 71.3 | 64.3 | 69.7 | 55.3 | 66.8 |
| Parametric CAD Bench | 85.1 | 87.5 | 73.5 | 86.2 | 73.8 | **91.5** |
| VLMsAreBiased | 43.8 | 61.2 | 74.1 | 59.8 | 36.6 | **88.3** |
| Dense200 | 20.8 | 31.1 | 69.7 | 55.3 | 60.7 | **87.0** |
**Where it clearly loses, so this isn't a highlight reel:** DeepSWE 1.1 (56.6 vs. GPT5.6 Sol's 73.0, Fable5's 70.0), SWE-bench Pro (67.7 vs. Fable5's 80.0), HLE (43.6 vs. Fable5's 53.3), HLE w/ tools (56.2 vs. Fable5's 64.5), MLS-Bench-Lite (41.0 vs. Fable5's 49.9), Toolathlon Verified (72.5 vs. Fable5's 77.9), WebArena-Verified (66.8 vs. Fable5's 71.3), and OSWorld 2.0, where Fable5's own partial score (66.1) is well clear of Qwen3.8-Max's 46.7. Several of these — DeepSWE, SWE-bench Pro, HLE w/ tools — are exactly the categories where the harness or judge asymmetries above cut in Qwen3.8-Max's favor elsewhere, which makes the clean losses more credible, not less.
The recurring problem across both tables is one [The harness effect](/articles/harness-effect) names directly: orchestration changes the number as much as the model does, so a table that scores different models on different harnesses is measuring two things at once and reporting only one. [Agent harnesses](/articles/agent-harness) makes the complementary point about what a harness actually *is* — tools, context management, control flow, an evaluator — which is exactly the layer these footnotes are quietly holding constant for some models and not others. None of that makes the underlying capability claims false. It does mean the honest reading of "Qwen3.8-Max leads Terminal Bench 2.1" is "leads it, evaluated differently than the models it's compared against" — a real result, with an asterisk that Alibaba, to its credit, discloses rather than hides.
## Getting it (or not)
Right now, Qwen3.8-Max is API-only, via QwenCloud. The API exposes a `reasoning_effort` parameter with three levels — `xhigh` (default, for demanding tasks), `medium`, and `low` — and `preserve_thinking` is on by default. The notable integration detail: QwenCloud's API is compatible with both the OpenAI and **Anthropic** protocols, so pointing Claude Code at Qwen3.8-Max is a matter of setting `ANTHROPIC_BASE_URL` and an auth token, no separate client needed. It also plugs into Codex, Qoder CLI, Qwen Code, and OpenClaw with similarly small config changes.
The open weights are, again, announced for "next week" — not this release. Until they land on Hugging Face and ModelScope, every claim in this article about what a 2.4T/95B model *is* rests on Alibaba's blog post and API behavior, not an inspectable checkpoint. That is a materially weaker evidentiary position than Kimi K3's, where the weights, a technical report, and a `config.json` all shipped together. Worth remembering the next time "first Qwen-Max-class open-weight model" gets repeated as though the weights were already out.
## The take
Strip away the marketing framing and what is left is genuinely interesting: a model that, by its maker's account, ran a coding project unsupervised for 16 days with a public commit trail, rebuilt and then beat a research paper's method from a bare citation, out-negotiated a game-theoretic supplier matrix for a simulated year, and found real architectural wins in a chip design 400 turns after the obvious ones were gone. If even most of that holds up, it is a meaningfully more concrete set of long-horizon-autonomy claims than "our model scored X on benchmark Y."
But the benchmark tables sitting next to those case studies are graded on a harness-by-harness, judge-by-competitor-model, in-house-benchmark basis that Alibaba discloses in footnotes rather than in the headline number — Terminal Bench 2.1 at avg@10 against everyone else's best-of-published, PaperBench judged by Opus 4.6, PLawBench judged by Gemini. That is not disqualifying. It is the same asymmetry every major lab's self-reported benchmark table has, and Alibaba's footnotes are, if anything, more forthcoming than most about exactly where the comparison stops being apples-to-apples. Read the case studies for what the model can apparently do unsupervised. Read the tables — and their footnotes — for how much weight the number itself can actually carry. They are not the same kind of evidence, and this release is unusually clear about which is which.
---
*Built from [Alibaba/Qwen's Qwen3.8-Max release post](https://qwen.ai/blog?id=qwen3.8) (2026-08-03). Figures 1–3 are the post's own images — the overall performance grid, the RL-scaling curve, and the cross-harness generalization chart — flattened onto white for dark-mode compatibility and capped near 1600px; not reassembled or relabeled. The gate-count staircase and case-study switcher are original interactive reconstructions of the source's own numbers, not independently measured. Benchmark tables are Alibaba's; footnote caveats quoted or closely paraphrased from the source's own numbered notes. Weights and technical report were not available at publication time — every architectural and infrastructure claim here is Alibaba's, unverified against a released checkpoint.*
---
# Recursive Harness Self-Improvement: beat your last harness, not a population of them
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/recursive-harness-self-improvement
> date: 2026-08-03
> tags: agents, harness-optimization, information-theory, llm, explainer
[Agent harnesses](/articles/agent-harness) argued that the loop wrapped around a model — tools,
context policy, control flow — matters as much as the model's own intelligence. [The harness
effect](/articles/harness-effect) showed that orchestration, not the model, is what actually sets an
agent's token bill. Both pieces treat the harness as a thing worth engineering carefully by hand.
[Recursive Harness Self-Improvement](https://arxiv.org/abs/2607.15524) (Lee, Xu, Seely, Lee, Zaharia,
Tang — Sakana AI and UC Berkeley) asks the next question: can the harness improve *itself*? Their
answer treats the harness as a single text prompt and updates it using nothing but a comparison
against its own immediately-previous version.
## The idea, and the objective it can't afford
The harness a coding agent runs under — roles, instructions, and the workflow connecting them — is,
in RHI's framing, just a string $H$ drawn from a space of harnesses $\mathcal H$. Optimizing it
against a broad population of competitors is the obvious move, and it's what most prior work does:
$$
H^*_x \in \arg\max_{H \in \mathcal{H}} f_x(H), \qquad
f_x(H) = \mathbb{E}_{H' \sim \mu,\; y \sim \mathcal{A}(H,x),\; y' \sim \mathcal{A}(H',x)}
\big[\mathbf{1}\{y \succ y'\}\big]
$$
$\mu$ is a distribution over competitor harnesses, $\mathcal A(H,x)$ is the agent running harness $H$
on task $x$, and $y \succ y'$ means an LLM judge preferred output $y$. The problem is cost: a
population of size $m$ needs $m$ fresh agent executions and $\binom{m}{2}$ pairwise judgments per
iteration — $\Theta(m^2)$ — before you can even take one optimization step. For a user continually
specializing a harness to a new task, that's not a research inconvenience, it's prohibitive.
RHI's relaxation replaces the population with a point mass on the harness's own previous version:
$$
\tilde{f}_x^{(i)}(H) = \mathbb{E}_{y \sim \mathcal{A}(H,x),\; y^- \sim \mathcal{A}(H_x^{(i-1)},x)}
\big[\mathbf{1}\{y \succ y^-\}\big]
$$
One new execution, one comparison, cached forever after. $\Theta(1)$ per iteration, independent of
how large a population you'd otherwise have wanted.
## Why comparing to yourself is still principled
The obvious objection: isn't comparing only to your immediate predecessor a much weaker signal than
comparing to a whole population? RHI's answer is a Bradley-Terry argument. Assume there's a latent
task utility $u_x : \mathcal H \to \mathbb R$ and a link function $\sigma$ (strictly increasing,
$\sigma(0) = \tfrac12$) such that $\Pr(H \succ H') = \sigma(u_x(H) - u_x(H'))$ — the standard
pairwise-preference model. Then both objectives are monotone in the *same* latent utility:
$$
f_x(H) = \mathbb{E}_{H' \sim \mu}\big[\sigma(u_x(H) - u_x(H'))\big], \qquad
\tilde{f}_x^{(i)}(H) = \sigma\big(u_x(H) - u_x(H_x^{(i-1)})\big)
$$
So any revision that beats $H^{(i-1)}_x$ with probability greater than one-half also increases the
ideal, population-level objective. RHI performs **noisy local ascent** on the same utility ordering a
much more expensive search would climb — it just takes a smaller, cheaper step each time, using the
accumulated preference history as the only signal for which direction is up. There's no proof this
converges, or how fast; it's a directional argument, not a guarantee.
The algorithm this licenses is short. At iteration $i$: run the agent under $H^{(i)}$, get an output.
Compare it against the cached output from $H^{(i-1)}$. Save the preference. Feed the accumulated
preference history to an LLM harness optimizer, which writes $H^{(i+1)}$.
Critically, the harness optimizer never sees the evaluation prompt $x_{eval}$ directly — only the
preference history, which was itself generated by a judge conditioned on $x_{eval}$. Alignment with
the actual evaluation criteria happens indirectly, through the accumulated comparisons, not because
the optimizer was told what's being graded.
## What actually gets rewritten
RHI decomposes the harness into **agent design** (roles and instructions for each candidate agent)
and **agent workflow**, which splits further into **contracts** — what information passes between
subagents and the orchestrator — and **hops** — the interaction structure and control flow. The
optimizer's own prompt is explicit about where to spend its edits: prioritize contracts and hops over
roles and instructions.
The hypothesis behind that priority: a task-specific contract tells the orchestrator and subagents
what to pass along instead of making them condition on the entire interaction history, which is
"conceptually analogous to imposing a task-dependent sparsity pattern on inter-agent information
flow" — sparse attention for agent communication, in effect. Better contracts should mean less
redundant context, better cache efficiency, and lower cost, for free, alongside better task
performance.
## Does it work, and what does it cost
Across 30 synthetic ML-research tasks (finance, robotics, pharma), a few RHI iterations
substantially raise the ceiling that test-time scaling alone can reach. With Opus-4.7, one iteration
is enough to beat both `xhigh` and `max` reasoning-effort settings. With Opus-4.8, two iterations beat
`xhigh`, `max`, **and** the provider's own built-in dynamic multi-agent harness, `ultracode` — a
user-constructed, prompt-level harness beating a vendor's dynamic scaffold.
The gains aren't from longer outputs — normalized token usage stays roughly flat across iterations
for Sonnet-4.6 and Opus-4.8 while win rate climbs (Opus-4.7's data can't separate the two hypotheses;
only two iterations were run and its token count rose alongside performance, which the paper states
plainly as inconclusive). What actually improves is cost, largely through less redundant cache
read/write from better-managed context:
The 60% figure is the abstract's headline, and it's the comparison against the provider's own
dynamic multi-agent harness — not against a same-family reasoning-effort setting. A companion
ablation (Appendix A) found something the paper didn't have to report: the provider's built-in
multi-agent mode scores a **lower** Elo than running single-agent, despite costing far more —
the vendor's own dynamic scaffold failing to pay for itself on this benchmark, stated without
softening.
## The information-theoretic account
Section 6.3 goes further than "it works" and proposes *why*: RHI implicitly maximizes task
information in the components it's told to prioritize (contracts, hops) while minimizing
task-conditional redundancy across all components. Formalized,
$$
J(g_i) = \underbrace{\sum_{hc \in \mathcal{C}_{ext}} \frac{1}{K^{hc}_{Xi}} \sum_{k=1}^{K^{hc}_{Xi}}
I\big(z^{hc,(i)}_{Xk}; X\big)}_{f_{ext}} \;-\; \beta \underbrace{\text{TC}\big\{z^{hc,(i)}_{Xk}\big\}
\big|X}_{f_{int}}, \qquad \beta > 0
$$
$f_{ext}$ is mutual information between the *externally-emphasized* components (contracts, hops) and
the task $X$; $f_{int}$ is the task-conditional total correlation — redundancy — across *all* four
component types. The hypothesis: RHI is implicitly raising the first term and lowering the second,
estimated here with canonical-correlation mutual information and total correlation over PCA-whitened
sentence embeddings.
The paper is careful about how much weight this deserves: it "does not prove that RHI optimizes a
unique scalar objective," should be read as "an embedding-based proxy," and is explicitly
"correlational rather than causal" — not a claim about what the optimizer LLM is actually doing
internally, just a consistent pattern in what its edits produce.
**Two honesty gaps worth stating plainly, because the paper doesn't headline them.** First, in the
Opus-4.8 experiments, one of the two LLM judges is **`opus-4.8-xhigh` itself** — the same model
family whose harness is being evaluated also serves as one of its own judges, scoring `opus-4.8` runs
against other `opus-4.8` baselines. The paper does average across two judges (the second is an
independent `gpt-5.5-max`), which dilutes the self-judging influence, but the paper does not discuss
it as a potential bias source. Second, every agent tested belongs to one vendor — Sonnet-4.6,
Opus-4.7, Opus-4.8 — with no cross-family test on GPT or Gemini as the agent under improvement (those
families only ever appear as *judges*). And there is no empirical comparison against any of the
roughly 15 directly competing methods the paper discusses in its own related work — Meta-Harness,
Self-Harness, TTHE, ADAS, GPTSwarm, AFlow, GEPA, DSPy, and others. RHI is only benchmarked against
same-family reasoning-effort scaling and the provider's built-in multi-agent mode, not against the
alternatives it explicitly positions itself against.
The benchmark itself is also self-constructed: 30 tasks synthesized from real job postings by an
LLM, evaluated by the same lab that designed the method being tested on them. None of this means the
result is wrong — the plain admission that Opus-4.7's token-length claim is inconclusive, and the
decision to run and report the ablation showing the vendor's own multi-agent mode underperforms
single-agent, are both the kind of finding a paper trying to look better than it is would have left
out.
## The take
The part of RHI I'd actually reuse is the trajectory-local relaxation itself: comparing only against
your immediately-previous version turns an intractable population search into something you can run
continuously, cheaply, and the Bradley-Terry argument for why that's still principled — not just
convenient — is genuinely nice. The information-theoretic account of *why* it lands on contracts and
hops is a good hypothesis, stated with the right hedges. What I'd want before trusting the magnitude
of any specific number: a comparison against even one of the population-based methods it explicitly
argues against, and an evaluator that isn't sometimes the model being graded. This reads as the first
half of a real idea — the paper says so itself, calling the harness-to-model feedback loop "the
second half" left to future work — and the half that's here is worth having, with its gaps named
rather than papered over.
---
*Source: [Recursive Harness Self-Improvement](https://arxiv.org/abs/2607.15524) (Hyunin Lee, Jinglue
Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang — Sakana AI, UC Berkeley), arXiv:2607.15524.
Figures 1, 2, and 3 are reproduced from the paper for commentary; the interactives are mine, built on
the paper's own reported formulas and measured endpoints.*
---
# Sol-Attn: deciding which attention blocks to skip while you're already streaming them
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/sol-attn
> date: 2026-08-03
> tags: diffusion, attention, sparse-attention, inference, video, explainer
Video generation has an attention problem that language models mostly don't. A few seconds of video at a useful
resolution is a very long token sequence, attention is quadratic in it, and diffusion runs the whole stack dozens
of times per clip. So attention stops being *a* cost and becomes *the* cost.
[Sol-Attn](https://arxiv.org/abs/2607.24027) — "Sparsifying online attention", from the SANA team at NVIDIA and
collaborators — is a training-free way to skip most of that work. The idea I like is not that it sparsifies
attention; everyone does that. It is *where* the decision happens.
## The problem with picking blocks
Training-free sparse attention works block-wise: score each block of keys with something cheap (a proxy), keep
the promising ones, skip the rest. The question is how you pick, and the paper's Figure 1 shows why the two
standard answers both misbehave — on two different attention-logit distributions, one peaked and one nearly flat.
Read the sparsity numbers across the rows:
- **Top-k** keeps a fixed fraction, so it reports **70.3% on both rows**. It cannot tell a peaked distribution
from a flat one. On the peaked row it is leaving free sparsity on the table; on the flat row it is throwing
away blocks that mattered.
- **Top-p** keeps blocks until their cumulative proxy mass hits a target, which is adaptive but wildly so:
**96.88%** sparsity on the peaked row and **21.9%** on the flat one. Budgets swing by a factor of four
between distributions, which is miserable for a kernel that wants predictable work per tile.
- **Sol-Attn** thresholds at **μ + βσ** — the mean of the block scores plus β standard deviations. It adapts
(75.0% vs 67.2%) but stays in a controlled band.
That statistical threshold is the whole trick, and its virtue is *computability*. A mean and a variance can be
maintained as a running summary while you stream blocks. A top-k ranking cannot: you have to see every score,
materialize them, and sort. Which brings us to where the decision gets made.
## Folding the decision into online softmax
Flash-attention-style kernels already stream. They walk key/value tiles, keep a running maximum and a running
sum, and rescale as they go — that is what makes softmax computable without ever holding the full attention
matrix. Conventional sparse attention bolts a *separate* pass in front of this: score all blocks, build a proxy
map in memory, rank it, then run the sparse kernel on the survivors.
Sol-Attn puts the decision inside the loop that was already running.
The outer loop computes proxy scores from mean-pooled keys and compares them against that tile's threshold,
producing a `1/1/0`-style mask. The inner loop then runs exact attention on the surviving tiles and skips the
rest. Because the threshold is a statistic rather than a rank, no proxy map is ever materialized — the budget
comes out dynamic *and* controllable, which is the combination neither top-k nor top-p manages.
## Not dropping, approximating
The second idea is smaller and does more work than it looks. Standard block-sparse attention treats an
unselected block as if it contributed nothing. Under aggressive sparsity that assumption is exactly where the
quality goes.
Sol-Attn has already computed a proxy score for every block, including the losers, since that is how it decided.
So instead of discarding them it **reuses those scores to approximate the skipped blocks' contribution** — a
correction term that costs nothing extra, because the information was a by-product of routing. Routing, sparse
computation and approximation correction all happen in a single online-softmax pass.
That is why the accuracy curve degrades gracefully rather than falling off a cliff: the tail is attenuated, not
deleted.
## What it actually buys
The paper reports **2.1× end-to-end for video generation** and **2.3× for video editing**. The more useful chart
is the cumulative breakdown, because it shows what is attributable to what:
Those headline multiples are **cumulative**, so it is worth doing the subtraction. On HunyuanVideo, Sol-Attn takes
328.4 s down to 170.6 s — a **1.92× marginal** gain on top of the other two techniques. On Wan2.1-14B it takes
217.6 s to 161.8 s, a **1.34× marginal** gain. Real, and the largest single contributor in the Hunyuan case, but
not 5.08×. Anyone quoting the total as an attention result is quoting three techniques.
## The engine around it
Sol-Attn does not ship alone. It is one of five composable techniques in the
[`sol-engine` branch](https://github.com/NVlabs/Sana/tree/sol-engine) of NVlabs/Sana (Apache-2.0), described as
"an efficiency-oriented inference codebase for high-resolution video diffusion, built on SGLang's
`multimodal_gen` runtime".
The five: **caching** (reuse or skip denoising-step outputs, TeaCache/EasyCache-style), **quantization**
(TransformerEngine NVFP4 4-bit, applied step-selectively), **kernel fusion** (memory-bound DiT ops — norm,
activation, precision conversion), **sparse attention** (Sol-Attn), and **token pruning** (dropping low-salience
video tokens during refinement steps). Reported end-to-end speedups, all on GB200 with warmup excluded:
| Model | Speedup |
|---|---|
| Wan2.2 TI2V-5B | ~2.89× |
| SANA-Video (2B) | ~2.77× |
| LingBot-Video (30B) | ~2.60× |
| LTX-2.3 (22B) | ~2.38× |
| Cosmos3-Super (64B) | ~2.27× |
| Wan2.2-A14B (14B MoE) | ~2.17× |
The consistency across 2B to 64B, dense and MoE, is the interesting part — these are mostly memory-movement and
redundancy wins, so they do not evaporate as models grow.
There is also an **agent-native workflow**: the repo is set up so a coding agent (Codex or Claude Code) does
environment setup, weight fetching and inference, troubleshooting as it goes. Worth noting on a site whose own
content is written this way — treating "an agent will be the one running this" as a first-class install path is
still rare.
## Where this sits
The same lab has been attacking this cost from the other end. [SANA-Video 2.0](/articles/sana-video2) makes
attention cheap *architecturally* — linear attention for three of every four layers, a deep-compression VAE to
shrink the token count before attention ever runs. That requires training the model that way. Sol-Attn is the
training-free counterpart: take a model somebody already trained and skip work at inference. SANA-Video is
literally the second row of the engine's own benchmark table, so both ends compose.
Against the site's other sparse-attention coverage — [MiniMax's approach](/articles/minimax-sparse-attention) —
the contrast is that most sparse attention is *trained*, with the model learning to live within a sparsity
pattern. Sol-Attn assumes no cooperation from the model at all. And where
[MrFlow](/articles/mrflow-diffusion-acceleration) attacks diffusion cost along the *step* axis, Sol-Attn attacks
the per-step cost; the engine's caching module is doing the step-axis job alongside it.
**Caveats.** (1) Every number here is **self-reported**, on a paper posted 2026-07-27 with no third-party
replication. (2) The engine's speedups are **GB200, warmup-excluded** — best-case hardware, and warmup is a real
cost you pay once. (3) The headline 5.08× is **cumulative across three techniques**; Sol-Attn's marginal
contribution is 1.92× and 1.34× on the two models shown. (4) The paper claims quality is preserved but I have
not seen an independent quality evaluation, and "preserved" is doing real work in a domain where the failure
mode is subtle temporal artifacts rather than a metric drop. (5) `sol-engine` is a **branch**, not a release.
## The take
The mechanism is the part worth keeping. Routing decisions in sparse attention are usually treated as a
preprocessing step — score, rank, select, then compute. Sol-Attn's claim is that the ranking was never necessary:
a statistical threshold gets you an adaptive budget from a running summary, which means the decision can live
inside the streaming loop the kernel already runs, which means the proxy map never has to exist. And once you
are computing proxy scores anyway, throwing them away for the skipped blocks is wasteful — reusing them as an
approximation is close to free.
Both ideas come from asking where the information already is, rather than adding machinery. That tends to be the
sign of a good systems result.
---
*Sources: [Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification](https://arxiv.org/abs/2607.24027)
(Haopeng Li, Yitong Li, Junsong Chen, Tian Ye, Haozhe Liu, Jincheng Yu, Duomin Wang, Ruihua Zhang, Zeke Xie,
Enze Xie, Song Han; arXiv 2607.24027, 2026-07-27) for the method and Figures 1–3, and the
[`sol-engine` branch](https://github.com/NVlabs/Sana/tree/sol-engine) of NVlabs/Sana for the engine, the five
techniques and the per-model speedup table. All figures are the paper's own, served locally. Marginal-speedup
arithmetic is mine, derived from the latencies in Figure 3. The interactives are mine and illustrate the
mechanism; they are not measurements.*
---
# ADR: the agent that watches your agents
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/uber-adr
> date: 2026-08-03
> tags: security, agents, mcp, llm, explainer
The first thing to get out of the way: [Uber's ADR](https://github.com/uber/ADR) has nothing to do with Architecture Decision Records. **ADR = Agentic Detection and Response** — an enterprise security framework for the AI coding agents your engineers already run. The tagline says it plainly: "ADR secures enterprise AI agents through observability, security benchmarking, and threat detection." It's infra, not a model — though it ships an LLM-based detector and a red-team benchmark, and it's been running inside Uber for **over ten months**, watching **7,200+ unique hosts** and **10,000+ agent sessions a day**.
The reason it exists is a gap that's easy to miss if you haven't operated one of these agents at scale: your existing security tooling watches file writes and process spawns. It has no idea *why* a file got written. An AI coding agent with shell access, file access, and a dozen MCP servers plugged in is a new kind of actor on your endpoints, and the thing that makes its actions dangerous or benign — the reasoning that led to the tool call — is exactly the part a traditional EDR agent can't see.
## The threat model
Three problems, named directly in [the paper](https://arxiv.org/abs/2605.17380) (Chenning Li, Pan Hu, Justin Xu, et al., accepted MLSys 2026 Industry Track):
1. **Limited observability** — "existing Endpoint Detection and Response (EDR) tools see file writes but not the agent reasoning, prompts, or causal chains linking intent to execution."
2. **Insufficient robustness** — static, rule-based defenses don't generalize across attack techniques.
3. **High detection cost** — running an LLM as a judge on every one of 10,000+ daily sessions is expensive if you do it naively.
The concrete shape of the threat is **indirect prompt injection through a tool a developer already trusts.** Uber emulated the industry's Agent Flayer incident (Zenity Labs, 2025) to validate ADR against it, and the attack graph is worth internalizing because it's the whole genre in five steps: a malicious email creates a Jira ticket with hidden instructions; a developer's Cursor session, connected to Jira through an MCP server, fetches that ticket as ordinary context; the hidden instructions coerce the agent into reading a local secrets file; and the agent — still just doing what its context told it — makes an outbound HTTP call that exfiltrates the credentials. Nobody typed a malicious command. The attacker never touched the developer's machine. The whole thing rides in on a support ticket.
Uber also ran an internal capture-the-flag exercise that names the second broad category: a custom shell tool integrated into the assistant, then a deceptive prompt instructing the agent to run `curl | python3` "under the guise of sandbox testing." Same underlying issue — the agent's own tool surface, or a malicious MCP server sitting on it, is the attack surface. ADR-Bench, the benchmark this whole thing gets scored against, formalizes that: it's built on **133 MCP servers** (78 benign, 25 deliberately backdoored, the rest environment-emulation), and one of the backdoored ones — a credential-harvesting tool called `location_harvester` — is registered under a spoofed identity, `weather_forecast_service`, specifically to test whether a detector falls for supply-chain impersonation. That's the taxonomy: prompt injection through trusted tool output, malicious or impersonating MCP servers, and an agent that ends up exfiltrating something it was never asked to touch.
## Why you can't just run an LLM judge on everything
The naive fix — point a capable model at every agent session and ask "is this suspicious?" — doesn't survive contact with the volume. At 10,000+ sessions a day, an LLM call per session that does real reasoning (reads code, checks policy, thinks about intent) is expensive and slow enough that latency alone rules it out for most of the traffic, almost all of which is a developer editing files and running tests.
ADR's answer is a **two-tier pipeline**, and the design is unapologetically borrowed from how a human security team already works: a triage analyst does a fast first pass on everything, and only escalates the ambiguous cases to someone senior enough to actually investigate.
**Tier 1** is a cheap, single-call triage model (`gpt-4o` by default) reading the full session transcript, tuned for high recall with an explicit escalate-on-ambiguity bias — anything that isn't clearly benign defaults to suspicious rather than being waved through. If triage says benign, the pipeline returns immediately. That early return, not a second model being smarter, is the entire cost-control story.
**Tier 2** only runs on the sessions Tier 1 flagged, and it is not an API call. Reading the released code (`Detection/guardrail/adr_agent/adr_baseline.py`) turns up something more interesting than a bigger judge model: it shells out to the **Claude Code CLI as a subprocess** — literally `subprocess.run(["claude", "--model", "claude-sonnet-4-6", "--mcp-config", ".mcp.json", ...])` — and gives that CLI session live MCP tool access to three purpose-built servers: one that fetches a flagged tool's actual source code, one that looks up known techniques for the suspected tactic, and one that checks policy violations. The reasoning agent doesn't just re-read the transcript harder. It *investigates* — it can go pull the suspicious MCP tool's own source before it commits to a verdict, the same way a human analyst would open the code instead of trusting the tool's name. Output is forced JSON (`is_threat`, `confidence`, `explanation`), with a documented retry path for the case where Claude itself flags the detector's own prompt as a possible injection attempt.
Step through both paths on the two real scenarios above — an ordinary edit-test-commit session, and the Agent Flayer emulation — and watch what each tier actually sees, calls, and costs:
The numbers back the design up directly. In Uber's own ablation, the triage layer alone resolves **40.7% of tasks** at **$0.017/task, 2.3s** latency; the escalated path costs **$0.029/task, 29.7s** — about 13x more expensive per task, which is the honest price of running an LLM twice, with tool calls, inside a detection loop. Removing the triage layer entirely and sending everything to the reasoning agent *does* raise recall (0.667 → 0.805) and F1 (0.800 → 0.868), but it also introduces false positives (precision drops to 0.943) and costs 29% more — a clean demonstration that the two-tier split isn't just an engineering shortcut, it's load-bearing for precision too. And in a second ablation that validates the "fetch the tool's own source" mechanism specifically: pull the source-code MCP server away from the reasoning agent and F1 falls from 0.800 to 0.676 — the single largest drop of any component removed, bigger than losing threat-intel lookup (0.800 → 0.716) or policy verification (0.800 → 0.806, barely a dent). The part of the mechanism that sounds like the interesting flex — an agent reading a tool's own code before judging it — is also, empirically, the part doing the most work.
## The benchmark, and where the baselines fall apart
**ADR-Bench** is 302 tasks (42 malicious, 260 benign) derived from real enterprise telemetry, spanning 17 named attack techniques across 5 tactics and those 133 MCP servers. The class imbalance — 13.9% attack rate — is the point: an academic benchmark that's roughly half attacks doesn't tell you how a detector behaves against the traffic mix a real enterprise actually has, which is overwhelmingly benign. ADR is also evaluated on **AgentDojo**, a public, mostly-balanced prompt-injection benchmark (93 tasks), giving two very different stress tests of the same detector.
The pattern that falls out is the honest headline: on AgentDojo, the baselines (ALRPHFS, GuardAgent, LlamaFirewall) look reasonable — ALRPHFS even gets within striking distance of ADR's precision. Move to ADR-Bench's enterprise-realistic imbalance and every baseline collapses to 30–40 false positives out of 260 benign tasks. That's not a rounding difference — at real traffic volume, that false-positive rate is what "unsuitable for production" cashes out to: a human review queue that drowns. ADR holds **1.000 precision, 0.667 recall, 0.800 F1, zero false positives** on ADR-Bench, and **0.927 precision, 1.000 recall, 0.962 F1** on AgentDojo — complementary strengths the paper is upfront about: precision-first where class imbalance is severe, recall-first where missing a prompt-injection attack is the costlier failure.
## What happened in production
The credential-exposure numbers are where this stops being a benchmark story. Deployed across Uber's MacBook fleet (Intel and ARM), ADR has "sustained reliable detection... uncovering hundreds of credential exposures across 26 categories," which fed a shift-left prevention layer — a regex-and-entropy pre-prompt hook running in Cursor and Claude Code — that hit **97.2% precision**: 206 true positives correctly caught, only 6 false positives, across 212 unique credentials seen in hundreds of thousands of sessions. Worth noting in passing: that prevention layer is explicitly *not* LLM-based. "Simple non-LLM checks work well for known attacks with static patterns... but fail for attacks requiring reasoning about tool semantics, causal context, and enterprise policies" — which is exactly the job Tier 2 exists to do instead.
One production detail is genuinely candid: of the sessions ADR's Tier 1/Tier 2 routed to a human review queue, only 34% were confirmed malicious true positives. Another 17% were true positives that turned out non-malicious — internal red-team and pentest activity that looked identical to an attack until a human checked. The paper doesn't paper over this: it explicitly flags the need to model *user personas* and intent, because a security team doing its job on purpose looks, from telemetry alone, a lot like an attacker.
## What's honest and what's missing
A few things worth stating plainly before treating any of this as settled:
- **Every production number is self-reported.** The 7,200 hosts, 10,000+ sessions/day, and 97.2% precision figures come from Uber's own paper — there's no third-party audit. A tweet from an Uber engineer separately claims 50,000+ sessions/day, a larger and more recent figure than the paper's own — treat that as informal color, not a citable number.
- **The baseline comparisons in Table 2 aren't independently re-runnable.** ALRPHFS and GuardAgent's code was stripped from this repo for licensing reasons; the paper's own comparison numbers are reproduced as-is, documented candidly in `docs/BASELINE_REPLICATION.md`, but you can't regenerate them yourself.
- **The open-source release is the detection half of a four-part system.** Per the architecture, ADR Explorer (pre-deployment red-teaming) and ADR Prevention (blocking unsafe actions in real time) are both explicitly excluded — "not included in the current open-source release. Stay tuned." What's shipped is the Sensor, the benchmark, and the dual-agent detector baseline — real, but not the complete production stack Uber runs internally.
- **The reasoning tier is unusually coupled to one vendor's CLI.** Shelling out to `claude --dangerously-skip-permissions` as a subprocess is a legitimate way to get a tool-using agent for free, and it's the most narratively interesting part of the design — but it's also a portability limitation worth naming as exactly that.
- **An LLM in the detection path is not free**, and ADR doesn't pretend otherwise: $0.024/task blended average cost and 18.5s average latency on ADR-Bench, with the escalated path alone running $0.029 and nearly 30 seconds. That's the real, ongoing bill for the precision this design buys.
- **ADR-Bench is Uber's own benchmark**, built from Uber's own telemetry and scored by Uber. It's a genuinely useful stress test — the class-imbalance framing is a real and underused idea — but it isn't a neutral third party's yardstick, and the paper is explicit that the benchmark's attack rate doesn't mirror real production incidence.
None of that erases the result: a two-tier detector that costs cents and seconds on the common case, escalates to an agent that can go read a suspicious tool's own source before it decides, and has been running against real attacks in production for the better part of a year.
## The take
ADR is worth reading past its confusing name because it's a concrete answer to a question [Lilian Weng's harness framing](/articles/agent-harness) leaves open: if the harness — the loop, the tools, the context policy — is where an agent's real capability lives, then the harness is also exactly where you'd go looking for an attack, and exactly where a defense has to watch. ADR watches the causal chain a traditional EDR tool can't see, and its own reasoning tier is built the same way the agents it's watching are: a model with tool access, investigating rather than just classifying. The cost/precision tradeoff it makes explicit — cheap triage on the routine 40.7%, an expensive investigating agent on what's left — is the same lesson [Antares](/articles/antares) makes from the opposite direction, with a much smaller model doing a narrower security job: the interesting engineering in agentic security right now is less about a bigger judge model and more about routing the right amount of reasoning at the right moment.
---
*Sources: [uber/ADR](https://github.com/uber/ADR) (Apache 2.0; the vendored AgentDojo benchmark under `Detection/benchmark/agentdojo/` is MIT); the paper, ["ADR: An Agentic Detection System for Enterprise Agentic AI Security"](https://arxiv.org/abs/2605.17380) (Chenning Li, Pan Hu, Justin Xu, Baris Ozbas, Olivia Liu, Caroline Van, Manxue Li, Wei Zhou, Mohammad Alizadeh, Pengyu Zhang, KK Sriramadhesikan, Ming Zhang; MLSys 2026 Industry Track). Figures reproduced from the paper for commentary, cropped from the arXiv PDF committed at `docs/adr-paper.pdf` in the repo and flattened onto white. The interactive pipeline trace and benchmark scatter are my own, built from the paper's own reported numbers — not measured traces.*
---
# Macaron-V1: four 1B adapters on a frozen 744B base
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/macaron-v1
> date: 2026-07-27
> tags: llm, lora, agents, open-weights, generative-ui, explainer
[Macaron-V1](https://huggingface.co/collections/mindlab-research/macaron-v1) is Mind Lab Research's agent
model family, released 2026-07-21 under MIT. The flagship, **Macaron-V1-Venti**, is described as a 748B model.
That number needs an asterisk immediately: **744B of it is a frozen GLM-5.2**, and Mind Lab's contribution is
**four 1B LoRA adapters** — about half a percent of the artifact.
That is the whole idea, and it is a genuinely different bet from how most post-training is done. Rather than
fine-tuning one monolithic model, take a strong open base, freeze it, and attach a small number of tiny
specialists. They call it **Mixture of LoRA (MoL)**.
## Mixture of LoRA
Four adapters, each 1B: `l0` Chat, `l1` Agent, `l2` Coding, `l3` GenUI. The routing detail is the neat part —
**`l0` is both the conversational backbone and the router**. It sees each new user request and dispatches it to
whichever specialist fits.
Note what routes and when. In a Mixture-of-Experts model, a router fires on *every token* at *every MoE layer*,
and the experts were built during pretraining — they are inseparable from the model. MoL routes **once per
request**, at the adapter level, and the thing being routed between is four swappable files sitting on a base
someone else trained. Ongoing reasoning and tool interaction stay inside the selected LoRA for the duration;
when a specialist finishes, its work is passed to the next as a concise summary rather than shared state.
The practical consequences are real. Specialization costs 1B parameters instead of a full fine-tune. Adapters
can be swapped, added, or updated independently. And when GLM-5.2 improves, you re-fit adapters rather than
retrain a 744B model. The cost is equally real: you inherit the base's ceiling, its licence obligations, and
its failure modes, and a request-level router cannot change its mind halfway through a turn the way per-token
routing implicitly can.
**Update, 2026-08-03: the Tall size conflict looks like two different counts, not a contradiction.** Mind Lab's
blog calls Macaron-V1-Tall **50B**; the model card, the Hugging Face listing and Novita all say **36B**. The
Hugging Face API settles half of it — `safetensors.total` for `Macaron-V1-Tall` is exactly **35,951,822,704
(35.95B)**, which is the 36B figure and is the *base checkpoint on its own*. The blog's 50B is its own
decomposition of base plus adapters: 35B + 4 × 3.7B ≈ 49.8B. So the two numbers are measuring different things,
and Novita hedges the gap as "10~50B".
What I could not verify is the per-adapter figure. The published 35.95B total does not appear to include the
adapter weights, so I cannot confirm 3.7B each from the metadata — that number is Mind Lab's, not something I
measured. It is worth flagging because it implies Tall's adapters are roughly **3.7× the size of Venti's 1B
ones** on a base twenty times smaller: on Venti the specialization is about half a percent of the artifact, on
Tall closer to a tenth. If that holds, "Mixture of LoRA" means something quite different at the two scales, but
the evidence for it is currently a single line in a blog post.
## What the adapters actually buy
Most of the release table compares Macaron to Claude, GPT and Gemini. That is the least informative comparison
available, because it confounds the adapters with GLM-5.2's own strength. The controlled experiment is sitting
right there in the same table: **each variant against the frozen base it was built on**. That isolates the only
thing Mind Lab actually changed.
The pattern is consistent and modest: roughly **+3 to +6 points** across chat, agent and coding work. Two rows
are inside noise — `T3-Bench` at +0.2 and SWE Atlas QnA at +0.6. And then UI4ABench jumps **+20.7** on Venti and
**+25.4** on Tall.
That outlier is the honest crux of the release. It is simultaneously the strongest evidence that a 1B adapter
can teach a frozen base a genuinely new skill, *and* the result most exposed to selection effects — UI4ABench is
Mind Lab's own benchmark, measuring generative UI, which is exactly the capability they built a dedicated
adapter for. Both things are true at once.
## Read the evaluation table twice
The published benchmark figure includes its own methodology notes, and they change how several rows should be
read.
Three things stand out:
- **The judges are other models, and one of them is the base.** ChatBench is scored by "a privately deployed
GLM-5.2 judge" — and Venti *is* GLM-5.2 plus adapters. A model's own base evaluating its output is a conflict
worth naming. Elsewhere the judge is a competitor: Claude Opus 4.6 on LivingBench, Claude Haiku 4.5 on
PinchBench, GPT-5.4 on ClawGym, GLM-5.1 on VitaBench, Gemini 3.5 Flash scoring UI4ABench rubrics.
- **Several rows are best-of-N, not single-shot.** PinchBench reports "the best observed score". DeepSWE allows
"up to three attempts, and report the best one". SWE Atlas QnA is pass@3. SWE-Bench Verified permits up to
three retries on evaluation errors and reports the best successful attempt. Those are legitimate protocols,
but they are not comparable to a single-trial number from another lab's report — and some competitor cells are
marked as taken from leaderboards or the models' own reports.
- **Macaron does not lead everywhere.** Claude Opus 4.8 wins SWE Verified (88.6 vs 85.6) and SWE Atlas QnA (57.3
vs 49.5). GPT-5.5 wins ClawGym (82.5 vs 77.7) and DeepSWE (70.0 vs 58.4). Qwen 3.7 Max wins VitaBench (61.2 vs
60.0) and Gemini 3.1 Pro wins VitaBench2 (50.2 vs 46.0). Mind Lab says as much in its own post: "Coding is
where we currently sit close to, rather than ahead of, the frontier."
The rows where Gemini, Qwen and Minimax collapse to 10.0–22.6 on DeepSWE and SWE Atlas QnA are almost certainly
harness incompatibility rather than capability — those evaluations run through Claude Code as the agent harness,
which is not neutral ground for every model.
## The infrastructure claims
Three systems are named, none with a technical report behind them yet:
- **MinT** — the post-training platform, claimed to support models up to a trillion parameters via adapter-only
handoffs and a "million-scale adapter catalog". Adapter-only handoff is the load-bearing idea: if
specialization is always a small file, you never move a 744B checkpoint between training stages.
- **MindForge** — an agentic RL framework built around discovery, expansion and update cycles against
production harnesses.
- **LongStraw** — million-token RL, which works by evaluating a shared prompt once into a reusable resident
state and then replaying only the response branches. For agentic RL where many rollouts share a long prefix,
that is the obvious win, and it rhymes with the external KV-cache pooling in
[Kimi K3](/articles/kimi-k3)'s RL infrastructure.
Macaron-V1 also ships a serving story: a
[Mixture-of-LoRA harness](https://github.com/MindLab-Research/Mixture-of-LoRA-Harness) that keeps an
OpenAI-compatible endpoint while adding the L0 router and same-request switching into the selected specialist,
plus [Macaron Artifacts](https://github.com/MindLab-Research/macaron-artifacts), a local WebUI and plugin that
runs inside Claude Code, Codex or Kimi Code.
## The take
The interesting claim in Macaron-V1 is architectural, not competitive: that request-level routing across a few
1B adapters on a frozen base is enough to build a specialized agent model, and that you can therefore treat a
frontier open-weight model as infrastructure rather than as something to fork. The base-versus-tuned comparison
supports a weaker version of that claim than the headline table does — a few points nearly everywhere, and one
large gain on the capability they purpose-built an adapter for.
What would settle it is the technical report, which the model card lists as "coming soon", along with the full
benchmark methodology. Until then this is a self-reported release with no third-party replication, several
best-of-N protocols, and its own base model sitting on the judging panel. The idea is worth watching. The
numbers are worth waiting on.
---
*Sources: the [Macaron-V1 collection](https://huggingface.co/collections/mindlab-research/macaron-v1) and the
[Macaron-V1-Venti](https://huggingface.co/mindlab-research/Macaron-V1-Venti) model card (architecture, adapter
roles, parameter counts, benchmark table), and Mind Lab's
[Introducing Macaron-V1](https://macaron.im/mindlab/research/introducing-macaron-v1) post (MinT, MindForge,
LongStraw, variant sizes). All benchmark numbers are Mind Lab's own, with per-benchmark judge models and
retry policies as annotated in their published table; no technical report has been released. The interactives
are mine.*
---
# Neutrino-1: quantization is a training decision, not a deployment one
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/neutrino-1
> date: 2026-07-27
> tags: llm, quantization, ternary, bitnet, inference, efficiency
Fermion Research shipped three models on July 27, 2026: **Neutrino-1 8B**, **Neutrino-1 0.6B**, and a 0.6B-Chat
variant, all on Hugging Face under Apache 2.0, all built on a ternary weight format — every projection weight is one
of exactly three states, minus, zero, or plus. That part of the pitch is not new; I wrote about
[Ternary15M](/articles/ternary15m) doing the same thing at 15M parameters a few days ago. What Neutrino-1 adds is
scale (about 500x more parameters) and, more importantly, a controlled comparison that the smaller model never ran:
what happens if you take the *same* ternary format and round a model into it after training, instead of training
inside it from the start.
The answer is not "a few points worse." It's a cliff.
## The cliff
Two 8B-class checkpoints, each about 3 GB, rounded into a ternary format after full-precision training, score **24.2**
and **24.7** on 5-shot MMLU. Chance, on a four-choice test, is **25.0**. Three models whose training ran *inside* the
ternary constraint from the start — at the same roughly 2-3 GB artifact size — score 47.24, 65.75, and **72.1**
(Neutrino-1 8B itself). Fermion's own line on this, and it's a good one: "rounding after training lands on the chance
line; every model trained inside its own format clears it by twenty-two points or more."
The mechanism is worth sitting with, because it explains why this isn't a smooth tradeoff curve. Round a trained
weight to the nearest of three values and the individual errors don't cancel — Fermion describes them as
uncorrelated, so they "accumulate along the row as a random walk instead of cancelling. The feature that leaves the
layer is not a noisier version of the right answer; it is a different number, and 36 layers compound the difference."
Their framing: "the failure is not noise. It is amnesia, and its signature is the cliff: scores do not degrade toward
chance, they arrive there." Train inside the constraint instead, and the optimizer never learns a solution the format
can't store in the first place — there's no gap between the model that was trained and the model that ships.
This is the flip side of what Ternary15M already showed at tiny scale: quantization-aware training with a
straight-through estimator gets hard-ternary inference to within +0.01 nats of the latent model, because the network
never experienced anything else. Neutrino-1 is the same bet, replayed at 8B, with the missing control group finally
run: skip the QAT and round instead, and the model doesn't degrade gracefully — it falls through the floor.
## What's actually ternary
Not everything. Of the 8B's 8.19 billion parameters, 6.95 billion (all 252 projection matrices — seven per layer,
attention and feed-forward, across 36 decoder layers) are ternary. The token embedding table and the output head stay
int8; the RMSNorm gains stay fp32. The reasoning Fermion gives is about where rounding error can hide: "a linear
layer's output feature sums hundreds of three-state contributions, so individual state errors cancel inside the sum;
that summation is what makes the format survivable." An embedding lookup returns one row verbatim — there's no sum to
average the error away — and the output head "decides tokens by small logit margins," where a rounding error can flip
the argmax. Both stay out of the ternary lane for the same reason: no averaging effect to hide behind.
The scale mechanism is also more granular than Ternary15M's. Ternary15M ternarizes each output channel around one
number: `absmean(W)`, the mean absolute weight for that whole row. Neutrino-1 groups weights into fixed-size blocks
along the input dimension and gives *each block* its own higher-precision scale — Fermion's phrase is "state times
scale." Smaller blocks track the underlying weights more closely at the cost of more stored scales; one scale for an
entire row (Ternary15M's approach) is the coarsest, cheapest case. The toy below runs the actual arithmetic — round
`clamp(w / scale, -1, 1)` per weight, scale is each block's mean absolute value — so you can see the tradeoff move:
## Where the bytes go
This is the honest caveat behind "4.2 times fewer bytes than fp16 at bf16 on the same memory system," and it's worth
being precise about it: that ratio is a **whole-artifact** number, not the per-weight ternary compression ratio. Log2
of 3 states is about 1.58 bits, which against a 16-bit float is closer to a 10x reduction — but the embedding and
output tables (int8, one byte per value, no ternary discount) and the norm gains (fp32) drag the average down. At 8B,
those non-ternary tensors are only 15.2% of the parameters but 32.8% of the bytes. At 0.6B the effect is worse: the
same un-ternarized vocabulary is 47.5% of the file, because a smaller transformer has fewer projection weights to
amortize a fixed-size vocabulary table against. It's the identical finding Ternary15M made at 15M parameters, where
the FP32 embedding table was over 60% of that model's total parameters and dominated its 43 MB footprint — the
direction of the effect is the same at both ends of a 500x scale range: the *smaller* the model, the more its
un-quantized vocabulary — not its ternary matmuls — decides the file size.
On disk, Neutrino-1 8B is 3.88 GB; it downloads at 2.56 GB because the ternary lane compresses further in transit
(Fermion reports 0.516-0.569 of raw bytes, layer-dependent). Neutrino-1 0.6B downloads at 328 MB.
## Sparsity is learned, not imposed
Across the 8B's 6.95 billion ternary weights, the split is **62.63% zero, 18.68% plus, 18.69% minus** — remarkably
close to balanced between the two nonzero states, and remarkably far from an even three-way split. Fermion's framing
is the one worth keeping: "most of the mass on zero: the format sets how much of each tensor falls silent, and the
learned weights decide which connections go." A float layer can only make a connection small; a ternary layer,
trained natively, can delete it outright and the training decides which ones.
That spike is the interesting part, because nothing about the format explains it — the format sets *how much* falls
silent on average, not *where* it clusters by depth. The four attention projections sit in a tight 61.84-63.51% band
at every one of the 36 layers, almost boring in its consistency. The feed-forward `down` and `gate` projections are
the exception: `down` reaches 72.47% zero at layer 3, `gate` reaches 70.48% at layer 4 — roughly ten points denser
than the rest of the network — and both settle back to the ~62% baseline by layer 5. The single densest tensor in the
whole model is the layer 1 `down` projection at 60.59% (its local minimum, immediately before the spike). Scrub
through the real per-layer numbers below:
The state statistics are also stable across scale in a way that argues they're a property of the format and the
training recipe, not of size: at 0.6B, fourteen times fewer parameters, the split is 62.26% zero / 18.87% plus / 18.86%
minus — within half a point of the 8B on every axis.
## How much of Qwen3-8B does it keep
Neutrino-1 8B is measured, on Fermion's own harnesses, against **Qwen3-8B at bf16** — described in the post as "the
full-precision base it was built from," which is itself worth flagging: this isn't an independently trained
architecture being compared to an unrelated baseline, it's a model built from Qwen3-8B's own weights and then
retrained natively in ternary. At 4.2x fewer bytes, Neutrino-1 8B holds:
Report the weakest number, not the flattering one: tool calling retention is 79%, and it's worse than that headline
suggests once you look at the breakdown by category (BFCL v3, macro-averaged to 68.9 overall). Held-out, textbook
function signatures score well — 82.3% simple, 83.5% multiple — but signatures drawn from real-world APIs in the wild
score much lower: 61.6% live-simple, 52.0% live-multiple. Non-Python languages are worse still: 54.0% JavaScript,
43.0% Java. "Tool calling: 79%" is an average that buries a 40-point spread between the easy and hard slices of that
same axis.
The MMLU headline (72.1) and the "96% general knowledge" retention figure are **not directly comparable** in
Fermion's own post. The retention percentages are computed against Qwen3-8B's score on an unnamed "general knowledge"
suite — Fermion never states Qwen3-8B's own 5-shot MMLU number anywhere in the piece, and never confirms that
"general knowledge" and "MMLU" are the same benchmark. The MMLU chart above only compares Neutrino-1 8B against
*other* ternary and rounded models at similar artifact sizes, not against its own full-precision progenitor. So
while the rounding-vs-native-training gap (24.2 vs 72.1) is well anchored, the honest answer to "what's the gap
between 72.1 and Qwen3-8B's own MMLU" is: **the source doesn't say, and you can't back it out from what's published.**
Every number in this section is self-reported by Fermion, on their own harness, with no third-party replication.
For what it's worth, the one place Neutrino-1 8B is reported to exceed the reference is answer-format discipline —
Fermion's explanation is that discipline is a trained *behavior*, not a bulk statistical property of the weights the
way knowledge is, so the format doesn't cap it the way it caps knowledge retention. At the small end, Neutrino-1 0.6B
is compared directly to Qwen3-0.6B on ARC-easy: 53.45 vs 60.82, 87.9% retention, at one-eighth the precision and a
238 MB vs 1.50 GB download.
## Serving it: the Neutrino Engine
The inference side ships as its own artifact — a `pip install fermion` package, a CUDA-enabled `llama.cpp` fork, and
an MLX pack — with one stated design constraint: output has to be **token-identical** to a full-precision reference on
every backend. Fermion gates every release on that: the speculative-decoding path (Neutrino-1 0.6B drafting for the
8B) was checked token-by-token against the undrafted path across 27,648 consecutive tokens before any drafted
throughput number was published, and they report zero divergences.
Measured numbers: **33.7 tokens/second** on a MacBook M5 (GPU path; 24.9 tok/s CPU-only), **30.7 tokens/second** on an
NVIDIA L4 at 4k context inside 4.68 GiB of VRAM (fits an 8 GB card), and **396 tokens/second** undrafted on an H100
80GB — rising to **763 tokens/second** with the 0.6B draft model, gated as above. Draft acceptance is prompt-dependent:
near 100% on counting/enumeration, 96.5% on factual recall, roughly 80% on prose, and roughly 50% on code — code is
where the smaller model diverges from the 8B's choices most often, so the speedup shrinks accordingly.
The core argument for why a smaller artifact is faster at batch size 1 is straightforward memory-bandwidth
accounting: single-stream decode reads every weight once per token, so bytes-per-token divided by memory bandwidth
sets a hard floor on latency that no kernel can negotiate around. Neutrino-1 8B's 3.88 GB artifact against roughly
16 GB for the same weights at fp16 puts that floor about four times lower before a single kernel runs.
That fp16 comparison is a **size** argument (3.88 GB vs ~16 GB), not a measured one. Every throughput number Fermion
publishes for the Neutrino Engine is ternary-format-versus-ternary-format — against a "reference stack" running the
same public 27B ternary model (105.15 vs 97.80 tok/s), against `bitnet.cpp` running BitNet b1.58-2B (102.4 vs 89.0
tok/s on an M5), against a lookup-table CPU kernel (T-MAC, 2.12x), against an int4 GEMV kernel (+12-13%). I could not
find a measured fp16 Qwen3-8B throughput number on the same M5, L4, or H100 hardware anywhere in the post. The
tokens/second figures are real and gated for correctness, but the *speedup over full precision* claim is anchored to
an artifact-size ratio, not to a same-hardware fp16 benchmark run.
One more honest number: KV cache growth is indifferent to weight format and scales with context regardless — 0.60 GB
at 4,096 tokens, up to 6.04 GB at the model's full 40,960-token window. Past roughly 26,000 tokens the cache alone
outweighs the 3.88 GB model artifact, so the memory story stops being about weights and starts being about context
length.
## What you can actually download
All three models are live on [huggingface.co/FermionResearch](https://huggingface.co/FermionResearch) as of today,
Apache 2.0, no waitlist:
| Model | Size | Role |
|---|---|---|
| Neutrino-1 8B | 2.56 GB download / 3.88 GB on disk | The frontier model |
| Neutrino-1 0.6B | 328 MB download | Draft model for speculative decoding, and usable standalone |
| Neutrino-1 0.6B-Chat | — | Conversational small model |
They're new enough that download counts were in the single digits at the time I fetched the org page — this is a
same-day release, not an established artifact with a track record.
## The take
Neutrino-1 is Ternary15M's bet — train inside the constraint instead of rounding into it, keep a shared scale next to
the signs, let signed accumulation replace multiplies — replayed at roughly 500x the parameters, with a block-wise
scale instead of one absmean per channel, a real inference engine with a correctness gate, and agentic/tool-use evals
that a 15M TinyStories model has no business running. The MMLU cliff (24.2-24.7 at chance versus 72.1 trained in
format) is the cleanest piece of evidence I've seen that quantization format is something you commit to before
training starts, not a knob you turn afterward — and it's consistent with, not contradicted by, Ternary15M's own
finding that training-aware ternary costs almost nothing when the network never knows another way to compute.
None of this is independently verified. Every number in this piece — the MMLU scores, the retention percentages, the
sparsity statistics, the tokens-per-second figures — is self-reported by Fermion Research on their own harnesses,
comparing their model to their own full-precision progenitor. The MMLU headline and the "96% general knowledge"
figure use two differently-named metrics that are never reconciled in the source. The speed claims are anchored to an
artifact-size ratio, not a measured full-precision baseline on the same silicon. And the model being celebrated for
"holding" Qwen3-8B's knowledge was built starting from Qwen3-8B's own weights, not trained from scratch as an
independent check on the method. The cliff is real and the mechanism is coherent; the specific numbers around it
deserve the same scrutiny you'd give any single-lab benchmark table until someone else reproduces them.
---
*Sources: Fermion Research, "[Intelligence at one-eighth the bits](https://www.fermionresearch.com/research/one-eighth-the-bits/)," "[Introducing the Neutrino-1 models](https://www.fermionresearch.com/research/neutrino-8b/)," and "[The Neutrino Engine](https://www.fermionresearch.com/research/the-neutrino-engine/)" (all July 27, 2026); model weights at
[huggingface.co/FermionResearch](https://huggingface.co/FermionResearch). All figures and quotes are self-reported by
Fermion Research with no third-party replication I could find. The two figures embedded above are reproduced from
Fermion Research's own per-layer and per-tensor measurements published in "One-eighth the bits"; the three interactive
components are mine. Related: [Ternary15M](/articles/ternary15m), the from-scratch 15M-parameter version of the same
bet.*
---
# Kimi K3: a 2.8T open model that turns compute into intelligence 2.5× better
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/kimi-k3
> date: 2026-07-17
> tags: llm, mixture-of-experts, linear-attention, kimi, scaling, explainer
Moonshot's [Kimi K3](https://www.kimi.com/blog/kimi-k3) is the largest open model anyone has shipped:
**2.8 trillion** parameters, **104B active** per token, a **1-million-token** context, natively multimodal. The
[weights are now out](https://huggingface.co/moonshotai/Kimi-K3) under the Kimi K3 License, along with a
[47-page technical report](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf) — so the interesting
claims can finally be checked against a `config.json` instead of a blog post.
**Updated 2026-07-27** after the weights and technical report landed. The most important change: K3 activates
**104B parameters per token**, not the ~50B figure circulating at announcement. That is a 3.2× jump over K2's 32.6B,
and it reframes the efficiency story — K3 is not a cheap-to-run model that punches up, it is a genuinely large model
that converts its compute unusually well.
The interesting part is not the parameter count. It is *how the parameters are spent*. K3 is built on two attention
changes — **Kimi Delta Attention (KDA)** and **Attention Residuals (AttnRes)** — that rework how information flows across
sequence length and across depth, and it scales up MoE sparsity hard: it activates **16 of 896 routed experts** per token
(plus 2 shared), inside a **Stable LatentMoE** framework. Together with refined training and data recipes, those
structural changes yield a measured **2.5× improvement in overall scaling efficiency** over K2. This piece is a
first-principles tour of each piece, why it is new, and what a model like this actually costs to build.
Read the block bottom-to-top: the hidden state passes through the attention sublayer (Gated MLA + KDA), Attention
Residuals reach back to earlier depths, and Stable LatentMoE routes the token to 16 of 896 experts before the block
emits its output. Four ideas, each doing a specific job. Take them one at a time.
## What the weights actually say
Before the mechanisms, the ground truth. Here is K3's shape as recorded in the released `config.json` and the report's
model summary:
| | |
|---|---|
| Total / activated parameters | **2.78T / 104.2B** |
| Layers | **93** (1 dense, 92 MoE) |
| Attention composition | **69 KDA + 24 Gated MLA** — 3:1 per block, plus a final global layer |
| Hidden dimension | 7168 · 96 attention heads · head dim 128 |
| Routed experts | **896**, **16 active** per token, **2 shared** |
| Latent MoE dimension | 3584 (half of hidden) · per-expert hidden 3072 |
| MLA compression | `kv_lora_rank` 512 · `q_lora_rank` 1536 |
| AttnRes block size | **12** layers (8 blocks, 9 counting the embedding) |
| Positional encoding | **none (NoPE)** |
| Activation | SiTU-GLU (`hidden_act: "situ"`) |
| Router | sigmoid scoring, `noaux_tc` — auxiliary-loss-free |
| Vision encoder | MoonViT-V2 · 401M · 27 layers · patch 14 |
| Context / vocabulary | 1,048,576 tokens · 163,840 |
| Quantization | MXFP4 weights, MXFP8 activations (QAT) |
Two of these are worth pausing on. The **3:1 KDA-to-MLA ratio** is not approximate — the config lists exactly which
layers are which, and the full-attention layers land on 4, 8, 12, … 92, then 93. The last layer is *always* global
attention, so whatever the linear layers summarized, the model gets one final unrestricted look at the whole sequence.
And `attn_res_block_size: 12` pins down the AttnRes design: 93 layers partitioned into blocks of 12.
Set against K2, the shape of the bet becomes clear:
## The official architecture diagram
## Kimi Delta Attention: constant-size memory over a million tokens
Ordinary softmax attention keeps a **KV cache** that grows by one entry per token. At a 1M-token context that cache is
the whole ballgame: decoding is [memory-bound on a cache that scales with sequence length](/articles/how-llm-inference-works),
and it only gets heavier as the context fills.
KDA is a **gated delta-rule linear attention**. Instead of a growing cache it keeps a **fixed-size recurrent state**
$S_t$ that each token updates in place: it *erases* a little of the old state (a gated decay) and *writes* the new
key/value association (the delta rule). The report's exact form applies a channel-wise decay before the delta update:
$$
S_t = \left(I - \beta_t k_t k_t^{\top}\right) \mathrm{Diag}(\alpha_t)\, S_{t-1} + \beta_t k_t v_t^{\top},
\qquad \tilde{o}_t = S_t^{\top} q_t
$$
where $\alpha_t \in (0,1)^{d_k}$ is the **channel-wise** one-step retention factor (the erase, per channel rather than
per head) and $\beta_t \in (0,1)$ controls the delta-rule write strength. The state $S_t$ is a fixed $d_k \times d_v$
matrix — its size does not depend on how many tokens came before. Queries and keys are produced by a short convolution
followed by Swish and L2 normalization; values by a short convolution and Swish. Scrub the recurrence and watch the
state stay constant-size while a softmax cache piles up:
That constant-size state is what makes a genuine 1M context tractable, and it is why Moonshot reports **up to 6.3× faster
decoding** in million-token contexts. It is not free — a linear-attention state is a lossy summary, not a perfect record,
so K3 interleaves KDA with full-attention layers (via Gated MLA) to keep exact recall where it matters. KDA also breaks
the assumptions of conventional prefix caching, so Moonshot contributed a KDA implementation to the vLLM community to make
serving practical.
### FlashKDA: the equation, shipped as a kernel
Moonshot also open-sourced the kernel underneath. **[FlashKDA](https://github.com/MoonshotAI/FlashKDA)** (MIT) —
"Flash Kimi Delta Attention" — is a set of CUTLASS kernels for exactly the recurrence above, and its exposed signature
is a direct read-out of that math. The main kernel, `flash_kda.fwd`, takes query, key and value plus a **gate**
tensor and **beta logits passed through a sigmoid** — the gate is $\alpha_t$, the channel-wise decay; beta is
$\beta_t$, the delta-rule write strength — with an optional recurrent state in and out, which is $S_t$ itself: the
fixed-size matrix that is the whole reason a 1M context stays tractable. It also batches variable-length sequences
via cumulative sequence lengths and runs mixed precision (bf16 activations, fp32 params).
FlashKDA needs SM90+ (Hopper and newer), CUDA 12.9+ and PyTorch 2.4+, ships benchmarks for H20 and GB200, and
auto-integrates with the `flash-linear-attention` library (v0.5.0+) through a `chunk_kda` op — the same KDA lineage as
the vLLM contribution above, now available as a standalone package rather than only living inside a serving
framework.
### NoPE: no positional encoding at all
Here is a detail the weights make unambiguous, and it is one of the quietly radical choices in K3: **there is no
positional encoding**. Not RoPE, not ALiBi, not learned embeddings. K2 used [RoPE](/architectures); K3 applies **NoPE**
to every MLA layer and lets the KDA layers carry position implicitly through their gating and decay — a recurrence is
inherently order-sensitive, so position falls out of the mechanism rather than being added to it.
The payoff is at the context frontier. Extending a RoPE model to 1M tokens normally means rescaling the frequency base
or applying YaRN-style interpolation, and every such trick is a place quality can quietly degrade. With NoPE there is
nothing to rescale: the report says K3 **extrapolates directly to 1M-token contexts without any positional-encoding
modification**. The 3:1 hybrid earns its keep here — KDA supplies position-sensitive, recency-aware mixing, while the
NoPE-MLA layers supply unrestricted global content interaction, and the two jobs stay cleanly separated.
## Attention Residuals: selective retrieval across depth
The second change is about *depth*, not length. A plain residual stream compresses all prior information into a single
state as it climbs — a bottleneck the report pointedly compares to an RNN over time. Transformers already solved that
problem along the sequence axis by replacing recurrence with attention; **Attention Residuals** applies the same move to
depth: each layer *selectively retrieves* representations from preceding layers rather than accumulating them uniformly.
Toggle the two modes and scrub the current layer:
Mechanically, each layer $l$ carries a learnable pseudo-query $q_l$; the keys and values are the outputs of all earlier
layers (plus the token embedding), and attention weights come from a softmax kernel with an RMSNorm inside — the norm
stops layers with large-magnitude outputs from dominating the read. Because depth is modest ($L < 100$), the full
$O(L^2 d)$ form is affordable in arithmetic; the real cost is the $O(Ld)$ memory of keeping every layer output alive.
**Block AttnRes** is the fix, and it is what K3 actually ships. The 93 layers are partitioned into blocks of 12; within
a block, layer outputs are summed into one representation, and full attention runs only over the ~8 block-level
representations. Memory and cross-stage communication drop from $O(Ld)$ to $O(Nd)$, and the block structure bounds the
inference-time state so inter-block results merge with intra-block partial sums via online softmax. The report notes
$N \approx 8$ recovers most of the benefit — which is exactly the 8 blocks of 12 the config encodes.
The payoff Moonshot reports is concrete: about **25% higher training efficiency at under 2% additional cost**. That ratio
is the tell — a cheap structural change that improves gradient flow and lets the stack go deeper without the usual
degradation, which is exactly the kind of lever that compounds into the headline 2.5× scaling number.
Alongside these, the attention sublayer uses **Gated MLA** — Multi-head Latent Attention with an input-dependent,
channel-wise **full-rank** output gate, letting each token modulate which channels it reads from global attention. The
MLP nonlinearity is a **Sigmoid Tanh Unit (SiTU-GLU)**, whose gate branch is a $\tanh$ bounded by a constant, so
activations cannot blow up. Small pieces, but at 2.8T scale "bounded" is load-bearing.
## Stable LatentMoE: 16 of 896, and why that is hard
Here is the aggressive part. K3's feed-forward is a mixture of experts with **896 routed experts**, of which only **16**
fire for any given token (plus 2 always-on shared experts) — a sparsity of **56**. The experts are **latent**: rather
than each selected expert receiving the full 7168-dimensional token, routed experts operate in a compact **3584-wide**
latent space, half the model width, while the shared experts keep a full-width path. That separation is what makes the
expansion affordable — in a conventional MoE, communication and expert-weight traffic grow with routing multiplicity, so
going to 16 active experts would be punishing at full width. Scrub a few tokens and watch the selected 16 change:
At this sparsity, two problems that are mild in a denser MoE become first-order. **Exploding activations:** the routed
path composes a down-projection, a gated multi-branch expert FFN, and an up-projection into a chain of nearly four
consecutive matmuls — ill-conditioned at 2.8T scale, which is what the normalization and the bounded SiTU-GLU are there
to contain. **Load balance:** balancing nearly a thousand experts per layer exceeds the regime where existing
auxiliary-loss-free schemes hold up. If a few experts hog the tokens, the rest never train, and the effective model
collapses to something far smaller than 2.8T.
### Quantile balancing: no auxiliary loss, no knob
K3 stays auxiliary-loss-free: balancing is done by adding a per-expert bias $b_j$ to the router score *before* Top-$k$
selection, and then omitting that bias from the mixture weights — so it steers dispatch without touching the router's
gradients. The standard version nudges $b_j$ by a fixed step in the direction of the load error, which forces a
trade-off between slow adaptation and load oscillation.
**Quantile Balancing** replaces the nudge with a direct solve. Routing runs Top-$(k{+}1)$ instead of Top-$k$: the first
$k$ entries are the routes actually taken, and the $(k{+}1)$-th is the **cutoff** a competing expert would have had to
beat. Each expert's next bias is then read off as a quantile of its *margins* (score minus cutoff) across the batch —
specifically the $(1 - k/n)$-quantile — which by construction hands every expert exactly its target load of $mk/n$
tokens. No auxiliary loss, no balance coefficient. Drag the quantile and flip to the aux-loss regime to see the
imbalance it removes:
At training scale those margins number in the millions and are scattered across ranks, so an exact quantile is not
computable. K3 estimates it from a **histogram**: each rank bins its own margins, a single all-reduce sums the bin
counts, and the quantile is recovered from the pooled histogram. Because counts are additive, the estimate reflects the
true whole-batch quantile up to the bin width — at a communication cost of a few hundred bins per expert. The bias is
frozen at inference.
The systems half matters just as much. K3 uses **perfectly balanced expert-parallel training with static shapes and no
host synchronization**. Variable expert loads normally produce variable tensor shapes, which force recompilation and
host-side synchronization that stalls a large cluster. Quantile balancing gives every expert the same load, so the shapes
are static, so the expert-parallel pipeline runs without host sync — the difference between 16-of-896 routing being a
nice idea and being trainable at 2.8T.
With all four pieces on the table, here is the module-level picture redrawn: the **Stable LatentMoE** and **KDA**
blocks in full detail on the left, and on the right the **Block Attention Residuals** backbone — where each module's
output flows through an `α` gate that can read *every* earlier block and the embedding, not just the layer below it.
### MoonEP: the same static-shapes claim, from the communication side
**[MoonEP](https://github.com/MoonshotAI/MoonEP)** (MIT) is Moonshot's expert-parallel communication library, and it
is the static-shapes claim from the quantile-balancing section above, attacked from the other direction. Quantile
balancing makes every expert's *load* equal before dispatch; MoonEP instead guarantees every rank *receives* exactly
$S \times K$ tokens — $S$ input tokens per rank, $K$ routed top-$k$ per token — no matter how skewed the actual
routing is. The mechanism is a small number of redundant experts, planned online from the current router outputs by
a near-optimal GPU planning kernel and prefetched before expert computation runs, with their gradients reduced back
to their home ranks on the backward pass. Because those redundant experts absorb whatever skew is left, every rank
ends up with an identical, statically-known token count — "statically known shapes eliminate per-layer MoE host
synchronization" is not a paraphrase of that claim, it is MoonEP's own description of what it buys. Tokens land
directly in their expert-grouped positions on remote ranks through zero-copy buffer views, so only a fixed
$S \times K$ buffer is needed per layer, with no per-layer host synchronization to stall the pipeline.
Moonshot's own benchmarks against DeepEP v2 on H20 make the case concrete: MoonEP's communication time stays close to
flat as imbalance (maxvio) grows, while DeepEP v2 degrades steadily and eventually OOMs under high imbalance — and
MoonEP's iteration time holds flat across the same range. It targets NVIDIA GPUs today, with Zhenwu PPU support listed
as under review, and credits DeepEP, Echo and UltraEP as inspiration.
## Native vision, trained from scratch
K3 is natively multimodal — text, images and video share one backbone and one context, with no post-hoc alignment stage.
The notable choice is how the vision tower was trained. Standard practice, including K2.5's own, initializes the encoder
from a contrastively pre-trained model like SigLIP. K3 instead trains **MoonViT-V2** (401M params, 27 layers, patch 14)
**entirely from scratch with next-token prediction**.
The reason given is stability, and the report shows the receipts: the SigLIP-initialized tower ran persistently higher
gradient norms with frequent spikes, while the from-scratch tower stayed smooth. Training under the language-modeling
objective also shapes visual features by what the LLM actually needs — fine-grained text and structure — rather than the
global semantics a contrastive loss rewards. The conclusion is the interesting bit: MoonViT-V2 **matched** the
SigLIP-initialized baseline on vision evals, so at this scale contrastive pre-training simply was not necessary.
## Turning compute into intelligence
Stack it up — KDA's cheap long-context memory, AttnRes's cheap depth, LatentMoE's extreme-but-stable sparsity, plus
refined training and data recipes — and the headline is a **~2.5× improvement in overall scaling efficiency** over K2.
This is not a vibe: it is a fitted scaling-law comparison on held-out out-of-distribution validation data, with
hyperparameters (batch size, learning rate, tokens-per-parameter, model shape) re-tuned independently for each family so
neither is handicapped by the other's settings.
Read the gap horizontally: pick any loss level and the red curve reaches it about 2.5× further left on the FLOPs axis.
Drag the capability marker to see the same trade in the other direction:
That is the number that actually matters. "2.8 trillion parameters" is a spec-sheet figure; "2.5× more capability per
FLOP" is an engineering result. A side note from the same study, useful to anyone tuning their own runs: under
independently optimized hyperparameters, **cosine decay consistently beat WSD** — the two schedules have very different
optimal peak learning rates and batch sizes, so comparisons that share one hyperparameter set tend to be unfair to
whichever schedule they fit worse.
## What it would take to train it
So what does building a 2.8T-A104B model actually cost? Sparsity still helps: training compute for an MoE scales with
the **active** parameters, not the total, so K3's per-token training FLOPs are those of a ~104B model rather than a
2.8T one. The standard estimate is
$$
C \approx 6 \, N_{\text{active}} \, D
$$
with $N_{\text{active}} \approx 104\text{B}$ and $D$ the number of training tokens. Moonshot still has not published K3's
token budget; for reference, K2 was trained on **15.5T tokens**. Plug in a frontier-scale budget and pick a cluster:
Three things make that estimate *achievable* rather than merely large:
- **Per-Head Muon.** K3 extends the Muon optimizer so that Newton–Schulz orthogonalization is applied to each attention
head's momentum block *separately* rather than to the whole Q/K/V projection. Full-matrix orthogonalization lets
large-gradient heads dominate the shared update direction; per-head equalizes the update scale across heads, which
improves stability at scale — and is slightly cheaper, since the iterations run on tall thin blocks.
- **MXFP4 / MXFP8 quantization-aware training.** From the SFT stage onward, K3 trains with **MXFP4 expert weights and
MXFP8 activations**, while attention projections, latent-MoE projections, shared experts and routers stay in higher
precision. The model is trained to be low-precision-native, which is why the full 2.8T weights fit in roughly
**1.4 TB** and why it is servable at all without a quality cliff.
- **Static-shape expert parallelism.** As above — quantile balancing plus static shapes and no host synchronization is
what keeps a large cluster busy instead of stalling on dynamic routing.
The context window is built up rather than trained flat: pre-training starts at **8K** and extends to **64K**, then the
cooldown phase walks **256K → 1M**. Concentrating the expensive long-sequence compute into a small slice of the budget is
what makes a 1M-token model economical. Length alone does not confer long-range ability, so Moonshot also *synthesizes*
long-context data by permuting and concatenating documents and sub-tasks such that the embedded task can only be solved
by attending across the full window — otherwise attention quietly degenerates into local patterns.
## Post-training: nine experts, then one
The pre-training story is where the architecture lives, but K3's post-training has a structure worth drawing. It is a
three-stage funnel: SFT for a cold-start policy, then RL that trains **nine separate experts** — three domains crossed
with three reasoning-effort levels — then **Multi-Teacher On-Policy Distillation** to collapse all nine back into the
single shipped checkpoint.
A few mechanisms make that work at 1M context:
- **Partial rollout.** In long-horizon RL, a handful of straggler trajectories can hold up an entire iteration.
Generation instead pauses once a fraction $\lambda$ of trajectories finish; the rest are enqueued and resumed at the
start of the next iteration, backed by persistent sandbox state. That means a single trajectory can span several
iterations, so the algorithm has to tolerate badly stale off-policy data — which it does via a per-token
regularization that keeps updates in a local neighborhood.
- **A white-box harness, not *the* harness.** Training against one fixed agent scaffold teaches the model that
scaffold's conventions. Moonshot's RL environment represents a harness as composable modules — tools, system prompts,
context management, skills, memories, subagents — and can instantiate Kimi Code, Claude Code, Codex, OpenClaw and
Hermes, mixing configurations across task groups so the model generalizes across harnesses rather than overfitting
one.
- **Deployment-aware training.** QAT runs through the *entire* post-training stage, and during RL the rollout and the
training pass share the same quantization scheme — eliminating the train/inference mismatch that usually shows up when
a model is quantized after the fact. Separately, K3's pre-trained multi-token-prediction layer is fine-tuned into an
EAGLE-3-style **draft model** for speculative decoding, optimized directly against the acceptance rate rather than a
KL surrogate.
### AgentENV: the sandbox layer, and it is open source
The piece that makes all of the above physically possible is the sandbox. Long-horizon agentic RL means running an
enormous number of real machines that agents can break, and **[AgentENV](https://github.com/kvcache-ai/AgentEnv)** —
built by Moonshot with partners, and released under MIT — is the microVM runtime they built for it.
The motivation is refreshingly blunt. Container-based sandboxes were not enough: in early experiments, aggressive agent
exploration caused **kernel panics and deadlocks**. And clamping down is the wrong fix, because hard tasks need a
sandbox close to a real machine — agents should be able to mount disks, run containers, even launch VMs. So AgentENV
runs each sandbox as an isolated **Firecracker** microVM, buying isolation and fidelity a container cannot.
On top of that it adds three lifecycle operations tuned specifically for RL:
- **Pause / resume.** A paused sandbox consumes no memory or CPU. This matters more than it sounds: the sandbox spends
as much as **98% of its lifetime** just waiting on the model's next inference result. Pausing that window is the
difference between renting an idle fleet and not.
- **Fork.** Branch a new sandbox from the *exact* state of a running one while the original keeps going — which is how
you run a reward judge against a trajectory without any side effects leaking back into it.
- **Snapshot.** Periodic checkpoints for error recovery.
The engineering is in the latencies: incremental checkpointing saves only pages dirtied since the last checkpoint,
giving **133 ms checkpoint and 49 ms resume**. Images use OverlayBD with a custom `ublk` driver, storage-layer sharing
and P2P transport, so tens of thousands of sandboxes with distinct images launch in **under a second**; copy-on-write
memory and page-cache tuning push memory overcommit to **6.5×** in real workloads.
The scale number is the one worth sitting with. Across K3's training and evaluation, Moonshot created
**51,219,741 sandboxes** spanning **1,505,678 distinct images**. That is what "agentic RL" costs when you actually run
it — and it is the part of the frontier stack that almost never gets published, let alone open-sourced.
AgentENV is one of three pieces of that stack Moonshot has now open-sourced: **MoonEP** (expert-parallel
communication, covered under Stable LatentMoE above) and **FlashKDA** (the attention kernel, covered under Kimi Delta
Attention above) are the other two — sandbox, communication and kernel, all MIT-licensed.
## The benchmarks
On coding, K3 is a clear #2-or-#3 behind Fable 5 and GPT-5.6 Sol, and ahead of everything else open or closed that
Moonshot tested — with a few outright wins.
On FrontierSWE it sits second, close behind Fable 5 and well ahead of the rest:
On Terminal Bench 2.1 it is effectively tied for first, and on the long-horizon SWE Marathon and Program Bench it is
first outright:
The agentic and visual picture is similar — competitive across the board, and #1 on browsing:
The pattern is consistent: K3 wins where the task is long-horizon and tool-heavy (SWE Marathon, Program Bench,
BrowseComp, Automation Bench, SpreadsheetBench 2), and comes second to Fable 5 or GPT-5.6 Sol on the single-shot,
knowledge-dense ones (GDPval and AA-Briefcase Elo, DeepSWE).
### Independent numbers
The obvious objection to everything above is that it is Moonshot grading its own homework. The report also collects
third-party leaderboards, which is the more useful evidence:
| Leaderboard | Kimi K3 | Rank | Best proprietary |
|---|---|---|---|
| Artificial Analysis Intelligence Index v4.1 | 57.1 | #4 / 580 | Fable 5 — 59.9 |
| Vals Index | 74.7 | **#2 / 39** | Fable 5 — 75.1 |
| WebDev Arena (Elo) | 1,678 | **#1 / 99** | Fable 5 — 1,634 |
| Text Arena (Elo) | 1,486 | #8 / 200 | Fable 5 — 1,507 |
| Agent Arena | 9.1 | #4 / 37 | Fable 5 — 12.7 |
An open model holding **#1 on WebDev Arena** and **#2 on the Vals Index** — 0.4 points off Fable 5 — is a materially
different claim from a vendor bar chart. Text Arena at #8 is the honest counterweight: general chat preference is not
where K3 shines.
## What it costs to serve
The sparsity that makes K3 cheap to train makes it cheap to run. API pricing is **$0.30 / MTok** on cache-hit input,
**$3.00 / MTok** on cache-miss input, and **$15.00 / MTok** output — and Moonshot reports cache-hit rates **above 90%** in
coding workloads, so the effective input price is closer to the cheap number than the expensive one. MXFP4 weights keep
the footprint at ~1.4 TB. The weights ship with deployment recipes for **vLLM**, **SGLang** and **TokenSpeed**.
The report's cost-efficiency comparison is the most quotable result in it:
Concretely: on **BrowseComp**, K3 takes the best score (91.2%) at **$2.03 per task** — half the cost of GPT-5.6 Sol and
an order of magnitude cheaper than the Claude models at max effort. On **Kimi Code Bench 2.0** it is 4.0 points behind
Fable 5 at **38% of the cost**, and at *high* effort it already matches Opus 4.8's *maximum*-effort score at roughly a
third of the price. On **GDPval-AA v2** it is within 50 Elo of GPT-5.6 Sol at 13% lower cost, and 2.6× cheaper than
Fable 5.
## What it can actually build
The case studies are where the long-horizon claims get concrete, and they are unusually ambitious:
- **GPU kernel optimization.** Given a sandbox and up to 24 hours per task, K3 cut **AttnRes** kernel latency from
283.6 ms to **114.4 ms**, cut DSA and KDA runtime by **55.1%** and **73.6%**, and reached over half of peak TFLOPS on
MLA — matching Fable 5 and beating Opus 4.8, GPT-5.6 Sol and GPT-5.5. Moonshot notes an early K3 checkpoint was
already doing most of their kernel-optimization work during late development.
- **A GPU compiler.** K3 built [MiniTriton](https://github.com/MoonshotAI/minitriton), a Triton-like compiler with a
tile-level Python frontend, an MLIR annotation layer and a PTX codegen pipeline, plus a dual-mode tensor library with
reverse-mode autograd and NCCL distributed primitives. On an L20 it beats PyTorch eager and `torch.compile` in
geometric mean, its from-scratch tensor-core matmul reaches ~90% of the measured machine roof, and it trains a GPT
end-to-end with gradients matching torch autograd to within torch's own fp32 rounding error.
- **A chip.** In a single **48-hour autonomous run**, K3 designed, optimized and verified an inference-chip prototype
([nano-kpu](https://github.com/MoonshotAI/nano-kpu)) using open-source EDA tools and the Nangate45 cell library.
Inside a 4 mm² budget it closes timing at 100 MHz for an RTL-simulated **8,700+ tokens/s** decode, with 1.46M standard
cells, 0.277 MiB of SRAM and an INT4 MAC array with fused dequantization.
**Read the caveats.** (1) K3 still **trails Fable 5 and GPT-5.6 Sol** on overall capability and on user-experience
polish — Moonshot says so directly, and flags sensitivity to thinking-history preservation and over-proactiveness in
ambiguous situations. (2) The headline charts are **Moonshot's own suite**, with opponents' fallbacks (Fable 5 hit
fallbacks on 35% of SWE-Marathon tasks) and cyberguards (GPT-5.6 Sol) noted; several suites also run K3 in *its own*
harness (Kimi Code) against competitors in Claude Code or Codex. The third-party leaderboards above are the better
evidence. (3) SWE-Marathon and PostTrainBench were run on **H20 GPUs** against an H20-recalibrated task branch, not the
official H100 setting. (4) The training-cost estimate is a **first-principles calculation**, not a disclosed figure —
Moonshot has published neither K3's token budget nor its cluster. (5) The license is a **custom Kimi K3 License**, not
Apache or MIT — read it before commercial use.
## The take
Strip away the size record and what is genuinely new in K3 is a coherent set of efficiency bets: **KDA** buys a real 1M
context with constant-size memory; **NoPE** means that context needs no rescaling tricks to reach; **AttnRes** buys depth
almost for free; **Stable LatentMoE** with **Quantile Balancing** buys 2.8T of capacity at 104B of active compute *and*
makes that extreme sparsity trainable without an aux-loss knob or host-sync stalls; **Per-Head Muon** and **MXFP4/MXFP8
QAT** make the whole thing converge and fit. The sum is the number that matters — **~2.5× more capability per FLOP than
K2**, measured on fitted scaling curves — delivered in the open at 2.8T.
It does not top the frontier, and it does not pretend to. What it proves is that the gap between open and closed is now
measured in scaling *efficiency*, not in whether an open lab can build at frontier scale at all — and with the weights,
the config, and a 47-page report on the table, that claim is now something anyone can go audit.
---
*Sources: the [Kimi K3 technical report](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf) (architecture,
scaling law, post-training, infrastructure, evaluations, case studies), the released
[model weights and card](https://huggingface.co/moonshotai/Kimi-K3) (config, deployment, license), the
[AgentENV repository](https://github.com/kvcache-ai/AgentEnv) (sandbox runtime), the
[MoonEP repository](https://github.com/MoonshotAI/MoonEP) (expert-parallel communication), the
[FlashKDA repository](https://github.com/MoonshotAI/FlashKDA) (attention kernels), and the
[Kimi K3 tech blog](https://www.kimi.com/blog/kimi-k3) (pricing). Figures 3–5 here are the report's Figures 2, 7 and 13,
reproduced for commentary. Benchmark numbers are Moonshot's except where marked third-party; the training-cost figures
are a first-principles estimate from $C \approx 6\,N_{\text{active}}\,D$ with clearly labeled assumptions, using K2's
15.5T-token budget as a reference. Interactive diagrams are mine; the routing, loop and cost visuals are illustrative.*
---
# BTL-3: a rank-32 LoRA that turns Qwen3.6-27B into a tool-use agent
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/btl-3
> date: 2026-07-24
> tags: agents, tool-use, code-generation, fine-tuning, open-weights, explainer
Most "new models" are not new weights. **BTL-3**, from **Bad Theory Labs**, is a clean example: it
is not a from-scratch 27B model but a **frozen rank-32 PEFT LoRA adapter** — about **934 MB** of
weights — post-trained on top of a pinned revision of **Qwen3.6-27B**. Load the base, apply the
adapter, and you get an agent tuned for coding, repository work, and structured tool use. The raw
capability is Qwen's; what BTL-3 contributes is **behaviour** — how the model runs an agent loop and,
notably, when it decides *not* to act.
That framing matters for reading the numbers honestly, so keep it in mind: this is a post-training
result on a strong open base, released under **Apache-2.0**, with the base checkpoint pinned to an
exact revision for reproducible loading. The card labels this frozen release **"RL-0013,"** and the
maximum RL sequence length (65,536 tokens) tells you the adapter was shaped by reinforcement learning
over long, multi-step trajectories — not just supervised fine-tuning on completions.
## The loop it was tuned to run
An agent model earns its keep inside a loop, not on a single completion. BTL-3's stated job is to
"reason, act, inspect tool results, recover from failures, and stop when no action is required." That
last clause is the interesting one. A tool-happy model that always reaches for a function call is easy
to train and annoying to deploy; the harder behaviour is the **route** decision — recognising when a
question needs a tool and when it just needs an answer. Pick a scenario and watch which path the model
takes:
The four scenarios line up with the four things BFCL (the Berkeley Function-Calling Leaderboard) v4
actually measures: a **single** call, **parallel** calls fired at once, recovery when a call **fails**,
and **irrelevance** — correctly declining to call anything. BTL-3's headline highlight is that last
one: a self-reported **91.2%** on knowing when to stay its hand. The loop is the product here; the
adapter's whole point is to make Qwen3.6-27B move through it reliably.
## The tool-use profile
Break BFCL v4 down by category and the shape of the model shows. It is strongest on the
straightforward cases and gives ground exactly where you'd expect — when it has to compose *several*
tools that each take *several* arguments:
The aggregate is **88.5% BFCL v4 AST** (1097/1240 on the full official set). The 70.0% on
parallel-multiple is the honest soft spot: issuing several correct calls at once, each with the right
arguments, is where structured tool use is genuinely hard, and a fifth of those cases still slip. The
**91.2% irrelevance** number is the one worth internalising — it is the difference between an agent you
can leave in a loop and one that invents work.
## Coding, and where it falls off
On standard code-generation benchmarks in **thinking mode**, BTL-3 posts strong pass-rates — and then
drops sharply on the hardest composite tasks. That gap is the useful part of the picture, not a number
to bury:
HumanEval at **95.12%** (156/164) and LiveCodeBench v6 at **88.1%** (170/193) are the flattering
figures — well-scoped "write this function" problems. **BigCodeBench-Hard Instruct at 26.35%**
(39/148) is the sobering one: strict pass@1 on tasks that chain many library calls into one correct
program is a different sport, and here the model solves roughly one in four. (BTL reports a softer
**59.25%** at the individual *test* level on the same suite — useful context, but a test-level score is
not a solved-task score, so read the strict 26.35% as the real one.) These are different benchmarks at
different difficulties, not a like-for-like ladder — the labels carry that.
## The Compact edition
Alongside the adapter, Bad Theory Labs ships **BTL-3 Compact**: the complete text model packed into a
single **8.39 GB** native file — smaller than an 8B model stored in FP16, which works out to an
effective **under 2.5 bits per parameter**. The claimed cost of that compression is measured on a
"fresh private 100-turn tool-contract gate": Compact retained **83 of the 90 behaviours** the full
model completed correctly, which BTL reports as **92.2% conditional tool-behaviour retention**.
Read that metric for exactly what it is. It is a *private* gate that BTL defined and ran, conditioned
on cases the full model already passed — so it says "Compact reproduces most of what the full model got
right," not "Compact loses only 8% overall." It's a reasonable internal check and a genuinely useful
artifact (a 27B-class agent in 8.39 GB is easy to self-host), but it is not an independent quality
measurement.
## Running it
Because BTL-3 is a LoRA adapter, deployment is "load Qwen3.6-27B at the pinned revision, then apply the
adapter" — a few lines with PEFT and Transformers, or `vllm serve` with `--enable-lora` and
`--max-lora-rank 32` and the Qwen XML tool parser for structured calls. The architectural context
window is **262,144 tokens** (inherited from Qwen3.6's hybrid attention), though the published
benchmarks were run at a **32,768-token** launch context. BTL recommends **thinking mode** for coding
and reasoning, which is also the mode every headline score was measured in.
The model card itself is direct about this: **run generated code and tool calls in a sandbox**, and
require **explicit confirmation before destructive, privileged, financial, or otherwise high-impact
actions**. An agent that scores 88.5% on tool calls still gets more than one call in ten wrong — that
residual is exactly where an unsandboxed loop does damage.
## The honest read
Every number here is **self-reported by Bad Theory Labs** — there is no independent evaluation yet. The
underlying **capability is Qwen3.6-27B's**; BTL-3 is a **post-training / RL result**, so credit the
adapter for the *loop behaviour* (tool routing, irrelevance, recovery), not for raw reasoning power.
Scores were measured in **thinking mode** at a **32K context** on protocols BTL chose, the coding wins
sit next to a **26.35% BigCodeBench-Hard** floor, and the Compact edition's **92.2% retention** is a
private, conditional gate, not a public benchmark. The model card ships **no figures or diagrams** — the
loop diagram above is my own reconstruction of the described behaviour. No training-data disclosure is
provided.
## The take
BTL-3 is a modest, honest kind of release: take a strong open base, spend an RL budget teaching it to
behave inside an agent loop, and ship the ~934 MB of difference under Apache-2.0. The most interesting
claim isn't a coding score — it's the **91.2% irrelevance**, the tuned instinct to *not* call a tool.
For anyone assembling a private, self-hosted coding agent, that plus the 8.39 GB Compact build is a
concrete, deployable proposition. Just hold the framing straight: the intelligence is Qwen's, the
discipline is BTL's, and until someone outside Bad Theory Labs runs the suite, every figure is a
vendor's own.
---
*Source: the [BTL-3 model card](https://huggingface.co/badtheorylabs/BTL-3) and
[BTL-3 Compact](https://huggingface.co/badtheorylabs/BTL-3-Compact) on Hugging Face, plus the
[runtime source](https://github.com/Badtheorylabs/BTL-3). The card ships no figures, so the diagram
here is mine; all benchmark numbers are Bad Theory Labs' self-reported values.*
---
# Token-level RL is a first-order approximation to the reward you actually want
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/first-order-rl
> date: 2026-07-24
> tags: reinforcement-learning, llm, mixture-of-experts, post-training, explainer
RL for reasoning models rests on a mismatch nobody had really justified. The **reward**
is assigned to a *whole response* — you sample a full chain of thought, check the final
answer, and hand back one scalar. But the **optimizer** — REINFORCE, GRPO, and the rest —
works one *token* at a time. We reward the sequence and update the tokens, and we mostly
just trust that closing the loop this way improves the thing we scored.
The Qwen team's [Stabilizing Reinforcement Learning with LLMs](https://arxiv.org/abs/2512.01374)
(Zheng et al., arXiv:2512.01374) takes that trust and makes it a theorem with fine print.
Their claim: the token-level objective is a **first-order approximation** to the true
sequence-level reward — exact in the limit, and valid only when two specific gaps are
small. The nice part is what falls out of it. Importance-sampling correction, clipping,
and Routing Replay for Mixture-of-Experts models — a grab-bag of stabilization tricks that
each arrived with its own justification — turn out to be the *same move*: keep the
approximation valid. One lens, and the whole toolbox lines up behind it.
## The objective you can't optimize
Write the thing we actually want to maximize — expected reward over responses the current
policy would generate:
$$
J^{\text{seq}}(\theta) \;=\; \mathbb{E}_{x\sim\mathcal{D},\; y\sim\pi_\theta(\cdot|x)}\big[R(x,y)\big].
$$
There's an immediate wrinkle: we don't sample $y$ from the policy we're training. Responses
come out of a fast **inference engine** (vLLM, SGLang) running policy $\mu_{\theta_{\text{old}}}$,
while gradients are taken in a **training engine** (Megatron, FSDP) holding $\pi_\theta$.
The standard fix is an importance-sampling reweight onto the rollout policy $\mu$:
$$
J^{\text{seq}}(\theta) \;=\; \mathbb{E}_{x\sim\mathcal{D},\; y\sim\mu_{\theta_{\text{old}}}(\cdot|x)}\!\left[\underbrace{\frac{\pi_\theta(y|x)}{\mu_{\theta_{\text{old}}}(y|x)}}_{\text{sequence-level IS weight}} R(x,y)\right].
$$
This is correct and completely impractical. A sequence likelihood is a product of hundreds
or thousands of per-token probabilities, so the ratio $\pi_\theta(y|x)/\mu_{\theta_{\text{old}}}(y|x)$
swings across an enormous dynamic range with brutal variance. Its gradient is technically
right and numerically hopeless. Nobody trains on it directly.
## The surrogate everyone actually uses
So instead we optimize the **token-level** objective — sum the per-token IS ratios instead
of multiplying them:
$$
J^{\text{token}}(\theta) \;=\; \mathbb{E}_{x\sim\mathcal{D},\; y\sim\mu_{\theta_{\text{old}}}(\cdot|x)}\!\left[\sum_{t=1}^{|y|}\underbrace{\frac{\pi_\theta(y_t|x,y_{
The intuition to keep: the surrogate isn't wrong, it's *truncated*. As long as the policy
you're optimizing stays close to the policy that generated the data, the truncation is
negligible and improving the cheap objective improves the real reward. Let them separate and
the surrogate starts optimizing something that isn't the reward anymore — which, in practice,
is exactly what a training collapse looks like.
## Two gaps, and the tricks that close them
"Keep $\pi_\theta$ close to $\mu_{\theta_{\text{old}}}$" sounds abstract until you factor the
token ratio into its two honest sources:
$$
\frac{\pi_\theta(y_t|\cdot)}{\mu_{\theta_{\text{old}}}(y_t|\cdot)} \;=\; \underbrace{\frac{\pi_{\theta_{\text{old}}}(y_t|\cdot)}{\mu_{\theta_{\text{old}}}(y_t|\cdot)}}_{\text{training–inference discrepancy}} \;\times\; \underbrace{\frac{\pi_\theta(y_t|\cdot)}{\pi_{\theta_{\text{old}}}(y_t|\cdot)}}_{\text{policy staleness}}.
$$
- **Training–inference discrepancy** is numerical. The same weights produce slightly
different probabilities in the training and inference engines — different kernels,
and inference deliberately disables batch-invariant kernels for throughput, so even one
engine isn't self-consistent. This is the gap between $\pi_{\theta_{\text{old}}}$ and
$\mu_{\theta_{\text{old}}}$.
- **Policy staleness** is procedural. To use more compute per rollout, we split a big batch
of responses into mini-batches and take several gradient steps, so later mini-batches are
optimized by a $\pi_\theta$ that has already drifted from the $\pi_{\theta_{\text{old}}}$
that generated them. Asynchronous frameworks make it worse.
Now the stabilization toolbox reads as one idea — shrink these two gaps so the first-order
approximation holds:
- The **IS weight itself** is not an optional variance trick; it *is* the first-order term.
Drop the training–inference correction and you're no longer approximating the sequence
objective at all.
- **Clipping** (the PPO move) stops gradients on tokens whose ratio has run too far,
directly capping policy staleness.
- **Routing Replay**, for MoE models, closes both — and it needs its own section, because
MoE breaks the story in a way dense models don't.
## Why MoE breaks it, and how Routing Replay repairs it
In a Mixture-of-Experts model, the probability of a token depends on *which experts the
router activated* for it. That turns the token ratio into a comparison over possibly
*different active parameters*: the inference engine routes the token to expert set $e^\mu$,
the training engine to $e^\pi$, and when those sets disagree the ratio
$\pi_\theta(y_t|\cdot)/\mu_{\theta_{\text{old}}}(y_t|\cdot)$ stops measuring "a small change
in the policy" and starts measuring "two different subnetworks." The $\delta_t$ are no longer
small, and the first-order approximation collapses. Routing is entangled with *both* gaps —
the engines can route differently (discrepancy) and the router's choice can shift as weights
update (staleness). Toggle the fix:
**Routing Replay** ([Zheng et al., 2025](https://arxiv.org/abs/2507.18071); Ma et al., 2025)
pins the routed experts during optimization so the token is scored over one fixed subnetwork
— the model is optimized like a dense one, and the ratio means what it should again. The
paper formalizes two flavors, differing only in *whose* routing you replay:
| | replays | closes | first mini-batch |
|---|---|---|---|
| **R2** — Vanilla Routing Replay | the training engine's rollout experts $e^\pi_{\text{old}}$ | policy staleness | target policy **unaltered** |
| **R3** — Rollout Routing Replay | the inference engine's experts $e^\mu_{\text{old}}$ | discrepancy **and** staleness | target policy altered |
There's no free lunch here, and the paper is careful to say so. Fixing the experts restores
the approximation but **biases the target policy** — you're now optimizing a model whose
routing is frozen to a past choice, not the routing it would pick itself. R2 leaves the first
mini-batch's target policy untouched; R3 alters it from step one but kills more of the
discrepancy. Which bias is worth paying turns out to depend on how off-policy you run — a
question only experiments can settle.
## MiniRL: the smallest honest baseline
To test the formulation instead of a pile of confounded tricks, the authors strip RL down to
**MiniRL** — REINFORCE with the token-level IS weight, group-normalized advantages (subtract
the per-prompt mean reward), and PPO-style clipping. That's it. It's deliberately the minimal
algorithm whose gradient stays faithful to the surrogate the theory justifies, which makes it
the right probe: if the formulation is real, the things that preserve the approximation should
be the things that stabilize MiniRL.
The setup is a genuine stress test. A 30B MoE (cold-started from Qwen3-30B-A3B-Base), **FP8
inference against BF16 training** — deliberately mismatched precisions to *inflate* the
training–inference discrepancy — on 4,096 verifiable math problems, scored as average accuracy
over 32 samples on HMMT25, AIME25, and AIME24. Hundreds of thousands of GPU hours, roughly
5–6 GPU-hours per gradient step. They track not just reward but two diagnostics: policy
**entropy** and the **training–inference KL divergence**, since a collapse announces itself as
a KL spike before the score falls.
## What the experiments say
**On-policy** (one gradient update per batch), the ablation lands exactly where the theory
predicts:
Three reads, each a prediction of the formulation:
- **MiniRL wins.** The plain first-order-faithful objective is the most stable and scores
highest.
- **Removing the training–inference IS correction collapses training** almost immediately —
the green curve nose-dives, entropy crashes, KL explodes. The IS weight was never optional;
it's the approximation's load-bearing term.
- **Length normalization is stable but worse.** Dividing the objective by response length is
common (GRPO and CISPO both do it), but it *invalidates* the first-order approximation — the
gradient no longer lines up with the true sequence objective — and the benchmark score pays
for it. Notably, **Routing Replay does not help on-policy**; with the gaps already small,
its bias is all cost and no benefit.
**Off-policy** (split the batch into $N$ mini-batches for $N$ updates), staleness enters and
the picture changes. Now clipping *and* Routing Replay both become necessary — drop either and
training collapses early:
The nuance the paper draws out: at **small off-policiness** ($\text{gbs}=2\times\text{mbs}$),
**R2 beats R3** — R2's lighter bias wins when the approximation is only mildly stressed. At
**larger off-policiness** ($4\times$, $8\times$), **R3 wins** — R2 can't hold training
together and R3's stronger discrepancy-killing earns its bias back. The recipe isn't "always
use X"; it's "match the replay to how off-policy you're willing to run."
The benchmark numbers in the bar chart above are **read off the training curves in Figure 1**,
not a reported results table — treat them as approximate. And the whole study is one task
(verifiable math), one model family (Qwen MoE), and a deliberately harsh FP8-inference /
BF16-training setup chosen to *amplify* the discrepancy. The mechanism is clean; how the exact
crossover points transfer to other rewards, modalities, and precision regimes is not something
one paper can settle.
## The result that reframes the field
The finding I keep coming back to isn't a trick — it's about what *matters*. Take one base
model, cold-start it three different ways (distilling from Qwen3-Max-Thinking-Preview,
DeepSeek-R1-0528, and gpt-oss-120b), then run the same stable recipe. They converge to the
**same place**:
Once training is stable, *how you started barely matters* — prolonged RL washes out the
cold-start differences and even on-policy and off-policy runs reach comparable peaks. The
implication is pointed: the field spends enormous effort curating cold-start data, and this
says that effort is mostly erased by enough stable RL. The lever that actually moves the
ceiling is **stability**, not initialization.
## The take
- **One approximation, one toolbox.** Token-level RL is the first-order truncation of the
sequence objective. IS correction, clipping, and Routing Replay aren't three unrelated
patches — they're three ways to keep the truncation valid by shrinking the
training–inference discrepancy and policy staleness.
- **The IS weight is structural, not cosmetic.** It's the linear term itself; removing it
doesn't add variance, it changes what you're optimizing, and training collapses on contact.
- **MoE needs Routing Replay, and it's a real trade.** Pinning experts restores the ratio but
biases the target policy. Use R2 when you're near on-policy, R3 when you push off-policy —
the paper's clearest practical recipe.
- **Stability is the scaling lever.** Different cold-starts, on-policy vs off-policy — once
stable, they land in the same place. The honest limits: one task, one model family, a
stress-test precision setup, and headline benchmark values read from curves rather than a
table. The formulation is the durable part; the exact numbers are a single, if very large,
data point.
---
*Built on [Stabilizing Reinforcement Learning with LLMs: Formulation and
Practices](https://arxiv.org/abs/2512.01374) (Zheng et al., Qwen Team, Alibaba). Routing
Replay's two flavors trace to [GSPO](https://arxiv.org/abs/2507.18071) (R2) and Ma et al.,
2025 (R3, arXiv:2510.11370). Figures are the paper's; the interactive diagrams are mine.*
---
# FLUX 3: when an image model decides to become a world model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/flux-3
> date: 2026-07-24
> tags: diffusion, image-generation, video-generation, flow-matching, black-forest-labs, explainer
Every FLUX before this one was an image model. FLUX.1 was a 12B rectified-flow text-to-image
transformer; FLUX.2 scaled the same recipe to a ~32B image-and-editing model. [FLUX 3](https://bfl.ai/blog/flux-3),
announced by [Black Forest Labs](https://bfl.ai) on 2026-07-23, is a different kind of object. It is a
single **multimodal foundation model** that learns jointly from images, video, and audio in one
architecture — and then generates all of them, plus robot *action*, from the same backbone. The framing
in the title is deliberate: "Real World Models." BFL is no longer trying to make prettier pictures; it is
trying to build a model of how the world looks, moves, and sounds.
The argument for one model is an information argument, and it is worth stating in BFL's own terms. No
single modality is a complete description of reality — each is a lossy projection captured by a different
sensor. Images fix spatial structure at one instant; video restores time and reveals physical dynamics;
audio exposes causal links between mechanical events and the sounds they make; language ties all of it to
goals and instructions. Learn from one projection and you get a good model of that projection. Learn from
all of them at once and their **mutual constraints** teach you more: the sound has to match the impact,
the motion has to obey the mass, the future has to follow from the past. The modalities stop being
separate problems and start being evidence about one underlying reality.
That is the pitch. Because FLUX 3 shipped as an Early Access *capabilities* announcement rather than a
tech report — no parameter count, no full architecture, benchmarks that are BFL's own — the honest way to
cover it is to explain the mechanism it is built on from first principles, show the evidence BFL did
publish, and keep the caveats visible. So: two mechanisms first, then the numbers.
## Mechanism 1: flow matching, and why few steps is the whole game
FLUX has always been a **flow-matching** model, and FLUX 3 builds on an approach BFL calls
[Self-Flow](https://bfl.ai/research/self-flow). Flow matching is the cleaner cousin of diffusion, and the
idea is simple enough to hold in your head. You want to turn a sample of pure noise into a sample of data.
So you define a path between them — for training, literally the straight line
$x_t = (1-t)\,x_{\text{noise}} + t\,x_{\text{data}}$ — and you train a network to predict the **velocity**
$v_\theta(x, t)$ that points along that path. Generation is then just solving an ordinary differential
equation: start at noise, and take steps in the direction the velocity field tells you, until $t = 1$.
The subtlety is *how many steps*. Each step is a full forward pass of a huge transformer, so steps are the
cost. And here the geometry of the path matters enormously. If the learned field is **straight** — a
constant velocity, which is what "rectified" flow aims for — then a first-order Euler integrator lands on
the data exactly, no matter how few steps you take. If the field is **curved**, few steps cut the corner
and miss the target; the error only closes as you add steps (and cost). Drag the step count and flip the
field to feel it:
This is why "straighten the transport paths" has been the central obsession of the whole FLUX / rectified-flow
lineage: a straighter field means a few-step sampler that still lands, which means a 20-second video is
merely expensive instead of impossible. The trajectory geometry above is the same one the
[Mage-Flow](/articles/mage-flow) piece leans on for its 4-step Turbo — few-step generation is a property of
the *field*, not a trick bolted on afterward.
## Mechanism 2: one sequence, joint attention across modalities
The second mechanism is how four modalities fit in one model. FLUX 3 sits in the **MMDiT** lineage — the
multimodal diffusion transformer that the [Mage-Flow explainer](/articles/mage-flow) walks through in
detail — and the load-bearing idea is that every modality is tokenized into a *single* sequence, and one
attention operation runs over the whole thing. A query token is not confined to its own modality: an audio
token can attend to the video frame that produced the sound; a video token can attend to the text that
describes the scene. Pick the query's modality below and flip joint versus per-modality attention:
Per-modality attention gives you four models in a trenchcoat. **Joint** attention is what lets the mutual
constraints actually do their work — it is the mechanical form of "the modalities are evidence about one
reality." BFL's own diagram makes the shape concrete: shared per-modality encoders feed a single
multimodal transformer, shared decoders read it back out, and *action* is added as one more
lane — explicitly marked extensible.
Self-Flow is BFL's method for aligning generation and understanding inside that one backbone, and the one
quantitative claim they put behind it is that unifying the two lifts *both* at once. Their ablation
reports lower generation error (Fréchet distance) per modality against a flow-matching baseline, and —
more interestingly — faster and higher-climbing success on downstream robot-control tasks when the
backbone is finetuned:
Read that right panel carefully, because it is the real bet: the same weights that generate video are a
**dynamics-aware prior** for physical control. That is the bridge from "content tool" to "world model,"
and it is the part most worth being skeptical of until there is a paper.
## What it actually does: video, with sound
The headline capability in Early Access is video. FLUX 3 generates up to **20 seconds in a single pass**,
and — the part competitors mostly don't have — every clip comes with **native audio**, generated jointly
rather than dubbed on afterward. The capability list is broad: text-to-video, image-to-video (animate a
still or use images as visual references), video-to-video (carry a character from a reference clip into a
new scene), keyframe-to-video for controlled transitions, multilingual dialogue, and *agentic chaining*
of clips into multi-shot sequences minutes long with consistent characters. Here is a representative shot
from BFL's own reel — a single continuous take of a galloping horse, the kind of coherent physical motion
the "world model" framing is really about:
The clip is muted and trimmed to keep the page light; the point it carries is temporal coherence — the
horse's gait, mane, and the blown debris stay physically consistent across the shot. BFL's full reel runs
image, video, and audio together; native sound is one of the model's stronger claims.
## The numbers — and exactly what they are
For the preliminary evaluation, BFL generated 10-second text-to-video clips at 720p with audio and ran
pairwise human-preference comparisons against a spread of current video models. These are the results
they published — read as "share of comparisons where a rater preferred FLUX 3 over the named model":
Hold these loosely. **50% is a tie**, not a loss — so the 52% against Seedance 2.0 and Gemini Omni Flash
is essentially even, while the 93% against Luma Ray 3.2 is a rout. The two highlighted bars (Runway,
Luma) are the comparisons BFL leads with in its post; I highlighted *their* emphasis, not mine. Every
number here is **vendor-reported**, from BFL's own harness, on short 720p clips, for a model BFL calls
"still in development" — there is no independent third-party evaluation yet, and the Grok figure is quoted
by BFL as "up to 69%." This is a preview signal, not a settled ranking.
## Image and action
Image generation and editing are coming a little behind video (Early Access "in the following weeks").
BFL says even mid-training FLUX 3 is a clear step over earlier FLUX on complex-prompt handling and
high-accuracy multilingual **text rendering** — the two things that have defined the modern image race.
The sample grid spans photographic, product, painterly, and graphic styles:
The most unusual branch is **action**. FLUX 3's world understanding is meant to extend to predicting what
happens next — and BFL takes two routes to it: native action prediction folded into the model, and using
the pretrained video backbone as a dynamics-aware foundation that specialized robot-control models finetune
from with little task-specific data. The first partner is [mimic robotics](https://bfl.ai/blog/flux-3),
with whom BFL built **FLUX-mimic**, a video-action model for dexterous manipulation reportedly tested on
production tasks at Audi. The bet, again, is that content creation and physical AI run on the *same*
foundation — the claim the right-hand panel of the Self-Flow chart is quietly staking out.
## The variants, and what "open" means this time
Everything ships from one underlying multimodal flow-matching model, rolled out in phases behind
Early Access gates for safety testing:
| Model | Covers | Access | Status (Jul 2026) |
|---|---|---|---|
| **FLUX 3 Video** | video + audio, gen & edit | API + private weights | Early Access (now) |
| **FLUX-mimic / FLUX 3 Action** | action prediction | research & commercial partners | rolling out (mimic robotics) |
| **FLUX 3 Image** | image, gen & edit | API + private weights | "in the following weeks" |
| **FLUX 3 Dev** | image + video + audio + action backbone | **open weights** | promised, not yet released |
That last row is the one to watch. FLUX.1 and FLUX.2 earned their standing partly because BFL shipped
open `dev` weights the community could actually run; FLUX 3 Dev promises the same for a *multimodal*
backbone — but as of the announcement it is a promise, and the near-term reality is API and private-weight
access. "Open" is on the roadmap, not on the table yet.
## The take
FLUX 3 is the most ambitious repositioning in the open-ish image world this year: from a best-in-class
image model to a single flow-matching backbone that treats image, video, audio, and action as one
learning problem. The mechanism is sound and well-motivated — joint attention over a unified token sequence
is the honest way to let modalities constrain each other, and rectified flow is what makes generating 20
seconds of it tractable. The evidence is thinner than the ambition: preliminary vendor evals on short
clips, a Self-Flow ablation without a paper behind it, an open release that is still a promise, and the
boldest claim — that a video generator is also a robot-control prior — resting on one chart. If it holds
up, "image model" will look like a strangely narrow way to have described what BFL was building. Worth
watching the Dev weights and the tech report; until then, admire the direction and keep the caveats.
---
*Source: [FLUX 3 — Real World Models](https://bfl.ai/blog/flux-3) (Black Forest Labs, 2026-07-23) and
[Self-Flow](https://bfl.ai/research/self-flow). Architecture, Self-Flow, sample, and benchmark figures are
BFL's own, shown for commentary; all evaluations are vendor-reported and preliminary. The flow-matching
and joint-attention interactives are mine. Related: [Mage-Flow](/articles/mage-flow) on the MMDiT backbone
and few-step Turbo sampling.*
---
# GEPA: optimize anything you can score and describe
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/gepa-optimize-anything
> date: 2026-07-24
> tags: optimization, prompt-optimization, evolutionary, agents, explainer
Most ways to make an LLM system better are expensive. Reinforcement learning like GRPO needs
thousands of rollouts to squeeze a scalar reward into the weights; hand-tuning a prompt needs a human in
the loop. **GEPA** (Genetic-Pareto) makes a different bet: language is a *richer* learning signal than a
number. Instead of a policy gradient from a sparse reward, GEPA reads the system's own execution traces,
**reflects on them in natural language** to work out what went wrong, and writes a better prompt — keeping
a **Pareto frontier** of candidates so specialist wins are never averaged away. The result, per its paper,
is a prompt optimizer that turns *a few* rollouts into large gains.
That core has now been generalized twice. `optimize_anything` drops the "prompt" assumption entirely — it
optimizes any text artifact you can **score and describe** — and the new **omni** release composes GEPA
with agent-based optimizers into a single meta-optimizer. This is a walk through the mechanism, then what
each layer adds, honestly labelled.
Two sources, two sets of numbers. The GEPA mechanism and the RL comparison are from the paper *GEPA:
Reflective Prompt Evolution Can Outperform Reinforcement Learning* (Agrawal et al., arXiv 2507.19457).
The `optimize_anything` / omni interface and the Frontier-CS results are from the GEPA team's July 2026
blog post. Every number below is **author-reported** on the authors' own tasks and harness; I have not
re-run them. The interactive diagrams are illustrations of the mechanism, not measured traces.
## The core: reflect, don't reward
GEPA treats a prompt (or a whole multi-prompt system) as the thing to evolve. One turn of its loop is
small and legible — sample a candidate, run it, read what happened, and let an LLM rewrite it:
The move that makes this work is stage 4. A GRPO update collapses an entire trajectory into one scalar
advantage and nudges billions of weights; almost all of the information in *why* the run failed is thrown
away. GEPA does the opposite: it hands a **reflective LLM call** the full trace — the reasoning, the tool
calls, the tool outputs — together with the evaluator's **textual feedback**, and asks it to diagnose the
failure and propose a fix in words. "You mis-parsed the date format on the third call" is a far denser
lesson than "reward = 0.2", and it applies after a single rollout instead of thousands.
Because the feedback is language, the same loop works on any system with one or more LLM prompts, and the
gains show up fast. Across six tasks the paper reports GEPA **outperforming GRPO by 6% on average and by up
to 20%, while using up to 35x fewer rollouts** — and beating **MIPROv2**, the previous leading prompt
optimizer, **by over 10%** (for example, +12% accuracy on AIME-2025). The claim isn't that reflection is
magic; it's that natural language is a higher-bandwidth channel than a reward scalar when your optimizer is
itself a language model.
GEPA comes out of the DSPy lineage — it ships as a DSPy optimizer, and the baseline it beats, MIPROv2, is
DSPy's earlier prompt optimizer. If you already express your system as a DSPy program, GEPA is a
drop-in optimizer over its prompts; the omni post also references a "DSPy full program evolution" tutorial
that evolves the program itself, not just its instructions.
## Why a Pareto frontier, not the best-so-far
The "Genetic" half of the name is evolution — mutate, evaluate, keep the good ones. The subtle part is
*which* ones you keep. A greedy search keeps the single best-scoring candidate and mutates that. GEPA keeps
the whole **Pareto frontier**: every candidate that is the best found so far on **at least one task
instance**. Toggle between the two policies and watch what greedy throws away:
Averaging is lossy. A prompt that nails a hard sub-case but is middling overall looks worthless to a
greedy optimizer and gets discarded — along with the one trick it had figured out. By keeping the frontier
and **sampling parents from it**, GEPA preserves those specialists so their lessons can be recombined later,
and it keeps the search from collapsing into a single lineage that stalls in a local optimum. It's the same
instinct as maintaining a diverse population in a genetic algorithm, made concrete on per-instance scores.
## Optimize anything: score + describe
Once you notice that nothing in the loop actually requires the artifact to be a *prompt*, the interface
generalizes. `optimize_anything` asks for exactly two things: a **seed** text artifact, and an
**evaluator** that returns a score and some feedback. Everything else — objective, background context — is
plain English the reflective LLM reads.
```python
from gepa.optimize_anything import optimize_anything, OptimizeAnythingConfig
def evaluate(candidate: str) -> tuple[float, dict]:
score, feedback = run_judge(candidate) # your metric + any text you want the LLM to see
return score, {"Feedback": feedback}
seed = open("seed_solution.py").read()
task = dict(
evaluator=evaluate,
objective="Maximize the score for this competitive-programming problem.",
background="A sandboxed judge runs hidden tests and returns a 0-100 score.",
)
result = optimize_anything(seed, **task, config=OptimizeAnythingConfig(engine="gepa"))
```
The artifact can be a prompt, a code file, a system message, a config, a plan — anything expressible as
text and gradeable by a function. The `feedback_dict` is the whole trick: whatever diagnostic text you put
in it (a failing test's stderr, a judge's rationale, a lint report) becomes the material the reflective LLM
reasons over. Score tells the search *whether* a candidate is better; feedback tells it *why*, which is what
makes the next mutation informed instead of random.
## Omni: no single optimizer wins, so compose them
The July 2026 post starts from an inconvenient observation: on hard, open-ended problems, **no single
optimizer dominates**. GEPA's reflective mutation, an autonomous coding agent, and an agent-based proposer
each win on different problems — and each one eventually **plateaus**. The `engine=` argument turns this
into a lever: the same `optimize_anything` task can be dispatched to any of three families.
| engine | family | how it proposes |
|---|---|---|
| `gepa` | LLM-based optimizer | one reflective LLM call mutates a parent drawn from the Pareto frontier; the framework owns the loop |
| `meta_harness` | agent-based | a coding-agent proposer mutates the candidate; the framework still owns the loop |
| `autoresearch` | autonomous agent | a long-horizon agent session owns the *entire* loop — selection, proposal, and orchestration |
Across ten problems the winner is unpredictable — the paper's per-problem tally is GEPA 3, AutoResearch 3,
Meta-Harness 4 — so betting on one engine is betting wrong 60–70% of the time. Worse, when an engine
plateaus, *seeding a different engine from the stuck candidate usually breaks through*: on one problem GEPA
stalled at 54.4 after about \$1.3 of budget and switching to AutoResearch lifted it to 62.7; on another,
AutoResearch stalled at 50.0 and both other engines climbed from there to a perfect 100.
**omni** turns that into a strategy. It splits a fixed budget in two phases: **explore**, running all three
engines in parallel on a small slice (about \$5 each) and keeping the best candidate; then **continue**,
seeding a *fresh* optimizer instance with that winner and spending the rest (about \$5) to push past the
plateau — all capped at \$20 total per problem.
The composition itself is exposed as a small kit of primitives — `optimize_best_of` (parallel, keep the
top), `optimize_sequential` (chain engines), `optimize_vote` (fair cross-engine comparison), and
`optimize_adaptive_sequential` (auto-switch on plateau detection) — so omni is one policy you can write, not
a hardcoded pipeline.
## Results on Frontier-CS
The benchmark is **Frontier-CS**: ten open-ended competitive-programming problems, a \$20 budget each, using
Claude Sonnet 4.6 with medium thinking. A single zero-shot LLM call averages **7.72**, so there is real
headroom. Under a matched \$20 budget, omni tops every standalone optimizer:
The per-engine story is that omni lifts *every* base optimizer, not just the strongest — the biggest jump is
GEPA's, which nearly doubles once it stops having to break through on its own:
| optimizer | standalone | as omni | lift |
|---|---|---|---|
| GEPA | 43.8 | 61.8 | +18.0 (+41%) |
| AutoResearch | 55.4 | 63.2 | +7.8 (+14%) |
| Meta-Harness | 50.9 | 59.3 | +8.4 (+16%) |
The mechanism behind those lifts is the plateau-break — a stuck candidate handed to a different optimizer
keeps climbing:
Read these as a promising engineering result, not a settled benchmark. Frontier-CS is ten problems on the
authors' own harness with an LLM judge, and the scores are averages over a small set with real variance —
the per-problem winner already swings widely. omni also spends its budget on *three* engines plus a
continue phase, so the fair comparison is the matched \$20 cap, which the post does hold. The headline is
narrow and honest: under that budget, composing beat every single optimizer they tried — not that omni is
optimal.
## The take
GEPA is a clean idea executed in layers. The core is that **language is a denser training signal than a
reward** when the optimizer is an LLM: reflect on the trace, keep a Pareto frontier so specialists survive,
and a few rollouts go a long way — up to 20% over GRPO at up to 35x fewer rollouts, on the paper's tasks.
`optimize_anything` strips away the "prompt" assumption and leaves a genuinely general interface: *any text
artifact you can score and describe* is now optimizable, with the feedback string doing the heavy lifting.
And omni's contribution is an honest one — since no single optimizer wins and each one plateaus, **explore
across engines, then continue from the best**, which on Frontier-CS beat every standalone optimizer at a
matched budget. The caveats are the usual ones for a fresh result: small benchmark, LLM judge, provider
numbers. But the shape is compelling, and because the whole thing is `pip install` and a scoring function,
it's unusually easy for others to check on their own artifacts.
---
*Sources: [GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning](https://arxiv.org/abs/2507.19457)
(Agrawal et al., 2025) for the mechanism and the RL/MIPROv2 comparisons, and the GEPA team's
[optimize_anything goes omni](https://gepa-ai.github.io/gepa/blog/2026/07/22/optimize-anything-omni/) post
(Tan, Agrawal, Lee, Zhang, Klein, Sen, Dimakis, Zaharia, 2026) for the interface and Frontier-CS results.
GEPA is developed in the [open-source repo](https://github.com/gepa-ai/gepa) and integrates with
[DSPy](https://dspy.ai). Figures are reproduced from the post for commentary; the interactive diagrams are
mine. All benchmark numbers are author-reported.*
---
# Lanyon: proving a PDE solver correct before you run it
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/lanyon-neurosymbolic
> date: 2026-07-24
> tags: neurosymbolic, theorem-proving, pde, scientific-computing, verification, explainer
When a frontier model writes a numerical PDE solver, it does two hard things at once: it *derives* the
scheme (the math) and it *implements* the scheme (the code), and nothing forces those two to agree. The
proof — if there even is one — can quietly be about a different object than the code that runs. [Lanyon](https://lanyon.ai/research/linear-benchmarking/)
is a **neurosymbolic** system built to close exactly that gap: it emits a solver and a machine-checkable
proof from *one* domain-specific specification, and a symbolic engine type-checks the proof and compiles
the kernels **before any simulation runs**.
Lanyon has published the first in a promised series of benchmarking posts, comparing itself on **simple
linear PDE solvers** — one-dimensional linear advection and the Maxwell equations — against five frontier
models: **Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, GPT-5.5, and Kimi K3**. The claims are large:
**20–250× fewer output tokens**, wall-clock measured in seconds rather than the frontier models' minutes,
and that — unlike every model tested — Lanyon's proofs respect the IEEE-754 floating-point axioms.
Two things are true at once here, and this piece keeps both in view. The *idea* — proving numerical code
correct against real floating-point semantics, before you trust its output — is genuinely important and
under-served. And the *numbers* are Lanyon's own, from an initial benchmark it designed and graded, on a
deliberately easy slice of the problem space, with no independent replication yet. Read the mechanism as
plausible; read the multipliers as vendor claims.
## The loop: propose, then prove before you run
The architecture is a tight loop between a neural **proposer** and a symbolic **engine**. The proposer
writes a candidate solver together with its formal spec in a domain-specific language (DSL); because the
proof and the code are expanded from the *same* DSL specification, Lanyon's pitch is that the classic
autoformalization failure — proving one thing while implementing another — is designed out rather than
caught after the fact. The engine then asks a question you can answer without a single time-step:
**does the expanded proof type-check, and do the expanded kernels compile?** If not, the specification is
wrong, and the error routes straight back to the proposer.
Drag the stage control and flip the candidate between *clean* and *has bug*. The point of the diagram is
the contrast, not the animation: the symbolic engine is a **gate** that rejects a bad solver before it
ever produces a number, whereas a pure-LLM path writes solver code token-by-token straight to an
unverified output. Lanyon frames this as a *tighter* reinforcement-learning loop than agentic
approaches — it can verify a specification *before* execution, where a frontier agent "can only verify
*post hoc*," by running the code and eyeballing whether the plots look right.
## Why a symbolic engine can be far cheaper
The token-efficiency claim is the one most worth understanding mechanistically, because it is the most
plausible. An LLM that solves a structured math problem end-to-end pays for the *derivation* in tokens:
it expands algebra, tracks indices, and reasons about stability step by step, in natural language and
scratch work, and it re-does that reasoning on every run. A symbolic engine does the exact algebra once,
in a representation built for it, and the neural component only has to *propose the specification* — the
expensive, exact, repeatable computation is offloaded to machinery that does it in closed form and
verifies it, rather than re-deriving it token by token.
For well-posed, structured problems — and a linear PDE solver is about as structured as scientific
computing gets — that division of labor is exactly where a hybrid should win. The honest flip side, which
we return to below, is that it is *also* exactly the regime where the symbolic half has the most to
exploit; the argument gets weaker as problems get less clean.
## Proving code correct under IEEE-754, not the reals
This is the genuinely interesting engineering, and it is subtler than "we use a theorem prover."
Floating-point arithmetic is not real arithmetic. The most important way it differs: addition is
commutative but **not associative** — $(a \oplus b) \oplus c \neq a \oplus (b \oplus c)$ in general,
because each $\oplus$ rounds its result. Distributivity fails for the same reason. So a proof that a
numerical scheme is stable or conservative, if it is carried out with ordinary **real-number** algebra,
can be a proof about a program that *does not exist* — an idealized version of your kernel that never
rounds. The code that actually runs obeys the weaker, messier IEEE-754 algebra.
Lanyon's answer is an internal, **Lisp-based** symbolic theorem-prover (not Lean) that builds on
Gorard and Hakim's 2025 work on formal verification of PDE solvers within finite-precision arithmetic.
Its symbolic layer is constrained to properties that actually hold under IEEE-754 — commutativity yes,
associativity no. For this benchmark, where every system's output is audited as Lean proofs plus C code
so the comparison is apples-to-apples, Lanyon's translation to Lean is deliberately careful: expressions
are **parenthesized** so any algebraic manipulation stays consistent with the IEEE-754 axioms, and where
there is any ambiguity it reaches for `simp only` instead of `simp`, and restricts `ring_nf` and
`field_simp` to cases where a commutative (semi)ring or field structure can be *safely* assumed. Those
Lean tactics quietly assume the real-number identities — associativity, distributivity — that floats
break, so using them freely is how a "proof" drifts away from the code. Lanyon's claim is blunt: **none
of the frontier models follow this more restrictive discipline.**
## What the errors look like
Because Lanyon graded every run against a rubric, the failure taxonomy is concrete. The verdicts:
| Verdict | Meaning |
|---|---|
| **Faithful** | The Lean proof matches the C formulas; the proofs are substantive |
| **Partial** | Honest proofs, but limited to a subset / dead code / a degenerate regime |
| **Misformalized** | Verification leans on escape hatches or invalid (vacuous, true-by-construction) theorems |
| **Disconnected** | The Lean proof is about a different object than the C actually implements |
The specific behaviors Lanyon reports observing in the frontier runs, especially under *terse* prompts:
- **Escape hatches.** `sorry`, `admit`, `native_decide`, vacuous hypotheses, and "true-by-construction"
theorems dressed up as substantive results.
- **Degenerate demos.** Running one-dimensional advection only at `CFL = 1.0`, where the scheme collapses
to an *exact shift* and verification becomes trivial — in one case despite a terse prompt explicitly
asking for a more general solver.
- **Algorithm substitution.** A GPT-5.6 Sol Maxwell run that used a different finite-volume method than
the one requested.
- **Incomplete verification.** Leaving limiters, the two-dimensional extensions, and time-dependent
properties like stability unproven while presenting the result as verified.
- **Timeouts.** Kimi K3 failed to finish two of three detailed Maxwell trials inside a two-hour window.
The through-line Lanyon draws is *misformalization*: "the proof not matching the code is the precise
failure mode run to run of other agents, especially under ambiguous prompts." Notably, under the
**detailed** prompts every model that finished was graded Faithful — the divergence shows up when the
prompt is terse and the model is left to decide what "verified" means.
## The numbers (vendor-reported)
Everything below is Lanyon's own measurement, on a benchmark it authored and graded, over two simple
linear PDE problems with three trials each. Lanyon reports **20–100×** (linear advection) to **50–250×**
(Maxwell) fewer output tokens than the frontier models, wall-clock in **seconds** versus their **minutes**,
and that its cost stays roughly flat from 1D to 2D while the frontier models' rises 1.5–2×. To reduce
self-bias each model reviewed every (anonymized) run — a reasonable control — but the benchmark is
self-selected, the rubric is Lanyon's, and none of it has been independently replicated. Treat the
multipliers as claims, not facts.
The output-token spread is the visual that carries the token-efficiency argument. These are the frontier
models' reported output tokens on the **Maxwell** solver under the detailed prompt; Lanyon reports its own
solver "takes seconds to generate" with orders-of-magnitude fewer tokens (the 50–250× figure), and does
not publish its exact count in the post:
Under the **detailed** prompts, every model that finished produced a faithful proof — the differences are
in cost and wall-clock, not correctness:
**Linear advection — detailed prompt**
| Model | Output tokens | Cost / trial | Wall-clock | Verdict |
|---|---|---|---|---|
| Claude Opus 4.8 | 153k | $7.03 ± 2.21 | 30.5 min | Faithful ×3 |
| Claude Fable 5 | 101k | $7.39 ± 1.17 | 20.9 min | Faithful ×3 |
| GPT-5.6 Sol | 23k | $2.56 ± 2.23 | 8.7 min | Faithful ×3 |
| GPT-5.5 | 22k | $1.39 ± 0.29 | 5.4 min | Faithful ×3 |
| Kimi K3 | not reported | $2.26 ± 0.38 | 37.9 min | Faithful ×3 |
**Maxwell equations — detailed prompt**
| Model | Output tokens | Cost / trial | Wall-clock | Verdict |
|---|---|---|---|---|
| Claude Opus 4.8 | 219k | $13.00 ± 5.16 | 49.0 min | Faithful ×3 |
| Claude Fable 5 | 149k | $11.03 ± 0.63 | 30.9 min | Faithful ×3 |
| GPT-5.6 Sol | 30k | $2.28 ± 0.79 | 12.3 min | Faithful ×3 |
| GPT-5.5 | 35k | $2.63 ± 0.91 | 9.0 min | Faithful ×3 |
| Kimi K3 | not reported | $5.43 ± 3.01 | 92.0 min | Faithful (1/3 finished; 2 DNF) |
The rubric bites under the **terse** prompts, where the model has to decide for itself what a "verified"
solver means. This degradation table is really the substance of Lanyon's correctness claim:
| Model | Advection (terse) | Maxwell (terse) |
|---|---|---|
| Claude Fable 5 | Faithful ×2, Partial ×1 | Misformalized ×1, Partial ×2 |
| Claude Opus 4.8 | Partial ×3 | Partial ×3 |
| GPT-5.6 Sol | Faithful ×2, Partial ×1 | Faithful ×2, Partial ×1 |
| GPT-5.5 | Partial ×2, Misformalized ×1 | Partial ×3 |
| Kimi K3 | Misformalized ×2, Faithful ×1 | Partial ×3 |
## The skeptic's read
Take the mechanism seriously and the numbers skeptically.
- **The domain is the easiest possible.** Simple *linear* PDEs with well-posed, structured solutions are
precisely where a symbolic engine has the most to exploit and an LLM has the least edge. Nonlinear,
stiff, shock-forming, or turbulent problems — where numerical analysis actually gets hard, limiters
matter, and closed-form structure evaporates — are exactly the regime this benchmark does not touch.
Lanyon calls this "the first in a series"; the interesting posts are the later ones.
- **Self-selected and self-graded.** Lanyon chose the problems, wrote the rubric, and defined what
"Faithful" means. The cross-model anonymized review is a real mitigation against self-bias, but it does
not make the benchmark neutral, and a rubric that centers *formal verification discipline* is one Lanyon
is built to win by construction.
- **The "errors" need replication.** The escape-hatch and degenerate-demo findings are specific and
falsifiable — which is good — but they are single-digit trial counts from one evaluator. "Frontier
models game terse prompts" is a claim that should be independently reproduced before it is repeated as
fact, not least because prompt phrasing is doing a lot of work here (the detailed prompts were all
Faithful).
- **Lanyon doesn't show its own homework.** It reports the frontier models' tokens and times but not its
own exact figures, so the headline multipliers are computed against a number ("seconds," "far fewer
tokens") the reader can't inspect.
None of that undermines the core idea. Verifying that a numerical kernel satisfies its spec *under
IEEE-754 semantics* — not under an idealized real-number fiction — is the right thing to want, and doing
it *before* the simulation runs is a real advantage over "run it and see if the plot looks physical." That
instinct is the same one behind verifier-gated systems like [Leanstral](/articles/leanstral-formal-proofs),
which grade with a checker built to reject `sorry` and `native_decide` outright; Lanyon points the same
discipline at floating-point numerical code.
## The take
Lanyon's contribution, stripped of the multipliers, is a stance worth taking seriously: derive the solver
and its proof from one specification, prove the proof against the arithmetic the hardware actually uses,
and reject the program before it ever runs if the two don't line up. That is a cleaner story than any
single benchmark number, and it is the part that would still matter if the numbers were half as large.
The numbers themselves are early, self-reported, and drawn from the friendliest possible corner of
scientific computing. Twenty-to-two-hundred-fifty-fold is the kind of figure that demands independent
replication and harder problems before it means anything durable — and the honest version of the excitement
is not "Lanyon is 250× better," it's "a neurosymbolic system that offloads exact computation and
float-faithful verification to a symbolic engine *should* be dramatically more efficient on structured
math, and here is the first, self-graded evidence that one is." Whether that holds when the PDEs stop being
linear is the whole question — and exactly what the promised follow-up posts have to answer.
---
*Source: Lanyon's [linear-PDE benchmarking write-up](https://lanyon.ai/research/linear-benchmarking/)
(Lanyon, 2026), which builds on Gorard & Hakim (2025) on formal verification of PDE solvers in
finite-precision arithmetic. All benchmark numbers, cost/token figures, error findings, and speed/efficiency
multipliers are Lanyon's own self-reported results on a benchmark it designed and graded; they have not been
independently verified. The interactive diagram is mine.*
---
# Ling-3.0-flash: a 124B open MoE that runs like a 5B and reaches for 1M tokens
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/ling-3-0-flash
> date: 2026-07-24
> tags: llm, mixture-of-experts, linear-attention, hybrid-attention, open-weights, explainer
[Ling-3.0-flash](https://huggingface.co/inclusionAI/Ling-3.0-flash), released 2026-07-23 by **inclusionAI**
(Ant Group's Ling/Bailing MoE family), is a **124B-parameter** open Mixture-of-Experts that activates only
**~5.1B parameters per token** — about **4%**. It ships with a **256K native context** that is designed to
extend to **1M tokens**, a hybrid-reasoning ("thinking") mode, and a claim that lands harder than the spec sheet:
with roughly **1/8 the total** and **1/12 the active** parameters, it *matches or beats inclusionAI's own 1T
flagship* on most of the benchmarks it was launched against.
The interesting part is not the sparsity ratio on its own — it's the **attention stack** that makes a cheap
million-token context tractable. Ling-3.0-flash is built on **hybrid linear attention**: it interleaves **KDA
(Kimi Delta Attention)** — a gated delta-rule linear-attention layer with constant-size memory — with **Gated
MLA** full attention, at a **5:1 ratio**, and feeds each block into a **512-expert MoE** that fires just **8
experts** plus **1 shared expert** per token. This is a first-principles tour of each piece, why it matters, and
where the launch numbers put it.
Read the center stack bottom-to-top: tokens are embedded, pass through **7 groups** of blocks — each group is
**5×** (KDA + MoE) followed by **1×** (Gated MLA + MoE) — then a final norm, a 157k-vocab output projection, and
a multi-token-prediction head. The two right-hand panels expand the attention modules. Take the ideas one at a
time.
## KDA: constant-size memory over a very long context
Ordinary softmax attention keeps a **KV cache** that grows by one entry per token. At a 256K–1M context that cache
*is* the cost: decoding becomes [memory-bound on a cache that scales with sequence length](/articles/how-llm-inference-works),
and it only gets heavier as the context fills.
**KDA** avoids that. It is a **gated delta-rule linear attention**: instead of a growing cache it keeps a
**fixed-size recurrent state** $S_t$ that each token updates in place — it *erases* a little of the old state (a
per-channel gated decay) and *writes* the new key/value association (the delta rule). A compact way to write the
family is
$$
S_t = \mathrm{Diag}(\alpha_t)\, S_{t-1} + \beta_t\, k_t v_t^{\top}, \qquad o_t = S_t\, q_t
$$
where $\alpha_t$ is the gated decay (the erase), $\beta_t\, k_t v_t^{\top}$ is the written delta, and $o_t$ reads
the state with the query. The state $S_t$ is a fixed $d \times d$ matrix — its size does **not** depend on how many
tokens came before. That is exactly the structure the right-hand KDA panel draws: queries and keys go through a
short conv and an L2 norm, values through a conv, and learned $\alpha$/$\beta$ gates (softplus $\varphi$ and sigmoid
$\sigma$) control the decay and write, with a final sigmoid output gate.
Because the state is constant-size, KDA runs in **linear time and constant memory** in the sequence length — so
five of every six layers pay *no* growing-cache cost at all. Only the single Gated-MLA layer per group keeps a
real KV cache, so the model's total long-context memory grows at roughly **1/6 the slope** of an all-attention
model. Drag the context length and watch the gap open up:
This is the lever behind "1M-token context" being a design target rather than a marketing number. It is not free —
a linear-attention state is a **lossy summary**, not a perfect record — which is exactly why Ling keeps full
attention in the mix.
## Gated MLA: one full-attention layer per group for exact recall
Every group's sixth block is a **Gated MLA** layer — **Multi-head Latent Attention** with **RoPE** and a learned
sigmoid gate. MLA compresses keys and values into a low-rank latent before attending, which shrinks the KV cache
of the full-attention layers themselves; RoPE gives the positional structure that makes long-range *exact* recall
work. The left panel's callout says it plainly: the **1M-token context** rides on RoPE-equipped MLA, while KDA
carries the **linear time complexity**.
The division of labour is the whole point. KDA is cheap but forgetful; full attention is exact but expensive. A
**5:1** interleave keeps one exact-recall layer in every group so the lossy linear layers have something precise to
anchor to — you get most of full attention's fidelity at a fraction of its memory. This is the same bet Moonshot
made in [Kimi K3](/articles/kimi-k3), and the resemblance is not subtle (more on that below).
## The MoE: 8 of 512, plus a shared expert, kept balanced
Ling-3.0-flash's feed-forward is a **fine-grained MoE**: **512 routed experts**, of which only **8** fire per
token — an activation ratio of **1/64** — plus **1 shared expert** that runs on every token to carry the common,
always-useful computation. That extreme sparsity is what lets a 124B model spend only ~5.1B parameters per token:
the compute is that of a ~5B model while the *knowledge capacity* is that of a 124B one.
At 1/64 activation two problems that are mild in a denser MoE turn first-order. **Routing** has to be learned well
or most of the capacity is wasted, and **load balance** matters even more — if a few experts hog the tokens, the
rest never train and the effective model collapses to something far smaller than 124B. Ling's answer is
**ALF-LB (adaptive load balancing)**: rather than a single brittle auxiliary-loss coefficient, it adapts the
balancing pressure per expert so utilisation stays even without destabilising training — the same *spirit* as K3's
aux-loss-free balancing, aimed at keeping all 512 experts alive.
Two more structural details worth stating:
- **The first 2 blocks use a dense FFN instead of MoE.** Early layers do broad, low-level feature mixing where
routing buys little and can hurt stability, so Ling keeps them dense and only switches to sparse experts deeper
in the stack — a common, load-bearing choice at this sparsity.
- **Multi-Token Prediction (MTP).** The training objective is next-token prediction **plus** an auxiliary
multi-token-prediction head (the node at the top of the stack). MTP densifies the learning signal per step and
doubles as a **self-speculative decoding** draft head at inference, which is part of how a "flash" model earns
the name.
Rounding out the sheet: a **157k-token vocabulary**, an **embedding dimension of 2,560**, and the 7-group hybrid
stack above.
## The Kimi K3 resemblance
If this all sounds familiar, it should. [Kimi K3](/articles/kimi-k3) is built on the same three bets — **KDA**
for constant-size long-context memory, **full attention interleaved** for exact recall, and an **extreme-but-stable
sparse MoE** with a balancing scheme that avoids a brittle aux-loss knob. Ling-3.0-flash runs the same playbook at
a very different scale: **124B/5.1B** for a fast production model versus K3's **2.8T/~50B** frontier system, and
**8-of-512** routing versus K3's 16-of-896. The convergence is the story — two independent open labs arriving at
the *same* architecture for efficient long-context reasoning strongly suggests this hybrid-linear + sparse-MoE
recipe is where open models are settling.
## The benchmarks
On its launch suite, inclusionAI compares the **Ling-3.0-flash(RC3)-Thinking** build against a field of thinking
models: its own 1T **Ring-2.6-1T**, **MiniMax-M2.7**, **Step-3.7-Flash-high**, **Deepseek-v4-flash-max**,
**Nemotron-3-Super-120B**, **GPT-5.4-mini-high**, and **Claude-Sonnet-4.6-maxthink**. The full grid:
The headline result is coding. On **SWE-Bench Pro** the 124B model edges the entire field — including the 1T
sibling and every frontier opponent it was tested against:
Long-context and instruction-following are where the hybrid stack should pay off, and it does. On **MRCR-128k** it
sits second, just behind Claude and comfortably ahead of the 1T Ring and everything else — while GPT-5.4-mini,
MiniMax and Step fall off a cliff:
On **SysBench** (system-prompt adherence) it's in a three-way near-tie at the top with Claude and Deepseek:
It is not a clean sweep. On **Terminal-Bench v2.1-AA** — long-horizon, tool-heavy agent work — Claude-Sonnet-4.6
is well clear and Deepseek leads the open pack; Ling lands mid-field, ahead of its own 1T sibling but not the
frontier:
The pattern is consistent with the architecture: Ling is strongest where **structured recall and instruction
adherence** dominate (SWE-Bench Pro, MRCR-128k, SysBench, IFBench), and merely competitive on the longest-horizon
agent loops where a top proprietary model still pulls ahead. For a 124B model activating ~5.1B parameters, being in
that conversation — and beating a 1T model at 1/12 the active compute — is the result.
**Read these as vendor numbers.** (1) Every score above is **inclusionAI's own launch report** for the
**RC3-Thinking** build, run against opponents at their listed settings (e.g. `Claude-Sonnet-4.6-maxthink`); treat
cross-lab comparisons as directional, not audited. (2) The **1M-token context** is a stated design target extending
a **256K native** window — long-context quality at the far end is not established by MRCR-128k alone. (3)
Architecture details (the 5:1 KDA:MLA interleave, E512A8 + shared expert, ALF-LB, dense first-2-blocks, MTP) are
drawn from inclusionAI's release and community write-ups; exact per-layer counts may differ in the final tech
report. (4) The diagram is a **faithful recreation** of the launch architecture figure in our house style, not the
original image.
## The take
Strip away the "beats a 1T model" headline and what's genuinely useful about Ling-3.0-flash is a **coherent,
reproducible recipe**: **KDA** buys a linear-time, constant-memory path to very long context; a **1-in-6 Gated MLA**
layer buys back the exact recall linear attention loses; an **8-of-512 + shared-expert** MoE with **ALF-LB** buys
124B of capacity at ~5B of active compute and keeps all the experts trained; and **MTP** plus dense early blocks
make the whole thing converge and decode fast. That it is essentially the [Kimi K3](/articles/kimi-k3) architecture
at 1/22 the size is the most telling part — the frontier recipe for efficient long-context reasoning is now open,
and it runs on a single node.
---
*Sources: the [Ling-3.0-flash model card](https://huggingface.co/inclusionAI/Ling-3.0-flash) and inclusionAI's
launch materials (architecture, hybrid KDA/MLA interleave, MoE configuration, benchmarks), the
[Kilo announcement](https://blog.kilo.ai/p/announcing-ling-30-flash-free-on) (124B/5.1B, 256K→1M context), and
inclusionAI's launch benchmark chart (reproduced above). Benchmark numbers are quoted from inclusionAI's own
report for the RC3-Thinking build. The architecture diagram is a house-style recreation of the launch figure; the
KV-vs-state chart is illustrative (order-of-magnitude, to show the shape).*
---
# MAI-Image-2.5-Pro and MAI-Voice-2-Flash: Microsoft builds its own
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mai-image-2-5-voice-2
> date: 2026-07-24
> tags: image-generation, tts, microsoft, multimodal, explainer
For most of the last three years, "Microsoft's AI" mostly meant OpenAI's models wearing a Copilot badge.
[This announcement](https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/) is
the other Microsoft — **MAI**, Mustafa Suleyman's Microsoft AI group — shipping two of its *own* frontier
models into public preview: **MAI-Image-2.5-Pro**, a text-to-image model tuned for quality, and
**MAI-Voice-2-Flash**, a speech model tuned for speed. Neither is a wrapper. Both are trained in-house,
and both are already swapped into products you use.
The headline isn't a leaderboard score — it's a supply-chain move. Microsoft has spent years renting its
image and voice capability from third parties; MAI is now making that capability itself, on its own data,
and serving it into Microsoft's product surface at a large discount. That's the frame worth reading these
two releases through.
## Two tracks, one strategy
The pair is deliberately split by objective. **Pro** is the quality lane — hero imagery, detailed edits,
precise in-image text, priced like a premium model. **Flash** is the throughput lane — fast, cheap speech
for high-volume voice, where responsiveness beats everything. What ties them together is where they land:
each has been dropped into Microsoft products in place of an outside model, and each ships with a
self-reported serving win.
Read the fan the way Microsoft wants you to: these aren't demos looking for a home. Bing Image Creator is
now **100% in-house** on MAI-Image-2.5; PowerPoint's image-to-image runs on it at a claimed **84% lower
GPU cost than GPT-Image-2**; OneDrive made it the default editor. On the voice side, MAI-Voice-2-Flash
powers Dynamics 365 Contact Center at a claimed **89% GPU-cost reduction** and feeds Azure Voice Live.
The numbers are Microsoft's own, but the direction is unambiguous — every one of these was previously a
place a third-party model would have run.
## MAI-Image-2.5-Pro: quality, and its own data
MAI-Image-2.5 is the model line; **Pro** is its high-fidelity tier, the one you reach for when the output
is the deliverable rather than a thumbnail. When the base MAI-Image-2.5 model debuted on
[LMArena](https://lmarena.ai) it landed at **No. 3 for text-to-image and No. 2 for image editing** — a
notch behind OpenAI's image models but, per third-party arena coverage, roughly level with Google's
Nano Banana 2. For a first fully in-house image model, that's a real result.
The sample reel leans hard on the two things generators historically fumble: **product photography** and
**legible in-image text**. Brand lockups, packaging copy, poster typography — the kind of output where a
single wrong glyph gives the game away.
That emphasis shows up in the self-reported Arena breakdown. Against the prior MAI-Image-2, Microsoft
reports a **+75 overall Elo gain**, and the two categories that moved most were exactly the hard ones:
These are *deltas versus the previous generation*, not absolute scores against rivals, and they're
Microsoft's own Arena tallies — read them as "where the team pushed," not as a competitive ranking. The
one claim that is genuinely strategic rather than aesthetic sits in the fine print: MAI says the model is
trained on **"clean, traceable, enterprise-grade data, without distillation from third-party models."**
For an enterprise buyer nervous about provenance and copyright, "we didn't distill someone else's model
and we can trace our data" is a feature, not a footnote — and it's a pointed contrast to the murkier
lineage of much of the field.
Pricing tells you which lane Pro is in: **$5 / 1M text-input tokens**, **$8 / 1M image-input tokens**,
and **$106 / 1M image-output tokens** — priced as a premium generation model, not a commodity one.
## MAI-Voice-2-Flash: the throughput lane
The voice release is smaller in ambition and clearer in purpose. **MAI-Voice-2-Flash** is a distilled,
speed-first sibling of MAI-Voice-2: Microsoft reports it is **2× faster** and **32% cheaper** while
keeping "the natural prosody and high acoustic quality" of the parent. It's priced at **$15 / 1M
characters** — the kind of number that only matters at contact-center volume, which is exactly the
target.
The MAI-Voice line has been a speed story from the start: its first model was pitched on generating a
full minute of audio in under a second on a single GPU. Flash extends that lineage in the direction that
matters for the deployment above — a call-center agent that has to respond *now*, thousands of
conversations in parallel, where a half-second of latency is the difference between natural and robotic.
Pairing "good enough prosody" with "cheap and instant" is the entire product thesis, and it's why the
Dynamics 365 and Azure Voice Live integrations lead the voice half of the announcement rather than a
quality benchmark.
Microsoft frames this as a *family*, not a single model: a **Pro/quality** tier and a **Flash/speed**
tier per modality, so a product team picks the point on the cost–quality curve it needs. That's the same
"pick your lane" packaging the rest of the industry has converged on (Pro vs. Flash, Opus vs. Haiku) —
Microsoft is now doing it with models it owns end-to-end.
## Why in-house, and why now
Strip away the model cards and the strategic logic is a spreadsheet. Every image or utterance Microsoft
generates from a third-party API is marginal cost it doesn't control and margin it doesn't keep. Owning
the model turns that into an internal transfer — and the reported serving wins (**−84%** GPU cost in
PowerPoint, **−89%** in Dynamics 365, **2.5× efficiency** with a **25%** P95-latency cut and a **26%**
higher save rate in OneDrive) are the payoff, measured across products that run at Microsoft scale. At
that volume, a double-digit-percent cost cut on a capability embedded in Office and Azure is a very large
number.
It's also insurance. MAI already builds its own [text models](https://microsoft.ai) and voice models;
adding a competitive image model means Microsoft can staff Copilot, Bing, Office and Azure from its own
frontier lab if it ever needs to — reducing dependence on any single outside provider. Two public-preview
models are a small headline; "Microsoft no longer *has* to rent its image and voice stack" is the actual
one.
Keep the caveats attached. Every number here is **vendor-reported**: the Arena deltas are Microsoft's own
tallies, and the GPU-cost and efficiency figures are Microsoft's internal measurements against its own
baselines, not independently reproduced. There's **no technical report** — no architecture, parameter
count, or training detail was published, only capability claims and prices. Both models are in **public
preview**, which means the quality bar and the pricing can still move.
## The take
MAI-Image-2.5-Pro and MAI-Voice-2-Flash are not the most capable image and voice models in the world, and
Microsoft doesn't claim they are. What they are is *sufficient* — a top-three image model and a fast,
cheap voice model, both good enough to swap into the real products where Microsoft used to pay someone
else. That's the whole move: not winning a leaderboard, but owning the supply chain and pocketing the
GPU-cost delta at Office-and-Azure scale, on data Microsoft says it can trace. It pairs naturally with the
research-lab counterpart from the same company, [Mage-Flow](/articles/mage-flow) — a 4B efficiency bet —
and with [Qwen-Image-3.0](/articles/qwen-image-3), another vendor deciding its image model should be
*useful* infrastructure rather than an art toy. The frontier that's being contested here isn't quality.
It's who owns the model behind the button.
---
*Source: [Introducing MAI-Image-2.5-Pro and MAI-Voice-2-Flash](https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/)
(Microsoft AI, 2026-07). LMArena placements and the "level with Nano Banana 2" comparison are from the
earlier [MAI-Image-2.5 launch](https://microsoft.ai/news/introducing-mai-image-2-5/) and third-party
arena coverage. All benchmark, cost, and efficiency numbers are Microsoft's own; the sample image is the
announcement's, shown for commentary. The interactive is mine.*
---
# SANA-Video 2.0: keeping video attention linear without losing the picture
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/sana-video2
> date: 2026-07-24
> tags: diffusion, video-generation, linear-attention, efficient-inference, nvidia, explainer
Video generation has a scaling problem that images mostly dodge. After a VAE compresses a clip, a single
1080p video still spans *tens of thousands* of latent tokens — and a standard video diffusion transformer
runs full 3D softmax attention over all of them, at every layer, at a cost that grows with the *square* of
the token count. Double the resolution or the duration and the attention bill doesn't double, it
quadruples. That $O(N^2)$ wall is why most open video models cap out at a few seconds and lean on clusters.
[SANA-Video 2.0](https://arxiv.org/abs/2607.21553), from NVIDIA, is an argument that you can walk around
the wall instead of paying to climb it. It generates high-quality video up to 720p **on a single GPU** —
the 5B model renders a 720p, 5-second clip in **13.06 seconds** on one H100, which the paper clocks at
roughly **120× faster** than Wan 2.2-A14B — while claiming quality parity with much larger full-softmax
systems. It does this by being disciplined about where quadratic attention is actually worth its price.
The organizing idea is the same "spend cost where the signal is" instinct that the
[Mage-Flow explainer](/articles/mage-flow) applied to image tokenizers — but pointed at *attention* and
stretched across *time*. To see why it works, start with the family it comes from.
## The SANA lineage: efficiency as a house style
SANA has always been the efficiency line. The original **SANA** image model made three bets that each cut a
different cost: a **deep-compression autoencoder** (DC-AE) that squeezes an image by 32× per side instead of
the usual 8×, so the transformer sees far fewer tokens; a **Linear Diffusion Transformer** that replaces
quadratic self-attention with linear attention; and a **decoder-only text encoder** — a small Gemma LLM in
place of the heavy T5 — for conditioning. Together they let a laptop-class GPU generate 1K and even 4K
images. **SANA-Video 1.0** carried the recipe to video with a **2B, pure-linear** DiT and posted big
speedups at competitive quality.
But pure linear attention pays for its $O(N)$ speed with **expressiveness**. Linear attention compresses all
of the past into a single fixed-size state matrix $S \in \mathbb{R}^{d \times d}$; that state simply cannot
encode every token-to-token interaction, and for video — where precise spatiotemporal correspondence and
fine detail matter — the missing interactions show up as softness and drift. This is the exact tension the
2.0 paper sets out to resolve, and it borrows its fix from a place you might not expect: recent large
language models. Qwen3-Next and Kimi-Linear keep a mostly-linear stack but insert a few **softmax anchors**
— periodic full-attention layers that restore exact interactions at fixed depths — and route information
across depth with **attention residuals**, a combination Kimi K3 runs at trillion-parameter scale.
SANA-Video 2.0 asks whether the same trick unlocks long, high-resolution video. It answers yes.
## Two levers, and why they compound
SANA's speed isn't one trick; it's two multiplicative ones. The first is upstream of the transformer
entirely: a **high-compression VAE** (SANA-Video 2.0 uses LTX-VAE 2.3, with a stride of **8×32×32** —
32× on each spatial axis, 8× in time) turns a clip into a modest sequence of latent tokens before the DiT
ever runs. A 5-second 720p clip that is millions of pixels collapses to on the order of ten thousand tokens.
The second lever is the one this paper is about: given that token sequence, **how does attention cost scale
as the clip grows?** Drag the length and flip the resolution:
The point isn't the exact numbers — it's the *shape*. Because attention is quadratic in the token count, a
full-softmax DiT's cost curls sharply upward as clips lengthen and sharpen; the linear-dominated hybrid's
stays low, so the gap *widens* precisely in the long, high-resolution regime where video is most expensive.
The paper's compiled profiling puts the DiT forward pass at **1.55× faster at 5 seconds rising to 3.2×
faster at 60 seconds** versus a matched full-softmax baseline at 720p, and 2.01× faster even at 1080p/121
frames. Compression gives you few tokens; linear attention makes each token cheap; the two savings multiply.
## The mechanism: 25% softmax, and residuals to carry it
So how much softmax do you actually need? SANA-Video 2.0's answer, established through reduced-resolution
proxy studies, is **one layer in four**. Its backbone is a stack of **Hybrid Linear–Softmax Attention**
layers at a **3:1 ratio**: three gated-linear-attention layers — cheap $O(N)$ mixing — for every one
**gated-softmax anchor** that restores the full-rank interactions the linear layers can't represent. Then a
second mechanism, **Block Attention Residuals (AttnRes)**, groups the layers into blocks and routes each
*completed* block's summary forward into later linear layers, so the anchors' refreshed representations
propagate across depth instead of decaying — worth about a **12% lift in deep-layer effective rank** in the
paper's probes. Flip the regime and toggle the residuals:
Two design choices are worth flagging as honest engineering, not magic. First, SANA-Video 2.0 is trained
**from scratch** as a hybrid — it is not a pretrained softmax model that was later "linearized," a shortcut
that usually leaves quality on the table. Second, the 25% figure is a *measured* trade-off point, not a
round number: fewer anchors and quality slips; more and you're paying for softmax you didn't need. The
paper's architecture diagram lays out both pieces — the hybrid layer, the 8-layer blocks, and the shared-query
router that does the aggregation:
The two models share this design at different sizes: the **5B** is a 32-layer, width-2,560 backbone; the
**14B** is 40 layers at width-4,096 (14.25B parameters). Both operate on LTX-VAE 2.3 latents and draw text
features from **Gemma-2-2B-IT** — the decoder-only text encoder carried straight from the SANA lineage —
through cross-attention at every layer.
## What "Video2" adds: making it a real generator
A cheap backbone is only half a video model; the other half is the training pipeline that teaches it motion
and taste. SANA-Video 2.0 is trained with **flow matching** (the same few-step-friendly objective the
[FLUX 3 explainer](/articles/flux-3) walks through for video), then sharpened in stages: a **Self-Flow**
distillation that compresses the sampler, **Direct Preference Optimization**, and an online
**Reward-Feedback-Learning** RL loop. It generates 480p–720p at 81, 121, or 193 latent frames — multi-second
clips, extendable to 8 seconds after fine-tuning — and, because the whole stack was built to be
hardware-friendly, a final **Sol-Engine** pass (kernel fusion, caching, and sparse attention) squeezes out a
further **3.58×** end-to-end, which is what brings the 5B pipeline to that 13.06s figure. Here is a
representative clip from the project page — a surreal "world in a bottle" bobbing on the ocean, the kind of
shot whose value is in staying *coherent* across time:
The clip is trimmed and recompressed from the project page's 8-second 720p sample to keep the page light;
the source reel runs at full resolution. What it's meant to show is stability over time — the failure mode
pure-linear video models fall into (drift, flicker, softening detail) is exactly what the softmax anchors
are there to prevent.
## The numbers — and what they are
The headline is latency, and it is dramatic. Reading straight off the paper's one-H100, 720p/5s profile and
expressing each baseline as a multiple of SANA 5B's 13.06s, the field looks like this:
The efficiency claim only means something if quality holds, and here the evidence is a **VBench** score of
**84.30** for the 5B at 40 sampling steps — essentially level with the 14B's 84.23 and with the Wan 2.2
quality point the paper marks on its chart — reached in a small fraction of the latency. The paper's own
framing is the honest one: *match* full-softmax quality while keeping linear attention's long-sequence
scaling.
Hold these the right way. The speedups above are **derived from the paper's own** latency table (Figure 1b)
— a single-GPU H100 profile from NVIDIA's harness, on their chosen baselines (including in-house or renamed
systems like "Bernini-R" and "Lance"), with both sides compiled on their best kernels. The quality claim
rests on **VBench**, one automated benchmark that correlates only loosely with human preference; there is no
independent third-party evaluation yet. And the strongest numbers stack two separate wins — the hybrid
*architecture* (the 3.2× DiT-forward gap) **and** the Sol-Engine *systems* pass (a further 3.58×) — so the
"120×" is an end-to-end pipeline figure, not the attention mechanism alone. It's a strong, well-instrumented
result; it is not a settled head-to-head ranking.
## Honest limitations
The ceiling is real: **720p** is the top resolution and clips are **seconds**, not minutes — this is not yet
a long-form or 1080p+ model, and the 720p/8s operating point comes from a small supervised fine-tuning stage
($\sim 10^4$ clips), so the highest-resolution, longest-duration quality is the least battle-tested part.
The 25% ratio is validated at reduced-resolution proxy scale and then trusted at full scale. The VAE that
does so much of the compression work is a **licensed external component** (LTX-VAE 2.3), not SANA's own
DC-AE — worth noting because it means the headline contribution here is squarely the *attention* design, not
the tokenizer. And as always with a fresh tech report, every number is the authors'. None of this undercuts
the core result; it just sizes it.
## The take
SANA-Video 2.0 is a clean, well-argued answer to the question that has quietly bounded open video
generation: *do you have to pay quadratic attention to get softmax-quality video?* The answer is no — keep
three layers in four linear, spend softmax only at periodic anchors, carry the anchors' work forward with
residuals, and feed the whole thing from a high-compression VAE so the token count is modest to begin with.
The savings compound exactly where video is most expensive, which is why the gap grows with length and
resolution rather than shrinking. It pairs naturally with [FLUX 3](/articles/flux-3), which bets on *scale*
and joint multimodality to reach video, and with [Mage-Flow](/articles/mage-flow), which makes the same
co-design argument for images: efficiency is an architecture problem, not only a compute one. Worth the
usual caveats on vendor benchmarks and the 720p/seconds ceiling — but as a demonstration that linear
attention can carry real video without visibly losing the picture, it's the most convincing one so far.
---
*Source: [SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation](https://arxiv.org/abs/2607.21553)
(Chen et al., NVIDIA, 2026), the [project page](https://nvlabs.github.io/Sana/Video2/), and the
[SANA repository](https://github.com/NVlabs/Sana). The teaser and architecture figures and the sample clip
are the authors', shown for commentary; all benchmarks are paper-reported. The attention-scaling and
hybrid-stack interactives are mine. Related: [Mage-Flow](/articles/mage-flow) on efficient tokenizers and
[FLUX 3](/articles/flux-3) on flow-matching video.*
---
# Scaling agentic RL: 365,000 environments behind one contract
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/scaling-agentic-rl
> date: 2026-07-24
> tags: reinforcement-learning, agents, environments, infrastructure, prime-intellect, explainer
The reinforcement-learning recipe for agents is, by now, boring in the good way: put the agent in an environment, let it act, check whether it succeeded, and reward it for succeeding. The hard part was never the algorithm. It is the phrase "check whether it succeeded." At scale you need *hundreds of thousands* of tasks that each come with a sandbox the agent can act in and a **grader that produces a clean, reproducible reward** — and the open-source ecosystem, for all its great datasets, does not ship that. Every SWE benchmark, every terminal corpus, every search eval invents its own harness, its own image conventions, its own grading scripts, and its own failure modes. They do not compose.
[Prime Intellect's post](https://www.primeintellect.ai/blog/scaling-agentic-rl) (Daniel Auras and team, July 2026) is an engineering answer to exactly that: they took **23 agentic tasksets across three domains and put them behind one taskset API** — roughly **365,000 tasks** (~198,000 software-engineering across 20+ languages, ~28,600 terminal, ~137,600 search), each with a prebuilt sandbox image, each grader withheld until scoring, many re-uploaded only after gold-validation. This piece walks the thesis, the design, the catalog, and the caveats.
This is a **company engineering post**. Prime Intellect sells the sandbox/compute platform (Prime Sandboxes, `prime-rl`) that this catalog runs on, so read the framing as a product argument, not a neutral survey. The task **counts are their displayed, shipped figures**; several shrink after validation (the [validation table](#gold-validated-then-re-uploaded) below shows by how much). And there is **no independent benchmark of training quality** here — the post ships a training *config* (GLM-4.5-Air on `scaleswe_v1`, 6 H200 nodes, 2 days) but no accuracy numbers, so treat the value as *infrastructure*, not a SOTA result. What is genuinely useful is the design pattern, which stands on its own.
## The bottleneck is the environment, not the algorithm
Here is the shape of the problem the post is solving. Each upstream taskset made reasonable choices for its *own* harness, and those choices don't compose: SWE-bench applies test patches inside a generated eval script; R2E-Gym bakes tests into the image and compares against expected outputs; every search benchmark invents its own judge. If you want to train **one** agent across all of them, you have to normalize those lifecycles without breaking each taskset's own scoring semantics — because the scoring semantics are the whole point. A reward you can't trust is worse than no reward.
Prime Intellect frames it in one line: *"A taskset row is only useful if it can produce a clean reward signal."* And a "surprising fraction" of open agentic data fails that precondition — broken images, network-dependent tests, expected outputs that drifted, and tasks that score as solved without touching the code at all. Scaling RL, in this telling, is far less about the loss function (see [Ring-Zero](/articles/ring-zero-trillion-scale-rl) and [frontier RL economics](/articles/frontier-rl-cheaper) for the loss/systems side) and far more about **manufacturing verified, reproducible environments in bulk**.
## One environment, three layers
The enabling idea is [verifiers v1](https://www.primeintellect.ai/blog/verifiers-v1), which decomposes an environment into three independent layers:
- **Taskset** — the data and its scoring logic (what problem, what counts as solved). *This post is the taskset layer.*
- **Harness** — how the agent is driven (Codex, their own harness, or yours).
- **Runtime** — where it executes (a Prime sandbox, local Docker, …).
Because those layers are independent, one command can run any taskset in any harness on any runtime:
```bash
# ScaleSWE, in the Codex harness, on Prime Sandboxes:
uv run eval scaleswe-v1 --harness.id codex --harness.runtime.type prime -n 3
```
The taskset-layer packaging format is called **Harbor**: SWE-bench Verified "runs the Harbor Hub packaging against the official instance images," and Terminal-Bench 2 is wrapped "through the same Harbor taskset, so the eval suite and the training corpora share one task format and one scoring contract." That last clause is the whole design in miniature — *the thing you evaluate on and the thing you train on speak the same format and the same scoring contract.*
## A rollout, end to end
Concretely, every taskset exposes the same handful of hooks, and a rollout walks them in the same order. The one beat worth internalizing is the **integrity** move: during a rollout the agent lives inside the *same* sandbox as the grading machinery, so "anything readable in the container is fair game for a reward hack." So the grading material — the test patch, the expected outputs, the grader — is **withheld until scoring**, then restored only to compute the reward. Step through it:
The contract that makes this uniform is small — a typed data schema plus four hooks:
```python
import verifiers.v1 as vf
class MyTaskData(vf.TaskData):
base_commit: str # state the sandbox resets to
test_patch: str # grading material — withheld until scoring
gold_patch: str # reference fix — used only by `validate`
class MyTask(vf.Task[MyTaskData]):
async def setup(self, runtime): ... # prepare the repo in the task's image
async def finalize(self, trace, runtime): ... # capture the agent's diff into the trace
@vf.reward
async def solved(self, runtime) -> float: # restore tests, apply test_patch,
... # run the taskset's own upstream grader
async def validate(self, runtime) -> bool: # gold patch must score 1.0;
... # the no-op (setup-only) run must not
```
Some upstream authors deliberately ship the tests readable — R2E-Gym keeps its grading tests at `/r2e_tests`, Multi-SWE leaves grading scripts and `test.patch` under `/home`. That is fine for an attach-at-eval harness, but not for a *live RL sandbox* under optimization pressure, so Prime Intellect's integrations hide those artifacts and restore them only for scoring.
## Four decisions that make tasksets compose
The integrations keep every taskset's original grading path — upstream log parsers, upstream report generation, upstream test commands — and normalize everything *around* it. Four decisions do the work:
- **One API.** Every taskset loads from a typed config (dataset, split, filters), provisions a sandbox from the task's image, and scores with the taskset's own logic. Swapping splits or adding a `filter_fn` works the same everywhere.
- **One image registry.** Task images live in Prime's own registry, co-located with the sandboxes — **~135,000 prebuilt open-source task images**, which they claim is the largest such catalog hosted by any sandbox provider. The point is operational: no Docker Hub rate limits "when running a thousand concurrent rollouts," and reproducibility from a pinned per-task image rather than a build step that can drift.
- **One integrity standard.** Grading material withheld until scoring, as above — the reward-hack defense.
- **One validation bar.** Before a dataset earns a default slot, it runs gold-patch and no-op validation, and a cleaned version is re-uploaded with exclusions preserved. That is the next section.
## The catalog: 23 tasksets, three domains
Pick a domain to see its tasksets and their (shipped) task counts. The three domains are lopsided — software engineering and search dominate the raw count; terminal is smaller but denser in verified benchmarks.
The domain split of the ~365,000 total:
### Software engineering
Real repositories, real diffs; the reward is whether hidden tests pass after the agent's patch. Counts are the displayed shipped totals; parenthetical notes flag the gold-validated re-upload sizes where they differ.
| Taskset | Tasks | What it is |
|---|---|---|
| SWE-bench Verified | 500 | Human-filtered GitHub issues in major Python repos; the canonical benchmark |
| SWE-bench Multilingual | 300 | The canonical set across C, C++, Go, Java, JS/TS, PHP, Ruby, Rust |
| SWE-bench Pro | 731 | Harder successor; large-scale diffs from license-friendly repos |
| SWE-smith | 83,519 | Bugs *injected* into healthy repos, keeping the tests that catch them (8 languages) |
| R2E-Gym | 4,578 | Executable envs from real commits with synthesized issues (4,522 gold-validated) |
| Multi-SWE | 6,835 | Containerized RL + eval instances across 7 languages (2,232 in the validated RL set) |
| SWE-rebench-V2 | 32,079 | Continuously mined fresh PRs, 20 languages, decontaminated by recency (6,275 verified) |
| Scale-SWE | 17,202 | Python tasks with test patches applied just before eval (from 20,181 raw) |
| SWE-Lego | 15,903 | SWE-bench-style training data at scale; tests applied only at scoring |
| OpenSWE | 36,884 | Tasks paired with per-task eval scripts kept out of the sandbox until scoring |
| Senior SWE-Bench | 50 | Investigation/design tasks from 12 production repos; pytest/vitest + optional LLM rubric |
### Terminal
Give the agent a shell and a goal; a hidden pytest grader checks the end state. Smaller in raw count, but this is where the community-standard evals live.
| Taskset | Tasks | What it is |
|---|---|---|
| TMax | 14,600 | Terminal tasks, each pinned to a prebuilt image; all 14,600 boot-and-setup verified |
| Terminal-Lego | ~13,800 | Docker-verified Terminal-Bench-style tasks built from real StackOverflow issues |
| OpenThoughts-TBLite | 100 | High-signal 100-task terminal-agent benchmark; hidden grader |
| Terminal-Bench 2 | 89 | Community-standard eval; 89 rigorously verified tasks |
### Search
The search tasksets share one design decision: they are **harness-agnostic and tool-free**. The taskset ships questions and scoring *only* — the harness brings its own search tool (the Codex harness's built-in web search, Prime's search skill, or yours). The same tasks then train and evaluate any search-capable agent without the environment prescribing a retrieval pipeline.
| Taskset | Tasks | What it is |
|---|---|---|
| PaperSearchQA | 59,907 | Biomedical deep-research QA (54,907 train + 5,000 test); judge-graded |
| WideSeek | 44,632 | WideSearch-style table compilation; scored by item-level cell F1 |
| S1-DeepResearch | ~15,000 | Multi-hop resolution questions with gold answers; judge-graded |
| OpenSeeker | 11,677 | Web-research QA with the original judge prompt |
| DeepDive | 3,250 | Hard multi-hop research (2,234 RL + 1,016 SFT); strict boxed-answer judge |
| BrowseComp | 1,266 | OpenAI's browsing benchmark, in its Explanation/Exact-Answer/Confidence format |
| REDSearcher | 1,000 | Long-horizon web-research questions |
| BrowseComp-Plus | 830 | BrowseComp re-grounded in a fixed 100,195-doc corpus, with a controlled BM25 `search` tool |
BrowseComp-Plus is the one exception to bring-your-own-search: because it serves the benchmark's own BM25 retriever over a fixed corpus, the retriever becomes a *controlled variable* and runs are reproducible — evidence recall is tracked alongside accuracy.
## Gold-validated, then re-uploaded
This is the part that separates the catalog from a link farm. For each dataset, Prime Intellect ran the **gold patch through the full scoring path** in fresh sandboxes, **retried failures up to 10×** to separate flaky from deterministically broken, ran **independent second passes** to catch noisy rows, and ran **multiple no-edit passes** to drop tasks that score `1.0` with no fix at all. The two-sided precondition is simple: *gold patch applied → tests pass; no patch → tests fail.* Every dropped row is persisted in the re-upload so you can audit the exclusion.
The shrinkage is not cosmetic — for the noisiest sources, most of the raw rows do not survive:
| Verified re-upload | Raw | Verified | What dropped |
|---|---|---|---|
| R2E-Gym-Subset-Verified | 4,578 | 4,522 | 56 network/timing-sensitive `aiohttp`/`tornado` tests |
| SWE-Lego-Real-Data-Verified | 4,432 | 4,323 | flaky rows, via two independent passes |
| Multi-SWE-RL-Verified | 4,703 | 2,232 | a no-edit filter caught tasks gradeable as solved with zero edits |
| SWE-rebench-V2-Filtered-Verified | 32,079 | 6,275 | wholesale-broken images; inline GitHub issue/PR references scrubbed |
| SWE-Bench-Verified-Quick | 500 | 468 | the slowest examples, for quick online-evals |
Two of these deserve a callout. **SWE-rebench-V2** goes from 32,079 to 6,275 — an 80% cut — and its design goal is worth stealing: it *"continuously mines fresh GitHub PRs into tasks… naturally decontaminated by recency."* If your tasks are always newer than any model's training cutoff, benchmark contamination stops being a worry by construction. And **Multi-SWE**'s no-edit filter is the quiet hero: a task that grades as solved before the agent does anything is pure reward-hack fuel, and it takes a dedicated pass to find them. The same tooling ships publicly — `uv run validate ` is the model-free sibling of `eval`, running the gold check and the setup-only no-op check in independent runtimes.
## Why this matters for RL at scale
Strip the product framing and the reusable lesson is a data-engineering one. RL at scale does not fail on the gradient; it fails on **thousands of tiny reward bugs** — a flaky test, a drifted output, a container that won't boot, a task solvable without work — each of which quietly poisons the learning signal. The contribution here is treating environments as a *manufactured, versioned, validated artifact*: one task format, one scoring contract, prebuilt per-task images for reproducibility, grading hidden until scoring for integrity, and a gold/no-op validation gate before anything is trusted. That is the same discipline data teams already apply to training corpora, finally applied to the *reward* side — which, for agentic RL, is where the actual difficulty lives.
## Where the reward signal still lies
The post is refreshingly candid that this is mitigation, not a solved problem. A reward signal can lie in two directions, and Prime Intellect names both.
**False positives — reward hacks.** As long as grading runs *where the agent lives*, "a policy under RL pressure will eventually find whatever seam is left" — an editable test file, tamperable grading state, an artifact that leaks the answer. Withholding grading material raises the bar significantly but is not a guarantee; the structural fix they put on the roadmap is **grading in isolated sandboxes**, so the environment the agent can touch and the one that scores it are separate.
**False negatives — correct-but-different fixes.** Validation cannot catch this one by construction. Tasks mined from merged PRs inherit that PR's tests, and those tests often assert *implementation details* rather than behavior — an exact error string, a private helper's name, a precise return shape. An agent that fixes the underlying issue a different but equally correct way still fails them, and the reward reads as a false negative. Gold-patch validation is blind to it (the original patch passes its own tests by definition). At RL scale these near-misses are noise that *punishes correct work*. Their mitigation — "Agentic Judging" — is announced but not yet detailed.
## The take
The headline number — 365,000 environments — is the least interesting thing here. The interesting thing is the **contract**: 23 datasets that each shipped their own harness, image conventions, and grader now load through one typed API, run on prebuilt per-task images, hide their grading material until scoring, and pass a gold/no-op validation gate before they are trusted — with the failed rows kept for audit. That is the unglamorous, correct answer to "how do you get reproducible reward at scale," and it is exactly the layer that has been missing while everyone argued about losses. Take the counts as shipped figures and the training value as unbenchmarked; take the design pattern as the real deliverable. If agentic RL is bottlenecked on verified environments — and the evidence says it is — then a validated, versioned, one-contract catalog is a more load-bearing contribution than another clever objective.
---
*Built on Prime Intellect's [Scaling Agentic RL: 365,000+ Environments for SWE, Terminal, and Search](https://www.primeintellect.ai/blog/scaling-agentic-rl) (Daniel Auras and the Prime Intellect Team, July 2026), with the taskset details drawn from the post and the [research-environments](https://github.com/PrimeIntellect-ai/research-environments) and [verifiers](https://github.com/PrimeIntellect-ai/verifiers) repos it links. All task counts and validation figures are Prime Intellect's own reported numbers. The two interactive diagrams are my redrawings of the mechanism (the rollout pipeline and the taskset map), not reproductions of the post's charts; the hero image is the post's own cover art. There is no independent benchmark of training outcomes in the source, and I have not run one.*
---
# Solar Open 2: Upstage's 250B-A15B hybrid-attention MoE
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/solar-open2-250b
> date: 2026-07-24
> tags: llm, open-weights, upstage, moe, long-context
**Solar Open 2** is Upstage's newest open-weight model: a **250B-parameter Mixture-of-Experts** that activates **15B per token**, ships on Hugging Face, and serves a **1M-token context**. Upstage frames it as an *agentic specialist* — built for tool calling, multi-step reasoning, and document-heavy officework — and as a sovereign-AI play, strong in Korean and Japanese as well as English. What makes it worth a close read is not the parameter count but the **architecture**: instead of a conventional softmax-attention transformer, Solar Open 2 runs a **hybrid stack** that interleaves one softmax-attention layer with three linear-attention layers, and drops positional encoding entirely.
I read the [model card](https://huggingface.co/upstage/Solar-Open2-250B) and the accompanying [technical report](https://huggingface.co/upstage/Solar-Open2-250B/blob/main/Solar_Open_2_Tech_Report.pdf) for this. Two framing notes before any benchmark chart: every number below is **self-reported** by Upstage on its own harness, and the comparison set (DeepSeek-V4-Flash, MiMo-V2.5, Command A+, and others) is Upstage's chosen field. I keep both caveats attached throughout.
All benchmark numbers here are **Upstage-reported** — its own eval harness, its own choice of comparison models and settings. Solar Open 2 tops **no** row against the strongest open model in its bracket: on most English benchmarks **DeepSeek-V4-Flash** (284B-A13B) leads and Solar Open 2 comes second. Read it as a **strong model for its 15B active size** and a genuinely interesting *architecture*, not a leaderboard winner. The daggered Korean rows (`Ko-AIME'25`, `KBank-MMLU`, `Ko-GDPval`) are Upstage **in-house** benchmarks; treat those as least comparable across vendors.
## The Solar name, and what this is not
If you know the **Solar** name, you probably know it for **depth up-scaling** — the 2023 trick behind Upstage's original SOLAR 10.7B, where you grow a model by duplicating and stacking layers from a smaller base rather than training a new shape from scratch. It's reasonable to expect the lineage to continue. It doesn't. **Solar Open 2's card describes no depth up-scaling.** What it describes instead is a **selective weight transfer**: the model is initialized from its predecessor, **Solar Open 1** (102B-A12B), but *"only the 2.3% of weights that survive the architectural change are carried over, and everything else is randomly initialized."*
That 2.3% is the honest headline of the training story. The architectural change is drastic enough — a new hybrid attention stack, no positional encoding, an expert pool grown from 128 to 320 — that almost none of the old model's weights fit the new shape. What transfers cleanly are the parts the two generations share: the token embeddings and output layer (same 196,608-token tokenizer), and the fragments of attention and MoE that survive. Everything else starts from noise. So this is not a re-skin of Solar Open 1 and not a depth-up-scaled Solar 10.7B — it is a **mostly-fresh 250B model** that borrows a running start.
## The hybrid attention stack
Here is the distinctive move, and the reason the 1M context is affordable. Solar Open 2 has **48 layers**, arranged as **twelve identical blocks**, each block being **one softmax-attention layer followed by three linear-attention layers** — the pattern the card writes as `[Softmax, Linear×3] × 12`. So **12 of the 48 layers are softmax; 36 are linear.**
The reason this matters is memory. A **softmax** layer keeps a **KV cache** that grows linearly with the sequence — every past token's keys and values must be stored so the next token can attend to them. A **linear-attention** layer does not: it folds the entire past into a **fixed-size recurrent state**, so its memory is constant no matter how long the context runs. By making three of every four layers linear, Solar Open 2 keeps a growing KV cache on only **12 of 48 layers** — *"holding long-context memory to roughly a quarter of an all-softmax model of the same shape,"* per the card.
Toggle between the hybrid stack and an all-softmax baseline, and drag the context length to watch the KV-cache footprint each one carries:
The KV-cache arithmetic is worth doing by hand, because it is the whole efficiency argument. Each softmax layer is **grouped-query attention** with **8 KV heads** and `head_dim` 128; storing K and V in fp16 costs, per token per softmax layer:
$$
2 \times n_{kv} \times d_{head} \times b = 2 \times 8 \times 128 \times 2 = 4096 \text{ bytes}
$$
Multiply by the number of KV-bearing layers and the context length. At **1M tokens**:
- **Solar Open 2 (12 softmax layers)** → about **48 GiB** of KV cache.
- **An all-softmax stack (48 layers)** → about **192 GiB** — exactly 4× more.
That 4× is the margin that turns a 1M-token window from "possible on a rack" into "fits alongside the weights." If linear attention is new to you, I built the mechanism up in [how transformers attention works](/articles/how-transformers-attention-works) and the sparse-attention variants in [MiniMax sparse attention](/articles/minimax-sparse-attention); the KV-cache side of the story is [how LLM inference works](/articles/how-llm-inference-works), and squeezing the cache further is [TurboQuant](/articles/turboquant-kv-cache).
Three details make the linear layers actually work at this depth, and the card is specific about all three:
- **NoPE — no positional encoding.** Because the linear layers *"encode token order intrinsically in their recurrent state,"* Upstage removes rotary encoding entirely. The upside the card claims: no RoPE extrapolation limit, so the trained window is not tied to a length distribution seen during training.
- **KDA with negative eigenvalues.** The linear layers use **Kimi Delta Attention** (the Kimi Linear lineage — see [Kimi K3](/articles/kimi-k3)), but with `allow_neg_eigval=True`, widening the state-transition write strength to $\beta = 2\sigma(\cdot) \in (0, 2)$. Standard linear cores restrict eigenvalues to $[0,1]$ (decay or persist only); allowing the sign to flip restores the ability to *erase* and self-correct — the card ties this to genuine state-tracking (parity, modular counting).
- **A sigmoid output gate on the softmax layers**, which the card says suppresses the "attention sink" pathology and improves long-context extrapolation.
One ordering detail separates it from its cousins: within each block the **softmax layer comes first** (`S-L-L-L`), unlike the linear-first ordering (`L-L-L-S`) of Kimi Linear and Qwen3.5. Upstage's own architecture figure lays the whole thing out — the 12× block on the left, and insets for the MoE, the GQA softmax layer, and the KDA linear layer, color-coded by which weights transferred from Solar Open 1:
## Where the parameters live
The sparsity is the economic argument, so account for it. Solar Open 2 is **250B total, 15B active** — a **6% activation rate**. Each MoE block holds **321 experts: 320 routed plus 1 shared**; the router keeps the **top-8 routed** experts per token, and the shared expert always runs, so 9 experts fire per token. There are **no dense layers** — every block is MoE. The backbone it inherits from Solar Open 1 is **48 layers, hidden size 4096, head dim 128, 64 query / 8 KV heads**, and the **196,608-token** vocabulary.
If MoE routing is unfamiliar, I built the router, the top-k gate, and the sparsity argument from nothing in [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch) — the same machinery that lets a 250B model serve at roughly the cost of a 15B dense one.
## Training
Upstage is thinner on the data story than on the architecture, and I won't pad it. The concrete figures: **~12 trillion pre-training tokens**, on **NVIDIA B200** GPUs, for **2M GPU-hours**, initialized by the 2.3% selective transfer above. The technical report adds that the run *"maximizes value per token over a globally deduplicated corpus"* and trains on *"purpose-built long-horizon agent scenarios"* spanning conversational tool use, coding, and officework — the agentic focus is baked into the data, not just the eval suite.
The tokenizer is the other inherited advantage. Solar Open 2 reuses Solar Open 1's **Korean-efficient** byte-level BPE tokenizer unchanged; the report claims global-model tokenizers spend **1.2–1.9× more tokens** on the same Korean text. In long agent trajectories, where the working context accumulates over many turns, fewer tokens per unit of Korean text translates directly into lower inference cost and a longer effective window — a real, if narrow, edge.
## The benchmarks
Upstage's headline figure bundles six benchmarks across knowledge/reasoning and agentic/professional work, with Solar Open 2 (dark violet) against Solar Open 100B and the sub-320B open field:
The clearest *win* is agentic tool-use in the **APEX-Agents** suite, where Solar Open 2 roughly doubles the rest of its bracket — a result consistent with the agent-scenario training:
On coding it edges the field on **LiveCodeBench v6** (92.4 vs DeepSeek-V4-Flash's 92.3) but sits behind on the harder **SWE-Bench Verified** agentic coding task — the "competitive, not leading" pattern in miniature:
The full English suite, with the best value in each row bolded (Upstage's own marking):
| Benchmark | Solar Open 2 250B-A15B | Solar Open 100B 102B-A12B | Command A+ 218B-A25B | Mistral Medium 3.5 128B | MiMo-V2.5 310B-A15B | DeepSeek-V4-Flash 284B-A13B |
|---|--:|--:|--:|--:|--:|--:|
| MMLU-Pro | **86.2** | 80.4 | 79.0 | 81.2 | 84.6 | 85.9 |
| GPQA-Diamond | 86.3 | 66.2 | 75.6 | 77.5 | 83.0 | **88.9** |
| HLE (no tools) | 28.8 | 11.5 | 11.4 | 12.8 | 24.3 | **32.3** |
| LiveCodeBench v6 | **92.4** | 56.5 | 86.1 | 84.9 | 89.1 | 92.3 |
| ArtifactsBench | 55.9 | 43.4 | 42.8 | 49.8 | 59.3 | **61.0** |
| HMMT 2026 | 93.9 | 68.9 | 73.5 | 62.9 | 61.4 | **94.7** |
| AIME 2026 | 95.7 | 87.7 | 96.0 | 89.0 | 92.3 | **97.0** |
| Multi-Challenge | 61.0 | 40.5 | 45.8 | 49.8 | 39.0 | **62.0** |
| IFBench | 80.0 | 57.7 | 73.9 | 69.0 | 67.1 | **80.3** |
| AA-LCR | 62.3 | 36.0 | 46.0 | 61.0 | 62.7 | **63.7** |
| SWE-Bench Verified | 70.4 | 15.4 | 14.4 | 69.6 | 73.0 | **73.8** |
| Terminal-Bench Hard | 28.3 | 2.3 | 25.0 | 33.3 | **41.7** | 34.1 |
| APEX-Agents | **16.6** | 2.4 | 1.6 | 6.1 | 13.4 | 13.2 |
| MCP-Atlas | 58.2 | 34.4 | 27.2 | 30.7 | **63.9** | 58.2 |
| τ³ (banking) | 19.6 | 7.4 | 5.8 | 5.8 | 8.7 | **22.3** |
| GDPval-AA v2 (ELO) | 1128 | – | 712 | 929 | 1145 | **1187** |
The shape is consistent: Solar Open 2 leads on **MMLU-Pro, LiveCodeBench, and APEX-Agents**, and is otherwise a close **second to DeepSeek-V4-Flash** — which, at 284B-A13B, is a comparable-scale sparse model. Against the smaller **Solar Open 100B**, the jump is large and uniform (SWE-Bench Verified 15.4 → 70.4, APEX-Agents 2.4 → 16.6), which is the more meaningful comparison since it isolates a generation of progress on one team's harness.
Where Solar Open 2 actually **leads** is Korean — unsurprising given the tokenizer and data focus. It tops **CLIcK, HAE-RAE, KBank-MMLU, KBL, and Ko-GDPval**, beating even the closed **GPT-5.4 mini** and **Claude Haiku 4.5** on several:
| Benchmark | Solar Open 2 | Solar Open 100B | MiMo-V2.5 | DeepSeek-V4-Flash | Claude Haiku 4.5 | GPT-5.4 mini |
|---|--:|--:|--:|--:|--:|--:|
| KMMLU-Pro | 78.4 | 64.0 | 69.1 | **78.9** | 67.9 | 78.1 |
| CLIcK | **90.7** | 78.9 | 78.4 | 89.2 | 53.5 | 89.6 |
| HAE-RAE v1.1 | **73.8** | 73.3 | 61.7 | 73.1 | 38.5 | 69.4 |
| Ko-AIME'25 † | 97.7 | 80.0 | 88.0 | **98.0** | 81.7 | 90.7 |
| HRM8K | 92.2 | 87.6 | 90.7 | **93.4** | 90.6 | 91.3 |
| KBank-MMLU † | **80.8** | 65.5 | 71.0 | 79.5 | 68.9 | 79.0 |
| KBL | **75.5** | 65.5 | 69.8 | 72.8 | 69.9 | 75.3 |
| KorMedMCQA | 93.0 | 84.4 | 87.7 | 94.1 | 87.0 | **94.2** |
| Ko-GDPval † | **86.8** | 3.4 | 81.0 | 85.0 | 68.3 | 59.4 |
*† Upstage in-house benchmarks — least comparable across vendors.* The report's boldest claim rides on the last row: on Ko-GDPval, a Korean officework-agent benchmark, it says Solar Open 2 *"essentially matches DeepSeek-V4-Pro (1.6T) at less than a sixth of its size."* That is an in-house benchmark and a self-comparison, so weight it accordingly — but the direction (a Korean-specialized 250B beating much larger generalists on Korean agentic work) is plausible and repeated across the daggered rows.
## Running it
The weights are ~250B in bf16, so this is multi-GPU territory: Upstage lists a **minimum of 4× H200** (141 GB) and **recommends 8× H200**. The supported production path is **vLLM** (an Upstage fork), with expert-parallel MoE:
```bash
vllm serve upstage/Solar-Open2-250B \
--served-model-name solar-open2-250b \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--moe-backend triton \
--reasoning-parser solar_open2 \
--tool-call-parser solar_open2 \
--enable-auto-tool-choice
```
Solar Open 2 is a reasoning model with a two-position knob: `reasoning_effort="high"` turns on chain-of-thought (a reasoning block capped at 131,072 tokens), and `reasoning_effort="none"` answers directly. Because the reasoning trace counts against `max_tokens`, Upstage recommends leaving room — up to 256K for the full response in high-effort mode — and preserving prior reasoning traces across turns.
```python
from openai import OpenAI
client = OpenAI(api_key="EMPTY", base_url="http://localhost:8000/v1")
resp = client.chat.completions.create(
model="solar-open2-250b",
messages=[{"role": "user", "content": "Prove that the square root of 2 is irrational."}],
reasoning_effort="high",
temperature=1.0,
top_p=1.0,
max_tokens=131584,
)
print(resp.choices[0].message.reasoning) # reasoning returned separately
print(resp.choices[0].message.content)
```
Two nice touches for agent builders. The same vLLM server exposes **both** an OpenAI-compatible `/v1` endpoint and an **Anthropic-compatible `/v1/messages`** endpoint, so **Claude Code** can point straight at it (`ANTHROPIC_BASE_URL=http://localhost:8000`) with no proxy, and MCP tools reach the model through the standard tool-calling interface. For smaller boxes, **NotaAI** publishes official quantized builds — INT4, NVFP4, and an INT4-GlobalPruned variant; the [NVFP4](/articles/nemotron-nvfp4) format is the same 4-bit floating layout NVIDIA pushes for Blackwell.
## License: open weights, with a name tax
Solar Open 2 is **open-weight, not fully open-source**. It ships under the **Upstage Solar License**, and the derivative terms are specific: any model you create, fine-tune, or distill from it must **prefix its name with "Solar"** (e.g. `Solar-MyModel-v1`), **prominently display "Built with Solar"** in public materials, and **include a copy of the license**. That is looser than a research-only license — commercial use and derivatives are allowed — but it is a **branded** license, not Apache-2.0. If you plan to build on it, the naming and attribution requirements are load-bearing, not boilerplate.
## The take
Solar Open 2's real interest is **architectural**, not positional. It is a clean, well-documented instance of the **hybrid linear/softmax** direction — three linear-attention layers for every softmax one, no positional encoding, KDA with negative eigenvalues — that makes a **1M-token context** affordable by keeping a KV cache on only a quarter of its layers. Paired with a 6%-activation MoE and a Korean-efficient tokenizer, that is a coherent systems story aimed squarely at long-horizon agents, and the honest, unusual init note (2.3% of weights transferred, the rest random) is a refreshing departure from the Solar brand's depth-up-scaling past.
The caveats are the standard open-weights ones, stated plainly. Every number is Upstage's own harness against a field it chose, and on that field Solar Open 2 is **a consistent second to DeepSeek-V4-Flash** on English work — leading its bracket on a few benchmarks (MMLU-Pro, LiveCodeBench, APEX-Agents) but not the pack. Its clearest edge is **Korean**, much of it measured on **in-house** benchmarks. And the license carries a name-and-attribution tax that Apache-2.0 models don't. For a team that wants an **open, agent-capable, genuinely long-context** model — especially one working in Korean — and can run 4–8× H200, Solar Open 2 earns a serious look. As the strongest open model at 250B, that title still belongs to the model it keeps finishing behind.
---
*Built from the [Solar Open 2 model card](https://huggingface.co/upstage/Solar-Open2-250B) and [technical report](https://huggingface.co/upstage/Solar-Open2-250B/blob/main/Solar_Open_2_Tech_Report.pdf) (250B-A15B, hybrid attention, 1M context, Upstage Solar License). All benchmark numbers are Upstage-reported; the two figures are reproduced from Upstage's model card and technical report for commentary. The interactive stack diagram is an illustration of the mechanism — the layer pattern, the KV-cache-bearing layers, and the fp16 KV-cache arithmetic use the published config; the linear-attention state memory is a small constant left out of the readout for clarity.*
---
# Ternary15M: a language model where every weight is −1, 0, or +1
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/ternary15m
> date: 2026-07-24
> tags: quantization, ternary, bitnet, efficiency, from-scratch, explainer
Most quantization is an afterthought: train a model in FP16, then squeeze the weights down to 8 or 4 bits for
deployment and hope the accuracy survives. [Ternary15M](https://github.com/brianbell-x/ternary15M) does the opposite,
and takes it to the extreme. It's a 15.19M-parameter, Llama-style language model where **every linear layer is
ternary** — each weight is one of exactly three values, `−1`, `0`, or `+1`, scaled by a single per-channel number.
The model is *trained that way from scratch*, not clipped down after the fact. It's a tiny, readable member of the
[BitNet b1.58](https://arxiv.org/abs/2402.17764) family, and it's a clean place to see why three-valued weights are
such an interesting bet.
The architecture is deliberately small and standard: dim 288, 6 layers, 6 query / 6 KV heads, a SwiGLU MLP with
hidden size 768, a 32k vocab, and a 256-token context. All 42 of its attention and feed-forward linear layers
(6 layers × 7 projections) are ternary; only the embedding table and the RMSNorm gains stay in full precision.
## The whole trick is the sign
Here's why anyone cares about `{−1, 0, +1}` specifically. A neural network is, underneath, a pile of dot products:
each output is a weighted sum of its inputs. In full precision, every one of those weights is an arbitrary float, so
every term in the sum is a hardware **multiply**. Multiplies are the expensive part — they dominate the energy and the
silicon area of a matmul.
Now make each weight a sign instead of a number. If the weight is `+1`, the term is just the input — **add** it. If
it's `−1`, **subtract** the input. If it's `0`, the input contributes nothing — **skip** it. The multiplies vanish; the
inner loop of the dot product becomes signed accumulation over the inputs. Toggle between the two below and watch the
multiply count collapse:
The single multiply that survives is the per-output-channel **scale** — one number `s` that rescales the whole
accumulated sum back to a sensible magnitude. So a ternary matmul is: signed adds across the row, then one multiply at
the end. On general-purpose GPUs the win is mostly memory-bandwidth (you move far fewer weight bytes); on hardware
designed for it, killing the multiplies is the point. Either way, the arithmetic is genuinely different from
"small floats."
## 1.58 bits, and where the bytes actually go
Three states carry log₂(3) ≈ **1.58 bits** of information each — that's where the "b1.58" name comes from, and why a
ternary weight is often quoted as costing ~1.58 bits versus 16 or 32 for a float. Ternary15M doesn't bit-pack that
tightly; its exported checkpoint stores each ternary weight as an `int8` (one byte) plus one FP32 scale per output
channel, which the author notes compresses to roughly 2 bits with basic entropy coding. The latent training checkpoint
is 182 MB; the deployed ternary model is **43 MB**.
But 43 MB is bigger than 1.58 bits × 15M would suggest, and the reason is worth sitting with:
At 15M parameters, the model is mostly its embedding table. The tied vocab embedding is 32,000 × 288 ≈ **9.2M
parameters** — over 60% of the model — and it stays FP32, so it alone is ~37 MB of the 43 MB file. The ternary trick
only compresses the ~6M weights in the 42 linear layers. The lesson generalizes: at small scale, quantizing the matmuls
buys you less than you'd hope because the un-quantized embedding dominates the footprint. Ternary pays off hardest on
*deep* models, where the linear layers, not the vocabulary, are the bulk of the weights.
## Teaching a network to live with three values
You can't train ternary weights directly, because rounding to `{−1, 0, +1}` has a gradient of zero almost everywhere —
nudging a latent weight from 0.31 to 0.32 doesn't change the rounded output, so ordinary backprop would see no signal
and learn nothing. Quantization-aware training gets around this with two ideas working together.
First, the **scale**. Each output channel's weights are ternarized around their own average magnitude — the mean of the
absolute weights in that row, `absmean(W)`. Dividing by that scale before rounding is what decides which weights round
to `±1` and which collapse to `0`, and multiplying the scale back afterward keeps the layer's outputs at roughly the
right size. It's computed per output channel, so every neuron sets its own threshold.
Second, the **straight-through estimator (STE)**. The forward pass uses the *ternarized* weights — so the network
actually experiences quantization while it learns — but the backward pass pretends the rounding was the identity
function and passes the gradient straight through to a full-precision **latent** copy of the weights. Those FP32 latents
are what the optimizer updates; the ternary weights are re-derived from them every forward pass. In PyTorch the whole
thing is a few lines:
```python
def forward(self, x: torch.Tensor) -> torch.Tensor:
weight = self.weight # FP32 latent weights
scale = weight.abs().mean(dim=1, keepdim=True) # absmean per output channel
safe_scale = scale.clamp_min(torch.finfo(weight.dtype).eps)
qweight = torch.round(torch.clamp(weight / safe_scale, -1, 1)) * scale
weight_ste = weight + (qweight - weight).detach() # STE: value = qweight, grad → weight
return F.linear(x, weight_ste)
```
The `weight + (qweight - weight).detach()` line is the STE in one expression: numerically it equals `qweight` (the term
you subtract is detached from the graph), but its gradient with respect to `weight` is 1, so the optimizer trains the
latent weights as if the quantizer weren't there. At export time the latents are thrown away and the weights are frozen
to `int8` in `{−1, 0, +1}` plus the FP32 scales:
```python
def ternary_components(weight: torch.Tensor) -> tuple[torch.Tensor, torch.Tensor]:
"""Return int8 {-1, 0, 1} weights and FP32 per-output scales."""
scale = weight.detach().float().abs().mean(dim=1, keepdim=True)
safe_scale = scale.clamp_min(torch.finfo(torch.float32).eps)
qweight = torch.round(torch.clamp(weight.detach().float() / safe_scale, -1, 1))
return qweight.to(torch.int8), scale
```
## Does it work?
The honest, useful result is that at this scale the quantization is nearly free. Trained on
[TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) (a synthetic corpus of simple children's stories)
for 655M tokens, the model's validation loss barely moves when you go from the latent weights to the hard-ternary ones:
The bars are almost identical on purpose — that *is* the finding. Freezing the network to pure ternary weights costs
only **+0.01 loss** over the latent model it was trained as. The quantization the model trained under is the same one it
ships with, so there's no distribution shift at export. And the deployed artifact is a fraction of the size:
The whole run cost about **$0.70** — ~50 minutes on a single L40S at ~200k tokens/second. That cheapness is a feature:
it makes ternary QAT something you can actually reproduce and poke at, not a claim you take on faith.
These numbers are the author's own, from a single training run on one small synthetic dataset, and TinyStories loss is
not a general capability benchmark. Treat them as a clean proof-of-concept that ternary-from-scratch *converges* at this
scale — not as evidence about how ternary trades off against full precision on a real, large model. BitNet's own papers
argue the gap stays small up to billions of parameters, but that's a separate, much larger claim than this repo makes.
## What a 15M ternary model can and can't do
It's worth being blunt about the ceiling. TinyStories exists precisely so that tiny models can learn *something*
coherent: the vocabulary and grammar are simple, the stories are short, and 256 tokens of context is plenty. Within that
box, Ternary15M does the job — it generates grammatical, on-topic little stories, and it does so from weights that are
almost entirely signs. That's the point of the artifact.
What it can't do is everything a real LM does. There's no world knowledge, no reasoning, no code, no long context, no
instruction following — 15M parameters and a children's-story corpus don't reach any of that, and ternary quantization
doesn't change the ceiling in either direction. The value here isn't the model's outputs; it's that the *training
recipe* — BitLinear layers, an absmean scale, an STE, and a from-scratch schedule — demonstrably works end to end and
lands within a hundredth of a nat of its full-precision-latent self.
## The take
Ternary15M is a good teaching artifact for a genuinely surprising idea: you can restrict every weight in a network to
one of three values and, if you *train* it that way rather than clipping after the fact, pay almost nothing in loss. The
mechanism is clean — signs replace floats, so dot products become signed accumulation with one scale multiply per
channel; the STE lets gradients flow to a latent copy the quantizer hides; the absmean scale keeps magnitudes sane. The
honest caveats are that this is a 15M model on TinyStories with self-reported numbers, and that at this scale the FP32
embedding table, not the ternary matmuls, dominates the file size. But as a from-scratch, $0.70, reproducible window
into how BitNet-style quantization actually works, it's about as legible as this idea gets.
---
*Source: the [Ternary15M repository](https://github.com/brianbell-x/ternary15M) (Brian Bell, MIT license) — its
`README`, `MODEL_CARD.md`, `RESULTS.md`, and `ternary15m/model.py`. The lineage is
[BitNet b1.58](https://arxiv.org/abs/2402.17764) (Ma et al.). Code snippets are from the repo; the interactive diagram
is mine. The repo ships no figures, so there are none to reproduce here — the numbers above are quoted from its
`RESULTS.md` and model card.*
---
# Nanbeige4.2-3B: looping a small model up to a big one's depth
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nanbeige-4-2-3b
> date: 2026-07-22
> tags: llm, agents, small-models, open-weights, explainer
[Nanbeige4.2-3B](https://huggingface.co/Nanbeige/Nanbeige4.2-3B) is a compact agentic model — **3B
non-embedding parameters** (4B total), Apache-2.0, bilingual EN/ZH — from the Nanbeige LLM Lab at Boss
Zhipin. Its claim is the one every good small model makes: it performs "well beyond its parameter
scale," reporting wins over **Qwen3.5-9B** and **Gemma4-12B** across tool-use, office-agent, code-agent,
and most reasoning benchmarks. The mechanism behind the claim is the interesting part — a **Looped
Transformer** — and with the [technical report](https://huggingface.co/Nanbeige/Nanbeige4.2-3B/blob/main/Nanbeige42_report.pdf)
now out, the config is no longer a mystery: it's a two-pass loop, **pretrained from scratch on 28T
tokens**, then shaped by a four-stage post-training pipeline. Here's the loop, the efficiency story, and
which of the numbers to lean on.
## The loop: depth without parameters
A standard transformer buys reasoning depth by stacking more distinct layers, each with its own
weights. A **looped** transformer buys it by running the *same* layers several times — so effective
compute depth is (physical layers × loops), while the parameter count stays at just the physical
layers. You pay for depth in FLOPs, not weights.
Three design choices in the report make this more than a slogan:
- **Two passes, not more.** The report studies the loop count directly and lands on **2** as the sweet
spot: it keeps roughly **75% of a standard Transformer's token efficiency** while adding real
capacity. More passes buy almost nothing and make training slower and less stable — so the model
loops exactly twice.
- **From scratch beats upcycling.** You *could* pretrain a normal transformer and then convert it into a
looped one ("upcycling"). Nanbeige compared both and found training the looped architecture from
scratch performs **significantly better** — the model needs to adapt its representations to repeated
layer reuse throughout pretraining, not have the loop bolted on afterward.
- **They kept the full KV cache.** Looping twice normally doubles the attention compute, so they tried
sharing the KV cache across passes to halve it. It consistently underperformed, so they **kept the
full, non-sharing loop** — a deliberate choice to spend inference memory on quality.
If that recurrent-depth bet sounds familiar, it's the same one as [LOTUS](/articles/lotus-latent-reasoning),
which loops a padded 3B Transformer to reason in its hidden states — and it's a cousin of the
architecture-over-scale thesis in [Motif 2.6B](/articles/motif-2-6b). You can explore the tradeoff in
the widget above; ×2 is what actually ships.
## A stronger base to start from
Before any agent training, the looped base model already leads its weight class. Pretrained from scratch
on a 28T-token corpus (larger and cleaner than Nanbeige 4.1's, with up-weighted math, code, and
synthetic-QA data — and a first taste of agentic trajectories mixed in), **Nanbeige4.2-3B-Base** beats
Qwen3.5-4B-Base, Gemma4-E4B-Base, and its own predecessor on *every* reported base benchmark: GSM8K
**92.7**, BBH **81.6**, MBPP **67.6**, SuperGPQA **35.2**, GPQA **53.3**. The knowledge gap is the
clearest — on MMLU-Pro the 3B looped base outscores a 4B Qwen base by twelve points:
That head start — the loop plus the refined 28T-token mixture — is what the post-training then turns into
agentic behavior.
## The results, and how to read them
Here's the headline chart from the report: the same benchmark suite against Gemma4 and Qwen3.5, with
Nanbeige4.2-3B in teal, essentially topping every agent and code panel and most reasoning ones.
The cleaner way to see the efficiency argument is to put score against parameters directly. On the
public code and agent benchmarks the teal point sits up-and-to-the-left — smaller *and* higher:
Now the honesty pass, because not all of these bars are the same kind of evidence:
- **The comparable, public ones** are the strongest signal: **SWE-Bench Verified 63.6** (vs Qwen3.5-9B
53.1), **SWE-Bench Pro 46.9** (vs 33.8), **Terminal-Bench 2.0 44.1** (vs 29.2), **LiveCodeBench-V6
72.5**, **HMMT-Feb-2026 82.8**. A 3B model at 63.6 on SWE-Bench Verified is genuinely notable if it
holds up in third-party harnesses. The report evaluates everyone under standardized protocols
(Table 3), which is a stronger footing than a model card's loose bars.
- **Treat the eye-catching ones with care.** GPQA-Diamond **87.4** for a 3B model is near-frontier and
surprising — it's self-reported in Think mode, the regime where small models gain the most and where
contamination is hardest to rule out. Several agent scores (GDPval, AgentIF-Oneday, OfficeQA-Pro) run
through Nanbeige's own agent stack, and **Recruit-Bench is Nanbeige's in-house benchmark** — all
reasonable to publish, none of it fully apples-to-apples.
- **Where it doesn't lead:** Gemma4-12B still wins **SciCode** (38.2 vs 35.6), **IF-Bench** (73.5 vs
54.6), and **Recruit-Bench** (69.4 vs 63.3) — strict instruction-following and some scientific coding
aren't the strong suit. The report is upfront about this: best on five of six reasoning benchmarks,
not all of them.
## Training: synthesize the environments, reward the process
The report's four-stage post-training pipeline is where the agent behavior is actually built.
**1. SFT with a STEM-to-agentic curriculum.** Starting from the pretrained checkpoint, supervised
fine-tuning runs in three stages that stretch the context window 64K → 128K → 256K while sliding the
mix of target tokens from reasoning toward agentic interaction — think first, then act:
The trajectories themselves come from large-scale environment *synthesis*: a repository-to-task pipeline
for software engineering (mine real repos, reconstruct a sandboxed container, keep only fail-to-pass
verified tasks), a hybrid real-plus-simulated pipeline for tool use (live MCP servers, Python-reconstructed
APIs, and LLM-simulated virtual tools), and an artifact-centric pipeline for office cowork (reports,
slides, spreadsheets). Crucially, the same task is solved by **multiple heterogeneous scaffolds** —
Claude Code, OpenHands, SWE-agent, Codex-style drivers — so the model learns scaffold-*invariant* repair
strategies rather than the quirks of one harness. A **turn-level loss mask** keeps bad intermediate turns
in context but out of the loss, so the model learns to recover from mistakes without being trained to
repeat them.
**2. Two-stage RLHF for hybrid thinking.** A pointwise reward model cleans up the failure modes a small
model is prone to — repetitive reasoning, cyclic reflection, delayed termination, malformed output. The
report's interesting finding is that this general-purpose RLHF **generalizes two ways**: *cross-task*
(fixing repetition and formatting also lifts math, code, and agentic scores — many agent failures are
generation loops, not reasoning errors) and *cross-mode* (behavior learned on Non-Think responses
transfers to Think mode). It's RLHF doing more than safety and style.
**3. Length-controlled reasoning RL.** A difficulty-aware penalty discourages over-long reasoning on
problems the model already solves reliably, while leaving still-hard problems room to explore — cutting
tokens without trading away correctness.
**4. Agentic RL with action-centric rubrics.** Finally, outcome rewards are combined with **process
rewards** — per-turn rubrics scoring tool-call accuracy and the information gained each step — for denser
credit assignment over long trajectories. For a model this small, the report finds it more stable to run
agentic RL on *easier* tasks (short trajectories, higher pass@8) than on the hardest ones. Across the RL
pipeline, accuracy rises while output tokens fall (e.g. AA-LCR 50.0 → 58.7 with average length dropping
19.5k → 6.7k tokens; PinchBench-V2 55.9 → 74.7).
## Small enough to live on your laptop
The payoff is deployment. At 3B non-embedding params the model is meant to run **locally** — the card
ships recipes for vLLM, SGLang, `llama.cpp`/GGUF, and Ollama (including MLX on Apple silicon), with a
configurable thinking mode (`enable_thinking`, `preserve_thinking`) and XML-format tool calls. Under the
**OpenClaw** agent framework — evaluated with the *same* scaffold and tools for every model — Nanbeige
reports beating both Qwen3.5-4B and 9B across all six daily, office, and deep-research benchmarks, with
the widest gaps on office workflows (GDPval 68.8 vs 38.0, AgentIF-Oneday 58.9 vs 32.1). The pitch is a
private, on-device assistant that can still carry multi-step tool workflows.
## The take
Nanbeige4.2-3B is another data point for a thesis this site keeps returning to: **architecture and
training, not raw scale, are increasingly what a small model needs to punch up a weight class** — the
same lesson as [LOTUS](/articles/lotus-latent-reasoning), [Motif 2.6B](/articles/motif-2-6b), and the
sub-1B security models in [Antares](/articles/antares). The looped-transformer bet is the genuinely
interesting bit, and now that the report shows the working — two passes, trained from scratch, full KV
cache — it reads less like a marketing line and more like a set of measured tradeoffs. The public
code-agent numbers are strong enough to take seriously; just keep the in-house scaffolds and
self-reported reasoning scores in the "promising, pending third-party replication" column — which is
exactly where an open-weight release lets anyone go check.
---
*Source: the [Nanbeige4.2-3B technical report](https://huggingface.co/Nanbeige/Nanbeige4.2-3B/blob/main/Nanbeige42_report.pdf)
(Nanbeige LLM Lab, 2026) and the [model card](https://huggingface.co/Nanbeige/Nanbeige4.2-3B).
Evaluations are self-reported, largely in Think mode, some using in-house scaffolds and benchmarks. The
performance figure is Nanbeige's; the interactives are mine.*
---
# Antares: a 1B model that hunts vulnerabilities like a person
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/antares
> date: 2026-07-22
> tags: security, ai, open-weights, agents, explainer
[Antares](https://blogs.cisco.com/ai/introducing-antares-the-most-efficient-open-weight-ai-models-for-vulnerability-localization)
is a family of small security models from Cisco's Foundation AI team, built for one narrow, expensive
job: **vulnerability localization** — given a vulnerability class and a codebase, pinpoint the source
files that actually contain the flaw. Two are open-weight today under Apache 2.0 — **Antares-350M** and
**Antares-1B** — with a 3B on the way, and all three are fine-tuned from **IBM Granite 4.0**. The claim
that makes them worth a look: they beat models tens to hundreds of times their size on this task, at a
fraction of the cost, and they're small enough to run **locally** so proprietary code never leaves the
building.
## The job: localize, don't fix
Antares doesn't patch anything, doesn't explain *why* a file is vulnerable, and doesn't emit exploits.
It answers one question — *which files should a human look at first?* — and it does it the way an
analyst would: as a **terminal agent**. Given only a CWE identifier and its generic description (no
advisory text, no file hints), with the repo mounted read-only, it issues shell commands — `grep`,
`find`, `cat` — reads the output, reasons, changes direction when a lead goes cold, and finally calls
`submit_vulnerable_files` with a ranked list. The catch that makes it hard: a budget of **15 terminal
calls per task**. No vector database, no retrieval index — just exploration. Scrub a run:
That "search, read, revise, backtrack" loop is the thing Cisco's earlier research argued you can
*train* into a small model — useful retrieval behavior from learned strategy, not from scale. Antares
is the test of whether it transfers to security.
## The numbers, and why the task is genuinely hard
To measure it they had to build a benchmark, because general code-search sets (SWE-Bench and the like)
test finding code relevant to a *dev task*, not localizing a *vulnerability* from a CWE description.
**VLoc Bench** is 500 tasks across 290 real repositories, 6 package ecosystems, and 147 CWE categories
(78% carry a real CVE); each repo is reconstructed at its pre-fix commit, and ground truth is the set
of files the actual security fix touched. The metric is **File F1** — the harmonic mean of how many
submitted files were right and how many of the right files were found.
Read the scores with the ceiling in mind: this is hard enough that **the best frontier model tops out
around 0.23**. Against that, a 1B open model at **0.209** is the story:
Antares-1B (0.209) clears **GLM-5.2 at 753B** (0.186) and **Gemini 3 Pro** (0.152), and the 3B
essentially matches GPT-5.5. Meanwhile several giants flail — Llama-3.3-70B lands at 0.012, plain GPT-5
at 0.048 — which tells you this isn't a capability that falls out of scale; it has to be trained in.
## Training, not scale
The cleanest evidence is the backbone itself. The same **Granite 4.0 1B** weights, untrained for this
task, score a flat **0.000** — they can't navigate a repo and submit useful files at all. Everything
Antares can do comes from a two-stage pipeline: **SFT** on cybersecurity reasoning, deep-research
traces, and terminal code-search trajectories, then **GRPO** — reinforcement learning over full
multi-turn agent trajectories with verifiable rewards for localization quality, valid submissions,
tool-use compliance, and exploration behavior. Pick a size and watch the build-up:
GRPO isn't cosmetic — it adds a real slice on top of SFT (+0.021 File F1 at 1B) by teaching the model
to *verify and stop* rather than imitate a trajectory. And it stacks down the size ladder: even the
**350M** GRPO model (0.135) beats a 753B open model.
## The economics: cheap enough to run on every commit
Accuracy-per-parameter only matters if it turns into accuracy-per-dollar, and this is where small wins
outright. The full 500-task sweep costs about **$0.71** for Antares-1B — versus **$12.50** for GLM-5.2
(15.2× more) and **$141** for GPT-5.5 (172× more) — and Antares-1B finishes it in **~13 minutes on a
single H100** with 16 parallel workers.
That's the unlock the researchers keep pointing at: as Stanford's Amin Saberi puts it, "near-frontier
accuracy on secure code reasoning at a fraction of the cost, fast enough to run on every commit." A
model this size runs on-prem, so — in NUS professor Reza Shokri's framing — "proprietary code never
leaves the machine," which matters most for the universities, public-sector teams, and smaller shops
that were priced out of token-heavy frontier models. It ships with a CLI that sweeps a read-only repo
snapshot and returns candidates as human-readable, JSON, or **SARIF** for CI/CD triage.
## Where it breaks
Cisco is refreshingly specific about the limits, and they follow directly from the design. The
15-command budget means performance **degrades on large repos** (>10MB) and on multi-file
vulnerabilities needing 5+ files of context. It's strong on **grep-able** patterns (CWE-843 Type
Confusion, CWE-1321 Prototype Pollution) and weak on ones that need real semantic understanding
(CWE-732 Incorrect Permissions, CWE-667 Improper Locking, CWE-401 Memory Leak). It has an April 2025
knowledge cutoff, and — by design — it tells you *which* files, never *why*. It's a first-pass triage
aid with a human in the loop, not a replacement for the security toolchain.
## The take
Antares is a clean demonstration of a claim that keeps getting more useful: for a **narrow, well-shaped
task**, the behavior that matters — search, verify, backtrack, know when to stop — can be trained into
a sub-1B model until it beats models 100–750× its size, cheaply enough to run always-on. It won't
generalize; it's not supposed to. It's part of Cisco's broader push (alongside its Foundry Security
Spec and CodeGuard efforts) to make AI security tooling something you can measure and deploy rather
than demo — and "frontier-adjacent accuracy at $0.71 a run, on hardware you own" is a genuinely
different offer than one more giant model behind an API.
---
*Source: [Introducing Antares](https://blogs.cisco.com/ai/introducing-antares-the-most-efficient-open-weight-ai-models-for-vulnerability-localization)
(Cisco Foundation AI, 21 July 2026), the [Antares model cards](https://huggingface.co/collections/fdtn-ai/antares),
and the [technical report](https://cisco-foundation-ai.github.io/antares/technical-report.pdf). Figures
are Cisco's; the interactives are mine.*
---
# Gigatoken: tokenizing at gigabytes per second
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/gigatoken
> date: 2026-07-22
> tags: systems, tokenization, performance, open-source, explainer
Tokenization is the step nobody thinks about. Before a language model sees a single token, terabytes
of raw text have to be turned into integer IDs — and if you've ever run that job over a real corpus,
you know it's slower than it has any right to be. [Gigatoken](https://github.com/marcelroed/gigatoken)
(Marcel Rød) is a drop-in tokenizer that does it at **gigabytes per second** — up to roughly **1000×
faster** than HuggingFace's `tokenizers`.
The detail that makes this interesting: it isn't a Python-vs-Rust story. HuggingFace `tokenizers` and
OpenAI's `tiktoken` are *already* multithreaded Rust. Gigatoken beats them by a factor of a thousand
anyway — which means the win is algorithmic, not a language swap.
## Why a regex was the bottleneck
A BPE tokenizer does two jobs. First it **pretokenizes** — splits the text into word-ish chunks,
almost universally by running a big regular expression over every byte. Then it applies the **BPE
merges** within each chunk. Everyone pictures the merges as the expensive part; in practice the
**regex pretokenization is the majority of the wall-clock**. Toggle it:
Gigatoken's two moves both target that. It replaces the regex with a hand-written **SIMD** scanner
that performs the exact same split at over 2 GB/s per thread, and it **caches pretoken→token
mappings**, so any word it has already encoded becomes a lookup instead of a re-run of BPE. Real text
is mostly repeated words, so the cache hits constantly. Add minimal Python round-trips and threads
that barely touch each other, and you get the numbers below.
## The numbers, per CPU
This is throughput encoding an 11.9 GB slice of OpenWebText, with the speedup over HuggingFace beside
each bar. Note the honest split baked into the colors:
**BPE tokenizers** — GPT-2, Llama, Qwen, DeepSeek, GPT-OSS, and friends — hit ~20+ GB/s on a big EPYC
and clear three-digit speedups. **SentencePiece-based** ones (Gemma, Mistral, CodeLlama) are the
weak spot the author flags openly: still faster, but ~10–20×, because Gigatoken hasn't optimized that
path. And the hardware matters as much as the tokenizer: a 144-core EPYC does GPT-2 at 24.5 GB/s, an
M4 Max laptop at 8.8 GB/s (its best speedups actually *exceed* 1000× because HF is slower there too),
and a single 8-core desktop Ryzen still lands around 100×.
## What "gigabytes per second" buys you
Numbers this large stop meaning anything without a yardstick, so here's one: at the EPYC's rate you
could tokenize **all of Common Crawl — about 130 trillion tokens, effectively the whole public
internet — in just under 6.5 hours.** The same job on HuggingFace's tokenizer runs for the better part
of a year. Drag the dataset size:
That's the real point. Tokenization is pure overhead on the path to training — you pay it every time
you change a vocab, re-shuffle a corpus, or add data — and a tokenizer that runs at disk speed turns a
multi-day preprocessing job into a coffee break.
## Using it
Two modes. **Compatibility mode** is the drop-in: wrap an existing tokenizer and it behaves like the
original, output matched exactly (`gt.Tokenizer(hf_tokenizer).as_hf()` or `.as_tiktoken()`) — a bit of
speed traded for bit-for-bit parity. The **Gigatoken API** is the fast path, letting the Rust side
read files directly and skip Python overhead entirely. You can benchmark any HuggingFace tokenizer
against your own data without installing anything:
```bash
uvx --with tokenizers gigatoken bench 'openai-community/gpt2' owt_train.txt \
--validate --doc-separator "<|endoftext|>"
```
It's honest about the edges, too: SentencePiece is under-optimized, WordPiece isn't supported yet,
Windows is untested (use WSL), and there's still ABI3 overhead the author expects to claw back another
2× from. The README even carries an AI-use disclosure noting most of the code was hand-written, with
AI help mainly for the user-facing API and porting SIMD strategies across AVX-512/AVX2/NEON.
## The take
Gigatoken is a clean reminder that "already optimized" is not the same as "optimal." A step everyone
had mentally checked off as solved — it's Rust, it's threaded, move on — was still leaving a **1000×**
on the table, because the actual hot loop (a regex nobody questioned) had never been rewritten for the
hardware. It won't change what your model learns. It will change whether the tokenizer is ever the
thing you're waiting on again.
---
*Source: the [Gigatoken README and benchmarks](https://github.com/marcelroed/gigatoken#benchmarks)
(Marcel Rød, 2026). Throughput measured on OpenWebText across EPYC 9565, Apple M4 Max, and Ryzen
9800X3D CPUs; the figure is the project's, the interactives are mine.*
---
# Mage-Flow: a 4B image model that bets on its tokenizer
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mage-flow
> date: 2026-07-22
> tags: image-generation, diffusion, open-weights, efficiency, explainer
The open image-generation frontier has been scaling backbones hard: Z-Image at 6B, Qwen-Image at 20B,
FLUX.2 at 32B, Hunyuan-Image-3.0 at 80B. [Mage-Flow](https://arxiv.org/abs/2607.19064), from Microsoft's
Mage team, bets the other direction — a **compact 4B stack** for text-to-image generation *and*
instruction-based editing that stays competitive with those much larger systems by **co-designing** its
three layers instead of just growing one. It's MIT-licensed and released as an open research baseline.
The organizing idea is "codec-aligned efficiency" — *spend representation capacity where the signal
is* — and it shows up as three co-designed components: a cheap tokenizer, a native-resolution backbone,
and a fused-kernel training system.
## The tokenizer that pays for everything
The VAE is the quiet tax on high-resolution diffusion: every image is encoded to latents and decoded
back, and at 2K that cost dominates. **Mage-VAE** is a lightweight pixel-diffusion tokenizer distilled
from the FLUX.2-VAE latent space, using one-step encode/decode with anchor-latent KL regularization. It
matches FLUX.2-VAE's reconstruction fidelity while doing far less work:
Make the tokenizer an order of magnitude cheaper without losing quality, and every downstream stage —
training and inference — inherits the saving. That's the lever the rest of the stack is built on.
## A backbone that doesn't crop
The generator itself is a **Native-Resolution Multimodal Diffusion Transformer**. Text prompts are
encoded by Qwen3-VL; images are turned into compact latents by Mage-VAE; and — the key move — images of
any resolution and aspect ratio are flattened into **variable-length token sequences** and packed
together with the text tokens in one batch. Per-sample 2D rotary embeddings and variable-length
FlashAttention let the 4B MMDiT process those packed sequences while preserving each image's native
spatial layout, so there are no fixed resolution buckets and no center-crops.
That's what makes one checkpoint span the whole range — 512² up to 2048², any ratio, out to an extreme
4:1 panorama:
## One backbone, three rungs
On top of that foundation Mage-Flow ships a family. A **Base** model trained with rectified flow
matching is aligned into the **RL** model with **Diffusion-NFT** (better prompt following, text
rendering, aesthetics, editing fidelity), then distilled into a **4-step Turbo** with Decoupled-DMD and
adversarial perceptual guidance. The same pattern produces the editing line. Watch the step count — and
the latency — fall:
The Turbo rung is the point: it turns a 30-step diffusion model into a **4-step** one, so at 1024² on a
single A100, generation drops from 4.37s to **0.59s** and editing to **1.02s** — interactive speed from
a model you can actually fit.
## The frontier that matters
Mage-Flow's case isn't that it tops any single benchmark — it's that it sits on a favorable
**quality–speed–memory** frontier. On GenEval (generation) and GEdit-Bench-EN (editing) it's
competitive with or ahead of much larger systems, while its peak GPU memory stays the lowest of the
field:
The concrete number is memory: across generation and editing, Mage-Flow's peak GPU memory stays around
**18–20 GB** — versus 58.8 GB for Qwen-Image, 65.5 GB for HiDream-I1, and a two-GPU 179.6 GB for
FLUX.2-dev. ~18 GB is a single desktop-class card, which is the whole pitch: a strong generation-and-editing
model that runs *locally*.
## The take
Mage-Flow is a clean argument that **image-model efficiency is a co-design problem, not a scale
problem**. The headline speed (0.59s at 1024²) comes from the tokenizer being cheap, the backbone
avoiding resolution buckets, the kernels being fused, and the sampler being distilled — each layer
pulling its weight so a 4B model can stand next to 20–80B ones. It pairs naturally with the far larger
[Qwen-Image-3.0](/articles/qwen-image-3): same task, opposite bet on where the capability should live.
Worth the usual caveat — the weights are MIT but released for research use, and the benchmark framing is
the authors' own — but the frontier it draws is a genuinely useful one.
---
*Source: [Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing](https://arxiv.org/abs/2607.19064)
(Zhang et al., Microsoft, 2026), the [Mage repo](https://github.com/microsoft/Mage), and the
[model collection](https://huggingface.co/collections/microsoft/mage). Figures are the paper's; the
interactives are mine.*
---
# NVIDIA Rubin: co-designing a GPU for the shape of agentic inference
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nvidia-rubin
> date: 2026-07-22
> tags: hardware, gpu, inference, systems, explainer
An agentic workload isn't a single prompt and response — it's sustained inference across many reasoning
steps that plan, call tools, verify, and revise over long contexts. That execution pattern stresses a
GPU differently than a chatbot does: it wants low *per-step* latency, high decode throughput, efficient
long-context attention, large KV-cache capacity, and the ability to spread a model across tightly
coupled GPUs. [NVIDIA's Rubin GPU](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/)
is a bet that the right response is to co-design for exactly that pattern — and the headline claim is
**up to 10× more agentic throughput per unit of energy than Blackwell.**
That "10×" is a vendor number on an internal 2-trillion-parameter MoE workload, so read it as a design
target rather than an independent benchmark. What's more interesting than the single figure is *how*
it's assembled — because it isn't one trick, it's a checklist of agentic bottlenecks each with its own
answer.
## The chip
Physically, Rubin is **two reticle-limited compute dies** fused into one package over a high-speed
inter-die link (NV-HBI): **336 billion transistors, 224 SMs, 896 Tensor Cores**, a third-generation
Transformer Engine that flexes precision across formats for **up to 50 petaflops of NVFP4**, up to
**288 GB of HBM4 at 22 TB/s**, and NVLink 6 at **3,600 GB/s** of scale-up bandwidth. Those are the raw
numbers; the architecture is about turning them into *sustained* utilization.
## One checklist, many bottlenecks
Here's the spine of the whole design. Each agentic-inference bottleneck gets a specific Rubin feature —
click through them:
Two are worth dwelling on. **Long-context attention** is where agentic runs spend their time, and Rubin
attacks it from two sides at once: it compresses the intermediate attention scores into a structured
**2:4 sparse** form so softmax and the second attention GEMM operate on fewer values, and it raises
exponential throughput so softmax — which becomes the bottleneck once the matrix math speeds up — keeps
pace.
The other is **MoE decode**: as expert counts climb, just locating and moving expert weights becomes
the cost. Blackwell tracks one memory descriptor per expert; Rubin keeps a single shared descriptor and
overrides the pointer and stride inline in the TMA instruction at runtime — less metadata bookkeeping,
more GPU time on actual matmuls.
## Generation over generation
The concrete comparatives NVIDIA gives are memory bandwidth and softmax (exponential) throughput. On
both, the Blackwell-Ultra-to-Rubin step is the large one:
Memory is the quiet star here. Decode — the token-by-token generation phase — is fundamentally
**memory-subsystem bound**, and agentic workloads spend more of their runtime there (long contexts,
big KV caches, interactive generation). HBM4 doubles the interface width of HBM3e for **2.8× the
bandwidth** of Blackwell, while 288 GB of capacity keeps trillion-parameter models and their KV state
resident instead of spilling to slower memory. Capacity and bandwidth do different jobs — one holds the
context, the other feeds the cores — and decode needs both.
## The data center as one unit of compute
The last move is to stop thinking about a GPU and start thinking about the **AI factory** as a fixed
power budget. Rubin's efficiency story is rack-scale: Intelligent Power Smoothing uses on-rack energy
storage to absorb the sharp power swings of AI workloads (about −10% average draw and −20% on 50 ms
peaks), and **DSX MaxLPS** turns that reclaimed headroom into more GPUs — up to **40% more** in the
same megawatt envelope.
All of this lands in the Vera Rubin NVL72 rack — third-gen MGX, cable-free trays, 45°C liquid cooling,
hot-swappable NVLink switch trays — designed so compute, networking, cooling, and power behave as one
execution domain.
## The take
Rubin is a useful lens on where inference hardware is going: not chasing a bigger peak-FLOPs poster
number, but **co-designing every layer around the execution pattern of agents** — sparse long-context
attention, distributed MoE decode, tighter kernel handoffs, fused scale-up communication, and power
treated as the real budget. It's the opposite end of the spectrum from running a model on a single
[DGX Spark](/articles/dgx-spark-batching), and the same underlying question — *how do you keep
expensive compute actually busy?* — answered at rack scale. Keep the caveat in mind: the marquee
numbers (10× throughput/watt, 40% more GPUs) are NVIDIA's own, on NVIDIA's workloads, for a platform
still rolling out. The architecture is real and specific; the multipliers are the vendor's to prove.
---
*Source: [Inside NVIDIA Rubin GPU Architecture](https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/)
(NVIDIA, 21 July 2026). Performance figures are NVIDIA's, several on an internal 2T-MoE workload; the
figures are NVIDIA's, the interactives are mine.*
---
# The harness is the generalizer: how a scaffold learns to solve longer, unseen tasks
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/harness-compositional-generalization
> date: 2026-07-21
> tags: agents, harness, llm, systems, explainer
Compositional generalization is the thing humans do without thinking: given a novel problem, break it
into familiar sub-problems, solve each, recombine. Transformers, famously, are *unreliable* at it — a
model that aces 32k-token tasks does not simply keep working at 2M tokens, and one trained to classify
Jeopardy questions does not automatically transfer to spam detection. The usual answer is *scale*:
more data, more parameters, and the ragged edges smooth out. [Alex Zhang and Omar
Khattab](https://alexzhang13.github.io/blog/2026/harness/) make a different argument, and it's a sharp
one: **the capacity for compositional generalization can live in the harness** — the program that wraps
the model — rather than in the weights.
Their one-line thesis: *"the primary job of the harness should be to carry a higher-level inductive
bias that can reduce unfamiliar and complex problems to compositions of simpler ones."* Scaling data
still matters most, they're careful to say — but *"the machinery that we feed that data into and its
inductive biases are what will determine the coefficients of that scaling."*
## The trick: keep every call locally in-distribution
Here is the mechanism that makes it work, and it's worth sitting with. A task can be wildly
out-of-distribution *as a whole* — no model trained on 32k-token inputs has ever seen a 2M-token one —
while every *individual* model call inside it stays comfortably in-distribution. Zhang calls this
property **locally in-distribution (LID)**. Drag the length multiplier and watch what it buys you:
A base Transformer reads the whole long task in one context window. Past the length regime it trained
on, that context is unfamiliar — the model degrades, the phenomenon people now call *context rot*. The
harness that stays LID never puts the model in that position: it decomposes the task so each call sees
only a short, familiar slice, and accuracy holds as the task grows.
## How the harness does it: RLM
The concrete harness they study is a **Recursive Language Model (RLM)**, and it earns LID with two
moves. The first is **context offloading**. Instead of appending each raw observation — a tool output,
a retrieved document, a sub-agent's answer — to the running context, the RLM stores it in a REPL
variable and passes only a tiny symbolic *handle*. The root LM's view stays a short, task-agnostic
prefix; the bulk data sits in the environment, peeked at through small probes. Drag the step count and
watch the two context sizes diverge:
The second move is **programmatic sub-agent calling**. Sub-agents behave like functions: they run,
and their output lands in a REPL variable rather than being spliced back into the caller's context.
Zhang stresses these are equal partners — *"programmatic sub-calling is equally as important as context
offloading"* — because together they're what keep the root context from bloating step over step, which
is exactly what would drag it out of distribution.
Standard agent patterns — ReAct, CodeAct, and by extension most coding agents — fail LID precisely
because they append everything to a growing history. The RLM is the same idea run in reverse: the
harness works to *keep the model's window small and familiar*, and lets the environment hold the state.
This is a different claim from the two harness pieces already on this site. [Agent
harnesses](/articles/agent-harness) is about *engineering the loop* — tools, context policy, the
self-improving outer loop. [The harness effect](/articles/harness-effect) is about *token economics* —
same model, cheaper orchestration. This one is about *generalization*: the harness as an inductive bias
you can train, so the scaffold itself learns to solve tasks it never saw.
## It generalizes — and the base model doesn't
The payoff is measured, not asserted. Zhang RL-trains the RLM on **short** tasks and evaluates on long
held-out ones, across six long-context benchmarks. The training signal is short-task reward; the
interesting question is whether the eval reward on much longer tasks *tracks* it.
It does. Training only on short tasks — 150 steps on `Qwen3-30B-A3B-Instruct` — the RLM generalizes to
tasks **8–32× longer**, with eval reward that *"more closely matches the train reward on shorter
tasks,"* while the base Transformer's eval stays flat even as its train reward rises. On MRCRv2,
GraphWalks, and OOLONG the trained 30B RLM approaches or exceeds a frontier `GPT-5.5` RLM. Zhang reports
roughly **10× the eval lift for the same train lift** versus a vanilla Transformer.
And it isn't only length. In a separate **strategy generalization** test, the RLM trained on one domain
transfers to a completely different one — Jeopardy-style TREC classification to spam/ham; essay
similarity to *math-problem* similarity; Twitter stance detection to error-detection in chat logs.
Again the RLM's train reward tracks its eval reward across the domain gap, and again the base
Transformer plateaus. The decomposition strategy the harness learns is the thing that transfers, not
the surface task.
It isn't free. The RLM runs **1.5–3× slower** than the base Transformer per sample — multiple LM calls
per step, sub-call latency — and for a couple of benchmarks (MRCRv2) it needed a light *"nudge to
decompose"* to converge on a generalizing strategy rather than a brittle one. Zhang's read: at scale no
supervision should be necessary, but a hint buys sample efficiency.
## The take
The reflex in this field is to push every capability into the weights and let scale sort it out. This
work is a reminder that *where* an inductive bias lives is a design choice. A harness that holds each
call locally in-distribution turns "solve a 2M-token task" into "solve a sequence of 32k-token tasks,"
and that reframing is learnable — you can RL-train it on cheap short tasks and watch it generalize to
long, unseen, even cross-domain ones. It fits the pattern the other harness pieces on this site keep
circling: the layer *around* the model is not glue. Here it's the part that generalizes.
---
*Source: [Language model harnesses are compositional generalizers](https://alexzhang13.github.io/blog/2026/harness/)
(Alex L. Zhang, with Omar Khattab), July 2026. Figures are the post's; the two interactives are mine.*
---
# Laguna S 2.1: an 8B-active model that won't give up
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/laguna-s-2-1
> date: 2026-07-21
> tags: llm, agents, coding, open-weights, explainer
[Laguna S 2.1](https://poolside.ai/blog/introducing-laguna-s-2-1) is poolside's new agentic coding
model: a **118B-parameter Mixture-of-Experts with 8B activated per token**, a **1M-token context** in
both thinking and no-thinking modes, and open weights on day one under OpenMDW-1.1. It went from the
start of training to launch in **under nine weeks**, and it is small enough to run on a single
[NVIDIA DGX Spark](/articles/dgx-spark-batching). The claim poolside makes is precise and, unusually,
falsifiable — they released full evaluation trajectories for every trial: it is *the most capable
agentic coding model in its weight class, by a wide margin.*
The interesting part is how they say they got there. Not by making the model smarter in the raw sense —
by making it **behave** better: verify more, take less for granted, stop declaring victory early, keep
going. It's the same family, and the same industrialized pipeline, as the models from
[Laguna's Model Factory](/articles/laguna-model-factory) — S 2.1 is the third release in three months.
## Punching above its weight class
Here's the whole pitch in one plot: Terminal-Bench 2.1 score against model size, log axis. Toggle to
**active params** and Laguna S 2.1 — at 8B active — sits above open models that activate five to seven
times as many parameters, and a dozen points under a closed frontier (GPT-5.6, Claude Fable 5) whose
models don't even disclose their size.
Read that honestly, the way poolside does: on Terminal-Bench 2.1 its **70.2%** is well short of the
88% ceiling that Kimi K3, GPT-5.6 Sol, and Claude Fable 5 share. Their own framing is that the top of
these benchmarks is saturating — "as the frontier advances, top scores cluster in the 70–90% range and
models that behave very differently end up no more than a few points apart." Where it actually leads is
**SWE-Bench Multilingual**, edging the much larger field:
Poolside is candid about the softer spots too — on **DeepSWE v1.1**, the least saturated of the set
(frontier models spread from 54% to 73%, and some 1T-plus models score under 10%), S 2.1 lands at
40.4%, mid-pack, and it reports that in its own harness rather than the leaderboard's. The takeaway
isn't "it wins" — it's "score-per-parameter," and on that axis nothing its size is close.
## The second axis: behaviors, not intelligence
> What we've done in this model is not necessarily add more intelligence, but improve the behaviors
> that lead to a more capable model: more verification, less taking things for granted, not declaring
> victory early, and being more persistent. — Pengming Wang, co-head of Applied Research
The concrete lever behind that is **test-time compute**. Of every model poolside has trained, S 2.1
has the largest gap between its no-thinking and max-thinking modes — its internal monologue is doing
real work, especially on the hard problems. Flip it:
There's no user-facing low/medium/high dial yet — it's off or max (default), with the model choosing
its own budget. Poolside says it has watched coherent, productive reasoning run for **hours and
hundreds of thousands of tokens**, which is also why the 1M-token context matters: long agentic
sessions genuinely accumulate that much working state.
## Seeing it work
Benchmarks are a proxy; the trajectories are the evidence. Poolside published three unedited runs.
**A browser engine from an empty folder.** Asked to build an HTML/CSS rendering engine in vanilla
JavaScript, Laguna S 2.1 worked one 50-minute session, 181 steps, no human intervention — building the
full pipeline (HTML tokenizer → DOM → CSS parser with specificity → cascade → box-model layout →
canvas-2D renderer). The resourceful part: with no vision of its own, it needed a way to *check* its
output, so it ran **headless Chromium to read the canvas back and compared screenshots numerically**
against a real browser.
The verbatim prompt, if you want to reproduce it:
```text
your job is it to build a simple browser engine (just html/css) in
javascript to demonstrate the capabilities of poolsides new "Laguna S"
model. the goal is to take render html snippets in a canvas like a real
browser. to demonstrate it the engine, build a self-contained single
page app that showcases a gallery of multiple html snippets and renders
them side by side (canvas with our render engine + iframe letting the
hosting browser render it for real for comparison). support for most
common layout and styling elements
```
**Optimizing poolside's own harness.** Pointed at the agent harness that trains and serves the models,
in an automated loop S 2.1 made it **5.2% faster with ~70% lower memory allocation** — finding an
O(n²) string-concatenation in token accumulation and swapping in buffers, then memoizing trajectory
materializations and pre-allocating slices. When speedups got marginal, it *kept going*, switching its
attention to allocations because those were still measurable. (Validated with Go's race detector and
`go vet` gating — real gains, not hidden race conditions.)
**Re-deriving Erdős problem #397, in Perl.** With no Python in the sandbox, it found Perl, did exact
prime factorizations there, conjectured a family, and proved it — a closed-form infinite family of
eight-index solutions to a problem open for over 50 years. It's a **re-discovery** (GPT-5.2 Pro solved
it in January 2026), and poolside says so plainly — but the construction is structurally different
(eight indices growing linearly vs the known six-index family), so it's a fresh derivation, not recall.
## What actually changed in training
S 2.1 is a **scale-up of the Laguna XS family, on exactly the same pre-training data as XS 2.1** — the
step up was scale, training-code fixes, and small recipe tweaks, not new data. Almost everything that
separates it comes from **post-training**, in two stages: an SFT stage (partly synthetic) that
bootstraps capability, then RL reserved for tasks the model can't already solve at a high pass rate.
It's also poolside's first model to run **RL in FP8 precision**.
The task corpus is the substance: **409k environments** — 83k terminal-focused, 168k standard
software-engineering — mostly grounded in real code history, the largest source reproducing **~38,000
real commits across ~17,000 repositories**, plus merged-PR reconstruction, injected-bug fixing, and a
new agentic step: given a repo, install every dependency and get its test suite running. And three
changes to the loop map straight onto the "persistence" story:
- **More generous rollout budgets** — longer timeouts, more tokens per turn, more turns per task than
any earlier model (likely why it keeps going).
- **A new sandbox** — background processes, selective network blocking to shrink the reward-hacking
surface, artifact caching.
- **Multi-harness rollouts** — the same prompts rolled out across several agent scaffolds, so it
learns behaviors that transfer instead of overfitting to one harness.
That cadence — M.1 and XS.2 in April, XS 2.1 in July, S 2.1 weeks later — is exactly what the
[Model Factory](/articles/laguna-model-factory) was built to enable: reproducible foundations, so each
release inherits the last one's work automatically. Poolside is upfront about the rough edges shipped
to move fast: some tool-schema slips in *third-party* harnesses (it leans on memory of its own tool
interface), invalid JSON in tools that expect array arguments, and occasional overthinking on
competition math.
## The take
Laguna S 2.1 is the clearest example yet of a thesis worth taking seriously: **how a model works is a
separate axis from how smart it is, and it's trainable.** Persistence, verification, and backtracking
aren't emergent gifts of scale here — they're the product of longer rollouts, a better sandbox, and
multi-harness RL, poured into an 8B-active model that then holds its own against giants. It won't top
the leaderboards, and poolside doesn't pretend it does. But "frontier behavior at a size you can run on
one desktop box" is a more useful thing to ship than another point of benchmark score — and it's the
[Model Factory](/articles/laguna-model-factory) that makes shipping it every few weeks look routine.
---
*Source: [Introducing Laguna S 2.1](https://poolside.ai/blog/introducing-laguna-s-2-1) (poolside,
21 July 2026); benchmark figures as published there (pass@1 averaged over 3–4 attempts), with full
trajectories at trajectories.poolside.ai. The browser-engine screenshot is poolside's; the
visualizations are mine.*
---
# LongStraw: fitting million-token RL onto a fixed GPU budget
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/longstraw-2m-rl
> date: 2026-07-21
> tags: rl, systems, long-context, llm, explainer
There's a widening gap in how long a context a model can *use* versus how long a context you can *train*
it on with RL. Inference servers now run million-token windows; RL post-training, as publicly described,
mostly sits at **256K tokens or below** and leans on length generalization at deployment. That gap
matters most for agents, whose observations, tool outputs, retrieved documents, and past decisions pile
up into the history that conditions the next action. [LongStraw](https://github.com/MindLab-Research/longstraw)
(MindLab and Fudan) sets out to close it — and it is unusually disciplined about framing the result as a
*systems feasibility* claim, not an accuracy one.
Read this as a systems paper. LongStraw shows you **can execute** million-token GRPO steps under a fixed
GPU budget — the memory transaction completes, on real hardware, with finite values. It does **not**
claim these runs improve reasoning accuracy, and it is careful to say so repeatedly. I'll keep that
distinction front and center, because it's the whole integrity of the work.
## Why GRPO is the hard case
Inference has an easy memory story: prefill the prompt, keep the state you need to decode, throw the
forward graph away. GRPO can't. It samples **G** responses for one prompt and updates the policy from
their *relative* rewards — so old-policy scoring, reference scoring, and every policy response all depend
on the *same* long history, and policy learning must also retain or reconstruct what backward needs.
Quadratic attention plus long-lived backward state make GPU memory the wall. The paper's framing: the
practical limit is **state lifetime, replay, and distributed ownership — not the attention kernel alone.**
## The move: change the graph boundary
LongStraw's core data structure is **resident state**: evaluate the shared prompt with *no autograd* and
keep only the minimal, model-native state later tokens need — recurrent state, KV pages, latent pages —
not the full prompt activation graph. Formally it stores z̄ₚ = stopgrad(zₚ(θ)). Its core algorithm is
**response replay**: restore that boundary, score the old and reference branches graph-free, rebuild
*one* policy response under autograd, backpropagate, and pop back to the boundary.
That last point is the honest heart of it. A full-sequence gradient has two terms — the direct response
gradient and the prompt-side term that flows back through zₚ(θ). LongStraw keeps the first and drops the
second. In the paper's own words, printed right on the figure: *"conditional-response gradient only …
full-sequence gradient parity is not claimed."* The measured objective is **response-only execution**.
## Serial replay makes G a schedule, not a memory multiplier
Because responses are replayed one at a time, the live autograd graph is bounded by a single response —
so the group size **G** becomes a scheduling/time dimension instead of a memory multiplier. This is the
lever that makes million-token steps fit. Drag G and watch the measured peak barely move:
On eight H20 GPUs, Qwen3.6-27B completes an exact-attention response-only GRPO step at **2,097,152
positions** for both G=2 and G=8 — and going from G=2 to G=8 adds only **0.208 GB** of peak allocated
memory per rank (97.503 → 97.711 GB). Holding all G response graphs, by contrast, would scale activation
memory with G and overrun the budget. That's the trade the design makes: G costs wall-clock (271 s per
member), not memory.
## Two incompatible architectures, same transaction
The design is instantiated for two structurally different models, which is most of the engineering:
- **Qwen3.6-27B** (8× H20, CP8) — 48 Gated-DeltaNet layers + 16 full-attention layers, dense FFNs. It
keeps compact recurrent state for GDN and right-sized CP8-sharded KV pages for full attention, composes
rank-local softmax statistics into the exact global attention output, replays response blocks in
reverse, and allocates K/V gradient pages only when a response backward touches them. NF4 QLoRA,
116.7M trainable parameters.
- **GLM-5.2** (32× H20, CP32/EP32) — 78 MLA/DSA attention layers (21 index + 57 IndexShare) and a
3-dense/75-MoE feed-forward stack routing each token to 8 of 256 experts (+1 shared). It holds
CPU-resident MLA latent pages and DSA index-key pages, reconstructs the sparse selection across
context-parallel owners, and replays the real Megatron router, EP32 all-to-all, and expert compute
under whole-layer checkpointing. (This is the same [GLM 5.2](/articles/glm-5-2) whose IndexShare makes
a million-token context cheap.)
Qwen solves *preserve and replay dense/recurrent history*; GLM extends it to *dynamic sparse selection,
routed experts, and cross-rank communication*.
## The numbers, and the ceiling
The headline is how far the training context moves — from the usual quarter-million to millions of
positions on fixed hardware:
At **4,456,448 positions**, one captured prefix supports **eight consecutive G=8 optimizer cycles** — 64
response replays — at **83.894 GB per rank**, comfortably under the H20's 150.755 GB. On 32 H20 GPUs, GLM
completes a deterministic 2M execution across all 78 layers with two full backward passes. The memory
plot shows how the operating points sit against the ceiling:
Note what the plot says and doesn't: 2M is *"an achieved operating point rather than a measured capacity
ceiling."* They ran it; they don't claim it's the max.
## The refresh knob
Capturing a million-token prompt is the dominant cost, so LongStraw reuses one capture across several
optimizer steps. But each update moves the parameters, and the cached prompt state goes stale. A 1M
fresh-prefix oracle measures exactly how stale — and turns "how long can I reuse a prefix" into a number:
Reuse is nearly free for a step or two (loss drifts ~0.04–0.12%) and clearly not by step four. So the
refresh interval becomes a measured control, not an assumption.
## The fine print (which is the point)
This is where LongStraw earns its credibility. It defines **four levels of validation** and states plainly
where each path lands:
**What is not claimed.** The runs establish *response-only execution*, not full-sequence gradient parity.
Qwen's global attention merge uses a BF16 numerator; its distributed **optimizer finalization is
incomplete** — the prototype all-reduces dQ but leaves page-owner K/V gradients rank-local, so the eight
AdamW instances are locally-applied, not replica-equivalent. The GLM 2M run predates restored gradient
finalization, so its stronger global-update claim is *unestablished*. And the 2M workloads are
**synthetic**: β=0, unclipped surrogate, old and reference scores coincide at step one, rewards and
advantages are synthetic. The real online sampling→reward→train loop (vLLM-DAPO-MATH) is validated only
in **short-context, archived external runs** — *"not a 2M online rollout, repeated policy learning, or
full-sequence gradient parity,"* and no evidence of long-context policy improvement.
Spelling that out is not a weakness of the paper — it's the substance. A lesser report would have shown a
2M run and let you assume it means a better model. LongStraw shows a 2M run and tells you exactly which
narrow, well-defined thing completed.
## The take
The useful reframing here is that **long-context RL is a state-lifetime and ownership problem, not an
attention-kernel problem.** Once you treat the long prompt as detached resident state and replay
responses serially, the memory transaction — not the kernel — is what sets how far you can train, and
million-token steps fit on inventory you already have. It's a real feasibility milestone, and it sits
next to the other systems-first takes on RL cost like
[frontier RL is cheaper than you think](/articles/frontier-rl-cheaper). What's left — and the paper is the
first to say it — is closing the distance from "the step executes" to "the gradient is exact and the
policy actually improves at 2M tokens." That's the next paper, honestly labeled.
---
*Source: [LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget](https://arxiv.org/abs/2607.14952)
(Zhou et al., MindLab & Fudan University), July 2026. Figures are the paper's; the two interactives are
mine, built on its measured numbers.*
---
# Qwen-Image-3.0: chasing “useful” instead of “good-looking”
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/qwen-image-3
> date: 2026-07-21
> tags: image-generation, multimodal, diffusion, qwen, explainer
Every image model since the first Stable Diffusion has been optimised, implicitly, for one thing:
making a picture you'd want to look at. [Qwen-Image-3.0](https://qwen.ai/blog?id=qwen-image-3.0) —
the third generation of Qwen's image line — says the flex out loud and then walks away from it. Its
whole pitch is one Chinese word, **实 ("Real")**: not *prettier*, but *useful* enough to put in a
production pipeline. Where 1.0's keyword was "Precision" and 2.0's was a five-word mouthful
("Precision, Variety, Completeness, Beauty, and Authenticity"), 3.0 collapses to a single claim and
splits it three ways — **Rich Content, Authentic Details, Deep Knowledge**.
Two things to set expectations first. This was a **capabilities announcement**, not a tech report:
Qwen published no architecture, parameter count, or benchmark table — only examples. And unlike the
open-weight 1.0 and 2.0 (Apache-2.0 releases the community could run), 3.0 arrived **closed**,
reachable through [Qwen Chat](https://chat.qwen.ai/?inputFeature=t2i) and Alibaba's API rather than a
weights download. So this is a piece about *what it claims to do* and how to reason about it — with
the model's own examples, the levers behind them, and an honest look at where it cracks.
## Rich Content: how much fits in one image
The demo that opens the announcement looks like a single tidy math slide. It isn't — it's one cell of
a **3×3 grid**, generated in a single pass, where every cell is a different dense infographic:
projectile motion, the Sylow theorems, a parasitology explainer, a bank internal-control diagram, a
cell-DNA comparison, and more. Nine unrelated technical posters, each with its own Chinese and English
text, formulas, and charts, laid out without bleeding into one another.
The thing actually being demonstrated isn't the grid — it's the **prompt budget**. Describing that
grid precisely takes about 3.7k tokens, and 3.0 raises the instruction ceiling to **4.5k tokens**,
several times what 2.0 would reliably follow. "How much you can draw" turns out to be gated by "how
long an instruction the model will still honour." Drag it:
There's a second axis to "rich": not just width, but **depth** — interfaces nested inside interfaces.
One instruction renders, outer to inner, a VSCode window holding a Qwen Chat window holding a WeChat
thread holding a pour-over-coffee poster, each preserving its own authentic chrome.
## Authentic Details: how finely it draws
If Rich Content is about *how much*, Authentic Details is about *how fine*. The headline number here
is **10px**: text small enough that most generators turn it into a texture that merely looks like
writing, which 3.0 claims to keep genuinely legible. Legible small text is the single hardest thing in
image generation — it's where the difference between "renders language" and "renders squiggles" lives.
The stress test is an academic paper: a full page of algebraic-geometry derivations with superscripts,
subscripts, curly braces, fraction bars, and multi-line aligned equations — the kind of layout where a
single wrong glyph is obvious.
The same precision shows up in **editing**, not just generation. Given a damaged traditional
ink-wash painting, the model restores the missing regions — matching brushwork, ink gradients, and
feather texture, removing mould spots — while leaving the original composition intact.
## Deep Knowledge: how broadly it draws
The third axis is coverage. Qwen claims native rendering across **12 languages** (Japanese, Korean and
Spanish are shown), 100-plus artistic styles, and a spread of real UI grammars — web pages, games,
livestream overlays — backed by enough world knowledge to build things like a scientific figure from a
photo. Given an insect photograph, the model keeps the subject and adds taxonomic labels, morphological
annotations, a magnified detail inset, and a scale bar: a publication-ready research figure.
It can even reach *outside* itself: the model connects to the web to pull current facts — the
announcement generates a weather-forecast card for a specific city and date — and composes with known
figures, e.g. staging Qi Baishi and Van Gogh co-hosting a livestream. This is the "productivity tool,
not toy" thesis in one line: newspapers, storyboards, exam papers, and UI mockups are the target, not
wallpaper.
## The catch
Here's where the honesty matters. Every image above is Qwen's own curated demo, and independent
testers have been blunter than the blog: reports describe it as roughly a **notch below** the best
proprietary generators (GPT-Image-class models, Nano Banana Pro), with **real gaps** once you leave the
reel — Korean text with typos, charts that break, and data-plotting tasks (a GDP chart) with points
placed in the wrong spots. Legible-at-10px and accurate-at-10px are different claims, and factual
layout — where the *numbers* have to be right, not just crisp — is exactly where it slips.
None of that erases the direction, which is the interesting part. Optimising an image model for
*deployable* output — long controllable prompts, small legible text, nested real UIs, factual layout —
is a more useful target than one more bump in aesthetic quality, even when the first release doesn't
fully hit it. The trade for it is openness: 1.0 and 2.0 were weights you could run; 3.0 is an API you
call.
## The take
Qwen-Image-3.0 is best read as a **repositioning**, not a benchmark win. "Real" is a good target —
image generation is far more valuable as a layout-and-typography engine for documents and interfaces
than as an art toy — and the three levers it leans on (a 4.5k-token instruction ceiling, a 10px text
floor, and world-knowledge grounding) are the right ones for that job. Just hold the demos and the
independent tests in the same hand: the ceiling is real, the floor is real, and the accuracy at that
floor is still catching up. It pairs naturally with the site's piece on
[Qwen Audio 3.0 TTS](/articles/qwen-audio-3-tts) — the same "3.0, make it deployable" push, one
modality over.
---
*Source: [Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge](https://qwen.ai/blog?id=qwen-image-3.0)
(Qwen team, 2026-07-16). Architecture, parameters and benchmarks were not published with the release;
availability and independent-testing notes via secondary coverage. All images are Qwen's official
examples, shown for commentary.*
---
# EAGLE-3: making the draft model scale, and a from-scratch build
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/eagle-3-speculative-decoding
> date: 2026-07-20
> tags: inference, speculative-decoding, llm, systems, explainer
Autoregressive decoding is the tax every LLM pays: one forward pass, one token, and each pass is
memory-bound — you stream all the weights through the chip to produce a single token. **Speculative
decoding** is the trick that beats it. A small, cheap **draft** model guesses the next few tokens; the
big **target** model verifies all of them in one parallel forward pass and keeps the longest prefix it
agrees with. The output is provably identical to normal sampling — you just get several tokens per
expensive pass instead of one. [EAGLE-3](https://arxiv.org/abs/2503.01840) (Li et al.) is the current
peak of the EAGLE family, and its contribution is subtle: it makes the *draft model itself* keep
getting better with more training data. Up to **6.5× faster**, losslessly.
## Draft and verify
The whole game is **acceptance**: how often the target agrees with the draft's guesses. A token only
counts if every token before it was accepted too, so its odds decay geometrically — which is why
deeper drafts hit diminishing returns and the real lever is raising the acceptance rate. Drag it:
The mean **accepted length τ** — tokens produced per target forward pass — is essentially the speedup.
Verification is exact: the target only ever emits a token it would have sampled itself (rejected
tokens are resampled from the corrected distribution), so nothing about quality changes. Speculative
decoding is a pure latency win with no accuracy cost.
## What EAGLE-3 changes
EAGLE's draft model was clever: instead of predicting tokens directly, it autoregressively predicted
the target's **top-layer feature** and reused the target's own LM head. But that came with a
feature-prediction loss that **constrained the draft** — and the authors found that scaling up its
training data barely helped. EAGLE-3 makes two changes.
First, it **drops feature prediction for direct token prediction**, and trains the draft the way it
will actually run — a technique they call **Training-Time Test (TTT)**. The problem it solves is
concrete: EAGLE's first drafted token got accepted often, but its *second* collapsed, because at step
two the draft feeds on its own step-one output, which drifts away from the features it was trained on.
TTT folds that multi-step rollout into training so the draft learns to consume its own predictions.
Second, freed from the feature constraint, the draft fuses the target's **low, middle, and high**
features instead of only the top layer — concatenated and projected down — giving it a richer basis
for predicting tokens two and three ahead.
## A scaling law for inference acceleration
Here's the payoff, and it's the reason the paper exists. With the feature constraint removed, the
draft model's speedup **keeps climbing as you give it more training data** — a relationship never seen
for EAGLE, which flatlines. EAGLE-3 was trained on roughly **8× more data** than EAGLE, and the curve
is still going up.
That reframes draft-model training as a data-scaling problem instead of a fixed architectural trick —
the same lesson the rest of the field keeps relearning.
## The numbers
On Vicuna-13B (greedy), EAGLE-3 averages **5.51× speedup** across five tasks, peaking at **6.47×** on
HumanEval with an accepted length of **7.54** — a clear step over EAGLE-2, and multiples over Medusa
and vanilla speculative sampling:
The part that matters for real serving is that the win survives **batching**. Most speculative methods
degrade past batch ~16 (the extra draft compute stops paying off when the GPU is already busy); in
SGLang, EAGLE-3 still delivers **+38% throughput at batch 64**, and at batch 1 it hits **373 tok/s vs
158 for plain SGLang** — a 2.36× serving speedup. It's also compatible with EAGLE-2's dynamic draft
tree, so the two stack.
## From scratch: tiny-speculators
If you want to see the machinery without the framework, [`tiny-speculators`](https://github.com/junuxyz/tiny-speculators)
(junuxyz) is a from-scratch EAGLE-3 **trainer**, verifier = Qwen3-8B, in five clear stages: prepare
ShareGPT data → extract the verifier's early/mid/late/final hidden states via vLLM → train the draft
with the **3-step TTT rollout** (using PyTorch **FlexAttention** for the non-causal TTT mask) → export
to vLLM's speculators format. The draft is a **single decoder layer** that fuses three hidden states
(projected `3H → H`), concatenates the sampled token's embedding, and predicts the next token.
It's honest about being a small educational build. Its 60k-sample checkpoint on HumanEval reaches a
**27.4% draft-token acceptance rate** and a **mean accepted length of 1.82**, cutting single-request
p50 latency from **2.61s to 1.78s** — while openly noting that at high concurrency it stayed *below*
plain serving. That gap between a from-scratch draft and the paper's 6.5× is exactly the value of the
scaling law: acceptance is a data problem, and EAGLE-3's whole point is that more of it keeps helping.
## The take
Speculative decoding was already the standard latency trick; the interesting move in EAGLE-3 is
turning the *draft model* into something that scales. Drop the constraint that stopped it learning
(feature prediction), train it on its own multi-step outputs (Training-Time Test), give it more of the
target's internal features to look at, and its acceptance — and therefore the speedup — rises with
data instead of plateauing. It pairs naturally with the site's pieces on
[multi-token prediction](/articles/multi-token-prediction) and
[DeepSeek's DSpark](/articles/deepseek-dspark): the same idea — predict more per step, verify exactly —
attacked from three directions.
---
*Source: [EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test](https://arxiv.org/abs/2503.01840)
(Li, Wei, Zhang, Zhang) and the [tiny-speculators](https://github.com/junuxyz/tiny-speculators) repo.
Figures are the paper's; the interactive is mine.*
---
# Motif 2.6B: differential attention and PolyNorm, trained at scale
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/motif-2-6b
> date: 2026-07-20
> tags: llm, architecture, pretraining, attention, explainer
Most "small strong model" reports are the same model everyone else ships — a dense pre-norm
Transformer with RoPE, GQA, and SwiGLU — trained on better data. Motif Technologies'
[Motif 2.6B](https://huggingface.co/Motif-Technologies/Motif-2.6B)
([arXiv 2508.09148](https://arxiv.org/abs/2508.09148)) is not that. It makes two architecture bets
that mostly live in papers, not products — **differential attention** and a learned polynomial
activation called **PolyNorm** — and it's the first model I've seen train *both* at real scale
(2.5T tokens). The result is a 2.6B model that goes toe-to-toe with 7–8B models on code and math.
Two mechanisms, a data schedule with a plan, and a couple of training tricks make it work.
## The block
Motif is a 32-layer, hidden-size-2048 dense Transformer — 16 attention heads, no GQA (16 KV heads),
a 219,520-token vocabulary, RoPE with θ = 500,000. Standard pre-norm skeleton. What's swapped in are
the two coloured boxes: the attention sublayer is Differential Attention, and the feed-forward
sublayer's nonlinearity is PolyNorm.
These weren't picked by taste. Motif ran controlled ablations at 0.6B, 1.8B, and 4.6B under a fixed
3e20-FLOP budget, testing QK-Norm, Cross-Layer Attention, and Normalized GPT (nGPT) alongside these
two — and differential attention plus the polynomial activation are what survived.
### Differential attention
Ordinary attention leaks. A softmax over the whole context puts non-trivial mass on tokens that have
nothing to do with the query — the more filler in the window, the more the signal gets diluted.
Differential attention (Ye et al.) fixes this by computing **two** attention maps per head and
returning their difference: `attn = [softmax(Q₁K₁ᵀ) − λ·softmax(Q₂K₂ᵀ)]V`. The second map is trained
to model the common-mode noise, so subtracting a λ-scaled copy of it cancels the leak. Drag λ:
It's the same move as a differential pair in analog circuits, or noise-cancelling headphones:
measure the noise separately, subtract it, keep the signal. Sparser attention shows up downstream as
better long-context retrieval, less hallucination, and cleaner in-context learning — and a GroupNorm
after the subtraction keeps the (now possibly negative) scores stable.
### PolyNorm
The other swap is the activation. Instead of committing the whole network to one fixed curve, PolyNorm
makes the nonlinearity **learned**: a degree-3 polynomial over normalized powers of the input,
`PolyNorm(x) = a₁·n(x) + a₂·n(x²) + a₃·n(x³)`. Each layer learns its own aᵢ, so it can bend toward
near-linear, saturating, or S-shaped as needed, and pick up higher-order interactions a single
activation can't. Mold it:
The per-power normalization is the load-bearing detail — it's what stops the x³ term from exploding
the activation scale, which is exactly why cubic activations don't normally survive contact with a
real training run. Motif's is capped at degree 3 on purpose.
## A schedule for the data
Here's the training idea I like most. Motif runs a **data-mixing scheduler** — a schedule for the
*corpus*, conceived like a learning-rate schedule. The 2.5T-token dataset is partitioned into eight
domain groups whose sampling ratios move **linearly** from a start mix to an end mix across training.
Drag the training-progress handle:
Early training is broad and web-heavy, teaching general language; the final, best-behaved tokens
pour into Korean, code, math, and reasoning — the dense skills the model will actually be graded on.
The base corpus is aggregated and filtered from **DCLM, TxT360, FineWeb2, and FineMath**, plus an
in-house Korean corpus (there wasn't a good open one).
Two more training details are worth stealing. The LR follows **WSD** (warmup-stable-decay): a peak of
`5e-4` held flat for the first 2T tokens, then annealed to 25% of peak over the final 0.5T. And
throughout, Motif does **checkpoint averaging** — every 8B tokens it takes a simple moving average of
the six most recent checkpoints and feeds the averaged weights straight back into the training loop,
a cheap, continuous smoothing that costs nothing at inference. A stage-2 anneal (~500B tokens) then
stretches RoPE from θ = 10,000 to 500,000 (ABF) and extends context 4K → 16K for the long-context
variant in the last 80B tokens.
## The finetuning stack
The post-training is where the reasoning gets sharpened. SFT is small — under 15B tokens, ~5M samples
— but heavily engineered.
A few of the moves are unusually specific. **Exam-CoT QA** synthesizes ~5M standardized-exam
multiple-choice items with step-by-step rationales. **EvolKit** (Auto Evol-Instruct, with Qwen3-8B)
rewrites existing SFT samples into harder ones. And **dataset fusion** compresses several samples into
one cohesive conversation — they found Qwen3-8B just concatenated the inputs, so they used GPT-4o to
actually fuse them, packing more knowledge per token. Rejection sampling against a reward model prunes
the weak generations. Then alignment is two-stage DPO — coarse-grained (Tulu 3 preference mixtures)
then fine-grained (MagpieLM + LMSys arena data).
## Punching above its weight
The scoreboard is the point of all of it. On code and math, a 2.6B model lands where 7–8B models do —
and sometimes past them:
On HumanEval it's within a point of Llama 3 8B and more than doubles Mistral 7B; on MATH it clears
Mistral 7B by 3×. GSM8K tells the same story — 75.7 (8-shot, maj@8) against Mistral's 52.2. The honest
caveat is knowledge: on MMLU the smaller model can't fake breadth.
MMLU rewards parameters you simply don't have at 2.6B, so Motif trails the 7–8B models there while
still beating Gemma 2 2B. That's the shape of the whole result: reasoning and code you can *train in*
with the right architecture and data schedule; raw knowledge still scales with size.
## The take
Motif 2.6B is a bet that the small-model recipe isn't finished — that there's still room to change the
architecture, not just the data. Differential attention buys cleaner attention, PolyNorm buys a learned
nonlinearity, the data scheduler front-loads language and back-loads skill, and WSD plus checkpoint
averaging smooth the ride. None of it is exotic in isolation; the report's contribution is showing the
combination survives 2.5T tokens and lands a 2.6B model in 7–8B territory on the things you can teach.
The pieces that "only work in papers" turn out to work in a model — which is the most interesting kind
of result.
---
*Source: the [Motif 2.6B technical report](https://arxiv.org/abs/2508.09148) (Motif Technologies) and
its [model card](https://huggingface.co/Motif-Technologies/Motif-2.6B). Figures are the paper's;
the interactive diagrams are mine. Differential Attention is from Ye et al.; the polynomial-activation
idea predates PolyNorm's use here.*
---
# Audex: audio, speech, and text through one decoder — without the text tax
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nemotron-audex
> date: 2026-07-20
> tags: audio, llm, multimodal, speech, explainer
There's a tax you pay when you make a text LLM multimodal. Bolt on a speech encoder, fine-tune for
audio tasks, and the model's text scores — reasoning, knowledge, instruction-following — tend to
sag. The audio ability arrives; some of the intelligence leaves. NVIDIA's
[Audex](https://huggingface.co/nvidia/Nemotron-Labs-Audex-30B-A3B) (Nemotron-Labs-Audex-30B-A3B,
[arXiv 2607.05196](https://arxiv.org/abs/2607.05196)) is a unified audio-text LLM whose whole point
is that it *doesn't* pay that tax. It does ASR, speech translation, audio understanding,
text-to-speech, text-to-audio, and direct speech-to-speech — and keeps the frontier text scores of
the model it was built on.
The design is almost aggressively simple, which is the interesting part.
## One decoder, one vocabulary
Most audio LLMs treat audio as a side channel: an encoder produces features, an adapter head or a
separate module consumes or emits them. Audex refuses the split. It is a **single Transformer
decoder** built on **Nemotron-Cascade-2-30B-A3B** — a hybrid Mamba–Transformer mixture-of-experts,
30B total parameters with ~3B active, a 1M-token context. Audio *inputs* are turned into continuous
embeddings by an audio encoder plus MLP adapters and dropped straight into the **text embedding
space**. Audio *outputs* are discrete tokens drawn from an **extended vocabulary** (reported at
205,312 entries) that the model predicts in exactly the same autoregressive stream as text.
The upshot is that "task" is not an architectural mode — it's just which token types show up in the
stream. Transcription is audio-in, text-out. TTS is text-in, speech-tokens-out. A spoken reply is a
single sequence that switches from text to speech tokens partway through. Flip between them:
At the far end, the discrete tokens hit dedicated detokenizers: a **speech decoder** (XCodec2, with
a causal variant for streaming) reconstructs the waveform for TTS and speech-to-speech, and an
**audio decoder** (XCodec1 with an enhancement VAE) handles general text-to-audio. Because the
whole thing is "just an LLM emitting tokens," it stays compatible with standard LLM training and
inference infrastructure — no bespoke serving path for the audio head.
## The text tax, measured
Here is the claim that matters, and it's a claim you can only make with a table. Scored on
**text-only** benchmarks against other recent audio LLMs, Audex doesn't just avoid regressing — it
sits at or near the top of nearly every one:
Read the paper's own table and the pattern is stark. Audex 30B-A3B posts **AIME 2025 91.2**,
**MMLU-Pro 78.9**, **GPQA-Diamond 74.9**, **ArenaHard v2 81.6**, **IFBench 77.8**,
**LiveCodeBench v6 85.3**, and a **1M-token context** with **99.4 / 83.4** on 256K/1M needle-in-a-
haystack — while audio models like Voxtral and MiMo-Audio show the tax plainly and even a strong
omni peer trails on reasoning and long context.
That "marginal or no regression" line in the abstract is the whole thesis. The audio ability is
additive, not a trade.
## What it does with the audio half
Keeping the text brain would be a hollow win if the audio were weak. It isn't. On the **OpenASR**
leaderboard (English), Audex 30B-A3B averages **6.82 WER** across eight test sets — 1.34 on
LibriSpeech clean, a table-best **1.76 on SPGI** — competitive with Whisper-large-v3 and the omni
models while being a single unified system:
Around that sit speech translation, audio question-answering, text-to-speech and text-to-audio
generation, and — the one that closes the loop — **speech-to-speech**: spoken input to spoken
output in one model, no cascaded ASR→LLM→TTS pipeline with its latency and error stacking.
## How you buy no-regression
The recipe is where the tax actually gets dodged. Audex is trained on **157.4B audio tokens and
320.5B text tokens** — note the text is still the majority — through multi-stage supervised
fine-tuning, then a **text-only** Cascade RL pass plus multi-domain on-policy distillation.
Two details do the work. First, the SFT is studied as **two curricula** — a multi-stage one that
adds capabilities one at a time (text SFT → audio warmup → audio-gen → audio-gen + understanding)
and a consolidated single-stage one that mixes everything at once. Second, and more important, the
reinforcement-learning stage that follows is applied in the **text domain**, the same Cascade-2 RL
the backbone already knew. Audio is learned as *additional* token vocabulary on a preserved base,
and the final polish happens where the text intelligence lives — so it's reinforced, not eroded.
## The take
Audex's bet is that you don't need a clever fusion architecture to add audio to an LLM — you need to
refuse to treat audio as special. Encode it into the text embedding space on the way in, emit it as
extra vocabulary on the way out, keep the majority of your training tokens textual, and do your RL
where the reasoning is. The reward is a model that hears, speaks, and generates audio while still
scoring 91 on AIME and holding a million-token context. The "unified" in unified audio-text LLM
turns out to mean *boring on purpose* — and that's the compliment.
---
*Source: [Unified Audio Intelligence Without Regressing on Text Intelligence](https://arxiv.org/abs/2607.05196)
(Zhifeng Kong et al., NVIDIA) and the [model card](https://huggingface.co/nvidia/Nemotron-Labs-Audex-30B-A3B).
Figures are the paper's; the interactive diagrams are mine.*
---
# Qwen Audio 3.0 TTS: an instructable LM-plus-flow-matching speech stack
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/qwen-audio-3-tts
> date: 2026-07-20
> tags: audio, tts, speech, generative, explainer
Modern text-to-speech has mostly converged on a shape: a language model predicts discrete speech
tokens from text, and a generative decoder turns those tokens back into a waveform.
[Qwen-Audio-3.0-TTS](https://funaudiollm.github.io/qwen-audio-3.0-tts/) — Alibaba's latest, in the
CosyVoice lineage — runs that shape hard and adds the things that make a TTS model actually usable:
instruction control, inline non-verbal events, 16 languages and 20 Chinese dialect regions,
one-pass long-form synthesis, and 48 kHz output. It currently sits at **#1 on the Artificial
Analysis Text-to-Speech leaderboard**. Here's how it's built.
## The stack
Two models do the heavy lifting. An **autoregressive LM** predicts a sequence of discrete speech
tokens from the text (and, for zero-shot cloning, a reference clip). A **flow-matching decoder**
turns those tokens into a mel-spectrogram. A **vocoder** reconstructs and super-resolves the
waveform to 48 kHz. Fronting all of it is a **12.5 Hz supervised speech tokenizer** — the piece
that quietly sets the model's latency. Click through the stages:
Splitting content (the LM) from voice-and-prosody (the flow-matching decoder) is what makes the
model *instructable*: you can change *what* is said and *how* it's said through different parts of
the system. It's also why the model is robust to a bad reference — a noisy or reverberant prompt
still conditions the LM, and there's no explicit denoising step to break.
## Why 12.5 Hz
The tokenizer's frame rate is the single most consequential number in an autoregressive TTS system,
because the LM emits one token per step and the steps are serial. Fewer tokens per second of audio
means fewer decode steps means less latency. Most neural speech codecs sit at 25–75 Hz;
Qwen's supervised tokenizer runs at **12.5 Hz**. Drag it and watch the step count move:
The word doing the work is **supervised**. A raw reconstruction codec at 12.5 Hz would throw away
too much to sound good; a *supervised* tokenizer is trained to keep exactly the content and speaker
information the LM needs, so it can afford the low frame rate. Short token stream, fast decode,
intact voice — that's the trade the tokenizer is engineered to win.
## Say what, and how
Controllability in most TTS models means "pick a preset voice." Qwen splits it into three
independent knobs. A **free-style natural-language instruction** sets role, emotion, speaking style,
rate, timbre, and accent for the whole utterance. **86 fine-grained inline tags** drop non-verbal
events — laughter, breathing, coughing, sighing — at the word level, inside the text. And the text
itself is the content. Switch the instruction and the delivery re-colors without a word changing:
On top of that: **16 languages and 20 Chinese dialect regions** (seven languages new this version),
**one-pass long-form synthesis up to three minutes**, a reproducible speaker fine-tuning protocol,
and vocoder **super-resolution to 48 kHz**. It also handles hard text-normalization cases and
degraded reference speech without a separate cleanup stage.
## The receipts
The project page leads with two radar charts across the **CV3-Eval** multilingual set — one for
content consistency (word-error rate) and one for speaker similarity — against MiniMax-Speech,
ElevenLabs v3, VoxCPM2, DotsTTS, and the Qwen3-TTS base. The shape tells the story: Qwen-Audio-3.0-
TTS holds a large, even envelope across all ~20 language axes, where the lighter baselines collapse
on the long-tail languages.
The paper reports state-of-the-art results across SEED-TTS-Eval, CV3-Eval, instruction-following,
long-form, and acoustic-robustness suites; the leaderboard #1 is the headline. (Exact WER/SIM
figures live in those radar charts rather than a table on the page — the shape is the claim.)
## The training, briefly
The two-model split has a matching two-track training recipe — **five progressive stages**: the LM
and flow-matching decoder are **pretrained independently**, then **jointly trained** with a
high-quality-data annealing phase, then the LM gets a **reinforcement-learning** pass, and the
decoder gets its own **robustness** stage and then its own **RL** stage. The robustness stage is
what lets the flow-matching decoder cope with degraded prompts; the separate RL passes are what
sharpen intelligibility and speaker fidelity without the two objectives fighting.
## The take
Qwen-Audio-3.0-TTS isn't a new paradigm — it's the LM-plus-flow-matching recipe executed with taste.
The 12.5 Hz supervised tokenizer keeps it fast, the content/voice split keeps it controllable, the
inline tags and instructions make it expressive, and the multilingual coverage is broad and even
rather than English-plus-a-long-tail. The interesting lesson is how much of "good TTS" is now about
the surfaces you expose — frame rate, instruction grammar, tag vocabulary — rather than the core
generative trick, which the field has largely settled.
---
*Source: the [Qwen-Audio-3.0-TTS project page](https://funaudiollm.github.io/qwen-audio-3.0-tts/)
(Alibaba / FunAudioLLM). The radar figures are theirs; the interactive diagrams are mine.*
---
# Cosmos 3: a world model that reasons and generates in one sequence
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/cosmos-world-model
> date: 2026-07-04
> tags: world-models, diffusion, mixture-of-experts, physical-ai, explainer
A language model can tell you what will probably happen if you tip a full mug. A **world model** has to
*show* you — render the next frames, consistent with gravity, contact, and the fact that the coffee ends
up on the table. That is the bet behind **physical AI**: robots and agents need a model that can both
**reason** about what to do and **imagine** the visual consequences of doing it. NVIDIA's **Cosmos 3**
(arXiv 2606.02800) is a world foundation model built around exactly that pairing — and the interesting
part is *how* it fuses the two, not that it does.
The usual way to bolt reasoning onto a generator is to run an LLM, get some text, and feed it to a
diffusion model as a prompt. Cosmos instead puts both in **one network and one token sequence**, so the
generator can read the reasoner's internal state as it works. The mechanism is a **Mixture-of-Transformers
(MoT)**.
## Two towers, one sequence
An MoT keeps **separate weights per modality** but runs **one shared attention** over a single sequence.
Cosmos uses two towers: an **autoregressive Reasoner** that does the planning, and a **diffusion
Generator** that produces video and world frames. Each token in the sequence is routed to its tower's
weights, but they all attend together. Step through it:
The load-bearing detail is the attention pattern. The Reasoner is autoregressive, so its queries attend
only **causally over its own keys** (`K_AR`) — standard next-token reasoning. The Generator is a diffusion
model, and its queries `Q_DM` attend over the **concatenation `[K_AR ; K_DM]`** — the reasoner's keys *and*
its own. That one cross-stream read is the whole idea: the generator can look directly at what the reasoner
decided, so a plan flows into pixels inside a single attention operation rather than through a text
bottleneck. This is the paper's dual-stream joint-attention design:
If you have read our [TwoTower explainer](/articles/nemotron-twotower), the silhouette will look familiar —
a frozen autoregressive tower feeding a diffusion denoiser — but the systems solve different problems.
TwoTower is a *language* model that splits one job (context vs. denoising) to speed up text decode; Cosmos
is a *physical-AI world model* whose two towers span different modalities and are trained jointly, with the
generator reading the reasoner online rather than cross-attending to a frozen copy. Same family tree,
different animal.
Two contrasts pin down what MoT *is*. A **dense** transformer would force one set of weights to be good at
both causal reasoning and bidirectional denoising — the conflict TwoTower also fights. A [Mixture-of-Experts
model](/articles/mixture-of-experts-from-scratch) routes *tokens* to expert FFNs but shares one attention
and one modality regime. MoT is the third option: **separate weights per modality/role (the two towers),
one shared attention** — so the split is by *what kind of token this is*, and the fusion happens in
attention. The diffusion half itself is a standard denoiser; if that machinery is unfamiliar, our
[diffusion](/articles/set-diffusion) and [diffusion-language-model](/articles/illada-diffusion-language-model)
pieces cover it.
## How the world becomes tokens
Before any of that attention can happen, five modalities have to become tokens in one sequence — and
Cosmos encodes each differently. Understanding tokens land in the Reasoner subsequence; generation
tokens land in the Generator subsequence. Trace a modality through its encoder:
The asymmetry is deliberate. The **understanding** path uses a ViT encoder trained *jointly* with the
backbone, so the reasoner sees vision the way its Qwen3-VL ancestor did. The **generation** path uses
*frozen* VAEs — a Wan2.2 video VAE (4× temporal, 32×32 spatial compression) and an audio VAE — so the
diffusion tower only has to produce latents a fixed decoder already knows how to render. And **actions**
from every embodiment collapse into one shared latent action space, which is what lets a single model be
a policy for arms, humanoids, and vehicles alike.
## Planning in pixel space: Action-CoT
Reasoning about *action* is not the same as reasoning in words. To act in the world you need a plan
expressed in terms of *where things move*. Cosmos's **Action-CoT** turns an instruction into a **2D motion
plan on the image plane** — a chain-of-thought drawn as a trajectory — before and while it generates. Pick
an instruction and scrub the plan into existence:
Concretely: the model predicts a path of waypoints across the frame (the gripper's route to the mug, the
block's slide, the drawer's pull), and that trajectory *conditions the diffusion tower*. The frames it
denoises then have to realize the motion, not merely look plausible — the chain-of-thought lives in the
same coordinate space the physics does. It is a neat answer to a real problem: language is a lossy way to
specify a manipulation, and image-plane motion is exactly what a downstream controller can consume.
## One backbone, six models
Because reasoning and generation share a backbone and everything is tokens, the *same* weights become six
different task models just by choosing which modalities go in and which come out. Route it:
The two dynamics models are the ones that matter for physical AI. A **forward dynamics** model predicts the
next video given the past frames and an action — that is the world model as a *simulator* an agent can plan
against. An **inverse dynamics** model recovers the action that connects two frames — useful for learning
control from unlabeled video. Both are the same network with the arrows reversed.
Cosmos comes in three sizes, all built on the dual-tower MoT: **Edge — 4B total on a dense 2B transformer
trained from scratch** (28 layers, a later release), **Nano — 16B total on a dense 8B**, initialized from
**Qwen3-VL-8B** (36 layers), and **Super — 64B total on a dense 32B**, from **Qwen3-VL-32B** (64 layers).
The Qwen initialization is telling — the reasoner tower inherits a strong pretrained VLM, while the generator
is a **flow-matching** diffusion tower (it predicts a constant velocity, `v* = ε − x₀`) grafted on and trained
to read it. Everything is openly released under **OpenMDW-1.1**: the Nano and Super checkpoints, the code,
five synthetic **SDG** datasets (physics, robots, driving, digital humans, warehouses), and the **Cosmos-HUE**
evaluation benchmark.
## The training pipeline
Two towers means two training tracks, joined where it counts. Click through the stages:
The Reasoner is extended from a VLM and fine-tuned on reasoning and Action-CoT data; the Generator is
pre-trained as a flow-matching denoiser, then **mid-trained jointly** with the reasoner — the stage that
actually wires up the dual-stream attention — before splitting into task-specific post-training for
text-to-image, image-to-video, and robot policy. Several headline results lean on **best-of-N sampling
against a learned reward model (WMReward)**, which is worth holding in mind when reading the numbers.
## The numbers, honestly
By the report's own tally, the post-trained models were the **best open-source Text-to-Image and
Image-to-Video models on Artificial Analysis**, and the **best policy model on RoboArena**, at the time of
writing — and Cosmos leads a physical-AI reasoning leaderboard **among open models**. Those are real, but
they're *open-model* rankings, and the scope is easy to lose in a press release. On the reasoning benchmark
where Cosmos Super posts its headline result, a closed frontier model still sits above it:
Read the SOTA claims narrowly. **Best *open* model is not best model** — Gemini 3.1 Pro (77.5) beats Cosmos
Super (73.7) on the reasoning benchmark, and Veo-3.1 leads on audio generation. Any "#1" leaderboard
position is a **dated snapshot** that moves as models ship. The text-to-image comparison uses a
**provider-selected harness** — a setup its authors chose, so treat the framing as favorable. And several
reported results use **best-of-N sampling with Cosmos's own reward model**, not single-shot generation;
that is a legitimate technique but not the same as raw one-shot quality.
None of that makes the work less interesting — it makes the *claim* precise. As an **open**, openly-licensed
omnimodal world model that fuses reasoning and generation in one attention operation and plans in image
space, Cosmos 3 is a genuinely new capability tier for people building on open weights. It just isn't the
best model in the world at everything, and the paper's own numbers say so if you read the parentheses.
## The take
The idea worth keeping is the **dual-stream joint attention**. Most "reasoning + generation" systems chain
two models and pay a text bottleneck between them; Cosmos makes the generator's queries attend over the
reasoner's keys inside one operation, so the plan reaches the pixels without being flattened into a prompt.
Pair that with **Action-CoT** — chain-of-thought as motion on the image plane — and you get a world model
whose reasoning is expressed in the same space its physics has to hold. The Mixture-of-Transformers is the
enabling structure: separate weights per modality, one shared sequence, fusion in attention. Whether the
scoped-SOTA numbers hold as the leaderboards churn is beside the point; the architecture is the
contribution, and it is a clean one.
---
*Built on **NVIDIA Cosmos 3** (arXiv 2606.02800; OpenMDW-1.1 license). The Reasoner/Generator MoT, dual-stream
joint attention, and Action-CoT are described in the paper (Figures 5 and 1); the interactive diagrams are
illustrations of the mechanism. Benchmark figures are quoted as scoped comparisons — best among open models,
with a closed frontier model (Gemini 3.1 Pro) still ahead — and some results use best-of-N with the
model's own reward model.*
---
# The harness effect: orchestration, not the model, sets your agent token bill
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/harness-effect
> date: 2026-07-18
> tags: agents, orchestration, token-economics, llm, explainer
The line I keep hearing is that per-token prices fall every quarter, so agent costs will
sort themselves out. The invoices say otherwise. This paper — [*The Harness Effect*](https://arxiv.org/abs/2607.06906)
from a Writer AI team (Muayad Sayed Ali et al., corresponding author Waseem AlShikh) — runs
the clean experiment I wanted someone to run: hold the model fixed, change only the
orchestration layer around it, and measure the bill. The result is that the orchestration
layer — the *harness* — moves cost per task more than switching between the cheapest and
most expensive model on the menu does.
One caveat up front, because it is load-bearing: the harness under test is Writer's own, and
Writer ran the evaluation. I read the numbers as directional, not as a neutral benchmark. What
makes it worth reading anyway is the mechanism — the paper formalizes *why* orchestration sets
token economics, and the formalization is provider-agnostic.
## Token maxing
The paper names the failure mode first. **Token maxing** is buying capability with tokens:
longer reasoning traces, more agent turns, wider tool payloads, larger replayed contexts — so
that tokens per task grow faster than task value. Falling per-token prices mask the pattern
without fixing it. Total spend rises anyway.
The bill for one agentic task is a sum over its `k` turns:
$$
C = \sum_{i=1}^{k}\left(p_{\text{in}}\,T^{\text{in}}_{i} + p_{\text{out}}\,T^{\text{out}}_{i}\right)
$$
where $p_{\text{in}}, p_{\text{out}}$ are the input/output prices per token and $T^{\text{in}}_i, T^{\text{out}}_i$
the tokens at turn $i$. The input side is the part the orchestration layer builds:
$$
T^{\text{in}}_{i} = \underbrace{S_i}_{\text{system}} + \underbrace{H_i}_{\text{history}} + \underbrace{G_i}_{\text{tool schemas}} + \underbrace{R_i}_{\text{retrieval}} + \underbrace{U_i}_{\text{user turn}}
$$
Here is the trap. If every turn replays the full transcript, the history term $H_i$ grows with
$i$, so the cumulative input over a task grows as $O(k^2)$. A harness that compacts and caches
history keeps it near $O(k)$. The gap between those two curves is spend that buys no quality —
and because the per-token price is falling the whole time, the total keeps climbing quietly.
Drag the horizon and watch it happen:
## The bill, and the one price that actually moves
The lever the harness pulls hardest is **prompt caching**. Providers serve a previously seen
prompt prefix from cache at roughly a tenth of the base input rate. If a fraction $h$ of input
tokens are cache reads billed at multiplier $\kappa$, the effective input price is
$$
p^{\text{eff}}_{\text{in}} = p_{\text{in}}\left(1 - h\,(1-\kappa)\right), \qquad \kappa \approx 0.1
$$
so a harness that keeps $h$ near 1 pays about a tenth of list price on the dominant input term.
The point the paper makes well: $h$ is not a model property and not a provider favor. It is a
function of how byte-stable your prompt prefix is across turns — which is set entirely by the
orchestration layer. On an identical-prefix call the harness served **99.9% of prompt tokens as
cache reads** (7,876 of 7,886). That is the whole game: shape the prompt so the expensive term
is almost always a cache hit.
## The controlled swap
The experiment is deliberately boring, which is why it is convincing. Twenty-two locked
evaluation tasks. Six foundation models — Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5,
Qwen 3.6, GLM 5.1, Palmyra X6. Each model runs the tasks twice: once under a frozen conventional
production loop, once under the Writer Agent Harness. Nothing else changes — same tasks, same
judges, same price table. Only the orchestration layer swaps. Flip it:
Blended across all six models and 22 tasks, replacing the loop with the harness cuts cost per
task 41% (`$0.21` → `$0.12`), median wall-clock 44% (48s → 27s), and tokens per task 38%
(14.2k → 8.8k) — with task-completion quality at parity (0.78 → 0.81, directional at this
sample size).
Two derived numbers make the parity concrete. Quality per dollar rises 82%. And throughput —
task-completions per million tokens — nearly doubles:
## Everyone gets cheaper
The efficiency win is not a quirk of one model. Under the swap, **every** model's cost and
latency fall — cost by 33% to 61%, latency by 33% to 55%. The effect is a property of the
orchestration layer, not of any model.
## Harness leverage
Now the finding that made me want to write this up. Efficiency is model-invariant, but the
**quality** gain from the same orchestration upgrade is not — it scales almost perfectly with a
model's baseline strength. Plot each model's mean quality gain against its baseline capability
and the points fall on a line: $r = 0.99$ over $n = 6$. The paper calls it **harness leverage**.
Stronger models extract more from the same harness. Scrub the models:
The honest edge of this: across 48 capability×model cells, 30 improve, 11 are flat, and **7
regress — all of them in the three smaller models**, concentrated in orchestration-heavy
capabilities (tool use over MCP, playbooks, presentations). Qwen 3.6 comes out net negative on
quality (−0.031). It is still 44% cheaper. So the harness is a strict efficiency win everywhere,
and a quality win that grows with the model you point it at.
## Six mechanisms behind the effect
The paper decomposes the harness into six mechanism families. None of them is exotic — they are
the unglamorous orchestration glue, which is exactly why they are easy to leave on the table:
1. **Cache-shape discipline — the two-zone prompt.** A byte-stable prefix (tool-schema catalog,
stable system prompt, append-only transcript) carries the provider's cache breakpoints;
everything volatile is confined to a tail that is rebuilt each turn and structurally excluded
from caching. This is what pushes $h$ toward 1 in the effective-price equation.
2. **Structured, incremental, cache-aware compaction.** Shrink history without breaking the
cache prefix — compact the middle, keep the front byte-identical.
3. **Context offload.** Tool outputs land in a store the model can reference, not in the prompt.
Tokens the model never pays to re-read.
4. **Zero-token waiting; durability as economics.** Durable execution so a pause, retry, or
long-running tool call does not replay the whole context to resume.
5. **Failure-spend governance.** Cap what a failing or looping run can burn before it is stopped.
Most runaway bills are failures, not successes.
6. **A model-agnostic floor.** The five above set an efficiency floor under *any* model — which
is what makes the savings a property of the layer, not the checkpoint.
## How other harnesses compare
The paper also scores six widely used agent systems on the same axes — vendor-integrated
clients, orchestration libraries, multi-agent conversation frameworks, and open personal
harnesses — from public documentation rather than head-to-head runs. The pattern is that most
frameworks implement some mechanisms and leave the rest "to the application to build and budget."
Cache-shape discipline and failure-spend governance are the two most often missing, and they are
two of the biggest levers. Treat that table as a design-time source study, not a measurement.
## What it is worth at fleet scale
The reason this matters past a single task: the per-task delta multiplies by volume, and by every
model you run. Apply the blended cost gap to monthly task volume and at **one million agent tasks
per month the harness is worth about `$90k`/month over the baseline — `$1.08M`/year** — and the
gap widens linearly with volume. An organization does not run one model; it runs a fleet, present
and future. The harness is the one component whose efficiency multiplies across all of them.
**Read the caveats.** (1) The sample is small — **22 tasks, 6 models**. The quality deltas are
directional at this size; the paper says so and calls its statistical posture "suggestive."
(2) It is the **vendor's own harness, evaluated by the vendor** (Writer), against a "frozen
conventional production loop" the vendor defined — a reasonable baseline, but not a neutral one.
(3) It is a **single workload**. The mechanisms generalize in principle; the exact 41% / 44% /
38% numbers are this task set, these price tables, these six models. (4) The `$0.21` → `$0.12`
and fleet-scale figures ride on current provider cache pricing ($\kappa \approx 0.1$); change the
price table and the arithmetic moves.
## The take
Strip the framing and the useful claim is narrow and testable: for agentic workloads, the
orchestration layer is a first-class cost object, and most of the cost is in prompt shape, not
model choice. The effective-input-price equation is the part I will actually use — it says the
expensive input term is a cache hit if and only if your prompt prefix is byte-stable across
turns, and that is an engineering property you control. Efficiency came out model-invariant
(every model 33–61% cheaper); quality came out capability-dependent (r = 0.99 with baseline
strength). I would want an independent harness and a second workload before trusting the exact
percentages. But the direction matches what I see in production: the token bill is set less by
which model you picked and more by how you assemble the context you hand it.
---
*Source: "The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise
Agentic AI" (Muayad Sayed Ali et al., Writer AI, 2026) —
[arXiv 2607.06906](https://arxiv.org/abs/2607.06906). Figures 1, 3, 4, and 6 are reproduced from
the paper for commentary. Benchmark numbers are quoted as reported; the interactive diagrams
illustrate the mechanisms and use the paper's headline values.*
---
# Diffusing blame: credit assignment under Dale's principle
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/diffusing-blame
> date: 2026-07-17
> tags: deep-learning, credit-assignment, biologically-plausible, backpropagation, theory, explainer
**Diffusing Blame** asks a sharp question: can a network learn useful representations while obeying the one rule real brains never break? That rule is **Dale's principle** — a neuron is either excitatory or inhibitory, and *every* synapse it makes carries that one sign. Backpropagation quietly violates a stack of constraints like this. The paper builds a network that respects them, trains it with a rule called **Error Diffusion**, and asks how far it gets. The answer: **96.7% on MNIST**, a **61.7% baseline on CIFAR-10**, and reinforcement-learning agents that hold their own against a backprop-free baseline — all while enforcing Dale's principle strictly. This is a walk through why that is hard, what the rule actually does, and the numbers, honestly labelled.
Everything below is from Yamada, Grillotti, Charakorn, Risi, Ha and Lange, *Diffusing Blame: Task-Dependent Credit Assignment in Biologically Plausible Dual-Stream Networks* (2026). Every accuracy, return, and ablation delta is **author-reported** on their own runs; I have not reproduced them. The two interactive widgets are my illustrations of the mechanism, not measured traces.
## Why backprop can't run in a brain
Backpropagation is the reason deep nets learn, and it is also the reason nobody thinks the brain runs it. Look at the backward pass. To update layer $\ell$, backprop needs the error signal $\delta_\ell$, and it computes it by pulling the layer above's error back through the transposed forward weights:
$$
\delta_\ell = \big(W_{\ell+1}^{\top}\,\delta_{\ell+1}\big) \odot \phi'(z_\ell).
$$
Read that literally as a circuit and three problems fall out.
- **Weight transport.** The backward pass uses $W_{\ell+1}^{\top}$ — the *same* forward weights, transposed. A synapse would have to read the value of the forward synapse it feeds and reuse it, exactly, on the way back. There is no known biological mechanism for a synapse to know its partner's weight.
- **Sign symmetry.** Because it is the same weight, the feedback carries the same *sign* as the forward connection. Forward and backward paths are locked together.
- **A separate error channel.** The $\delta$'s have to travel back through a network that is distinct from the forward one, layer by layer, without disturbing the forward activations.
**Dale's principle** makes all of this worse. In cortex a neuron's outgoing synapses are uniformly excitatory or uniformly inhibitory — the sign belongs to the neuron, not the synapse. A weight in a standard net has no such loyalty: one unit can push one target up and pull another down in the same step. Toggle between the two regimes and click a source neuron to see its fan-out:
Enforcing Dale's principle means the sign is frozen per neuron, so a net has to split into separate excitatory (E) and inhibitory (I) populations and coordinate them. That changes credit assignment at the root: you can no longer flip a weight's sign to fix an error, and the tidy $W^{\top}$ feedback path is off the table anyway. Prior biologically plausible rules — feedback alignment and its kin — dodge weight transport by sending the error back through a *fixed random* matrix $B$ instead of $W^{\top}$. That helps, but it trades one implausible object (the transpose) for another (a dedicated random feedback matrix), and historically these rules stall out past MNIST. The paper wants neither the transpose nor the random matrix.
## Dale's principle, in the weights
The architecture is dual-stream. Each layer carries a positive activation vector $\mathbf{p}$ and a negative one $\mathbf{n}$, and there are **four** weight matrices between consecutive layers — the within-stream pair $W_{pp}, W_{nn}$ and the cross-stream pair $W_{np}, W_{pn}$:
$$
\mathbf{p}_i = \phi_i\!\big(\mathbf{p}_{i-1} W_{pp} - \mathbf{n}_{i-1} W_{np} + \mathbf{b}_p\big), \qquad
\mathbf{n}_i = \phi_i\!\big(\mathbf{n}_{i-1} W_{nn} - \mathbf{p}_{i-1} W_{pn} + \mathbf{b}_n\big).
$$
The trick is in the signs. **Every learnable weight is constrained non-negative**, $W_{\bullet\bullet} \ge 0$ element-wise, and the minus signs in front of the cross-stream terms are *hardcoded into the wiring*, not learned. So the E stream always excites and the I stream always inhibits, structurally — Dale's principle holds by construction, and gradient descent can never sneak a sign flip past it. A readout that needs a signed output just subtracts the streams: $\hat{y} = y^{+} - y^{-}$.
This is a real cost. A standard dense layer is one matrix; this is four non-negative matrices with a fixed sign pattern, and the optimizer has to move the whole coordinated E/I system in lockstep. The question the paper answers is whether a learning rule can drive that system without any of backprop's illegal moves.
## Diffusing the error instead of transporting it
Error Diffusion's answer: don't route the error *back through the layers* at all. Route it *directly to every layer*. Take the output error $S$ (shape $B \times C$ for a batch of $B$ over $C$ classes) and broadcast it to the hidden units through a fixed routing matrix $M$, then form each layer's local update from presynaptic activity and the postsynaptic nonlinearity's derivative:
$$
R = S\,M^{\top}, \qquad U_p = \phi'(Z_p) \odot R, \qquad \Delta W_{pp} \propto A_p^{\top}\,U_p.
$$
$R$ is the routed error drive, $U_p$ scales it by the local activation slope $\phi'$, and the weight change is an outer product of presynaptic activations $A_p$ with $U_p$. No $W^{\top}$ appears anywhere. No random feedback matrix appears either — the routing $M$ is a fixed, structured broadcast, not a learned or random projection.
The original Error Diffusion was defined for binary classification. To go past that, the paper adds **modulo error routing**: hidden unit $i$ is assigned to output channel
$$
r(i) = i \bmod C,
$$
and learns from that channel's error $s_{r(i)}$. It is coarse — several hidden units share a channel, and unit $C$ wraps back to channel 0 — but it is deterministic, transport-free, and enough to spread class-specific blame across a wide hidden layer. Step through the forward pass, the diffusion, and the update, and flip between backprop and Error Diffusion to see the paths diverge:
The contrast is the whole point. Backprop's blame crawls back one layer at a time, each hop paying the $W^{\top}$ transport tax. Error Diffusion drops the error onto every hidden unit at once and lets each layer compute a local update. That is what makes it plausible — and also what makes it approximate, since a modulo-routed broadcast is a much blunter credit signal than the exact gradient.
## Three fixes that turn it into a learner
Error Diffusion out of the box does not learn much — the seed configuration lands at **50.4% on MNIST** and **11.6% on CIFAR-10** (barely above chance on ten classes). Three domain-specific fixes close most of the gap.
**Layer-specific sigmoid widths.** The activation is a temperature-controlled sigmoid,
$$
\phi_i(z) = \frac{1}{1 + e^{-2z/\alpha_i}},
$$
with a per-layer width $\alpha_i$. Why it matters: the update is scaled by $\phi'$, and a standard sigmoid's derivative is tiny once units saturate. The paper measures a **25x attenuation** of the surrogate gradient from the output down to the first hidden layer, so the early layers barely move. Widening the sigmoid (larger $\alpha$) keeps the derivative alive deeper in the stack. Their CIFAR-10 setup uses $\alpha = 3.0$ for convolutional layers and $\alpha = 6.0$ for fully connected ones; MNIST uses $\alpha = 6.0$ throughout.
**Batch-centered class error.** Instead of feeding raw one-vs-all errors, the class error is centered across the batch,
$$
\tilde{E}_{b,c} = E_{b,c} - \frac{1}{B}\sum_{b'} E_{b',c},
$$
so every class's error signal is zero-mean over the mini-batch. This removes a constant per-class bias that would otherwise push all units in a channel the same way regardless of the input.
**Asymmetric E/I initialization.** The excitatory matrices $W_{pp}, W_{nn}$ are scaled up by $1.5\times$ at init and the inhibitory $W_{np}, W_{pn}$ scaled down by $0.5\times$, a starting excitation-to-inhibition ratio of roughly **3:1**. That gives the network net-positive drive to begin with, and the paper shows the ratio relaxes toward a biological-like balance as training proceeds.
## What it scores
On the standard benchmarks, the constrained network learns — not to backprop's level, but well past chance, and well past unconstrained biologically plausible baselines that stall on MNIST. Direct Feedback Alignment (DFA), the backprop-free baseline that still uses a random feedback matrix, sits a few points ahead as the reference ceiling:
The most interesting result is not the headline accuracy — it is what the ablations reveal. Remove each fix and measure the accuracy drop, and the ranking **reverses between the two datasets**:
| removed component | MNIST Δ (pp) | CIFAR-10 Δ (pp) |
|---|---|---|
| layer-specific sigmoid widths | **−71.4** | −15.1 |
| batch-centered class error | −0.3 | **−47.9** |
| asymmetric initialization | +0.0 | −5.5 |
On MNIST the whole model lives or dies by the sigmoid widths — pull them and accuracy craters by 71 points, while the batch-centering does almost nothing. On CIFAR-10 it is the exact opposite: batch-centered class error is load-bearing (−47.9), and the widths matter far less. Same architecture, same rule — but the credit-assignment bottleneck is *task-dependent*, and a single-benchmark evaluation would have hidden that entirely. That is the paper's sharpest point: which fix carries the model is a property of the task, not the method.
## Into RL: ED-PPO
Classification is the easy setting — the error signal is a clean label. To test the rule where credit assignment is genuinely hard, the paper drops Error Diffusion into **PPO**, replacing the backprop gradient through the hidden layers of both the policy and value networks (the PPO objective still supplies the output-level error). Policy errors route to hidden units by action channel; value errors broadcast to all units. On Brax continuous control, ED-PPO is competitive with DFA and, on HalfCheetah, clears backprop:
That HalfCheetah result — ED-PPO at 5494±691, essentially matching DFA-PPO's 5581±359 and beating backprop's 3520±485 — is the strongest single number in the paper, but it does not generalize cleanly. On Humanoid, ED-PPO (6670±2592) trails backprop (8478); on the open-ended exploration task **Craftax**, ED-PPO edges out DFA-PPO (19.8±1.5 return) but sits below BP-PPO. The honest read is "competitive with the backprop-free baseline, still short of backprop on the hardest tasks" — which is exactly what the abstract claims, and worth stating plainly rather than cherry-picking HalfCheetah.
Keep the scale in view. These are small networks on MNIST, CIFAR-10, Brax and Craftax — not a scaling result, and not close to state of the art. Backprop still wins on accuracy on every classification task here (97.6% DFA and higher for standard backprop vs 61.7% on CIFAR-10), and beats ED on the harder RL environments. The contribution is not a better optimizer; it is a demonstration that representation learning is *possible at all* under strict Dale's principle, without weight transport or random feedback matrices — plus the finding that the binding constraint shifts with the task. Read it as biology-motivated evidence, not a drop-in replacement for backprop.
## The take
The idea is clean and the framing is honest. Real neural circuits obey constraints backprop ignores — a synapse can't read its partner's weight (no weight transport), and a neuron can't flip signs synapse by synapse (Dale's principle). Build a network that respects both, and credit assignment stops looking like a transpose and starts looking like a broadcast: Error Diffusion drops the output error straight onto every hidden layer, routes it by a modulo rule, and updates each layer locally. Three fixes — wider per-layer sigmoids to fight a 25x gradient decay, batch-centered class error, and a 3:1 excitation-to-inhibition initialization — are what turn a 50%-on-MNIST seed into a 96.7% learner and a 61.7% CIFAR-10 baseline. None of that is state of the art, and the paper doesn't pretend otherwise. What it earns is a real claim: you can learn representations under the brain's actual wiring rules, the gap to backprop is a few points rather than a chasm, and — the part I'll remember — *which* trick matters most depends on the task, a bottleneck you only see if you test on more than one benchmark.
---
*Built on Y. Yamada, L. Grillotti, R. Charakorn, S. Risi, D. Ha and R. T. Lange, [Diffusing Blame: Task-Dependent Credit Assignment in Biologically Plausible Dual-Stream Networks](https://arxiv.org/abs/2606.31700) (arXiv 2606.31700, 2026). Figures 1 and 2 are reproduced from the paper for commentary. The `DaleNetwork` and `ErrorDiffusion` widgets are my own illustrations of the mechanism, not measured traces; all accuracies, returns, and ablation deltas are author-reported and I have not independently reproduced them.*
---
# Intern-S2: a 397B model that reads the raw page
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/intern-s2
> date: 2026-07-17
> tags: llm, multimodal, scientific-ai, mixture-of-experts, explainer
Intern-S2-Preview-397B, from InternLM (Shanghai AI Lab), is a 397-billion-parameter
multimodal foundation model built for one thing: science. Not "science" as a benchmark
category bolted onto a general chatbot — science as the training objective, down to how
the model tokenizes a molecule and how it reads a figure off a paper.
The headline is a shape, not a single number. On general knowledge, math, and agentic
coding, Intern-S2 sits at **frontier parity** — a 397B model trading blows with much
larger closed systems, usually a hair behind. On specialized science — multi-omics,
molecular reasoning, material generation, protein-binder design — it **leads every
frontier model**, often by 4× or more. That gap is the whole story, and it's the payoff
of three specific design choices.
## The lineage
Intern-S2 is the third step in a line, and each step is worth naming because Intern-S2
inherits all of it:
- **Intern-S1** — a 235B mixture-of-experts on a Qwen3 backbone plus a 6B InternViT
vision encoder, continuously pretrained on 5T tokens, more than half of it scientific.
- **Intern-S1-Pro** — scaled to a trillion-parameter MoE with 512 experts and 8 active
per token, added Fourier Position Encoding (FoPE) and explicit time-series modelling.
- **Intern-S2-Preview-397B** — the most capable of the family. (There is also a
lightweight **Intern-S2-Preview-35B**, continued-pretrained from Qwen3.5 with a
shared-weight multi-token-prediction head, a KL loss, and chain-of-thought compression.)
The MoE math from S1-Pro is the standard sparsity trade. With $k$ of $N$ experts firing
per token,
$$
\theta_{\text{active}} \;=\; \frac{k}{N}\,\theta_{\text{expert}} \;+\; \theta_{\text{shared}},
\qquad k = 8,\; N = 512,
$$
so a trillion-parameter model pays for only ~22B activated parameters per token. FoPE and
time-series modelling let it ingest sequences from $10^0$ to $10^6$ points — the kind of
range a seismograph or a mass spectrometer actually produces.
## Three ideas that matter
Strip away the scale and Intern-S2 is three ideas working together:
1. A **vision-language pretraining paradigm** that learns directly from raw pages of
scientific literature — no OCR-and-parse preprocessing step in front of the model.
2. A **dynamic tokenizer** that natively represents molecular formulas, protein
sequences, and seismic signals as meaningful units rather than subword debris.
3. Large-scale **multi-task reinforcement learning** across more than 20 scientific
domains, trained jointly, which also happens to sharpen general reasoning.
Take them in order.
## Reading the raw page
A conventional document pipeline flattens a page to a string before the model ever sees
it: OCR recovers the words, a layout parser guesses the reading order, and the figures
and equations are dropped on the floor. The text model then learns from a transcript that
has already thrown away the thing you care about — how *this* curve relates to *that*
caption and *that* variable.
Intern-S2 skips the transcript. It "learns directly from raw pages of scientific
literature, jointly modelling symbolic semantics and visual relationships in a shared
representation space without intermediate parsing." The vision encoder maps text, figures,
and equations into one representation, and both a symbolic-semantics head and a
visual-relations head read off that same space.
The consequence is that a plot and the sentence that references it are learnable as a
single object, not two disconnected streams. For scientific literature — where the
argument often *lives* in the figure — that is the difference between a model that reads
the paper and one that reads a description of the paper.
## The dynamic tokenizer
A tokenizer is a vocabulary learned on a corpus, and a standard BPE vocabulary is learned
on natural-language text. Hand it a SMILES string or a protein sequence and it splits
where its merge statistics say to split — which has nothing to do with where the *meaning*
is. The aromatic ring in aspirin gets smeared across three tokens; a run of amino acids
gets merged into a chunk that no longer corresponds to any residue.
Intern-S2's dynamic tokenizer emits scientifically-meaningful units directly: an atom, a
bond, a residue, a waveform sample each become a token the model can address. This is not
cosmetic. If a residue's identity is spread across a token boundary, the model can't attend
to that residue cleanly — the representation is fighting the tokenizer. Native tokenization
is what lets Intern-S2 treat a molecular formula, a protein, or a time series as a
first-class input instead of a string that happens to look like one.
## Multi-task RL across 20+ domains
The last piece is post-training. Intern-S2 runs large-scale reinforcement learning across
more than 20 scientific domains **jointly**, rather than fine-tuning a separate model per
task. Training the domains together is what gives the model its leading general-reasoning
scores as a side effect: the same optimization that teaches it multi-omics and material
chemistry also rewards careful, multi-step reasoning that transfers.
It deploys on the usual high-throughput stacks — **LMDeploy, vLLM, and SGLang** — with a
256K-token context for text reasoning and 64K tokens for multimodal input. It's genuinely
strong at generative science: biomolecular interaction design and material-structure
generation, not just question answering.
## The benchmarks
Here is the shape, in one chart. Flip between the two task families and watch Intern-S2's
dot move from *just behind* the best competitor to *far ahead* of it.
### General tasks: frontier parity
On general benchmarks Intern-S2 rarely wins outright, but it rarely loses by much — which
is the remarkable part for a 397B model standing next to the largest closed systems. It
posts MMLU-Pro 89.75 (Gemini-3.1-Pro leads at 91.00), HMMT-2026 91.57 (GPT-5.5 at 97.06),
MMMU-Pro 80.46, and SWE-Bench-Multilingual 81.67 — effectively tied with GLM-5.2's 82.00.
On factual recall it clearly clears the open field even where it trails the closed leader —
SimpleQA-Verified is a good example:
The honest read: it trails the very top closed models on the hardest coding and knowledge
benches — TerminalBench 2.1 67.42 vs Claude-Opus-4.8's 84.60, SWE-Bench-Pro 61.56 vs
69.20. Parity, not conquest.
### Scientific tasks: dominance
Now the inversion. On specialized science the gaps stop being fractions of a point and
start being multiples. Biology-Instructions (multi-omics) is the clearest case: Intern-S2
scores 56.92 where the next-best frontier model manages 13.87, and most models land between
4 and 10.
Material-structure generation tells the same story: MP20 67.88 against a next-best of
16.75, with most models between 1.5 and 16. Molecular reasoning (Mol-Instructions 52.37 vs
GPT-5.5's 40.49), remote sensing (XLRS-Bench 51.97), microscopy VQA (MicroVQA 68.81), and
biomolecular interaction design (ProteinBinder-9 4.36 vs a best competitor near 2.4) all
land the same way. Roughly:
$$
\frac{56.92}{13.87} \approx 4.1\times, \qquad \frac{67.88}{16.75} \approx 4.05\times.
$$
It is not a clean sweep, and that's worth saying: on MolecularIQ, GPT-5.5 still leads
(76.41 vs Intern-S2's 61.49). But across the science suite as a whole, a 397B model beats
GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.8 — the payoff of the raw-page pretraining, the
dynamic tokenizer, and multi-domain RL compounding.
The caveats are real. This is a **Preview**, not a final release. On the hardest general
coding and knowledge benchmarks it still trails the top closed models. Several of the most
lopsided scientific wins — MP20, ProteinBinder-9 — are **internal benchmarks**, so treat
the exact multiples as InternLM's own measurement until third parties reproduce them. And
at 397B it is heavy to self-host: frontier-scale hardware, not a workstation.
## What I make of it
- **The specialization is the product.** Most "science" models are general chatbots with
a domain fine-tune. Intern-S2 pushes science into the tokenizer and the pretraining
objective, and the benchmark gaps show the difference that makes — 4× is not a
prompt-engineering delta.
- **Parity at 397B is the quiet achievement.** Matching Gemini-3.1-Pro and Opus-4.8 on
general tasks with a fraction of the (public) scale, while dominating science, is a
stronger statement than any single scientific score.
- **Trust the shape, verify the numbers.** The parity-vs-dominance pattern is convincing
and mechanistically motivated. The internal-benchmark wins want independent replication
before I'd quote the exact multiples as settled — but even halved, the lead holds.
---
*Sources: the [Intern-S2-Preview-397B model card](https://huggingface.co/internlm/Intern-S2-Preview-397B)
and the [Intern-S1 project](https://github.com/InternLM/Intern-S1) (InternLM / Shanghai AI
Lab). Benchmark numbers are quoted as reported on the model card; several scientific
benchmarks are internal.*
---
# LIMSSR: scoring actions when modalities go missing at training time
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/limssr
> date: 2026-07-17
> tags: multimodal, missing-modality, llm, action-quality-assessment, explainer
**Action Quality Assessment (AQA)** is the task of watching a video of an action — a
dive, a figure-skating program, a gymnastics routine — and predicting a *numeric quality
score* a judge would give it. The strong systems are multimodal: they read RGB frames,
optical flow, and sometimes audio, because form, motion, and rhythm all carry signal.
Which is fine until a modality goes missing. And in the real world, modalities go missing
not just at test time but *during training* — a dataset where some clips never had audio,
or flow was never computed. That is the case **LIMSSR** (Xu, Wu, Ke, and Peng; ICML 2026
spotlight) is built for, and it is meaningfully harder than the usual setup.
## Why training-time missingness is the hard case
Most missing-modality work makes a quiet assumption: you trained on complete data, learned
what a "normal" audio or flow stream looks like, and only lose one at inference. Then you
can reconstruct the missing stream, or distill a complete-modality teacher into an
incomplete-modality student.
Take the missingness back into training and both tricks weaken. You cannot reconstruct a
distribution you never fully saw, and there is no clean complete-modality teacher to distill
from if the training set itself is full of holes. The score is a scalar, so a wrong
imputation does not announce itself — it just quietly biases the number. To make "how much
does missingness hurt" precise, AQA is graded by **Spearman's rank correlation** $\rho$
between predicted and ground-truth scores: 1.0 is a perfect ordering of performances, 0 is
chance. It is a ranking metric, so it punishes exactly the systematic bias a bad guess
introduces.
## Impute in prompt space, don't reconstruct
LIMSSR's first move is to stop reconstructing. Each modality gets a frozen feature
extractor and a small projection; a missing modality is not zero-filled but replaced with a
**special token**, and the model is handed a **text prompt that names which modalities are
present and which are gone**. The LLM is then asked to *infer the missing modality's latent
role from the context that remains* — role inference, not feature reconstruction.
That distinction is the whole point. A zero-filled slot drags a fixed-slot fusion model
toward a wrong answer; a described-absence lets a language model reason about what the gap
means. Drop a modality on each side and watch the gap open:
The audio-only case is the tell. A naive fusion model, fed zeros for the two missing
streams, collapses to $\rho = 0.177$ — barely above chance. LIMSSR, told in words that
video and flow are gone and asked to reason from audio alone, holds $\rho = 0.687$ on the
same split.
## The pipeline
End to end: per-modality features → specific projection → **prompt-guided context-aware
modality imputation** (special tokens + the modal-condition prompt) → an **LLM-driven
multidimensional representation fusion** that packs everything into fusion tokens → a frozen
**large language model with LoRA** doing the sequence-to-score reasoning → a mask-aware
aggregation head that emits the score.
## Mask-aware dual-path aggregation: don't trust a lucky guess
Reasoning about a missing modality invites a failure mode: the model confidently
hallucinates the part it cannot see. LIMSSR's aggregation head is built to suppress exactly
that. It runs **two paths** off the same missingness mask $m$:
- **Cross-modal pattern recovery** — cross-attention, gated weighting, and a *learnable
confidence*. Strong when modalities are present, shaky when they are not.
- **Uncertainty-calibrated reasoning** — role-aware weighting and mask-aware refinement that
explicitly discounts low-confidence dimensions, so it degrades gracefully.
A learnable-confidence gate blends the two. As modalities drop, the recovery path has little
left to cross-attend over, its confidence falls, and the gate shifts weight onto the
calibrated path — the one that already distrusts what it cannot verify. Toggle the modalities
and watch the gate swing:
That shift is the anti-hallucination mechanism. On FS1000 the full gate cuts mean-squared
error from **18.18** (simple fusion) to **14.08**, at $\rho = 0.789$.
## Results
Across every available/missing combination, LIMSSR's predicted scores track the diagonal —
the ground truth — more tightly than the multimodal-expert baselines (MoMKE, MCMoE), and it
holds up in the settings where they fall apart:
The starkest number is the hardest split — audio only, the two visual streams gone:
With every modality present the margin is smaller but still real ($\rho$ 0.891 vs 0.819),
which is the shape you want: a method that helps most exactly where the problem is hardest,
and does no harm when the data is complete.
Keep the scope in view. (1) This is **AQA**, a narrow regression task on relatively small
datasets (FS1000 and friends), not a general multimodal benchmark — the gains are real but
domain-specific. (2) The interactive pipelines here are **illustrative**; the $\rho$ and MSE
values are the paper's reported FS1000 numbers, but the gate dynamics I animate are a
simplification. (3) Bolting a LoRA-tuned LLM onto a scoring head **adds parameters and
latency** versus a lightweight fusion model — you are paying for the reasoning that buys the
robustness.
## The take
The reframing is the idea worth keeping. Missing-modality learning has mostly been treated as
a *reconstruction* problem — rebuild the pixels or features you lost. LIMSSR treats it as a
*reasoning* problem: describe the absence in language, let an LLM infer what the missing
stream would have contributed, and gate the answer by how much you can trust it. That it
works under training-time missingness — the case reconstruction and distillation both
struggle with — and lifts audio-only $\rho$ from 0.177 to 0.687 is a good argument that, for
messy real-world multimodal data, telling the model what it is missing beats trying to fake
what it lost.
---
*Source: "LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete
Multimodal Observations" (Huangbiao Xu, Huanqi Wu, Xiao Ke, Yuxin Peng), ICML 2026 spotlight.
Numbers ($\rho$, MSE) are the paper's reported FS1000 results; the interactive diagrams
illustrate the mechanism.*
---
# Monolith 1.0: a 1.6T open MoE built for reasoning
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/monolith-1-0
> date: 2026-07-17
> tags: llm, mixture-of-experts, reasoning, long-context, explainer
Basalt Labs' [Monolith 1.0](https://huggingface.co/basaltlabsai/monolith-1.0) is a
**1.572-trillion-parameter** Mixture-of-Experts with **49.5B active per token**, released as
open weights under an MIT license — weights, tokenizer, and eval harness, commercial use
allowed. It is a decoder-only, reasoning-focused, Chinese-English model, and Basalt is blunt
about what it is for: "a 1.6T open Mixture-of-Experts foundation model for reasoning at scale."
The [tech report](https://basaltlabs.org/monolith) lays out the recipe.
The headline number is the total parameter count, but the number that governs everything else is
the **active** one. Monolith spends 1.57T parameters of capacity but pays for only 49.5B of
compute per token — a **32x sparsity** ratio. This piece walks the pieces that make that work: the
MoE routing, the two-stage context extension to a million tokens, the training recipe, and the
decode trick that makes a model this large servable. Full prose and math carry each idea; the
interactive diagrams are there to build intuition.
## The spec sheet
- **1.572T** total parameters, **49.5B** active per token — a **32x** sparsity ratio.
- **80 layers**, model dimension **8,192**.
- **Grouped-query attention**: 64 query heads, 8 KV groups, head dim 128.
- **128 routed experts** per layer (SwiGLU, intermediate dim 6,144) **+ 1 shared expert**; **top-2** routing.
- **RoPE base 5e6**, two-stage YaRN; context length **1,048,576** tokens.
- Byte-level BPE tokenizer, **151,936** vocab.
Two design choices do most of the work: how the experts are routed, and how the context is
stretched. Take them one at a time.
## Routing: top-2 of 128, plus one that never sleeps
Monolith's feed-forward block is a mixture of experts. Each layer holds **128 routed experts** and
**one shared expert**. For every token, a router scores the 128 and keeps the **top-2**; the shared
expert is always on. So **3 of 129** experts fire per token. Scrub the token index and watch the
routed pair swing while the shared expert stays lit:
The always-on shared expert is the part worth dwelling on. A pure top-$k$ router has to relearn the
common, every-token computation inside many different experts, which wastes capacity. Splitting off
one shared expert lets the routed experts specialize while the shared one carries the baseline work —
a pattern that has become standard in large MoEs because it stabilizes routing at high sparsity.
The sparsity is what makes the size affordable. Active parameters are the shared expert plus the
top-2 routed, so the per-token compute is that of a roughly 50B model while the knowledge capacity
is that of a 1.57T one:
$$
\text{sparsity} = \frac{N_{\text{total}}}{N_{\text{active}}} = \frac{1.572 \times 10^{12}}{4.95 \times 10^{10}} \approx 32\times
$$
That 32x is the whole bet: you get the memory footprint of a trillion-scale model but the FLOPs of a
mid-size one, provided you can route well and keep every expert busy.
## Attention: grouped-query, so the 1M cache fits
Before the context trick, one attention detail matters for it. Monolith uses **grouped-query
attention** — 64 query heads sharing just **8 KV groups**. The KV cache stores keys and values per
group, not per query head, so it is **8x smaller** than full multi-head attention at the same width.
At a million-token context the KV cache is the dominant memory cost of decoding, so an 8x reduction
there is the difference between a 1M window being a spec and being something you can actually hold in
memory.
## Long context: two YaRN stages to a million tokens
Monolith pretrains at a cheap **4,096-token** window, then extends to **1,048,576** tokens — a
**256x** stretch — using **YaRN** in two stages on a RoPE base of **5e6**. YaRN rescales the rotary
position frequencies so positions far beyond the training length stay in distribution instead of
aliasing into nonsense. Doing it in two 16x steps rather than one 256x leap keeps long-range
attention coherent. Step through the stages:
The extension factor is exactly
$$
\frac{1048576}{4096} = 256,
$$
and a two-stage split puts the midpoint at the geometric mean, $\sqrt{256} = 16$, so each stage is a
16x reach: $4096 \to 65536 \to 1048576$. (The 65,536 midpoint is my illustration of a
clean two-stage split; Basalt reports two YaRN stages without pinning the intermediate length.) The
reason to stage it is numerical: RoPE extrapolation degrades faster than linearly with the extension
factor, so two moderate stretches with a re-anchor in between hold up where one aggressive stretch
smears the attention over distant tokens.
## Training: 60T tokens, and a FLOP budget that checks out
Monolith is trained on **60T tokens** (a multilingual mixture) in **BF16 mixed precision with an
FP32 optimizer state**. Basalt reports a compute budget of about **1.8e25 FLOPs**. That number is
not arbitrary — the standard estimate for transformer training compute is
$$
C \approx 6 \, N_{\text{active}} \, D,
$$
with $N_{\text{active}}$ the active parameters and $D$ the token count. MoE training compute scales
with the **active** parameters, not the total, because only the active experts run per token. Plug in
$N_{\text{active}} = 49.5\text{B}$ and $D = 60\text{T}$:
$$
6 \times (4.95 \times 10^{10}) \times (6.0 \times 10^{13}) \approx 1.78 \times 10^{25}\ \text{FLOPs},
$$
which lands on the reported 1.8e25. The sparsity pays off twice: at 32x it means a 1.57T model trains
at the per-token cost of a ~50B one.
Post-training is a three-stage reasoning pipeline: **SFT**, then **DPO**, then **RLVR** — reinforcement
learning with verifiable rewards. RLVR is the reasoning-specific piece: instead of a learned reward
model, the reward comes from checking whether the answer is actually correct (a math result that
verifies, code that passes tests), which is a cleaner signal for training long chains of thought and
harder to reward-hack than a preference model.
## Serving it: self-speculative decoding
A 1.57T-parameter model is memory-bound at decode time, so Monolith ships **self-speculative
decoding**: the model drafts several tokens cheaply, then verifies them all in one forward pass,
keeping the longest correct prefix and correcting the first miss. Scrub the phase and flip the
domain:
Because the verify pass reproduces the base model exactly, this is **lossless** — the output is token
for token what greedy decoding would have produced, just fewer expensive passes to get there. Code is
more predictable than prose, so more drafts survive verification: Basalt reports **~2.1x** faster
decoding on natural language and **~2.7x** on code.
Even so, "open weights" here still means rack-scale hardware. Basalt targets **one GB300 NVL72 rack
(72 GPUs)** at FP8, or a **CloudMatrix-384** at native BF16. You can download the weights; running
them is another matter.
## The benchmarks
Here is where honesty has to lead. On Basalt's own harness, at maximum thinking effort, Monolith
posts numbers that are not just ahead of the field but near the ceiling of the tests themselves.
Compared against GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, and Kimi K2.6:
Read those bars, then read the next box before you form an opinion.
**These are single-lab, self-reported, own-harness numbers — treat them as directional, not settled.**
A **99.4** on Humanity's Last Exam and a **100.0** on AIME 2025 are effectively saturation: the model
is not beating the field by a few points, it is sitting at the top of the scale while every frontier
model on the same chart trails by 30 to 40 points. Numbers that clean from the lab that built the
model and ran the eval are exactly the ones to be skeptical of. Own-harness, maximum-effort results
select for the configuration that flatters the model; they say nothing about a neutral setup, a
contaminated-test check, or a different prompt. The thing to wait for is **independent third-party
evaluation** on held-out sets. And the practical caveat does not go away with better scores: "open
weights" for a 1.57T model still means a **GB300 NVL72 rack** to run it, so open here is a licensing
fact, not an accessibility one.
## What I make of it
- **The engineering is coherent and legible.** 32x sparsity via top-2-of-128 plus a shared expert,
an 8x-smaller KV cache from grouped-query attention, a staged YaRN reach to 1M tokens, and a FLOP
budget that checks out against $6 N_{\text{active}} D$ — none of it is exotic, all of it is the
right lever for a trillion-scale reasoning model. The self-speculative decode is a real, lossless
serving win.
- **The MIT license is the genuinely useful part.** Weights, tokenizer, and eval harness, commercial
use allowed — that is more open than most models at this scale, and it means the benchmark claims
can, in principle, be checked by anyone with the hardware.
- **The benchmarks are the part to hold at arm's length.** Saturated, self-reported, own-harness
scores are a marketing artifact until someone independent reproduces them. I would love to be
wrong; I would rather wait for the third-party numbers than quote 99.4 as if it were settled.
The bet Monolith makes is that a trillion-scale open MoE, routed and staged carefully, can be a
frontier reasoning model in the open. The architecture is a credible version of that bet. Whether it
actually reasons at 99.4-on-HLE levels is a question its own harness cannot answer.
---
*Sources: the [Monolith 1.0 model card](https://huggingface.co/basaltlabsai/monolith-1.0) and the
[Basalt Labs tech report](https://basaltlabs.org/monolith) (architecture, training, deployment,
benchmarks). Benchmark numbers are quoted as reported by Basalt on their own harness at maximum
thinking effort. The training-compute figure is checked against $C \approx 6\,N_{\text{active}}\,D$
with the reported 49.5B active parameters and 60T tokens. The interactive diagrams illustrate the
mechanisms; the routing, context, and decode visuals are illustrative, and the 65,536-token YaRN
midpoint is my own clean two-stage split, not a disclosed intermediate length.*
---
# VideoChat3: a 4B video model that watches longer for less
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/videochat3
> date: 2026-07-17
> tags: multimodal, video-understanding, vision-language, efficiency, explainer
VideoChat3, from Nanjing University, Shanghai AI Lab, NTU, and Peking University
([arXiv 2607.14935](https://arxiv.org/abs/2607.14935),
[HF](https://huggingface.co/papers/2607.14935)), is a **4B-parameter video
multimodal LLM** with one clear thesis: you should not have to pick between a model
that generalizes across video types, a model that is cheap to run on long clips, and
a model you can actually reproduce. Most video MLLMs give you one of the three. This
one aims for all three, and ships the training data to prove the last.
The interesting part is not the leaderboard — it is *how* a 4B model stays coherent
over 2048 frames without the token count exploding. Two mechanisms carry that: an
inflated 3D tokenizer, and adaptive frame resolution. I'll walk both, then the
numbers.
## The problem: tokens grow with time
A video is frames, and every frame is a few hundred vision tokens. Feed a clip in
frame-by-frame and the sequence length grows linearly with duration — a few thousand
tokens for a short clip, tens of thousands for a long one. That is what makes long
video expensive: the LLM pays quadratic attention over a token budget set by the
tokenizer, and the tokenizer, if it treats each frame independently, has no reason to
be frugal.
Two failure modes fall out of that. Compute blows up on long clips. And a 2D image
tokenizer, bolted on frame-by-frame, never models *motion* — it sees a stack of
stills. VideoChat3 attacks both at the tokenizer.
## I3D-ViT: inflating a 2D tokenizer into 3D
The first move is the **Inflated 3D Vision Transformer (I3D-ViT)**. Start from a 2D
image encoder — a plain patch-and-attend ViT — and *inflate* it into a spatiotemporal
one instead of training a video encoder from scratch. The inflation is three steps:
1. **Chunk the frames.** Group `T = 4` consecutive frames into a chunk.
2. **Attend within the chunk.** Run self-attention across the chunk's tokens, so
spatial and temporal structure are modelled together — motion, locally.
3. **Pool the chunk down.** A temporal pooling (×4) plus a 2×2 spatial merge collapse
each chunk to a compact, motion-aware token slot.
Temporal ÷4 times spatial ÷4 is a **16× spatiotemporal compression** — the token
budget grows with the frame count, but 16× slower than the naive per-frame path.
Drag the frame count and watch where the tokens go:
The reason this works without wrecking accuracy: the compression happens *after* the
chunk has been attended, so motion is already encoded into the surviving tokens. You
are not throwing frames away — you are summarising each 4-frame window into one slot
that remembers what moved. Native resolution and aspect ratio are preserved through
absolute temporal embeddings, so the model still knows *when* each token happened.
## Adaptive frame resolution: spend pixels where the evidence is
Compression handles the *count* of tokens. The second move handles the *cost per
frame*. In a streaming setting most of a video is uneventful — a static room, a
held shot, dead air. Processing every frame at high resolution spends the same budget
on the boring frames as on the one that answers the question.
So VideoChat3 conditions the per-frame resolution on state. Routine moments are
perceived under a low **224²-pixel** quota. When a *Standby* cue fires — the signal
that an answer is about to appear — the following window is enlarged to a **448²**
quota, roughly 4× the tokens, to catch the detail. Click frames to promote them and
watch the budget move:
The framing is a three-state stream: **Silence** (nothing to report, low-res),
**Standby** (something is coming, stay ready), **Response** (answer now, high-res).
The budget follows the state instead of the clock. On a stream that is mostly
silence, that is most of the frames spent at a quarter of the token cost.
## The benchmarks
The headline: at **4B parameters**, VideoChat3 beats comparable open models
(Qwen3-VL-4B, Molmo2-4B) across temporal perception, long video, reasoning, temporal
grounding, and online tasks — the paper's cross-benchmark sweep is one figure:
Where it separates most is **temporal grounding** — answering *when* something happens,
not just *what*. Over Qwen3-VL-4B the gains run from **+9.7** (Charades) up to **+20.6**
(VUE-TR V2 in the TimeLens suite), depending on the benchmark. Charades makes the point:
The same ordering holds on ActivityNet (54.8 / 48.2 / 39.8) and QVHighlights
(67.1 / 58.7 / 58.7). It is not just grounding, though — the general video and
reasoning benchmarks land ahead too, if by smaller margins:
VideoMME 70.1 (vs 69.3 / 69.6), LVBench 56.7 (vs 56.2 / 53.9), TempCompass 75.6 (vs
70.8 / 72.8), StreamingBench 83.0 (vs 80.2). The deltas on the general suites are
single digits; the grounding deltas are where the tokenizer design shows up.
## The efficiency payoff
The point of 16× compression is latency, and it compounds with length. Same
hardware, same clip, VideoChat3 vs Qwen3-VL, end-to-end inference:
| Frames | Qwen3-VL | VideoChat3 |
|---|---|---|
| 512 | 3.84s | 3.60s |
| 1024 | 12.25s | 8.10s |
| 2048 | 44.45s | **20.41s** |
At 512 frames the gap is small — the tokenizer overhead is a rounding error. At 2048
frames it is **2.2×**: `44.4s → 20.4s`. The compression buys you the frames the
grounding benchmarks reward, at a latency that stays usable as the clip grows.
## Fully open: the data, not just the weights
The "fully open" claim is the part I'd flag to anyone who has tried to reproduce a
video MLLM. The bottleneck is never the architecture — it is the ~3M-sample
instruction mix nobody publishes. VideoChat3 releases three:
- **VideoChat3-Academic2M** — 2.27M caption/QA instances from six academic sources,
with evidence-grounded annotation enhancement.
- **VideoChat3-LV116K** — 116.2K long-form samples, mean durations 156s to ~1.3K
seconds.
- **VideoChat3-OL617K** — 617,183 streaming/online instances across 40 shards.
Trained through a four-stage curriculum: tokenizer pretraining → video-language
alignment → general instruction tuning → long/streaming tuning. Weights and data
both out, so the recipe is checkable end to end.
Read the comparison for what it is: a **4B-vs-4B** result. VideoChat3 beats *comparable
open* models at its size — it is not claiming to beat frontier closed video systems or
much larger open ones, and the general-suite margins (VideoMME +0.5 to +0.8) are
inside the range where mix and eval harness matter. The token math here is
illustrative (I use ~64 spatial tokens/frame to keep the diagrams honest about
*ratios*, not absolute counts); the 16× compression, the 224²/448² quotas, and the
latency numbers are the paper's. Adaptive resolution has a real failure mode too: set
the Standby threshold too tight and a fast event is only ever seen in 224².
## What I make of it
- **The tokenizer is the whole story.** I3D-ViT is a clean idea — inflate a 2D
encoder, attend within short chunks, pool 16×. It is *why* a 4B model can watch 2048
frames in 20 seconds, and *why* the temporal-grounding gaps are as large as they are.
Motion survives the compression; that is the trick.
- **Adaptive resolution is the right shape for streaming.** Spend the budget on the
evidence, not the clock. It maps cleanly onto a Silence/Standby/Response state
machine, and the savings are largest exactly where video is cheapest to skimp — the
dead air.
- **Open data is the contribution that outlasts the benchmarks.** Numbers age; a
released 3M-sample video instruction mix is something the rest of the field can build
on. For a model whose pitch is "generalist *and* reproducible," shipping
Academic2M + LV116K + OL617K is the part that makes the claim real.
---
*Source: "VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video
Understanding," Li, Zhu, Zeng, Dong, Wu, et al.
([arXiv 2607.14935](https://arxiv.org/abs/2607.14935)). Benchmark values read from the
paper's reported figures and tables; numbers quoted as reported. Interactive diagrams
are my own illustration of the mechanism — token counts are illustrative, ratios are
the paper's.*
---
# ZUNA 1.1: a channel-agnostic EEG foundation model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/zuna-1-1
> date: 2026-07-17
> tags: eeg, foundation-model, diffusion, signal-processing, neuroscience, explainer
ZUNA 1.1, from Zyphra, is an EEG foundation model: a 380M-parameter transformer
encoder–decoder diffusion autoencoder that reconstructs missing channels, denoises
corrupted ones, and upsamples sparse montages to electrode positions that were never
recorded. It is Apache-2.0, runs on a consumer GPU or a plain CPU, and — the part I
find interesting — it is **channel-agnostic**: the same weights read a 4-electrode
consumer headband or a 256-channel research cap, because it treats every electrode as
a token at a physical coordinate, not a fixed slot in a montage.
There is no arXiv paper; this is a release across the [Zyphra blog](https://www.zyphra.com/our-work/zuna1.1),
the [Hugging Face model card](https://huggingface.co/Zyphra/ZUNA1.1), and the
[GitHub repo](https://github.com/Zyphra/zuna) (`pip install zuna`, plus a browser-based
Cloud EEG Playground). It is a point release on ZUNA1 — same 380M parameters, a bigger
and cleaner training corpus, and a handful of design changes that matter more than the
version bump suggests.
## The problem: no two EEG setups agree
EEG is a mess to model across datasets, and the mess is structural, not incidental.
A clinical 10-20 montage has 19 electrodes; a sleep lab might run 6; a research cap
runs 64, 128, or 256; a consumer headband runs 4. The electrodes sit at different
scalp locations, sample at different rates, and are filtered differently. Worse, within
a single recording, channels die, drift, or go noisy for part of a session and recover
later. Most models paper over this by fixing a channel list and re-referencing every
recording onto it — which throws away recordings that do not fit and cannot exploit
the extra electrodes when they exist.
ZUNA's bet is to stop treating a recording as a fixed-width tensor and start treating
it as a **set of tokens, each tagged with where and when it was measured**. Once the
model keys on physical position instead of channel index, the montage stops mattering.
## Channel-agnostic: an electrode is just a coordinate
The mechanism is a 4D rotary positional encoding. Each 0.125-second segment of one
electrode becomes a token, and its position is the tuple $(x, y, z, t)$: the
electrode's 3D coordinate on the scalp plus a coarse time index. Attention applies
rotary phases over all four axes, so "nearby" means nearby in space *and* time — the
model learns that Cz and C3 covary because they are physically close, not because they
happen to be channels 10 and 9 in some file.
That single choice is what buys channel-agnosticism. There is no learned per-channel
embedding to run out of, so an arbitrary subset or superset of electrodes is just a
different set of coordinates. Drop a region and the decoder fills those coordinates
from the electrodes it still has; ask for coordinates that were never recorded and it
predicts them the same way. Switch montages below and watch the same model span a
headband and a dense cap:
## Inside the model: a diffusion autoencoder
The architecture is two transformer stacks. The **encoder** reads the clean context
and compresses it to a latent. The **decoder** cross-attends to that latent and
produces the signal at the requested coordinates. The latent is injected into every
decoder layer through **adaptive-RMS norm** — the latent sets the per-layer scale of
the normalization, which is a cheap, stable way to condition a deep stack on a global
summary (the same conditioning trick diffusion image models use for the timestep).
The decoder is trained with a **rectified-flow** objective rather than a plain
regression loss, and this is the right call for reconstruction. Filling a missing
channel is genuinely uncertain — many signals are consistent with the surrounding
scalp — so a model trained to minimize mean-squared error returns the blurred average
of all of them. A generative decoder instead returns a *sample* from the plausible
set. Rectified flow makes that sampling cheap: it learns a straight-line transport from
a noise draw to the data. With noise $x_0$ and target signal $x_1$, the interpolant and
its velocity are
$$
x_\tau = (1-\tau)\,x_0 + \tau\,x_1, \qquad \frac{dx_\tau}{d\tau} = x_1 - x_0,
$$
and the decoder regresses a velocity field $v_\theta(x_\tau, \tau)$ onto that constant
velocity $x_1 - x_0$. At inference you draw $x_0$ and integrate the field from
$\tau=0$ to $\tau=1$. Because the target path is a straight line, few integration steps
get you most of the way — which is why this runs on a CPU. Drag the scrubber:
The same decoder does all three jobs — reconstruct a missing channel, denoise a noisy
one, upsample to a new position — and only the conditioning changes. That is the payoff
of framing everything as "predict the signal at these coordinates given those."
## Training: corruption on purpose
The corpus roughly doubled over ZUNA1, from about 2M to **3.5M channel-hours** of
public EEG. Two things about how it was prepared are worth noting. First, quality is
scored **per channel, per second**, so a channel that is clean for most of a session
and noisy for a stretch is used where it is good instead of being dropped whole.
Second, each recording is kept in two filter variants — a bandpass at 0.1–45 Hz and a
minimally processed version (0.01 Hz high-pass plus a notch for line noise) — so the
model sees both heavily and lightly filtered signal. Inputs are variable length, 0.5 to
30 seconds, bucketed into four bins so short clips are not wasted padding a long window.
The interesting part is that the model is trained to reconstruct under four distinct
corruption patterns, not one. This is the whole reason it generalizes to messy
real-world recordings:
Whole-channel dropout teaches it to rebuild a dead electrode from its neighbors.
Full-time dropout — a gap across every channel at once — teaches temporal inpainting.
Channel-time dropout is the realistic case: a cluster of electrodes goes bad for a
window. Random-uniform scatter mimics muscle artifacts and momentary failures. Because
training mixes all four, the model handles almost arbitrary space-time masks at
inference, which is exactly what `reconstruct_fif()` exposes — it auto-detects MNE bad
channels and `BAD_` annotations and repairs them.
## Results: reconstruction as channels drop
The metric is normalized mean-squared error between the held-out true signal and the
reconstruction,
$$
\mathrm{NMSE} = \frac{\lVert \hat{x} - x \rVert_2^2}{\lVert x \rVert_2^2},
$$
where 1.0 is the trivial "predict zero" baseline and lower is better. The baseline to
beat is MNE's spherical-spline interpolation, the classical way to rebuild a missing
electrode from a smooth fit over the others. Zyphra publishes the comparison as
figures, not tables, so the numbers below are read off the plots and are approximate.
At 20% channel dropout everything is close — NMSE around 0.4 to 0.6 across the four
datasets, because with most electrodes present even a spline does fine. Push to 90%
dropout and the spline blows up (Berlin BCI reaches roughly 2.7, i.e. worse than
predicting silence) while ZUNA1.1 and ZUNA1 hold near 1.0 to 1.5. That widening gap is
the headline: the learned prior degrades gracefully as information disappears; the
classical interpolant does not.
ZUNA1.1 versus ZUNA1 is, honestly, a wash — ZUNA1.1 is a touch better on ANPHY-Sleep
and BCI2000 and marginally behind on Berlin BCI and AAD. That matches Zyphra's own
claim: better or essentially equal NMSE at the same 380M parameters, with the real
gains going to stability and the broader input regime rather than raw accuracy.
The more realistic test deletes a whole brain region and rebuilds it from the other
seven:
Here the two learned models track each other closely and both crush the spline in
frontal and temporal regions (spline around 0.8 to 1.0 NMSE, ZUNA around 0.35 to 0.6).
The three converge only over parietal cortex, where the field is smooth enough that a
spline is already a reasonable prior. Central electrodes are the easiest — every model
does well — because they are surrounded by neighbors on all sides.
## Where it breaks
A reconstruction is a generative prior, not a measurement. The model fills a missing
channel with signal that is plausible *given the rest of the scalp* — which is
precisely wrong when the thing you care about is a focal event that only the missing
electrode would have seen. For a sleep-staging or BCI pipeline that leans on spatial
redundancy, that is fine. For reading a possible focal spike off a dead electrode, a
low NMSE can hide a confidently hallucinated normal trace. Denoising and upsampling
carry the same caveat: the output is the model's best guess at a signal that is
*consistent*, not the signal that was actually there.
Two more honest limits. The evaluation is four datasets and F32 weights; generalization
past those recording conditions is asserted, not shown. And rectified-flow decoding is
iterative — CPU inference works, and it is cheap because the transport path is straight,
but latency still scales with how many sampling steps you take, so "runs on a CPU" and
"real-time" are not the same claim.
## The take
- **The reframing that pays off is spatial.** Making position, not channel index, the
thing the model keys on is the whole idea, and 4D RoPE over $(x, y, z, t)$ is a clean
way to do it. It is the same move that made vision transformers resolution-flexible,
applied to the scalp — and it is what lets one set of weights span a headband and a
256-channel cap and interpolate to electrodes it never saw.
- **The diffusion-autoencoder choice fits the problem.** Reconstruction is genuinely
uncertain, so a generative decoder that samples a plausible signal is more honest
than a regressor that returns the blurred mean. Rectified flow keeps that sampling
cheap enough to run without a GPU.
- **It is built to be used, not admired** — Apache-2.0, `pip install zuna`, a browser
playground, and an MNE-friendly `reconstruct_fif()` entry point. The win over
classical interpolation is decisive; the win over ZUNA1 is a tie, and Zyphra says so.
Both are worth saying out loud.
---
*Sources: the [ZUNA 1.1 release](https://www.zyphra.com/our-work/zuna1.1), the
[Hugging Face model card](https://huggingface.co/Zyphra/ZUNA1.1), and the
[GitHub repo](https://github.com/Zyphra/zuna). Figures are from Zyphra's release; NMSE
values are read off the published plots and are approximate, since exact tables were
not released. Released 2026-07-16 under Apache-2.0.*
---
# GRAPE: RoPE, ALiBi, and FoX are the same construction
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/grape-position-encoding
> date: 2026-07-16
> tags: llm, attention, positional-encoding, transformers, explainer
Self-attention has no idea what order its tokens came in — permute the sequence and the raw
attention scores are unchanged. So every transformer bolts on a *positional encoding*, and over the
years these have multiplied into a small zoo of unrelated-looking tricks. [Rotary embeddings
(RoPE)](/articles/attention-mechanisms) rotate queries and keys by angle-per-position. **ALiBi**
subtracts a linear penalty proportional to how far apart two tokens are. The **Forgetting Transformer
(FoX)** multiplies in a per-token forget gate. Each is derived on its own terms, with its own
intuition, and the folklore treats them as competing families: multiplicative *phase* versus additive
*bias*.
**GRAPE** — *Group Representational Position Encoding*, from Princeton, UCLA and Tsinghua's IIIS
(ICLR 2026) — makes the deflationary claim that all of them are the **same construction seen through
different generators**. A position $n$ acts on the query/key space through one group action,
$$\mathbf{G}(n) = \exp(n\,\omega\,\mathbf{L}),$$
and *which kind of matrix* you put in the generator $\mathbf{L}$ decides the family. A rank-2
skew-symmetric $\mathbf{L}$ gives a **rotation**, and RoPE falls out exactly. A rank-1 nilpotent
$\mathbf{L}$ gives a **shear** that injects a linear bias, and ALiBi and FoX fall out exactly. Scrub a
position below and flip the generator to watch the two behaviours emerge from one law:
The payoff of the unification is not a new record — it's a *design space*. Once RoPE is "the rotation
with the canonical basis and a log-uniform spectrum," you can ask what the *learned* basis does; once
ALiBi is "the rank-1 unipotent action with a fixed slope," you can ask what a *content-dependent* slope
does. GRAPE names and tries both.
## One law, two generators
The organizing principle is that a positional map should respect an **exact relative law**:
$$\mathbf{G}(t-s) = \mathbf{G}(s)^{-1}\,\mathbf{G}(t), \qquad \mathbf{G}(n+m) = \mathbf{G}(n)\,\mathbf{G}(m).$$
This is what makes attention translation-invariant: if you apply $\mathbf{G}(i)$ to query $i$ and
$\mathbf{G}(j)$ to key $j$, the score $\tilde{\mathbf{q}}_i^\top\tilde{\mathbf{k}}_j =
\mathbf{q}_i^\top\mathbf{G}(j-i)\mathbf{k}_j$ depends only on the offset $j-i$, never on absolute
position. Any one-parameter subgroup $\mathbf{G}(n) = \exp(n\mathbf{L})$ satisfies it automatically —
so the whole design question collapses to *which generator $\mathbf{L}$ to exponentiate*. GRAPE
identifies the two generator types that keep the exponential cheap and the geometry clean:
- **Rank-2 skew** $\mathbf{L} = \mathbf{a}\mathbf{b}^\top - \mathbf{b}\mathbf{a}^\top \in
\mathfrak{so}(d)$ exponentiates to an **orthogonal** map — a norm-preserving rotation in
$\mathrm{SO}(d)$. This is *Multiplicative GRAPE*.
- **Rank-1 nilpotent** $\mathbf{A}$ with $\mathbf{A}^2 = \mathbf{0}$ exponentiates, in one term, to a
**unipotent** map $\mathbf{G}(n) = \mathbf{I} + n\omega\mathbf{A}$ in the general linear group $\mathrm{GL}$
— a shear that translates a feature and shows up as an additive logit bias. This is *Additive GRAPE*.
## Multiplicative GRAPE: RoPE is a rotation with a fixed basis
Build the generator from two vectors $\mathbf{a},\mathbf{b}\in\mathbb{R}^d$. With
$\alpha = \|\mathbf{a}\|^2$, $\beta = \|\mathbf{b}\|^2$, $\gamma = \mathbf{a}^\top\mathbf{b}$ and
$s = \sqrt{\alpha\beta - \gamma^2}$, the rank-2 skew $\mathbf{L}$ squares to
$\mathbf{L}^2 = -s^2\,\mathbf{P}_{\mathcal{U}}$ on the plane $\mathcal{U} = \mathrm{span}\{\mathbf{a},\mathbf{b}\}$.
That single fact collapses the matrix exponential to a **Rodrigues-type closed form**:
$$\exp(\mathbf{L}) = \mathbf{I} + \frac{\sin s}{s}\mathbf{L} + \frac{1-\cos s}{s^2}\mathbf{L}^2,$$
a pure rotation by angle $s$ inside the plane $\mathcal{U}$, computable in $O(d)$ flops with no matrix
ever materialized. Stack $d/2$ of these on disjoint coordinate pairs with frequencies $\theta_i$ and
they commute, so
$$\mathbf{G}(n) = \prod_{i=1}^{d/2}\exp(n\theta_i\mathbf{L}_i)
= \mathrm{blockdiag}\big(\mathbf{R}_2(n\theta_1),\dots,\mathbf{R}_2(n\theta_{d/2})\big).$$
That block-diagonal of $2\times2$ rotations *is* RoPE — the paper's Proposition 3.1 states RoPE is
**exactly** commuting multi-subspace GRAPE-M with the canonical coordinate pairs and a log-uniform
spectrum. What GRAPE adds is the freedom RoPE gives up: the planes and spectrum can be **learned**
(commuting subspaces at $O(d)$ per head), or you can allow a compact **non-commuting** mixture (at
$O(rd)$ per head) so different feature subspaces can *couple* — geometry the fixed RoPE basis cannot
express.
## Additive GRAPE: ALiBi and FoX are shears in a lifted space
To get an *additive* bias out of a *multiplicative* group, GRAPE uses the classic trick of a
**homogeneous lift**: augment $\mathbf{x}\in\mathbb{R}^d$ to $\hat{\mathbf{x}}\in\mathbb{R}^{d+k}$ and
work in $\mathrm{GL}(d+k)$ with a nilpotent generator. Because $\mathbf{A}^2 = \mathbf{0}$, the
exponential is just $\mathbf{G}_{\mathrm{add}}(n) = \mathbf{I} + n\omega\mathbf{A}$. With an asymmetric
lift $\hat{\mathbf{q}}_i = [\mathbf{q}_i;1;0]$, $\hat{\mathbf{k}}_j = [\mathbf{k}_j;0;1]$ and the rank-1
generator $\mathbf{A}_h = -\beta_h\,\mathbf{e}_{d+2}\mathbf{e}_{d+1}^\top$, the score becomes
$$\hat{\mathbf{q}}_i^\top\,\mathbf{G}_{\mathrm{add},h}(j-i)^{-\top}\,\hat{\mathbf{k}}_j
= \mathbf{q}_i^\top\mathbf{k}_j + (j-i)\,\beta_h,$$
which is **exactly ALiBi** with head slope $\beta_h$. The nilpotent structure is not decoration: it is
what guarantees the exact relative law and clean streaming (cache the rotated keys once). GRAPE then
generalizes the *slope*. Replace the constant $\beta_h$ with non-negative softplus gates on the query
and key, and the bias becomes **content-dependent**:
$$\tilde{\mathbf{q}}_i^\top\tilde{\mathbf{k}}_j
= \mathbf{q}_i^\top\mathbf{k}_j + (j-i)\,\omega\big[\mathrm{softplus}(\mathbf{v}^\top\mathbf{q}_i/\sqrt{d})
+ \mathrm{softplus}(\mathbf{u}^\top\mathbf{k}_j/\sqrt{d})\big].$$
This is **GRAPE-A-QK**: a learnable, content-adaptive linear bias derived from first principles rather
than hand-set per head. Drag the gate below to see the fixed ALiBi head-fan give way to a
content-driven slope:
The **Forgetting Transformer** falls out of the same picture. FoX's per-token forget gates accumulate
a bias $b_h(t,j) = \sum_{\ell=j+1}^{t}\log f_{\ell,h}$, which is precisely a *path product* of unipotent
factors $\prod_\ell(\mathbf{I} + \log f_{\ell,h}\,\mathbf{E}) = \mathbf{I} + b_h(t,j)\,\mathbf{E}$. So
FoX is an exact instance of **Path-Integral Additive GRAPE (GRAPE-AP)** — the endpoint-dependent
version that keeps row-wise composition and prefix-sum streaming.
## The whole map
Put the pieces together and every named scheme is a leaf on one tree. Click through them — each is
either *recovered exactly* by a specific generator or sits just past a known method as a GRAPE
*extension*:
## Does the extra freedom help?
Here is where the honesty starts. GRAPE is validated at **small scale**: 353M and 770M models trained
on 50B tokens of FineWeb-Edu, context length 4,096, in a nanoGPT/Llama-style setup, evaluated 0-shot on
a standard NLU suite (ARC, HellaSwag, OBQA, PIQA, WinoGrande, SciQ). The training curves are close, but
GRAPE's additive variants hold a persistent small edge, and the authors note RoPE showed a training
instability at 770M that GRAPE did not:
On downstream average, the ordering is consistent but the margins are small. For the 353M models,
GRAPE-AP (path-integral) is the best of the eight variants, edging FoX and ALiBi, with plain RoPE last:
The 770M models tell the same story — GRAPE-AP first, RoPE last — again by roughly a point:
**Read the wins narrowly.** (1) *Small scale, standard benchmarks.* Everything is 353M/770M on 50B
tokens at 4K context, on ARC/HellaSwag-style tasks — there are **no long-context or
length-extrapolation experiments** (no RULER, no retrieval), which is striking given the paper motivates
itself with long-context and ALiBi's extrapolation. (2) *The rotational story didn't pay off
empirically.* The Multiplicative variants that generalize RoPE — GRAPE-M-ctx/nonctx — actually
**underperform RoPE** on the large models (54.7–54.8 vs 55.76 avg); all the downstream gains come from
the **additive** family, so the framework's practical dominance rests on the ALiBi/FoX side, not the
rotation side. (3) *Margins are ~0.3–1.5 average points* over strong baselines, and the ranking flips
under the KV-shift setting: with KV-shift enabled, FoX edges GRAPE-AP at 770M (57.09 vs 56.86). (4)
*No efficiency measurements.* The $O(d)$/$O(rd)$-per-head costs are stated, not timed — there are no
wall-clock or FLOP comparisons. (5) Baselines are the authors' own reimplementations; there is no
comparison to tuned production models. The contribution is the **unifying theory and design space**,
lightly validated — not a demonstrated accuracy or efficiency SOTA.
## The take
GRAPE's real product is conceptual compression. Positional encoding stops being a list of tricks and
becomes a single knob — *which generator do you exponentiate?* — with RoPE, ALiBi and FoX as three
specific settings and a labelled space of alternatives (learned rotation bases, non-commuting mixtures,
content-gated slopes, path-integral biases) in between. That is genuinely clarifying, and the exact
recoveries are proved, not hand-waved: RoPE as commuting rank-2 rotations, ALiBi as a rank-1 unipotent
action, FoX as its path integral. What the paper does *not* yet show is that the new freedom the map
opens up buys much at scale — the strongest empirical variant is a modest improvement on the *additive*
side, the rotation-generalizing side trails plain RoPE, and the long-context claims the framing invites
go untested. As a theory it's a clean unification worth knowing; as a recipe, GRAPE-AP is a small,
honest win over FoX-style biases, and the rest is an invitation to experiment.
---
*Built on [Group Representational Position Encoding](https://arxiv.org/abs/2512.07805) (Zhang, Chen,
Liu, Qin, Yuan, Xu, Yuan, Gu, Yao; Princeton / UCLA / Tsinghua IIIS, ICLR 2026). Equations, tables and
figures are quoted from the paper (353M and 770M models, FineWeb-Edu, 0-shot lm-evaluation-harness);
the interactive diagrams are illustrations of the mechanism, not measured data. Related reading:
[how attention works](/articles/how-transformers-attention-works),
[a tour of attention mechanisms](/articles/attention-mechanisms),
[MiniMax Sparse Attention](/articles/minimax-sparse-attention), and
[how LLM inference works](/articles/how-llm-inference-works).*
---
# Bonsai 27B: a 27B model at 1.125 bits, small enough for a phone
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/bonsai-27b
> date: 2026-07-15
> tags: quantization, inference-optimization, on-device, multimodal, explainer
Most of the work that makes a large language model *usable* on your own hardware is not a
better model — it's a smaller one that behaves like the big one. **Bonsai 27B**, from PrismML,
is a clean example: it takes a full-precision **Qwen3.6-27B** and re-encodes its weights at
close to one bit each, ending up small enough to load inside a phone's memory budget. The
intelligence is Qwen's. What Bonsai contributes is the **extreme low-bit representation** —
run end to end, not just on the easy layers — and the custom kernels that make it fast.
The headline is the collapse in size. A 27B model at FP16 is 54 GB; Bonsai ships two quantized
variants, and the smaller one is **3.9 GB** — "27B-class capability at a footprint smaller than
a full-precision 2B model," as PrismML puts it. Flip between the precisions:
## Two encodings, close to one bit each
The two variants differ only in how each weight is stored. The **ternary** model uses the
three-value set `{−1, 0, +1}` with an FP16 scale shared across a group of weights — that works
out to **1.71 effective bits per weight** and a 5.9 GB model. The **1-bit** model drops the
zero, storing `{−1, +1}` plus the group scale, for **1.125 bits** and 3.9 GB. (The theoretical
floors are log₂3 ≈ 1.58 and 1.0 bits; the group-wise scales are the small overhead on top.)
Everything else — the hybrid-attention architecture, the 262K-token context window, the
Apache-2.0 license — is inherited from the base.
## The hard part: every block, not just the MLPs
Quantizing a transformer to a couple of bits is not new. What usually happens is that the
*sensitive* parts — the token embeddings, the attention projections, the LM head — are kept at
higher precision, and only the big feed-forward MLPs get squeezed. That protects quality, but
it also means the footprint only partly shrinks: a model is not small until its embeddings and
head are small too. Bonsai's claim is that the low-bit representation "runs end to end across
the language network, embeddings, attention, MLPs, and the LM head," with a compact **4-bit
vision tower** alongside. Toggle between the two philosophies:
Pushing 1-bit weights through the parts everyone else keeps in FP16 is exactly where accuracy
usually falls off a cliff — which is why the interesting question is not the size, but what it
costs. This is the *inference-time* mirror image of [native FP4 training](/articles/nemotron-nvfp4):
there the goal was to keep the math stable during training while deliberately holding some layers
higher precision; here the goal is to serve an already-trained model with nothing held back. If
you want the mechanics of why low-precision inference is memory-bound in the first place, the
[how LLM inference works](/articles/how-llm-inference-works) piece sets that up, and
[TurboQuant](/articles/turboquant-kv-cache) covers the complementary problem of quantizing the KV
cache rather than the weights.
## What survives — and what doesn't
Here is the honest part, and it's the part a size-and-speed announcement tends to bury. PrismML
reports that ternary keeps **~95%** of full-precision quality and 1-bit keeps **~90%**, averaged
over a 15-benchmark suite in thinking mode. Both averages check out against their own table — but
the average hides a wide spread. Pick a category and watch the three precisions, then read the
per-category retention strip:
Math is remarkably robust: the 1-bit model holds ~96% of the full-precision score. But
**agentic tool-calling** falls from 80.0 to 66.0 and **vision** from 72.6 to 59.6 — roughly 82%
retention each, nearly a fifth of the capability gone. Long, multi-step tool use and multimodal
perception are precisely the workloads that lean on the fine-grained information that one-bit
weights throw away. The overall number:
## On the device
The point of all this is where it runs. Bonsai reports up to **163 tok/s** for the 1-bit variant
on an RTX 5090 (134 for ternary), and up to **87 tok/s** on an Apple M5 Max (58 for ternary) — and,
the flashiest claim, that the 3.9 GB model fits inside an iPhone 17 Pro's app-memory budget, making
it "the first 27B-class model to run on a phone." PrismML frames this as *intelligence density* — a
coined score-per-GB metric on which the 1-bit model scores 0.53/GB, which they call more than 10×
the full-precision baseline. It ships with weights on Hugging Face, an MLX path for Apple silicon and
CUDA for NVIDIA, and speculative-decoding support for lossless draft-and-verify acceleration.
Read the numbers for what they are. Bonsai's **capability is Qwen3.6-27B's** — this is a compression
and kernels result, not a new model. Every score above is **vendor-reported on PrismML's own
15-benchmark suite in "thinking mode,"** so treat the suite and mode as chosen, not neutral. The
"~90–95% retained" headline is a real average that **masks much larger, uneven drops**: math barely
moves, but agentic tool-calling and vision lose ~18% at 1-bit — so the right variant depends entirely
on your workload. "Intelligence density," "first 27B on a phone," and "10×" are marketing framings
(intelligence-density is a coined score-per-GB metric), and the throughput figures are specific to an
RTX 5090 and an M5 Max. No independent evaluation exists yet.
## The takeaway
Bonsai is a bet that for a large slice of real use — on-device assistants, privacy-sensitive tasks,
hybrid deployments that route only the hard cases to a frontier API — a model that keeps 90% of a 27B's
quality while fitting in 3.9 GB beats a bigger model you can't run locally at all. That bet is strongest
where quality degrades gracefully (math, general reasoning) and weakest where it doesn't (agentic, vision).
The genuinely impressive engineering is the end-to-end part: getting one-bit embeddings and a one-bit LM
head to work is what turns "quantized MLPs" into a model that actually fits on the phone in your pocket.
---
# Inkling: an open-weights multimodal MoE built to be adapted
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/inkling
> date: 2026-07-15
> tags: llm, mixture-of-experts, multimodal, reinforcement-learning, attention, explainer
Most model launches lead with a leaderboard. Thinking Machines Lab's **Inkling** does the opposite: the
announcement states plainly that it is **"not the strongest overall model,"** and is instead **"designed
for broad adaptation through fine-tuning."** That framing is the right lens for everything below. Inkling
is an **open-weights**, multimodal **[mixture-of-experts](/articles/mixture-of-experts-from-scratch)**
foundation model — **975B total parameters, 41B active** — with text, image and audio in one stack and up
to a **1M-token** context. A smaller companion, **Inkling-Small (276B total / 12B active)**, ships in
preview. The pitch is a customizable *base*, not a frontier trophy.
What makes it worth a close read is the mechanics: an attention design tuned for long context, an
encoder-free multimodal path, a reinforcement-learning run whose reward scaled *log-linearly* while the
model's reasoning got *shorter* on its own, and a knob that lets you dial how many tokens the model spends
per query. Let's take them in turn — and keep the honest caveats in view throughout.
## The backbone: sparse experts, hybrid attention
Inkling is a **66-layer** decoder-only transformer. Two forms of sparsity run through it. In the
feed-forward path, every layer is a mixture-of-experts: **256 routed experts plus 2 shared experts**, with
**6 routed experts active per token**. A **sigmoid-based router** with an **auxiliary-loss-free
load-balancing** bias decides which six fire — the same "drop the aux loss, use a bias term" trick that
has become standard for keeping expert utilization even without a loss that fights the main objective. The
2 shared experts are always on, giving every token a common backbone of computation. Net effect: only
**41B of the 975B** parameters do work on any given token.
The attention path is a **hybrid**: of the 66 layers, **55 are sliding-window (512-token) local and 11 are
global** — an interleaved **5:1 ratio** — with **64 query heads** tied to **8 KV heads**, over a **6144-dim**
residual stream. Five cheap local layers pass for every one exact global layer — the same local/global
bargain that makes long-context serving affordable in
[MiniMax's sparse attention](/articles/minimax-sparse-attention) and
[MiMo-V2-Flash](/articles/mimo-v2-flash).
Then a cluster of small but telling choices — the kind you only catch by reading the config, not the
launch post:
- **Relative position bias**, not [RoPE](/articles/how-llm-inference-works) (`d_rel=16`, `rel_extent=1024`)
— a learned bias on relative distance, chosen for cleaner extrapolation past the trained length.
- **Short depthwise convolutions** (kernel size 4) in *several* places — after the key and value
projections and on the residual branches — a cheap way to blend a little local context into each token
before attention even runs. Convs-inside-a-transformer is a recurring "free lunch" for stability.
- An easy-to-miss one: a **separate RMSNorm on the token embeddings**, applied *before* the usual
per-block RMSNorms (`use_embed_norm=true`). The residual stream is normalized at *entry*, not only inside
each layer — extra insurance on embedding scale that most decoder-only stacks skip.
None of these are headline features; together they read as the fingerprint of a team tuning the backbone
for stable long-context training rather than chasing a benchmark. Scrub the stack to see both sparsities at
once — which layers are global, and which experts a token lights up:
## Encoder-free multimodal
The multimodal design is deliberately minimal: **no separate vision or audio encoder**. Instead every
modality is turned into tokens the transformer reads directly. **Audio** becomes **discrete dMel
spectrogram** tokens; **images** are cut into **40×40-pixel patches** and lifted by a small **four-layer
hMLP** patch encoder; all modalities land in the **shared hidden space** and flow through the same experts
and attention. There's no bolted-on CLIP-style tower whose representation you have to align — the model
learns text, image and audio in one backbone. That is part of why it's pitched as an adaptation base:
fine-tuning touches one stack, not a federation of encoders.
## Controllable effort — the signature move
Inkling can vary how much it "thinks." The **system message plus a per-token cost** let you trade accuracy
for token spend: turn effort down and it answers tersely; turn it up and it reasons at length, approaching
its ceiling. The headline result is on **Terminal-Bench-2.1**, where the lab reports Inkling reaching
**Nemotron-3-Ultra-equivalent** accuracy at **roughly one-third the generated tokens**. Drag the effort
knob and read the tie line — the same score sits about **3× further right** on the reference curve:
The efficiency framing matters more than any single point on the curve. A model that lets the *caller*
choose the accuracy/latency trade-off, per request, is a different product from one with a fixed thinking
budget — especially for the fine-tuning-and-deploy audience Inkling targets, who care about tokens-per-task
cost at scale. (The curve shape above is illustrative; the ~63.8% Terminal-Bench-2.1 plateau and the
~1/3-token match are the real, vendor-reported anchors.)
## The RL story: log-linear reward, self-shortening reasoning
Post-training leaned on **large-scale asynchronous [reinforcement learning](/articles/ring-zero-trillion-scale-rl)** —
**over 30 million rollouts**. Two findings stand out. First, the **aggregate held-out eval reward rose
log-linearly** across those rollouts, climbing from **0.264** at the SFT-initialised checkpoint to
**0.356** at release — a straight line on a log-rollouts axis, i.e. more RL compute kept paying off
predictably rather than saturating. Second, and more surprising: with **no brevity objective** in the
reward, the model's **chain-of-thought became more concise on its own**, "dropping grammatical overhead
while remaining comprehensible." Reasoning compression emerged as a side effect of optimizing for correct
answers. Drag the marker to watch reward climb as thought-length falls:
This connects back to controllable effort: a model whose reasoning is naturally terser is cheaper to run at
any accuracy target, and the effort knob then lets you push that further.
## Training and release
Pretraining ran on **45 trillion tokens** of mixed text, image, audio and video, optimized with **[Muon](/articles/muon-optimizer)
for the large matrix weights and Adam for everything else** (weight decay coupled to the squared learning
rate), on **NVIDIA GB300 NVL72** systems. Alongside the standard weights, Thinking Machines released
**[NVFP4](/articles/nemotron-nvfp4)** weights for Blackwell — the same 4-bit format NVIDIA used to train
Nemotron. The release is genuinely open: **weights on Hugging Face** (both standard and NVFP4), an **API on
Tinker plus Together, Fireworks, Modal, Databricks and Baseten**, day-one **vLLM / SGLang / llama.cpp**
integration, and a public **Playground**.
## Results — read them as vendor-reported
Here are the headline numbers from Thinking Machines' own suite (at high effort). Reasoning first:
And the agentic / coding side, where the effort story is most relevant:
Multimodal and safety round it out: **VoiceBench 91.4%**, **MMAU 77.2%**, **MMMU-Pro 73.5%**,
**Global-MMLU-Lite 88.7%**; on safety, **FORTRESS Benign 95.9% / Adversarial 78.0%** and **StrongREJECT
98.6%**. Inkling-Small tracks the big model closely on several evals (HLE text-only **31.6%**, HLE with
tools **47.8%**).
**Scope the numbers.** Every score here is **vendor-reported on Thinking Machines' own evaluation
suite**, and the comparison set (GPT-5.6 Sol, Claude Fable 5, GLM 5.2, Nemotron-3-Ultra, and others) is
**provider-selected** — so treat this as a self-report, not a neutral head-to-head. The "~1/3 the
tokens," the log-linear RL scaling (0.264 → 0.356), and the emergent reasoning-compression claim are all
**their measurements on their evals**: real and interesting, but not independently verified. And this is
a **company blog and model card, not a peer-reviewed paper** — there is no external methodology to audit.
The genuine mitigant is that it's an **open-weights** release, so the architecture and the claims become
independently checkable over time in a way a closed model's never are.
## The take
Inkling's most refreshing feature is its honesty about what it is. Thinking Machines did not build the
model to top a leaderboard; they built a broad, open, multimodal base and tuned the *ergonomics of
adapting and running it* — controllable effort so callers own the accuracy/cost trade-off, a 5:1
local/global attention hybrid and relative positions so 1M-token context stays affordable, an encoder-free
multimodal path so fine-tuning touches one stack, and NVFP4 weights so it deploys cheaply on Blackwell. The
two research results worth remembering are the **log-linear RL reward** (evidence that the post-training
recipe kept scaling) and the **emergent reasoning compression** (shorter chains-of-thought with no brevity
reward) — both their own measurements, both the kind of thing open weights will let others probe. Judge it
not as "is this the best model" — the lab already answered no — but as a base you can take, fine-tune, and
serve. On that axis, an open 975B/41B MoE with these ergonomics is a substantial thing to hand the
community.
---
*Built on Thinking Machines Lab's [Inkling announcement](https://thinkingmachines.ai/news/introducing-inkling/)
and [model card](https://thinkingmachines.ai/model-card/inkling/), plus the
[Hugging Face release](https://huggingface.co/thinkingmachines/inkling). All benchmark and scaling figures
are vendor-reported; the interactive diagrams are illustrations of the mechanism, with real endpoints
noted inline. Related reading:
[mixture-of-experts from scratch](/articles/mixture-of-experts-from-scratch),
[NVFP4 training](/articles/nemotron-nvfp4),
[MiniMax sparse attention](/articles/minimax-sparse-attention),
[MiMo-V2-Flash](/articles/mimo-v2-flash),
[trillion-scale RL](/articles/ring-zero-trillion-scale-rl),
[the Muon optimizer](/articles/muon-optimizer), and
[how LLM inference works](/articles/how-llm-inference-works).*
---
# LOTUS: reasoning in the hidden states, not the token stream
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/lotus-latent-reasoning
> date: 2026-07-15
> tags: llm, reasoning, chain-of-thought, latent-reasoning, inference-optimization, explainer
The way a reasoning model earns its answer is by writing out its work: an explicit **chain of thought (CoT)**,
one token at a time, before it commits to a final answer. That is where the latency goes. Each of those
intermediate tokens is a full sequential decode step — a [memory-bound pass over the growing KV
cache](/articles/how-llm-inference-works) — so the more the model thinks, the slower it answers.
**Latent CoT** is the tempting alternative: do the multi-step reasoning inside the model's *hidden states*,
replacing decoded tokens with continuous representations, and skip the token-by-token bottleneck entirely.
The problem is that it has never quite worked at scale. Methods like Coconut, CODI, and SIM-CoT match
explicit CoT on small models, but **beyond 1B parameters no latent method keeps up on math reasoning, and
the gap widens as the backbone grows**. LOTUS — *Looped Transformers with parallel supervision on latents*,
from Ying Fan, Anej Svete, and Kangwook Lee — is, to the authors' knowledge, the first latent-CoT method
to close that gap at the **3B** scale, while cutting the thought phase by **2.5×–6.9×**.
It gets there by fixing the two things the authors argue were holding latent CoT back.
- **(P1) Sequential generation.** Coconut, CODI, and SIM-CoT still produce their latent tokens
*autoregressively* — the sequential bottleneck is still there, just moved into latent space.
- **(P2) No CoT grounding.** Without supervision that aligns each latent position to a specific gold
reasoning step, the latent trace drifts and destabilizes as the model gets bigger.
## The loop: one weight set, R passes, K blocks in parallel
LOTUS builds a **padded latent region** into the prompt. Between two learnable delimiters `⟨BoT⟩` and
`⟨EoT⟩` it inserts $K$ blocks of $c$ shared, learnable `⟨lat⟩` tokens — a fixed $K\cdot c$ latent positions
(the deployed config is $K=6$, $c=25$, so 150 positions). The question $Q$ sits before `⟨BoT⟩`; the answer
$A$ comes after `⟨EoT⟩`.
The reasoning then happens by **looping the base language model over that region**. Let $E$ be the learnable
latent embeddings and $f_\theta$ the ordinary LM backbone. Starting from the latents, LOTUS reuses the *same
weights* for $R$ iterations, adding the previous pass's output back in each time:
$$
h^{(0)} = f_\theta\!\big(E \mid C_{\text{pre}}\big), \qquad
h^{(t)} = f_\theta\!\big(E + h^{(t-1)} \mid C_{\text{pre}}\big), \quad t = 1,\dots,R
$$
where $C_{\text{pre}}$ is the reused KV cache of the question. This is a **recurrent-depth (looped)
Transformer**: it adds computation depth by reusing parameters, not by adding them. The crucial property is
that all $K\cdot c$ latent positions are refined **together** on each pass — so the whole thought phase is
$R$ sequential forward passes, not one pass per generated token. Scrub the loop and watch the latents sharpen,
then read out at the final iteration:
That parallelism is the answer to **(P1)**. Where an autoregressive latent method spends a forward pass per
latent token — the same shape of bottleneck that makes [multi-token prediction](/articles/multi-token-prediction)
and [diffusion language models](/articles/illada-diffusion-language-model) attractive — LOTUS spends only $R$
passes for the entire trace, regardless of how many tokens that trace would have been.
The other way to see this is as a **network**. A looped Transformer is a recurrent-depth
network: roll it up and it is one block with a loop-back edge; unroll it and it is an effective
$R$-deep stack of the *same* weights, with the latents $E$ fed back in at every pass. The depth is
real — each pass is a full forward through the backbone — but the parameter count never grows past
$1\times$. Toggle between the rolled and unrolled views, and drag the unroll depth:
That is the whole trick behind "add computation depth by reusing parameters, not by adding them": a
3B backbone reasons at a depth its parameter budget alone would not buy, because depth here is
$R$ passes through shared weights rather than $R$ times the weights.
## Parallel supervision: grounding each latent in its gold step
The loop alone is not enough; the latents need to be told *what to compute*. This is the answer to **(P2)**,
and it is the part that makes LOTUS more than "a looped model." After the final iteration, LOTUS reads each
post-loop latent position $h^{(R)}_{i,j}$ **through the base LM head** $f_{\text{head}}$ and trains it, with
cross-entropy, toward the gold CoT-step token that belongs in that slot. Each gold CoT step $i$ is tokenized
and padded/truncated to $c$ tokens $T_{i,\cdot}$, and every position is supervised at once:
$$
\mathcal{L}_{\text{step}} = \frac{1}{N_{\text{step}}}\sum_{i=1}^{K}\sum_{j=1}^{c}
\operatorname{CE}\!\big(f_{\text{head}}(h^{(R)}_{i,j}),\, T_{i,j}\big)
$$
This is direct, position-aligned supervision to real reasoning tokens — much like ordinary explicit-CoT
supervision — rather than the indirect hidden-state or KV-cache distillation used by earlier parallel-latent
methods (PCCoT, KaVa). A separate final forward pass then supervises the answer against the latents:
$$
\mathcal{L} = \mathcal{L}_{\text{ans}} + \lambda_{\text{step}}\,\mathcal{L}_{\text{step}}
$$
The paper frames why both losses are needed with a **Parallel Chain Likelihood (PCL)** view. Because the
step loss factorizes over positions rather than autoregressively, it induces
$$
p_\theta^{\text{PCL}}(T\mid Q) = \prod_{i=1}^{K}\prod_{j=1}^{c} p_\theta(T_{i,j}\mid Q)
$$
The two losses then split the work: $\mathcal{L}_{\text{step}}$ provides **coverage** — it puts probability
mass on the correct gold token at each position — while $\mathcal{L}_{\text{ans}}$ provides **selection** —
it forces the jointly-computed latents to actually support the right answer. The ablation makes the split
concrete: with only $\mathcal{L}_{\text{step}}$, the model's post-loop latents recover the gold top-1 token
just **9.1%** of the time; with only $\mathcal{L}_{\text{ans}}$, **9.4%**; with **both**, **70.9%**
(NLL 3.07 vs 9.29 and 5.97). Neither loss alone builds a readable, correct latent trace.
## Why it is fast
The efficiency story is simple once the loop is clear. Explicit CoT's thought-phase latency scales with **how
much it writes**; LOTUS's scales with $R$, which is fixed. So the win grows exactly when the rationale gets
verbose. On Llama-3.2-3B the paper measures the thought phase directly — drag the playhead and watch LOTUS
finish while explicit CoT is still decoding:
On compact math-expression CoT the thought phase drops from **338.8 ms to 133.0 ms** (2.5×), and total
latency from 384.2 ms to 181.2 ms (about 2.1× overall). Swap in verbose **natural-language** rationales and
explicit CoT balloons to **963.6 ms** while LOTUS barely moves to **140.8 ms** — a **6.9×** thought-phase
speedup, at essentially the same accuracy (68.13% vs 68.41%). The query-prefill and answer phases are nearly
identical across methods; only the thought phase moves.
## The numbers
Held against explicit CoT and the strongest latent baselines on GSM8K (Llama-3.2-3B, in-domain), LOTUS lands
within about a point and a half of explicit CoT and clearly ahead of the latent baselines:
The pattern holds across backbones — GPT-2 (LOTUS 44.1 vs explicit 42.7), Llama-1B (57.3 vs 58.4), Llama-3B
(70.0 vs 71.5) — so unlike prior latent methods the gap does **not** widen with scale. And on the
**out-of-domain** average (GSM-Hard, MultiArith, SVAMP), LOTUS actually edges ahead of explicit CoT, **63.9
vs 62.1**, led by a near-perfect 99.9% on MultiArith and 75.7% on SVAMP.
## How deep does the loop need to be?
Loop depth $R$ is LOTUS's compute knob, and reasoning turns out to need a real amount of it. Train the model
at increasing $R$ and accuracy climbs steeply — a shallow loop simply cannot fit multi-step arithmetic:
Two honest wrinkles live in this chart. First, depth is **not** free test-time compute you can dial up after
training: take the model trained at $R=6$ and run it at $R=7$ and accuracy *dips* to 69.3% — LOTUS reasons
best at the depth it was trained for. Second, the parallel **width** $c$ is nearly free where depth is not:
sweeping $c$ from 1 to 50 tokens per block (a 50× change in latent positions) moves the thought phase by only
about 30 ms — 110.9 ms to 141.2 ms — because those positions are processed in parallel. A single token per
step (c=1) is too narrow (51.4%), but moderate widths (c=25–30) saturate at 70%.
## Does it actually reason, or memorize?
Because LOTUS reads its latents through the ordinary LM head, you can literally decode the thought — the same
trick behind interpretability tools that [unembed intermediate activations](/articles/jacobian-lens). The
post-loop latents recover the gold CoT at **70.9% top-1 / 85.8% top-5**. More telling is a multi-path test:
for a question with a *trained* gold chain (G) and an *unseen-but-valid* alternative chain (U), the readout
orders their likelihoods $G \ll U \ll \text{random}$ (NLL 0.07, 4.28, 8.16), assigning graded probability to
valid-but-never-seen reasoning. That ordering is the paper's evidence that the latents encode reasoning
structure, not just a memorized string.
**Read the setup before believing the headline.** (1) *Scope is narrow:* every result is **math word
problems** (GSM8K-family, trained on GSM8k-Aug); the authors explicitly flag transfer to other domains as
open. (2) *The budget is fixed:* $K$, $c$, and $R$ are hyperparameters set to cover the expected step count —
chains **longer than $K$ steps fall back to autoregressive completion**, and making the budget adaptive is
listed as future work, not a solved problem. (3) *"Bridges the gap" means parity, not a win:* LOTUS is ~1.5
points **behind** explicit CoT in-domain at 3B (70.0 vs 71.5); it needs the LOTUS+CODI combo to get within a
point, and it leads only on the OOD average. (4) *Speedups are the paper's own measurements* on Llama-3B
(H-class GPU), and the flattering 6.9× is specifically the verbose natural-language regime; the compact-math
number is 2.5×. (5) *Baselines are author-selected* latent methods (Coconut, CODI, SIM-CoT, PCCoT, KaVa) and
one explicit-CoT reference — not a broad frontier-model comparison.
## The take
LOTUS's real contribution is a clean recombination: take a **looped padded Transformer** (depth from weight
reuse, all latent positions refined in parallel) and give it **direct, position-aligned supervision** to gold
CoT tokens through the base LM head. The loop kills the sequential bottleneck that every prior latent method
kept; the supervision kills the drift that made latent reasoning fall apart at scale. Together they do
something no earlier latent-CoT method managed — **stay on the explicit-CoT accuracy curve at 3B** — while
turning a variable, write-everything thought phase into a fixed $R$-pass one.
The honest frame is that this is a *parity-with-a-speedup* result on math, bounded by a thinking budget you
have to choose in advance. Whether a fixed $K\cdot c\cdot R$ box holds up when problems demand more steps than
you budgeted — and whether the story survives outside arithmetic — are the open questions the authors name
themselves. But as a demonstration that latent reasoning can finally keep pace with explicit reasoning at
scale, and answer 2.5–6.9× faster while doing it, LOTUS is the first latent-CoT method that clears the bar.
---
*Built on [Bridging the Gap Between Latent and Explicit Reasoning with Looped
Transformers](https://arxiv.org/abs/2606.31779) (Fan, Svete, Lee; 2026). All accuracy, latency, and ablation
figures are quoted from the paper (Llama-3.2-3B unless noted; GSM8K-family math benchmarks; deployed config
$K=6$, $c=25$, $R=6$). The interactive diagrams are illustrations of the mechanism; the gold-CoT tokens in
the loop diagram are illustrative.*
---
# Ring-Zero: what a trillion-parameter model learns from reward alone
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/ring-zero-trillion-scale-rl
> date: 2026-07-15
> tags: llm, reinforcement-learning, reasoning, mixture-of-experts, explainer
**Zero RL** is the stripped-down recipe behind the reasoning-model boom: take a pretrained base model, give it math problems whose answers can be *checked*, reward it for getting them right, and let reinforcement learning grow the chain-of-thought on its own — no supervised fine-tuning, no human-written reasoning traces, no reward model. DeepSeek-R1 made it famous. But almost every published zero-RL study runs on models small enough to fit a modest cluster, which leaves the interesting question open: **what happens when you do this to a genuinely large model?**
Ring-Zero is that experiment. The authors run zero RL directly on **Ling-2.5-1T-Base** — a **1-trillion-parameter** Mixture-of-Experts model with **63B active** parameters per token — and report what changes at scale. The paper frames its results as a vindication of the *bitter lesson*: with enough scale, hand-crafted heuristics for "good reasoning" become unnecessary because the model develops them itself. The contribution is less a single trick than a **stable, four-stage training pipeline** that survives trillion-scale RL, plus a careful account of the training dynamics and the behaviors that emerge.
## The pipeline is the method
There is no single loss here. Ring-Zero's real content is a **sequence of four stages**, each one repairing a failure mode the previous stage creates. Click through them:
The logic of the sequence is worth stating plainly. **First-stage RL** uses a *token-level* loss — the per-response loss is deliberately **not** divided by length — so a longer correct trace earns more total credit and the model learns to think at length. That works, but it also teaches the model to pad (more on that below). **Self-distillation** then samples from the stage-1 expert, keeps the *shortest correct* trace, self-filters redundant steps, and fine-tunes the base model on the result — compressing the bloat and, crucially, **resetting the gap between the training and inference engines**. **Second-stage RL** switches to a *sample-level* (length-normalized) loss so gradients no longer reward length, and drops the KL penalty now that the model is a strong starting point. **Third-stage RL** adds three difficulty tiers with their own prompts so one checkpoint can reason short or long on demand.
## The setup
The base is **Ling-2.5-1T-Base**, a hybrid MoE combining **MLA** (multi-head latent attention) and **Lightning Attention** layers, trained *from scratch with no SFT*. The smaller **Ling-2.5-flash-Base** (104B total, 7.4B active) is used as the scaling foil and — importantly — as the workhorse for most ablations. Training runs on **320 × H200** GPUs with Megatron for updates and SGLang for rollouts. Each step draws **G = 16** rollouts per question at temperature 1.0; the reward is dead simple and rule-checkable:
$$
r_i = r_{\text{acc},i} + r_{\text{format},i}, \qquad r_{\text{acc},i},\, r_{\text{format},i} \in \{0, 1\}
$$
where $r_{\text{format}}$ checks for well-formed `...` and `...` tags and $r_{\text{acc}}$ is rule-based matching early on, LLM-as-judge (Qwen3-Next-80B) later. That is the *entire* supervision signal — no human labels, no learned reward model. This is what "zero" means. It builds directly on ideas we have covered before: the MoE backbone (see [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch) and [Switch Transformers](/articles/switch-transformer)) and RL from verifiable rewards on reasoning (see [Leanstral](/articles/leanstral-formal-proofs)).
## The objective, precisely
Stage-1's policy objective is a **clipped-importance** RL loss with a group-normalized (GRPO-style) advantage:
$$
\mathcal{J}(\theta)=\mathbb{E}\!\left[\sum_{i=1}^{G}\sum_{t=1}^{|o_i|}\operatorname{sg}(\hat{\rho}_{i,t})\;\hat{A}_{i,t}\;\log \pi^{M}_{\theta}\!\left(o_{i,t}\mid q,\,o_{i,
This is the same disease diagnosed — from the routing angle — in [Rollout Routing Replay](/articles/rollout-routing-replay): when the engine that generates a rollout and the engine that computes the gradient disagree, the importance ratio blows up and MoE RL diverges. Ring-Zero attacks the numerical side of it (and adds a small KL leash, $\beta = 10^{-4}$, with the reference model refreshed every 400 steps, plus **mixed-precision control** — BF16 everywhere except FP32 in the attention softmax and the LM head, the two places rounding error is worst).
## From token-level to sample-level: killing length inertia
The token-level loss has a side effect the paper names **length inertia**. Because the loss is not normalized by length, the model discovers a lazy shortcut: emitting more tokens is mathematically safer, so responses inflate *even on easy problems it already solves on the first try*. The fix is the Stage-2 loss, identical to Stage-1 except for one factor:
$$
\mathcal{L}_{\text{II}}(\theta)=-\mathbb{E}\!\left[\sum_{i=1}^{G}\frac{1}{|o_i|}\sum_{t=1}^{|o_i|}\operatorname{sg}(\hat{\rho}_{i,t})\;\hat{A}_{i,t}\;\log \pi_{\theta}\!\left(o_{i,t}\mid q,\,o_{i,
## Does it work? The numbers
The headline is **scaling**. On the first stage of RL alone, the 1T model clears the 104B model by wide margins on every math benchmark. First-stage 1T scores (with the 104B flash model in prose for contrast): AIME 2024 **89.1%** (flash 71.2), AIME 2025 **83.3%** (63.5), AIME 2026 **84.2%** (65.3), HMMT Feb 2026 **66.2%** (50.3), IMOAnswerBench **59.3%**.
The full pipeline (second-stage RL, with a 2× YaRN context extension) pushes those to **94.1% / 92.3% / 93.2%** on AIME 2024/25/26. The scaling advantage is visible not just in the endpoint but in the *slope* — the 1T model learns faster per step:
Honesty check on where this lands. Ring-Zero's best AIME 2026 number (**93.2%**) is genuinely strong — but it still **trails the frontier** models the authors themselves list. This is a zero-RL-at-scale study, not a SOTA claim:
Where Ring-Zero does claim an edge is **CoT quality**, measured three ways. Its traces win LLM-as-judge *comprehensibility* comparisons against GLM-5.1, Kimi-k2.6, MiniMax-M2.7 and Qwen3.5-397B. They *reproduce* better under distillation: fine-tuning student models on only **100K** Ring-Zero traces beats distilling **800K** DeepSeek-R1 traces —
— and they are *efficient*: on problems both solve, Ring-Zero averages **6,368 tokens**, less than half its baselines' length.
## Two phases: discovery, then sharpening
The second finding is about *how* RL improves the model over time. Track two quantities: **pass@1024** (can the model solve a problem in *any* of 1024 attempts — a measure of coverage) and **pass@1** (does it nail it on the first try — reliability). They move on different schedules.
Coverage saturates early — pass@1024 flattens around step 800 — meaning RL has already surfaced essentially every reasoning pattern it will ever use (the **discovery** phase). But pass@1 keeps rising long after (the **sharpening** phase): the model is not finding new tricks, it is becoming *reliable* at the ones it has. The paper reads this as evidence for a sharper claim in its discussion — that zero RL **optimizes within a boundary set by pretraining** rather than expanding it. Which is exactly the honest ceiling in its limitations: RL cannot invent a proof technique the base model never saw.
## Behaviors nobody programmed
The third finding is the paper's "bitter lesson" payoff: with scale, the model **spontaneously develops** cognitive behaviors that smaller-model work usually has to elicit with hand-crafted prompts or rewards. The paper documents five:
- **Anthropomorphism** — traces narrate themselves ("I might have a brain fart here", "let me not wing it", "genius idea"), artifacts of the pretraining corpus surfacing as reasoning scaffolding.
- **Structured formatting** — spontaneous "Step 1: / Step 2: / Verify:" scaffolds appear with no formatting instruction, hinting at a higher-level action space above raw tokens.
- **Parallel reasoning** — the model branches into competing strategies within a single rollout, compares outcomes, and commits only when evidence converges (tree-of-thought, self-taught).
- **Self-verification** — it re-checks assumptions, substitutes answers back into the problem's constraints, learned because that is what secures the correctness reward.
- **Context anxiety** — approaching its token limit, the model *strategically aborts* deep reasoning to guarantee a well-formatted answer, revealing an implicit awareness that format compliance is also rewarded.
The claim is not that these are magic; it is that at 1T scale they arrive *for free*, making the elaborate reasoning-elicitation machinery of small-model RL redundant.
## What the ablations actually establish
Several design choices are backed by ablations — with one caveat that matters (see below): they are run on the **104B flash** model, not the 1T model.
- **RL algorithm.** Comparing GRPO / DAPO / CISPO / GSPO reveals a **speed-stability tradeoff**: amplifying low-probability tokens (CISPO, DAPO) learns fastest but its entropy collapses; GRPO is most stable but slowest. Ring-Zero's clipped-importance-plus-corrections scheme is the attempt to get both.
- **KL penalty.** Remove it in stage 1 and the log-prob gap diverges, entropy collapses, and reward crashes within ~2,000 steps. With $\beta = 10^{-4}$ it stays healthy.
- **Ratio correction.** The naive ratio collapses near step 800; a clip-only patch delays collapse to ~2,700 steps; the training-engine-numerator correction trains indefinitely. This is the single most important stability result.
- **Format reward.** A single opening `` tag lets length explode with no reward gain; requiring properly closed tags with an EOS token is what makes stopping — and therefore credit — well-defined.
- **Hyperparameters.** Robust to learning rate over $\{1,2,3\}\times10^{-6}$; $G=32$ is fastest per step but $G=8$ fastest in wall-clock; token-level loss grows length, sample-level keeps it flat — motivating the stage-1 to stage-2 switch.
**Read the scope before the headline.** (1) The efficiency and stability *ablations* — RL-algorithm choice, KL, ratio correction, format reward, hyperparameters — are run on the **104B flash** model, not the 1T model, for cost. The conclusions are *assumed* to transfer up. (2) On raw accuracy, Ring-Zero **trails the frontier** models it lists (GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8, Qwen3.7-Plus, Kimi K2.6 all sit at 95.7–98.3% on AIME 2026 vs Ring's 93.2%); the win it claims is CoT *quality*, not peak score, and those quality judgments lean on LLM-as-judge. (3) The five "emergent behaviors" are **qualitative** — quoted trace snippets and interpretation ("context anxiety"), not quantified frequencies. (4) The frontier comparison numbers are **as reported by the authors** for competitors. (5) Adaptive depth carries **negative transfer**: the jointly-trained High tier (93.2%) sits *below* the dedicated second-stage model (94.1%) — flexibility costs a little peak. (6) The model, weights, and 320×H200 infrastructure are **not released**, so the 1T result is not externally reproducible. (7) Training is capped at **64k context** by hardware; the paper expects longer windows to unlock more, i.e. this is not the ceiling of the recipe.
## The take
Ring-Zero's value is not a new loss function — it is a **demonstration and an engineering recipe**. The demonstration: run the simplest possible RL signal (right/wrong plus format) on a trillion-parameter base, and you get sharp gains, a clean two-phase learning dynamic, and reasoning behaviors that smaller models have to be coaxed into. The recipe: a four-stage pipeline where token-level RL grows reasoning, self-distillation compresses it and resets the engine gap, sample-level RL sustains it without length bloat, and tiered RL makes depth controllable — all held together by a training-inference ratio correction that is easy to overlook and, per the ablation, the difference between training and diverging.
The honest frame is the one the paper itself offers in its limitations: zero RL **sharpens the reasoning already latent in pretraining; it does not transcend it**. That is why coverage saturates while reliability climbs, and why a bigger base — not a cleverer reward — is what moves the ceiling. As a controlled study of what pure reward does at scale, it is unusually candid; as a frontier-accuracy claim, it is not one, and it does not pretend to be.
---
*Built on [Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning](https://arxiv.org/abs/2607.12395) (Tang, Cao, Liu et al., 2026; CC BY 4.0), the team behind the Ling / Ring models. All benchmark, efficiency, and ablation figures are quoted from the paper; the interactive diagrams are illustrations of the mechanism, not reruns of the experiments. Ablations are on the 104B flash model unless noted; the 1T model and infrastructure are not publicly released.*
---
# Mach-Mind-4-Flash: specialize, then integrate
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mach-mind-4-flash
> date: 2026-07-13
> tags: llm, reinforcement-learning, knowledge-distillation, mixture-of-experts, explainer
The dominant recipe for a better model is still *scale the pre-training* — more parameters, more
tokens, more compute. Mach-Mind-4-Flash, from Li Auto's Foundation Model team, is a bet on the other
axis. It starts from an existing compact base — **Qwen3.5-35B-A3B**, a Mixture-of-Experts model with
35B total parameters but only **3B active per token** — and pushes it toward the score band of
100B-class models using **post-training only**: reinforcement learning, expert fusion, and
inference-time efficiency. No extra pre-training compute.
The one idea to leave with is the shape of that post-training: **specialize, then integrate.** Rather
than run one big mixed-reward RL job — which tends to rob Peter to pay Paul across capabilities — the
team trains **more than ten domain specialists in parallel** (across Reasoning, General, and Agent
tracks), then fuses them into a single deployable generalist. The fusion is the paper's headline
contribution, **Multi-Teacher On-Policy Distillation (MOPD)**, and a second stage, **HMPO**, trims the
model's reasoning length without paying for it in accuracy.
## The fusion problem, and MOPD
If you train separate RL experts — a math expert, a code-agent expert, a safety expert — each is
excellent in its lane. The naive way to get one model with all of those skills is to mix every
domain's reward into a single RL objective. In practice that fails in a specific way the paper names
**see-saw degradation**: the gradients from different rewards collide, so a gain on one capability is
"routinely offset by regressions on others." You climb one hill by sliding down another.
MOPD sidesteps the collision. Every training sample carries a **routing key** $k$ that deterministically
selects the one frozen domain teacher $\pi_{T_k}$ that should supervise it. The student generates a
rollout under its *own* policy, and each token is pulled toward its routed teacher with a **token-level
reverse-KL**:
$$
\mathcal{L}_{\text{MOPD}}(\theta) = \mathbb{E}_{(x,k)\sim\mathcal{D}}\;\mathbb{E}_{y\sim\pi_\theta(\cdot\mid x)}\!\left[\frac{1}{|y|}\sum_{t=1}^{|y|} D_{\mathrm{KL}}\!\big(\pi_\theta(\cdot\mid x,y_{
MOPD sits inside a **unified RL/OPD objective**, $\mathcal{L} = \alpha\,\mathcal{L}_{\text{OPD}} +
\beta\,\mathcal{L}_{\text{RL}}$, so the same framework can run pure RL ($\alpha=0$), pure distillation
($\beta=0$), or a joint blend — and new teachers register as config nodes with "zero intrusion into
the framework's core logic." A pilot on the tool-agent domain shows the mechanism converging: the
teacher-student top-$K$ token-overlap rate climbs from **0.73 to 0.84** over training.
Fusion is not perfectly lossless, and the paper is candid about it. Across the three tracks it reports
three distinct outcomes: **capability anchoring** for Reasoning (the frozen expert prevents the
student from regressing during fusion), **full retention** for General, and **mixed results** for
Agent — where the fused model sometimes lands *below* its own expert teacher (SWE-bench Verified
71.1 after fusion vs. 73.8 for the standalone expert). Long-horizon agent behavior is the hardest
thing to distill without smoothing away.
## HMPO: pay for correct-and-short
The second stage attacks **overthinking** — reasoning chains far longer than the task needs, which
inflate latency and serving cost for no accuracy gain. **Hybrid Median-length Policy Optimization
(HMPO)** is a single-stage token-efficiency method with a neat trick for the length budget: don't set
a threshold, *measure* one. For each query the policy samples a group of $G=10$ rollouts, and the
budget $b$ is the **median length of the correct ones**:
$$
b = \operatorname{median}\{\,n_i \mid i \in \mathcal{C}\,\}
$$
where $\mathcal{C}$ is the set of correct rollouts. The token reward is a cosine decay that starts at
1 and fades toward $\lambda$ as a correct trace grows, then cliffs to zero the instant it runs over
budget — and any incorrect trace earns zero at any length:
$$
R_{\text{token}} = \begin{cases}\min\!\big(1,\ \cos(\tfrac{\pi n}{2b}) + \lambda\big) & \text{if correct and } n < b\\[4pt] 0 & \text{otherwise}\end{cases}
\qquad
R_{\text{final}} = R_{\text{acc}}\cdot R_{\text{token}}
$$
The multiplicative composition enforces a strict **correctness-first, length-second** hierarchy:
wrong or over-budget traces get exactly zero reward, so efficiency gradients never flow to bad
answers. And because $b$ is the group median, it **self-tightens** as the policy gets more concise —
an implicit curriculum with, in the authors' words, zero tuning. Drag the candidate length, the
training progress, and $\lambda$:
Trained on a compact set of ~6.5K math problems (group size $G=10$, $\lambda=0.8$), HMPO cuts
generation length by **19-46% with at most a 0.7-percentage-point accuracy drop** — and although the
training is math-only, the learned length control **generalizes** to unseen domains: code generation,
science QA, and instruction following. As a single-pass method it also costs **1.5-2.5× fewer
GPU-hours** than the multi-stage length-control baselines it replaces.
## The infrastructure that makes it cheap
None of this is free unless the training loop is fast, and the third contribution is the plumbing: a
unified RL/OPD framework with **operator-level acceleration** reported at a **17% end-to-end training
speedup**. The wins are Hopper-specific kernel work — a deep integration of *SonicMoE* into Megatron
that implements an efficient **Indexed Grouped GEMM** for the MoE MLPs (using TMA copy, warp
specialization, and multi-stage producer-consumer pipelines), a **gate-up fusion**, and a **segmented
fusion with the shared expert** that overlaps communication and computation by splitting the shared
expert into AllGather / compute / ReduceScatter stages staggered against the routed experts. It also
leans on [multi-token prediction](/articles/multi-token-prediction) and multi-dimensional hybrid
parallelism. This is the same category of problem as
[stabilizing MoE RL](/articles/rollout-routing-replay) and [cutting RL's
cost](/articles/frontier-rl-cheaper): the algorithm is only as good as the systems that let you run it
at scale.
## The numbers
The result is a 3B-active model that trades blows with much larger ones. On raw reasoning it is
strong but *not* the frontier — at AIME'26 it lands ahead of the 122B-active Qwen3.5 but behind the
309B MiMo-V2-Flash and the 1T-parameter Kimi-K2.5:
Where it genuinely *leads* the pack — beating both the 122B-active Qwen3.5 and the 1-trillion-parameter
Kimi-K2.5 — is on instruction-following, safety, tool use, and Chinese web search. These are the axes
the specialize-then-integrate recipe was built to lift:
For context on those: IFBench 82.8 vs. 76.1 (Qwen 122B) and 67.4 (Kimi 1T); Behavioral-SafetyBench
80.7 vs. 29.9 and 67.8; BFCL-v4 75.8 vs. 72.2 and 74.5. Elsewhere it is competitive rather than
dominant — GPQA-Diamond 83.1, LiveCodeBench-V6 80.9, SWE-bench Verified 70.6, $\tau^2$-bench 80.0 —
solidly in the mix for a model activating a fraction of its rivals' parameters. And the efficiency
story is where the whole thing pays off: on AIME'26, HMPO puts Mach-Mind-4-Flash at the **upper-left**
of the accuracy-vs-tokens frontier, matching frontier accuracy at far fewer tokens per trajectory than
models of much larger active scale.
Read the wins with their scope. **All numbers are the authors' own** — a single-vendor technical
report, not an independent evaluation. The "100B-class performance" framing is real but selective:
Mach-Mind-4-Flash reliably beats the *122B-active* Qwen3.5 it's compared against, yet it **trails the
1T-parameter Kimi-K2.5** on AIME'25 (92.1 vs. 96.1), GPQA-Diamond (83.1 vs. 87.6), LiveCodeBench (80.9
vs. 85.0) and SWE-bench Verified (70.6 vs. 76.8) — this is *not* a SOTA claim. Fusion has costs: the
Agent track shows "mixed results," with the fused model landing below its own standalone expert on
some agentic tasks. HMPO's compression is real but measured on **single-turn** reasoning and the
math-trained generalization is the authors' evaluation. The **17%** speedup is Hopper-specific
infrastructure, not a modeling result. The paper's own limitations are blunt: MOPD leaves "a small but
consistent gap on extremely long-horizon tasks such as repository-level software engineering"; HMPO
does not yet extend to multi-turn agentic trajectories; and persistent web browsing (DeepSearch) plus
long-context comprehension "remain the weakest axes for compact models."
## The take
Mach-Mind-4-Flash's contribution isn't a new architecture — it reuses an off-the-shelf 35B-A3B MoE —
it's a **post-training recipe that composes cleanly**. MOPD turns "train many experts, ship one model"
from a lossy averaging problem into a routed distillation where each capability keeps its own gradient,
and the see-saw that plagues mixed-reward RL mostly disappears. HMPO is the tidy companion: make the
length budget a measured group-median instead of a tuned hyperparameter, gate it behind correctness,
and reasoning gets shorter for almost free. Set against the field's default answer — spend more
pre-training compute — this is the argument that a lot of headroom is still sitting in the
*post*-training stack, reachable by a compact model that only lights up 3B parameters at a time.
Whether a 35B model can truly close the last gap to a trillion-parameter frontier on the hardest
long-horizon tasks is the open question the paper's own limitations point at — but as a demonstration
that specialize-then-integrate scales to a dozen domains without falling apart, it's a clean result.
---
*Built on the [Mach-Mind-4-Flash Technical Report](https://arxiv.org/abs/2607.09375) (Foundation Model
Team, Li Auto Inc., 2026; CC BY-NC-ND 4.0). All benchmark and efficiency figures are quoted from the
report; the MOPD and HMPO interactives are illustrations of the mechanism, and the capability curves in
the fusion diagram are illustrative, not measured. Related on this site:
[Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch),
[Rollout Routing Replay](/articles/rollout-routing-replay), and
[MiMo-V2-Flash](/articles/mimo-v2-flash), which also leans on multi-teacher distillation.*
---
# Soofi S: a sovereign 3B-active model that keeps its cache near-constant
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/soofi-s
> date: 2026-07-13
> tags: llm, mixture-of-experts, pretraining, inference-optimization, explainer
Most open-model releases are open in name only: you get weights and an aggregate token count, not the
data, recipe, or checkpoints needed to audit or rebuild them. Most general-purpose multilingual models
spread their capacity thinly across dozens of languages, leaving German underrepresented relative to its
economic weight. And most are dense full-attention Transformers, whose per-sequence [KV
cache](/articles/how-llm-inference-works) grows with context and drags throughput down exactly in the
long-context, high-concurrency regime that costs the most to serve. **Soofi S 30B-A3B** — from a German
consortium coordinated by the KI Bundesverband (DFKI, Fraunhofer IAIS/IIS, TU Darmstadt, and others),
funded by the German BMWE — sets out to close all three gaps at once, and to do it on sovereign European
infrastructure.
The design that ties those goals together is a **hybrid Mamba–Transformer Mixture-of-Experts**: 31.6B
total parameters, but only ~3.2B active per token, and a backbone that is mostly linear-time Mamba-2 with
attention in just 6 of 52 layers. That last choice is the whole serving story — decode throughput stays
nearly flat as context grows, where dense baselines fall off a cliff:
## The architecture: hybrid backbone, sparse experts
Soofi S reuses the openly published **Nemotron 3 Nano** reference design without modification — a
deliberate choice, so the effect of the German–English data recipe can be measured against an
architecture-identical control (the same-arch [Nemotron](/articles/nemotron-nvfp4) baseline). The
backbone is 52 layers: **23 Mamba-2** sequence-mixing layers, **23 granular MoE** layers, and only **6
Grouped-Query Attention** layers, distributed sparsely through the depth. Scrub the stack — and note how
few layers actually hold a cache:
The Mamba-2 layers carry most of the sequence mixing with a *fixed-size recurrent state*; the attention
layers give exact long-range recall but are the only ones whose cache grows. The capacity lives in the
MoE layers, and this is where "30B at the cost of 3B" comes from. Each MoE layer has **128 routed
experts** plus **2 shared experts**; a learned, sigmoid-gated router activates just **6 routed experts**
per token, with the 2 shared always on:
The exact config, for reproducibility (Table 1): model dimension 2688, 32 attention query heads over just
2 KV heads (head dim 128), Mamba-2 state dimension 128, expert dimension 1856, squared-ReLU MoE
activation, RMSNorm, no positional embeddings, untied embeddings. Total 31.6B parameters, ~3.2B active
per token (~3.6B including embeddings).
## Why the cache stays near-constant
Decoding is memory-bandwidth bound: every generated token must re-read the model weights **and** the
attention cache of every sequence in the batch. In a dense full-attention model that per-sequence cache
grows with context, so at tens or hundreds of thousands of tokens, served many-at-once, the KV reads come
to dominate and throughput decays. Soofi's hybrid backbone attacks exactly this. Only 6 of 52 layers keep
a KV cache, with 2 KV heads each, so the incremental attention-cache footprint is about **6 KB per token
per sequence** — the report puts that at **11–53× lower** than the dense models in its comparison. As
context grows, only that small attention component scales with length; the Mamba-2 recurrent state stays
constant-size.
The measured payoff: at 40K context and batch 32, Soofi sustains **4.82k aggregate decode TPS/GPU**, a
reported **9.2×** over Ministral 3 14B, while fitting the weights and all 32 sequence states on a single
GPU. Across 4K→256K the aggregate decode rate stays essentially flat (no point more than ~34% below the
4K value), where dense throughput decays with context. Among the comparison models only Qwen3.5 — itself
a Gated-DeltaNet hybrid — scales similarly, and its 35B-A3B variant still measures ~1.9× slower than Soofi
at 40K. The prefill side shows the same shape: time-to-first-token at 256K is 372.7s for Soofi versus
2,058.9s for dense Ministral 3 14B and 6,428.6s for a dense Qwen3 32B control.
## The data: ~27T tokens, German on purpose
Soofi S was pretrained on approximately **27 trillion tokens** (~26.68T actually consumed) under a
three-phase Warmup–Stable–Decay curriculum: ~20T of diverse, quality-tiered pretraining, ~6.58T of
high-quality annealing, and a ~0.10T long-context extension that pushes the usable window to 1M tokens.
The defining move is that **German is deliberately up-weighted** — to 7.2% of the stable phase and 15.32%
of the annealing mixture, more than triple the ~5% total multilingual share of the reference Nemotron
recipe, and concentrated in a single language rather than spread across dozens.
The corpus is documented at the granularity of individual source datasets — raw tokens, epoch multiplier,
effective tokens, and even sources that were evaluated and *excluded* — so the mixture can be audited and,
where licenses permit, rebuilt. German coverage combines naturally occurring web and document text (HPLT,
German Commons, Genios, German FinePDFs/FineWiki) with machine-translated and synthetic German, since
high-quality native German text is far scarcer than English. The report identifies that German data
pipeline as the principal bottleneck for further gains.
Just as notable is *where* it ran. Soofi S was trained end-to-end on the **Industrial AI Cloud** operated
by Deutsche Telekom in Munich — up to **512 NVIDIA B200 GPUs** (64 DGX B200 nodes), ~253,000 B200
GPU-hours from 24 March to 13 May 2026, on a facility powered by renewable energy and cooled with water
from the Eisbach canal. Training on German soil under European data-protection rules is itself part of the
"sovereign" claim.
## Results, and how to read them
Against a set of large open-source models (Alia 40B, EuroLLM 22B, Apertus 70B, Olmo 3 32B), Soofi S is
the strongest in the set: highest **German aggregate** and, among fully open models, the highest English
and German evaluation scores — ahead of Olmo 3 32B and Apertus 70B despite activating a fraction of their
parameters.
The point the authors most want to land is capability-per-active-parameter: Soofi matches dense 14–27B
models on English and German aggregates while activating only ~3.2B parameters per token.
It also posts the best code aggregates in that comparison (HumanEval 73.8, MBPP 70.2, HumanEval-DE 65.5,
MBPP-DE 84.2 — first on four of five code benchmarks), and leads the set on mathematics (GSM8K 86.1).
Read against the *open-weight* set, the picture is more measured and the report says so: Soofi is **not**
the top model on aggregate — Qwen3.5 35B-A3B leads (English 74.6, German 81.6), and Soofi's English
aggregate (70.1) essentially ties Gemma 3 27B and Ministral 3 14B (both 70.3). Its clearest, cleanest
result is the architecture-identical comparison: versus Nemotron 3 Nano 30B-A3B (same backbone, different
data), the German–English recipe lifts German aggregate +4.2, held-out English +6.7, GPQA-Diamond +9.6,
and German-language proficiency (GLP-DE) +15.1 — while English capability is preserved or improved, the
usual price of monolingual specialization avoided.
## The honest caveats
This is a **consortium tech report, not a peer-reviewed paper**. Every number is **author-reported**,
the "Capability Index" in Figure 1 is **author-defined** (an average of five benchmark groups, each
normalized to the best-plotted model), and the throughput figures use an **author-selected baseline set
and measurement protocol** (TP=1, one B200, batch 32, latency-subtraction). Soofi S is a **3B-active**
model: "matches dense 14–27B" is active-vs-total, not 30B-dense compute. "Best/highest among fully open"
and "outperforms every European sovereign baseline" are **scoped to their comparison and eval suite**,
and deliberate German up-weighting shapes the aggregates — so these are not unqualified SOTA claims.
The report is candid about its own limitations. **Competition-style math in German** is the clearest gap
to the frontier: Minerva MATH-DE 56.0 trails Qwen3.5 35B-A3B (76.5) and Gemma 3 27B (65.6). **Open-domain
factual recall** is capacity-limited — NaturalQuestions 79.0 trails the largest dense baselines (Gemma 3
27B 83.5), consistent with storing world knowledge in only ~3B active parameters (the authors expect
retrieval-augmentation to close this in practice). And on openness itself the report draws an explicit
line: it satisfies the OSI's OSAID 1.0 (weights, checkpoints, training and eval code, exact per-source
data accounting, all under permissive licenses), but falls short of the stricter "every training token
must be redistributable" bar on exactly one component — the commercially licensed Genios corpus (1.3% of
Phase 1) — so ~99% of the mixture, not 100%, can be independently reconstructed.
## The take
Soofi S's contribution is less a new mechanism than a **thesis about deployment cost, executed
transparently**. The near-constant cache is the load-bearing idea: by keeping only 6 of 52 layers as
attention and letting Mamba-2 carry the rest with a fixed-size state, decode throughput stops caring about
context length — which is where dense models bleed. Wrap that in a sparse MoE (3.2B active of 31.6B) and
you get a model that serves like a 3B but scores like a 14–27B dense on its target languages. Set the
knobs honestly — author-reported numbers, an author-defined capability index, a scoped baseline set, a 3B
active budget, real gaps in German competition math and factual recall — and what remains is genuinely
notable: a fully documented, per-source-audited, German–English pretraining run on sovereign European
B200 hardware, released with checkpoints and code. As a template for "open in substance, efficient by
architecture, and built at home," it is a clean and unusually legible bet.
---
*Built on the [Soofi S Pretraining Report v1.0](https://huggingface.co/Soofi-Project) ("A Sovereign,
Open-Source Foundation Model for German and English", the Soofi-Team; consortium coordinated by the KI
Bundesverband, funded by the German BMWE). Architecture, data, and benchmark figures are quoted from the
report for commentary; the interactive diagrams are illustrations of the mechanism, and the throughput
curves use the report's measured endpoints with an illustrative in-between shape. Related: the
architecture-shared [Nemotron in NVFP4](/articles/nemotron-nvfp4), [mixture-of-experts from
scratch](/articles/mixture-of-experts-from-scratch), [how LLM inference
works](/articles/how-llm-inference-works), and [large-scale
pretraining](/articles/megatrain-single-gpu-training).*
---
# Colibri: running a 744B model on a 25 GB machine by streaming experts from disk
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/colibri
> date: 2026-07-10
> tags: inference-optimization, mixture-of-experts, systems, explainer, llm
[GLM-5.2](/articles/glm-5-2) is a 744-billion-parameter model. Loaded the normal way, its
weights want hundreds of gigabytes of fast memory — data-center territory. [Colibri](https://github.com/JustVugg/colibri)
is a single-file, pure-C engine, zero external dependencies, that runs that same model on a
consumer machine with about **25 GB of RAM**. Not a distillation, not a smaller sibling — the
full 744B checkpoint, answering correctly, on a box that costs less than one H100's cooling fan.
Before the "how," the honest headline: **this is a feasibility feat, not a usable-speed setup.**
Cold, Colibri decodes at roughly **0.05–0.1 tokens per second** — that is *10 to 20 seconds per
token*. Warm, with every trick engaged, the best real community result is about **0.37 tok/s**,
still ~3 seconds per token. And "25 GB of RAM" is only true if you *also* have ~**370 GB of fast
NVMe** for the experts. Keep both numbers in the same sentence or the claim is misleading.
**Read the fine print before you get excited.** (1) It is *slow* — 0.05–0.1 tok/s cold is 10–20
seconds **per token**; the best community case, ~0.37 tok/s, is still ~3 s/token. (2) "Runs on 25 GB
RAM" **requires ~370 GB of fast disk** for the experts — the RAM number alone is misleading. (3) The
numbers here are **community-reported / single real test cases**, not official benchmarks — treat them
as existence proofs, not spec sheets. (4) It works only because GLM-5.2 is an **extremely sparse MoE**
*and* because of **int4/int8 quantization** — remove either and the trick collapses.
## Why this is even possible: extreme sparsity
Colibri does not compress 744B parameters into 25 GB. It exploits the fact that, at any moment,
almost none of those parameters are doing work. GLM-5.2 is a
[Mixture of Experts](/articles/mixture-of-experts-from-scratch): each MoE layer holds **256 experts**,
but a router picks only a small **top-k** of them per token. Across the model, only about **40B of the
744B parameters activate for a given token**, and of those, only ~**11 GB of weights** actually *change*
from one token to the next — the routed experts. Everything else (attention, shared experts, embeddings —
the "dense" ~17B) is used on *every* token.
That split is the whole design. Split the model by how often each piece is touched:
- **The dense core (~17B params)** is touched every token → keep it **resident in RAM**, int4-quantized
to **~9.9 GB**.
- **The routed experts (21,504 of them: 75 MoE layers × 256, plus the MTP head, ~19 MB each at int4)**
are touched rarely and unpredictably → leave them **on disk (~370 GB)** and fetch only the handful a
token actually routes to.
The second enabler is quantization. The original FP8 checkpoint is ~**756 GB**; Colibri's offline
converter requantizes it to int4 (the experts) so the dense core fits in ~9.9 GB of RAM and the expert
store shrinks to ~370 GB on disk. Sparsity says *you rarely need most experts*; quantization says *the
ones you do need are small enough to stream*. You need **both**.
## The three-tier memory, made visible
So the memory hierarchy has three tiers: the **int4 dense core resident in RAM**, an **LRU cache of hot
experts in RAM**, and the **full expert store on disk**. When the router picks an expert, one of two
things happens. If that expert is already in the RAM cache (or pinned), it is a **hit** — served at
memory speed. If not, it is a **miss**: a disk read of ~19 MB that sits **on the critical path** of that
token. The token cannot finish until the bytes arrive.
Route a token through one MoE layer and watch it resolve. Scrub the token, resize the cache, toggle
hot-expert pinning, and watch the LRU fill and evict:
This is the entire performance story in one picture. The router picks its experts; the cache absorbs the
ones you keep re-using; everything else is a disk fetch. A cold token — nothing cached — reads about
**11 GB from disk** (75 layers × ~8 experts × ~19 MB). At a typical NVMe rate of ~1 GB/s, that read alone
is ~10 seconds, which is exactly why cold decode lands at 0.05–0.1 tok/s. **Decode is disk-bound**, not
compute-bound: the CPU sits waiting on I/O. This is the same memory-wall intuition as ordinary
[LLM inference](/articles/how-llm-inference-works), pushed to its limit — except the "memory" the decode
waits on is a spinning queue of NVMe reads instead of GPU HBM.
Colibri softens the disk with the usual systems tricks: an **LRU cache** so repeat experts stay hot, an
**async readahead** that reads the *next* block of experts while the current one is still multiplying (so
compute and I/O overlap), the **OS page cache** acting as a free second-level cache, and a **RAM safety
budget** — the cache is auto-sized from `MemAvailable` at startup so it fills spare memory without ever
triggering an OOM kill. None of these change the physics of a cold miss; they just make misses rarer.
## Budget the RAM and the disk together
Because the "25 GB" number is the seductive, misleading one, it is worth drawing to scale. On a 25 GB
machine the resident footprint is the 9.9 GB int4 core plus an auto-sized LRU expert cache plus a safety
headroom — and the 370 GB of experts sit on disk, roughly **15× larger than the entire RAM**. Slide the
cache size and see how little of the model is ever in memory at once:
That 15× gap is the point. Colibri did not shrink GLM-5.2 to fit in RAM; it arranged for only the
in-use ~4% to be in RAM at any instant, and made disk the backing store for the rest. Take away the fast
NVMe and there is no engine — which is why quoting the RAM figure without the disk figure sells a fiction.
## Buying speed back: warm cache, MTP, and pinning
Cold decode is the floor, not the experience. Three things stack on top of it, and it is worth being
precise about what each one is.
**A warm cache** is just locality: real prompts re-route to the same experts, so after a few tokens the
hottest experts live in RAM and the miss rate drops. **Pinned hot experts** make that permanent — Colibri
records which experts your usage actually routes to (a `.coli_usage` file) and pins the hottest ones in
spare RAM, so the engine *literally gets faster the more you use it*.
**MTP** is the subtle one. GLM-5.2 ships a native [multi-token-prediction](/articles/multi-token-prediction)
head — a lightweight draft model that proposes several future tokens, which the main model then *verifies*
in one batched forward. Colibri runs it natively at **int8** (this matters: at int4 the draft head's
predictions are so degraded that acceptance collapses to 0–4% and speculation never engages; at int8 it
reaches ~39–59% acceptance, community-measured). The payoff is **2.2–2.8 tokens per forward** — but note
that is a *speculation/acceptance rate*, the number of tokens you get out of one main-model pass, **not**
tokens per second. It amortizes the fixed per-forward cost; it does not make the disk faster. On a *cold*
cache MTP can even be a net loss, because verifying extra draft tokens routes to *more* experts
(~660 → ~1100 expert-loads/token) — so speculation only pays once the cache and pins are warm.
Stack them, and honestly label which rungs are measured versus an illustrative split:
The endpoints are real community numbers; the per-factor decomposition in the middle is illustrative
(Colibri's README reports the endpoints, not the split). The takeaway is the seconds-per-token column: the
best real result, **0.37 tok/s on a Ryzen AI 9 Framework 13** with a warm cache, MTP and pinning, is still
about **one token every three seconds**. A different community machine — an M5 Max with **128 GB of RAM** —
reaches **1.06 tok/s**, but only because far more RAM lets far more experts stay resident, which just
confirms the thesis: *the disk is the bottleneck, and RAM buys you out of it.*
## Colibri is the engine, not the model
Keep the two things distinct. [GLM-5.2](/articles/glm-5-2) is the *model* — the 744B MoE, its routing, its
MTP head, its training. **Colibri is an *engine*** that runs that model's forward pass in pure C on tiny
hardware. The cleverness is entirely in the *systems* layer: how to lay out weights on disk, when to read
them, what to cache, how to overlap I/O with compute, how to quantize the head so speculation survives.
It reimplements GLM-5.2's forward pass faithfully — MLA attention with a compressed KV cache (~57× smaller
than dense), DeepSeek-V3-style routing, native MTP — but adds nothing to the model's *capability*. Same
weights, same answers; the contribution is fitting them onto a laptop.
## The take
Colibri is a lovely demonstration of a real principle: **a sparse MoE's active footprint, not its parameter
count, is what a runtime has to hold in fast memory.** Because GLM-5.2 fires only a few of its 256 experts
per layer per token, and because int4/int8 quantization shrinks both the resident core and each streamed
expert, the working set collapses from hundreds of gigabytes to ~10 GB of RAM plus on-demand disk reads.
That is genuinely clever, and it is pure C with zero dependencies, which makes it a beautiful object to read.
But be clear-eyed about what it buys. It buys *access*, not *speed*: you can hold a conversation with a
744B frontier model on a 25 GB machine, at the pace of a few seconds per token, provided you also own a
370 GB fast disk. The bottleneck is not going away — it is the physics of pulling ~11 GB across NVMe for
every cold token — and the caching, pinning and speculation only push against it. As an existence proof
that extreme sparsity plus aggressive quantization can put a frontier model on consumer hardware, Colibri
is compelling. As a way to actually *use* one at interactive speed, it is not there, and it is refreshingly
honest about that.
---
*Built from the [Colibri repository](https://github.com/JustVugg/colibri) (JustVugg; Apache-2.0) and its
README. All throughput and hardware figures are **community-reported single test cases**, not official
benchmarks: cold ~0.05–0.1 tok/s and MTP 2.2–2.8 tok/forward on the dev machine; 0.37 tok/s (Ryzen AI 9,
Framework 13), 1.06 tok/s (M5 Max, 128 GB), 0.11 tok/s (Core Ultra 7, 24 GB) from community runs.
Architecture figures — 744B total, ~17B dense, 9.9 GB int4 core, 21,504 experts, ~370 GB on disk — are from
the repository README. The interactive diagrams are illustrations of the mechanism, not measurements.*
---
# One small box, 64 users at once: continuous batching on a DGX Spark
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/dgx-spark-batching
> date: 2026-07-10
> tags: inference-optimization, llm, systems, explainer, moe, quantization
A [DGX Spark](/articles/deepseek-dspark) is not a datacenter. It's a small GB10 Grace-Blackwell
box with **128 GB of unified memory** — the kind of thing that sits on a desk. And yet one of them,
running [vLLM](https://github.com/vllm-project/vllm), will hold a conversation with **dozens of people
at the same time** on a 35B-class model. The surprising part isn't the model or the silicon. It's a
scheduling trick called **continuous batching**, and it's the single most important reason a small box
can feel like a shared server.
The instinct is to picture the users taking turns — one prompt finishes, the next begins. That's not
what happens. At every decode step the server bundles **all** the currently-active conversations into
**one** forward pass through the GPU, and appends exactly one new token to each. Sixty-four people, one
pass. Drag the concurrency and step the clock to see it move:
## Why decoding one token is a waste of a GPU
To see why batching helps so much, you have to look at what generating a single token actually costs.
LLM inference splits into two phases — the [prefill and decode](/articles/how-llm-inference-works) —
and it's the **decode** phase, one token at a time, that dominates a chat session.
Decoding one token for one user means: load the model's weights out of memory, multiply the single
current token's vector through them, read that user's entire [KV cache](/articles/turboquant-kv-cache)
to do attention, and emit one token. The arithmetic is tiny — one token — but you had to stream **all
the weights** through the chip to do it. That makes single-stream decode **memory-bandwidth-bound**:
the GPU's compute units sit almost idle, waiting on memory. You paid to load the weights and barely
used them.
Continuous batching is the fix that falls out of that observation. If you're loading the weights
anyway, run *more* tokens through them on the same load. Gather N users' current tokens, stack them,
and do one fused matmul against the weights you already fetched. The weight-load cost — the expensive
part — is now **amortized over N streams instead of one**. Aggregate throughput climbs steeply, and
keeps climbing until you run out of either compute or KV-cache memory.
That's the trade in two curves. **Aggregate** tokens/second rises and saturates; **per-user**
tokens/second falls the whole way. Both are true at once, and conflating them is the most common way
these numbers get oversold.
## "Continuous," not just "batched"
Plain static batching — wait for N requests, run them together, wait for all N to finish — would be
useless for chat, because requests arrive at random times and finish at wildly different lengths. One
long generation would stall everyone.
vLLM's batching is **iteration-level** (also called *in-flight* batching): the batch is re-formed
**every single decode step**. A request whose prompt just arrived joins the batch on the next step; a
request that just emitted its stop token leaves it and frees its memory immediately. The GPU never
blocks waiting for the slowest member — that's the "continuous" part, and it's what the join/leave
churn in the first diagram is showing. Streams flow through a batch that is constantly being rebuilt.
The enabling piece underneath is **PagedAttention**. Each user's KV cache is stored not as one
contiguous slab but as a list of fixed-size **blocks** (the paged rows in the diagram), allocated on
demand from a shared pool — exactly like virtual memory pages. Without it, fitting 64 independent,
different-length caches into one memory space would fragment badly and you'd waste most of it. With it,
64 caches pack tightly, and a finished request's blocks return to the pool for whoever's next. KV-cache
memory — not FLOPs — is what ultimately caps how many users fit.
## The numbers, and which ones I trust
Here's where honesty matters. The story above is mechanism, and it's solid. The specific figures need
sorting into what the public [spark-bench](https://github.com/Weschera/spark-bench) results actually
contain versus what's a reported run.
The model is **Qwen3.6-35B** — specifically an **A3B mixture-of-experts**: ~35B total parameters but
only **~3B active per token**. That's half of why it's fast (you only compute a fraction of the network
each step) and why it fits comfortably (the weights are also **quantized** — the committed throughput
runs use vLLM with **NVFP4**, a [4-bit format](/articles/nemotron-nvfp4), and an FP8 variant). A 35B
model serving this briskly on a 128 GB box is a *quantized MoE*, not a dense fp16 35B — and that
distinction is load-bearing, not a footnote.
The committed `spark_bench.csv` sweeps concurrency **1 → 16** and shows the batching curve directly. On
the NVFP4 build, aggregate decode throughput roughly triples from a single user to sixteen:
…while each individual user's stream slows by roughly the same factor — you're trading personal latency
for collective capacity:
**Read the caveats before quoting a headline number.**
- **700+ tok/s is _aggregate_, across all streams — not per user.** At 64 users that's ≈ **11 tok/s each**,
a modest personal reading speed. Continuous batching raises *throughput*, not single-stream latency;
it does not make any one user's tokens arrive faster (time-to-first-token actually rises with batch
size, from ~0.3 s single-stream to ~2.5 s at concurrency 16 in the committed runs).
- **The 64-user / ~700 tok/s / 32,768-tokens-in-54-seconds figures are a _reported_ run, not one I could
confirm in the repo.** The committed sweep for Qwen3.6-35B tops out at concurrency **16** (~217 tok/s
median, ~450 in the best single run). A ~723 tok/s aggregate *does* appear in the committed data — but
at concurrency **32**, and for a *different* model ([laguna](/articles/laguna-model-factory)), not this
one. So the 64-user number is consistent in magnitude with where the curve is heading, but treat it as
reported, not verified.
- **The ~38 W figure I could not verify at all** — there is no power column anywhere in the committed
results. The DGX Spark's whole-box envelope is well above 38 W under load, so whatever that number
measures (a sub-component? an idle draw?), take it as reported until someone publishes the methodology.
- **Quantization is doing real work.** These rates are for a **4-bit (NVFP4) / 3B-active MoE**, not a
dense fp16 35B. Different precision, different story.
## The take
Continuous batching is the quiet reason "local inference" and "serving other people" stopped being
mutually exclusive. The mechanism is honest and general: decode is memory-bound, a lone stream wastes
the GPU, so amortize the weight-load across as many live streams as KV-cache memory will hold, rebuilding
the batch every step so nobody waits on anybody. On a DGX Spark that turns a desk-sized box into a
small shared server — genuinely dozens of concurrent users on a quantized 35B MoE.
What it is *not* is a speedup for the person on the other end. Each user's tokens come at a human
reading pace, and that pace gets slightly worse as the room fills up. The DGX Spark headline — one small
box, many users — is real and impressive; it's just a statement about **aggregate** capacity and clever
scheduling, not about raw single-stream speed. Hold both halves of that at once and the number stops
being a magic trick and becomes what it actually is: good systems engineering.
---
*Mechanism (continuous / in-flight batching, PagedAttention) is standard
[vLLM](https://github.com/vllm-project/vllm). Throughput and latency figures are read from the committed
[spark-bench](https://github.com/Weschera/spark-bench) `results/spark_bench.csv` (Qwen3.6-35B-A3B, vLLM
NVFP4/FP8, concurrency 1–16); the 64-user / ~700 tok/s / ~38 W figures are a reported run and are labelled
as such above. The interactive diagrams illustrate the mechanism; their tok/s curve is a saturating fit
pinned to the committed points and the reported endpoint.*
---
# KAT-Coder-V2.5: training a coding model to live inside a repository
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/kat-coder-agentic-training
> date: 2026-07-10
> tags: llm, agents, reinforcement-learning, code-generation, explainer
A "coding model" usually means a next-token predictor you paste a function stub into. **KAT-Coder-V2.5**, from the Kwaipilot / Kuaishou team, is trained for a different job: to *act* — to open a real repository, run its tests, read the failures, edit files, and iterate until the tests pass, across dozens of turns. Once that's the target, the model stops being the hard part. The hard part is manufacturing enough **precisely specified, executable, objectively verifiable** tasks to train on, and building an RL loop that survives long, sparse-reward trajectories. This report is mostly about that manufacturing stack. It pairs with our pieces on [agent harnesses](/articles/agent-harness) and on [near-frontier code RL](/articles/swe-1-7); KAT-Coder is a full-stack answer to the same question those raise.
The one idea to leave with: **an autonomous coding agent is an environment-and-data problem before it is a modeling problem.** Everything below — AutoBuilder, the data flywheel, harness randomization, the asymmetric critic, multi-teacher distillation — exists to feed a policy verifiable practice inside realistic scaffolds, without letting it overfit any single one.
## AutoBuilder: turning real repos into verifiable sandboxes
The raw material is public repositories, but an issue title and a merged PR are not a task. Two problems have to be solved. First, **task mining**: raw issue/PR text is ambiguous and often misaligned with what actually got merged, so AutoBuilder regenerates a structured spec from the *golden patch* and *test patch* — a **problem statement** (from the golden patch), **requirements** (from the test patch), and **interface constraints** (from both) — then runs a clarity check to ensure it's self-contained.
Second, **environment construction**: a build agent writes a configuration script and a separate verification agent runs it in an isolated sandbox, accepting the environment only when it can collect **>90% of the expected tests** with reproducible fail-to-pass and pass-to-pass outcomes (exit codes and log-scraping are explicitly rejected as too easy to fool). Combining base images, language templates, and retrievable build recipes, this reaches a **57.2% environment-construction success rate** (up from a 16.5% starting point) and yields **>100,000 verifiable environments across 12 languages**. That corpus of executable tasks is the substrate everything else trains on.
## The data flywheel: recovering near-misses without leaking hints
Verifiable environments are necessary but not sufficient: on genuinely hard tasks the model's raw pass rate is near zero, so rollouts produce no learning signal. KAT-Coder's answer is a two-stage recovery loop. First, inject **process-level hints** to lift near-misses over the line, raising the pass rate from ~0% to **~20%**. But a trajectory that only succeeds *because it was handed a hint* teaches the model to expect hints it won't have at deployment — so the second stage **replays the same tasks hint-free** and keeps only the trajectories that still recover. The final data carries no hint leakage and stays faithful to the original distribution. Drag the hint strength and step the stages:
What survives then passes a **process-score filter** that scores each trajectory on exploration, localization, pre-edit reasoning, specification fidelity, repository conventions, patch minimality, verification quality, recovery, and honesty — down-weighting exploitative or unstable behavior even when the tests happen to pass. A parallel **harness-robustness** step randomizes tool names, argument conventions, and output formats and injects realistic perturbations (missing dependencies, transient failures, truncated outputs, noisy logs), so trajectories don't encode a single scaffold's quirks.
## KwaiClawEnv: general agentic tool use
Repository work isn't the only skill. **KwaiClawEnv** is a three-layer pipeline that manufactures general tool-use trajectories: a **Service layer** builds callable capabilities from human-authored Skills and LLM-generated Services (>90% generation success from open-source community Skills), a **Task layer** expands real task seeds into variants with configurable difficulty and tool-chain length, and an **Eval layer** converts rollouts into SFT-ready samples and feeds quality signals back upstream. It yields **>100,000 high-quality instances** with an **average of 15 tool calls** per task and the longest **exceeding 100 steps** — genuinely long-horizon agentic data, validated through a three-stage checker (service reachability, task-schema legality, sandboxed execution) and two-layer filtering (hard rules plus LLM-as-judge).
## Reinforcement learning, hardened for long horizons
### Harness randomization
An agentic policy trained inside one fixed scaffold overfits the *surface* of that scaffold, not the task. The report names three failure modes: **format overfitting** (anchoring to one action format, so parsing breaks when the protocol changes), **context-structure overfitting** (depending on how history is concatenated), and **control-flow overfitting** (relying on a fixed reflection/stop schedule). The fix is to train across many harnesses that vary along three axes — tool-invocation protocol, context management, and control flow — spanning **white-box** harnesses (like mini-swe-agent: simple, uncompressed, clean signal) and **black-box** production harnesses (Claude Code, Codex, OpenClaw, OpenHands: with compression and context reorganization). Flip the axes and watch the rendered action change while the task — and its reward — stay put:
Underneath, the sandbox itself had to be hardened: container-image and environment-variable bugs meant **~16% of trajectories** initially failed for reasons unrelated to the model. Fixing disk pressure (95% → 60% usage) and timeouts drove sandbox-related failures to **below 2%**, and invalid-rollout rates from 6–7% to under 1%.
### Asymmetric PPO with a hindsight critic
Long-horizon tasks are sparse-reward: the signal lands only at the end, when the tests pass or fail. That makes the critic's job — estimating the value of an intermediate state — brutally high-variance, and a noisy value function means a noisy advantage, which destabilizes training. KAT-Coder's move is an **asymmetric actor–critic**: the actor sees only the normal harness state $s_t$ (so it behaves identically in training and deployment), while the critic is given a *privileged* **hindsight context** $c_t$ — the eventual reward, test outcomes, coverage signals, patch-level diffs, trajectory statistics, and subsequent turns. Scrub the turn and toggle hindsight to see the value estimate tighten:
Concretely, the standard clipped PPO objective is optimized:
$$
\mathcal{J}_{\mathrm{PPO}}(\theta)=\mathbb{E}_{q,\,o}\left[\frac{1}{|o|}\sum_{t=1}^{|o|}\min\!\left(r_t\hat{A}_t,\ \mathrm{clip}(r_t,1-\epsilon,1+\epsilon)\hat{A}_t\right)\right], \quad r_t=\frac{\pi_\theta(a_t\mid s_t)}{\pi'(a_t\mid s_t)},
$$
with advantages from GAE, $\hat{A}_t=\sum_{l=0}^{T-t-1}(\gamma\lambda)^l\delta_{t+l}$ and $\delta_t=r_t+\gamma V'(s_{t+1})-V'(s_t)$. The asymmetry is entirely in the value function: the critic conditions on the hindsight context, so its regression target uses $V_\psi(s_t,c_t)$ rather than $V_\psi(s_t)$:
$$
\mathcal{L}_{\mathrm{critic}}^{\mathrm{asym}}(\psi)=\mathbb{E}_{(s_t,c_t,R_t)}\left[\left(V(s_t,c_t;\psi)-R_t\right)^2\right].
$$
Because $c_t$ contains information the actor can't see, the value estimate is far less of a blind guess — lowering advantage variance without ever contaminating the deployed policy. Reward itself is a two-part framework: a **rule-based** reward with a core task score (all fail-to-pass *and* pass-to-pass tests must pass for full credit), eight behavioral constraints (duplication, garbled output, tool-call accuracy and placement, redundant calls, parallelism, debug-artifact cleanup), and failure-path incentives (file-search $F_2$, unit-test pass rate); plus a **model-based** generative reward model scoring fault diagnosis, post-fix validation, and execution strategy. The reported SWE training curve rises stably throughout.
### Multi-Teacher On-Policy Distillation
Rather than run five separate RL specialists and hope they compose, KAT-Coder fuses them with **Multi-Teacher On-Policy Distillation (MOPD)**: the student generates on-policy, and for each domain $d$ its distribution is pulled toward that domain's expert teacher $\pi_{T_d}$ under a reverse-KL objective —
$$
\mathcal{L}_{\mathrm{MOPD}}(\theta)=\mathbb{E}_{(x,d)}\,\mathbb{E}_{y\sim\pi_\theta(\cdot\mid x)}\left[\sum_{t=1}^{|y|}w_t\,\mathrm{KL}\!\left(\pi_\theta(\cdot\mid x,y_{
The picture is not uniform. On tool-use and repo-level SWE, KAT-Coder is at or near the frontier; on **Terminal-Bench 2.1 it scores 60.7** — behind GLM-5.2 (77.9), Kimi-K2.6 (73.0), and Opus 4.8 (84.6) — and on **SciCode 50.3**, behind GLM-5.2 (50.5) and both Kimi-K2.6 and Opus 4.8 (53.5). On KAT Claw Bench it reaches 85.5, with GLM-5.2 (86.8) and Opus 4.8 (90.7) ahead. The efficiency of the *training stack* — not inference — is the story; the paper reports no wall-clock or cost comparison against the frontier panel.
Honest caveats. **The panel and two of the six benchmarks are author-chosen.** KAT Code Bench and KAT Claw Bench are the team's *own* new benchmarks; the comparison excludes other open agentic-coding methods and reports no ablations isolating what each component (harness randomization vs. the hindsight critic vs. MOPD) actually contributes — the "16% → <2%" sandbox and "0% → ~20%" hint numbers are engineering deltas, not controlled ablations. The wins are **real but scoped**: first on PinchBench, second on SWE-Bench Pro, but **behind every panel model on Terminal-Bench 2.1 and behind three of four on SciCode** — the report itself flags terminal and scientific tasks as open weaknesses. The whole stack (AutoBuilder, KwaiClawEnv, the benchmarks) is internal infrastructure, so transfer to non-standard or closed repositories is unproven, and long-horizon credit assignment is "partially addressed" but still called a challenge.
## The take
KAT-Coder-V2.5 is best read not as a model but as a **recipe for the surrounding machine**. Its most transferable ideas are architectural in the systems sense: verify environments by actually collecting tests (not scraping logs), recover near-misses with hints and then *strip the hints* so the data stays honest, randomize the harness so the policy learns the task instead of the scaffold, and hand the critic — but never the actor — a view of the future to tame sparse-reward variance. Set against [rollout-stability work like Routing Replay](/articles/rollout-routing-replay) and [cheaper frontier RL](/articles/frontier-rl-cheaper), the throughline is the same: at long horizons, most of the win is in the plumbing, not the loss function. Whether the fixed recipe generalizes past the team's own repositories and benchmarks — and closes the terminal and scientific gaps — is the open question; but as an account of what it takes to make a coding model *live inside a repository*, it's unusually complete.
---
*Built on the [KAT-Coder-V2.5 Technical Report](https://arxiv.org/abs/2607.05471) (Huang, Li, Xu et al.; Kwaipilot / Kuaishou, 2026). All benchmark values are quoted from the paper's Table 4 (unified Claude Code harness); the interactive diagrams are illustrations of the mechanism, not measured traces. PinchBench averages are the paper's, retrieved from pinchbench.com on 2026-07-02.*
---
# MusaCoder: teaching a model to write GPU kernels with execution-feedback RL
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/musacoder-gpu-kernels
> date: 2026-07-10
> tags: llm, reinforcement-learning, gpu, cuda, code-generation, explainer
Ask a language model to turn a PyTorch module into a hand-written CUDA (or [MUSA](https://en.wikipedia.org/wiki/Moore_Threads), Moore Threads' CUDA-alike) kernel and you hit a task that punishes it at every turn. The output has to compile against a real toolchain, launch without an illegal memory access, produce numerically-matching results across dtypes and shapes, **and** run faster than the reference — or it is worthless. Most first attempts fail outright, which is exactly what makes the obvious training recipe, execution-based reinforcement learning, so hard: when nearly every rollout scores the same failing reward, there is no gradient to learn from. Worse, the model quickly discovers it can "pass" a correctness check by quietly calling the very PyTorch operator it was asked to replace.
**MusaCoder** (Cheng et al., Moore Threads, 2026) is a full-stack answer to that. It is less a single trick than an assembled pipeline — data synthesis, supervised and rejection fine-tuning, and RL against an execution verifier — with three stabilization mechanisms that keep the RL from collapsing. The result: a 9B model built on Qwen3.5-9B that matches the frontier closed models the paper evaluates, and a 27B model (on Qwen3.6-27B) that tops them on the paper's KernelBench protocol.
## The reward is where correctness gets enforced
Everything downstream depends on one design choice: the scalar reward `s(c)` that MooreEval — the paper's distributed compile/execute/verify environment — assigns to a candidate kernel `c`. MooreEval first produces a structured verdict `V(c,x) = (compiled, correct, legal, speedup, category, detail)`, and collapses it into a **correctness-first** reward (Equation 2):
$$
s(c)=\begin{cases}
-1, & \text{extraction / compile / runtime fails},\\
-1, & \text{a disallowed PyTorch/aten::* fallback is detected},\\
-1, & q=0,\\
-0.5+0.5\,q, & 0
That single wall at zero is what keeps the model honest. Speed is never rewarded until correctness is banked, and forbidden fallbacks are punished as hard as a crash — so the fastest way to positive reward is to actually write the kernel.
## Data and fine-tuning: getting off the floor before RL
RL only works if the base model clears the bar *sometimes*. MusaCoder spends most of its pipeline manufacturing that starting competence. A three-stage data engine expands the PyTorch-to-CUDA/MUSA workload distribution (real modules, cleaned GitHub projects, and NNSmith-generated computation graphs across ~162 operators), injects tensor **shape/stride/contiguity** hints extracted with `torch.fx`/`torch.export`, and enforces a six-step structured-reasoning template before any code is written. Auxiliary corpora add GPU-kernel knowledge Q&A and a kernel-**reviewer** task that must emit `VERDICT: CORRECT` or `INCORRECT`.
Two fine-tuning stages follow. Multi-task **SFT** teaches canonical kernel patterns and error-diagnosis (with loss masking so feedback tokens are context-only). Then a **diversity-preserving rejection fine-tuning (RFT)** step deliberately breaks with convention: standard RFT keeps only the single fastest correct sample, which collapses entropy; MusaCoder instead retains a *heterogeneous* set of verified-correct implementations, preserving the exploration diversity that RL will need. That choice alone is worth 2.2 points of Pass@8 (SFT 84.8 → 82.6 without RFT, Table 2). If you have not met [`torch.profiler`](/articles/torch-profiler) — the same tool MusaCoder uses to diagnose which operator families the base model is weak on — it is worth a detour.
## The three stabilizers
RL runs in two stages: a single-turn warmup to establish basic execution understanding, then multi-turn feedback RL where a failed kernel gets MooreEval's error log appended and the model tries again. On top of a GRPO objective, three mechanisms keep it from falling over.
### PrimeEcho — anchor the reward to the turn that ships
In multi-turn RL the naive reward is the best score across all turns, `max_k s_k`. But the model only ever *deploys* its first-turn kernel, and rewarding best-of-turns teaches it to defer correctness — ship something broken, then "fix" it once the verifier hands it the error. PrimeEcho blends the two (Equation 9):
$$
R_{\tau} = \alpha\,s_{1} + (1-\alpha)\max_{1\le k\le K} s_{k} + b_{\text{early}}(\tau),
$$
with an early-success bonus `b_early = β₁·1[success at turn 1] + β₂·1[fail at 1, success at 2]`. Keeping α high anchors the reward to zero-shot quality while still letting later turns supply exploration signal. Slide α and watch a deliberately-late trajectory get *more* reward as the anchor weakens — the exact hack PrimeEcho suppresses:
### Buffered Dynamic Retry — rescue the all-failed groups
GRPO normalizes advantages within a group of `G` rollouts. When a task is hard enough that **all** `G` samples fail — `r_i = −1` for every `i` — the advantages are all zero and the sample contributes **no gradient**, so the hardest tasks teach nothing. Buffered Dynamic Retry (BDR) composes a *repair task* `x' = Compose(x, c⁻, f⁻)` from a failed kernel and its feedback, pushes it into a FIFO buffer `B`, and mixes buffered repair tasks back into training with probability `p_buf`. It turns a dead rollout group into a feedback-conditioned second chance. In the paper's isolated test (Table 3) BDR lifts Pass@8 from 59.6 → 62.4 on a Qwen3-8B checkpoint (~16% of previously-failed tasks recovered) and 73.2 → 74.4 on Qwen3.5-9B (~28% recovery).
### MirrorPop — catch the off-policy sequences that cancel
MusaCoder's rollouts are generated asynchronously, so the rollout policy drifts from the training policy and the per-token importance ratio `ρ_t` no longer sits at 1. Vanilla sequence-level masking scores a response by its *signed* mean log-ratio — but a response that is badly off-policy with roughly equal positive and negative deviations averages to ≈0 and slips through as if it were on-policy (the paper's Figure 11 "cancellation" case). MirrorPop instead uses the mean **absolute** log-ratio, which every token can only push upward, and masks the sequence when it exceeds a threshold δ (Equation 21):
$$
M_i^{\text{mirrorpop}} = \mathbf{1}\!\left[\frac{1}{L_i}\sum_{t=1}^{L_i}\big|\log \rho_{i,t}\big| \le \delta\right].
$$
Toggle the two responses below — an on-policy one and a drifted one whose ratios cancel — and see which filter catches the drift:
This is the same failure mode that [Rollout Routing Replay](/articles/rollout-routing-replay) fixes at its source for MoE routers and that [async frontier-RL setups](/articles/frontier-rl-cheaper) wrestle with generally — here it is handled at the masking layer. Of the three stabilizers, MirrorPop is the one whose removal hurts most.
## The numbers
MusaCoder is evaluated under its own strict MooreEval protocol on KernelBench, split into Level 1–3 by difficulty. **Pass@8** asks whether at least one of 8 samples passes verification; **Avg.@8** is the mean correctness rate across the 8; **Faster Rate** counts a candidate only if it is correct, legal, *and* beats the baseline by more than 1.1×. The headline is overall correctness — MusaCoder-27B-RL reaches **93.2 Pass@8 / 88.6 Avg.@8**, ahead of every model the paper tests:
The gap widens on the hardest tier. On **Level 3**, MusaCoder-27B-RL scores **72 Pass@8 / 65.75 Avg.@8** against Claude Opus 4.7's 54 / 39.25 and GLM-5.1's 54 / 38.50 — the RL model roughly doubles the average correctness of the frontier baselines on the tasks where kernels are hardest to get right. Notably the 9B model (77.2 Avg.@8) edges Claude Opus 4.7 (77.3 is essentially tied) despite being a fraction of the size, and the RL stage is decisive: MusaCoder-27B jumps from 79.4 (SFT) to 88.6 (RL) Avg.@8.
Correctness is the easy win; **speed is much harder**. Even a good kernel rarely beats a fused `torch.compile` baseline, so absolute Faster Rates are low across the board — but MusaCoder still leads:
Against `torch.compile` (a tougher bar), MusaCoder-27B-RL's Faster Rate is 9.2% vs Claude Opus 4.7's 7.5%. On the authors' ported **MUSA KernelBench** (Table 4), the 27B model leads on both correctness and speed (92.4 Pass@8 / 81.7 Avg.@8 / 12.5 Faster) over DeepSeek-V4-Pro (92.0 / 56.9 / 5.7) and GLM-5.1 (88.0 / 66.4 / 6.9).
The ablation (Table 2) confirms each stabilizer earns its place, measured as removals from the full RL model (93.2 Pass@8):
Dropping MirrorPop costs the most (93.2 → 86.0), consistent with off-policy drift being the dominant instability in asynchronous kernel-generation RL.
## The honest caveats
The comparison is **provider-run and provider-defined**. MooreEval, the strict verification protocol, the difficulty split, and the *MUSA* KernelBench variant are all authored by the same team as the model; the closed frontier baselines (Claude Opus 4.7, DeepSeek-V4-Pro/-ProMax, GLM-5.1, Kimi K2.6) were evaluated by the authors under that protocol, not self-reported. Read the numbers as "MusaCoder wins on the bench MusaCoder built," which is a real result but not a neutral one.
A few more things worth stating plainly:
- **The base models are already strong.** MusaCoder-27B starts from Qwen3.6-27B (67.2 Pass@8) and MusaCoder-9B from Qwen3.5-9B — the recipe adds a lot on top, but this is not a from-scratch capability.
- **Speed remains the weak axis.** A ~15% Overall Faster Rate means the *large majority* of even correct generated kernels do not beat the reference. The framing is correctness-first for a reason; treat the speedups as a bonus, not the story.
- **The reward's tuning knobs aren't disclosed.** The paper leaves `λ` (performance weight) and `ν_max` (speedup clip) as symbols; the values in the reward interactive above are illustrative, chosen to show the shape, not read from the paper.
- **No dedicated limitations section.** The paper does not enumerate its own failure modes or generalization limits, so the boundaries of the approach — how it holds up on operator families outside the ~162 synthesized, or on GPUs beyond the CUDA/MUSA pair — are left for the reader to infer.
## The take
MusaCoder's real contribution is not any one of its parts but the recognition that execution-feedback RL for kernel generation fails in *three specific, nameable ways* — sparse rewards, multi-turn reward hacking, and off-policy cancellation — and a targeted fix for each. The correctness-first reward makes cheating pointless; PrimeEcho keeps the model honest about the turn that ships; BDR rescues the hardest tasks from the dead-gradient zone; MirrorPop stops async drift from poisoning the update. It is careful RL engineering more than a new algorithm, and the payoff — a 9B model tied with the closed frontier and a 27B model ahead of it on the paper's bench — is the kind of result that only shows up when every stage of the pipeline is doing its job. Whether the lead survives a neutral, third-party benchmark is the open question the provider-run setup leaves on the table.
---
*Built on [MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU](https://arxiv.org/abs/2606.04847) (Cheng, Lu, Liao et al.; Moore Threads, 2026; CC BY-SA 4.0). All benchmark numbers are quoted from the paper's tables; the interactive diagrams illustrate the reward and stabilization mechanisms and use illustrative parameter values where the paper leaves them unspecified.*
---
# Switch Transformers: route every token to exactly one expert
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/switch-transformer
> date: 2026-07-10
> tags: mixture-of-experts, llm, architecture, deep-learning, explainer
Every modern giant open-weights model — [LongCat 2.0](/articles/longcat-2) at 1.6T
parameters, and the rest of the sparse-MoE zoo — runs on one idea: don't run all the
parameters on every token. Keep a big pile of experts, and for each token light up
only a few. The mechanism, built from a router and a sparse forward pass, is walked
through in [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch).
**Switch Transformers** (Fedus, Zoph & Shazeer, 2021) is the paper that made that idea
*simple* enough to scale — and it did so with one deliberately blunt move: route each
token to **exactly one** expert.
That sounds like a footnote. It was the whole contribution. The prior MoE recipe
(Shazeer et al., 2017) argued you needed to route each token to at least the **top-2**
experts — the reasoning being that comparing two experts gives the router a gradient
signal to learn *which* is better. Switch throws that out and keeps only the **top-1**,
the single argmax expert. Flip the toggle below and watch the routing collapse from two
connectors per token to one:
Why is top-1 worth a paper? Because the top-2 you save is not free compute you were
wasting — it is a second copy of every token that has to be **dispatched to another
device**. Experts are sharded across accelerators; routing a token to an expert means
sending its activation over the network. Halving the experts-per-token roughly halves
both the router's arithmetic and the all-to-all communication volume — the actual
bottleneck at scale. Switch's claim is that with the right guardrails, one expert per
token loses little quality while making the whole thing dramatically cheaper to run.
## The capacity buffer, and the tokens it drops
The catch with routing is that it is dynamic — you don't know until runtime how many
tokens will pick each expert, but hardware needs **fixed** tensor shapes. Switch solves
this by giving every expert a fixed buffer:
$$
\text{expert capacity} = \frac{\text{tokens per batch}}{\text{number of experts}} \times \text{capacity factor}
$$
If routing were perfectly uniform, a capacity factor of `1.0` would give each expert
exactly its fair share of slots. It never is uniform. When more tokens route to an
expert than it has slots, the overflow tokens are **dropped** — they skip the layer
entirely and pass through the residual connection unchanged. Drag the load imbalance and
the capacity factor and watch tokens overflow into red:
This is the core tradeoff, made concrete. A higher capacity factor (the paper tests
`1.0`, `1.25`, and `2.0`) means fewer dropped tokens — but every empty slot is compute
and memory spent on nothing. A lower factor is cheaper but throws away more tokens. The
whole point of good load balancing is to flatten the routing so a *small* buffer
suffices.
## The two tricks that make top-1 stable
Blunt top-1 routing would collapse — the router would learn to send everything to a
handful of experts, starving the rest. Two mechanisms hold it together:
- **Differentiable load-balancing loss.** An auxiliary loss added at every Switch layer,
scaled by $\alpha = 10^{-2}$, is minimized when tokens are spread **uniformly** across
experts. It's the product of the fraction of tokens dispatched to each expert and the
router's average probability mass on that expert, summed over experts — a smooth
penalty that pushes the router toward balanced assignment without hard constraints.
- **Selective precision.** Large sparse models train in `bfloat16` for speed, but the
router's `exp`/softmax over expert logits is numerically fragile — small perturbations
flip the argmax and destabilize training. Switch casts **only the router's internal
computation to `float32`**, keeping everything else in `bfloat16`. The fp32 stays local
to the router (it isn't communicated across devices), so it buys stability at no
bandwidth cost.
Add **expert dropout** at fine-tuning time — a higher dropout rate of `0.4` inside the
expert layers versus `0.1` elsewhere — to keep the huge sparse model from overfitting
small downstream datasets, and top-1 routing trains cleanly.
## What it bought: 7× at matched FLOPs
Held to the **same FLOPs per token** as a dense T5, Switch reaches the same pretraining
quality far sooner. The 64-expert Switch-Base hits T5-Base's quality in about
**one-seventh** the training steps; scaled up, Switch-XXL reaches T5-XXL's quality about
**4×** faster.
The other headline is raw scale. By stacking experts, the paper builds **Switch-C** with
**2,048 experts** and roughly **1.6 trillion** total parameters — while **Switch-XXL**
takes a different bet, only 64 experts but a much larger per-expert FFN, at ~395B
parameters. Both were among the largest models trained at the time.
## Distilling back to dense
A 1.6T-parameter sparse model is impractical to *serve* for many use cases — the
parameters have to live in memory across many devices even if each token only touches a
few. So the paper distills the sparse teacher back into a small **dense** student, and
finds you can compress the model by up to **99%** while still keeping about **30%** of the
quality improvement the sparse model earned over its dense baseline. Not all of it — but
a meaningful slice of the gains survives into a model you can run on modest hardware.
Read the wins precisely — none of them is a free lunch.
- **1.6T is sparse, not dense.** Switch-C activates a *single* expert's FFN per token, so
the FLOPs and activated parameters per token stay close to the dense T5 backbone —
nowhere near 1.6T of compute. The trillion parameters are capacity you *store and
communicate*, not compute you *spend* per token. Never read "1.6T" as dense-1.6T cost.
- **7× is sample-efficiency at matched FLOPs, not wall-clock magic.** It means reaching a
quality target in fewer steps at equal FLOPs-per-token — bought by spending far more
**memory and cross-device communication** on many more parameters. On different
hardware or with communication-bound serving, the real-world speedup shrinks.
- **Top-1 is not strictly better than top-2.** Routing to one expert can be lower-quality
per FLOP in some settings; Switch's contribution is *showing it works well* once you add
the capacity buffer, the load-balancing loss, and the fp32 router — not that fewer
experts is always better.
- **Dropped tokens are lost information.** At a low capacity factor, overflow tokens skip
the layer entirely. That's a genuine cost you trade against buffer waste — there's no
setting that removes it, only balances it.
## Why it still matters
Almost every technique in a modern MoE — [LongCat 2.0](/articles/longcat-2)'s 1.6T /
~48B-active split, the routing and load-balancing machinery in newer systems — is a
descendant of the choices made here. Switch didn't invent Mixture-of-Experts; it made it
**simple and stable enough to scale**, by proving that the aggressive top-1 route works
if you surround it with a fixed capacity buffer, a balancing loss, and a precision-safe
router. The interesting later work mostly *pushes back* on the simplifications — smarter
routing than pure argmax, softer handling than hard token drops — but they all start from
the Switch layer. It's the foundation the zoo is built on.
---
*Built on [Switch Transformers: Scaling to Trillion Parameter Models with Simple and
Efficient Sparsity](https://arxiv.org/abs/2101.03961) (Fedus, Zoph & Shazeer, 2021).
Figures are the paper's own (Figures 2 and 5), used for academic commentary; the
interactive diagrams are our illustrations of the mechanism. Numbers — the α=0.01 loss
coefficient, capacity factors, the 7× and 4× speedups, 2,048 experts / 1.6T parameters,
99% compression retaining ~30% of gains, and 0.4 expert dropout — are quoted from the
paper.*
---
# SWE-1.7: near-frontier code RL, and the async training loop behind it
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/swe-1-7
> date: 2026-07-09
> tags: explainer, agents, reinforcement-learning, systems, inference-optimization
**SWE-1.7** is Cognition's newest in-house software-engineering model — the SWE-1 family is the set of models behind Devin and the Windsurf editor. It's trained with reinforcement learning on real SE tasks, starting from a **Kimi K2.7** base that had *already* been through heavy RL post-training. That starting point is the headline claim: Cognition still pulled large gains on top of an RL-saturated base, which they read as evidence against a "post-training ceiling." The model is tuned for long-horizon, asynchronous engineering — the multi-hour tasks Devin runs — and it's served through **Cerebras at 1000 tokens/sec**. That speed is the other half of the pitch: near-frontier quality, cheap and fast, moving the cost-performance Pareto curve rather than the top of the leaderboard.
Every number here is **provider-reported**, run under Cognition's own harness (Claude Code for Anthropic models, Codex for OpenAI, Devin CLI otherwise; `timeout=4h`, max reasoning effort). On the published suite SWE-1.7 tops **nothing**: **Claude Opus 4.8** leads all three benchmarks and SWE-1.7 lands roughly at **GPT-5.5**. It is also **not open-weights** — there's no model card or download, only the served model in Devin. Read it as a strong, fast, cheap near-frontier model and a genuinely interesting training write-up — not a new SOTA.
## The numbers, in full
Three agentic coding benchmarks, pass rate (%). `FrontierCode` is Cognition's own eval; `SWE-Bench Multilingual` is the multi-language slice of the SWE-bench family; `Terminal-Bench 2.1` is agent-in-a-terminal.
| Benchmark | SWE-1.7 | Kimi K2.7 Code (base) | GPT-5.5 | Opus 4.7 | Opus 4.8 | GLM-5.2 | Composer 2.5 | SWE-1.6 |
|---|---|---|---|---|---|---|---|---|
| FrontierCode 1.1 Main | **42.3** | 30.1 | 43.0 | 38.5 | 46.5 | 24.5 | 25.6 | 9.4 |
| Terminal-Bench 2.1 | **81.5** | 72.7 | 84.2 | 83.0 | 86.9 | 81.0 | 76.0 | 39.7 |
| SWE-Bench Multilingual | **77.8** | 73.5 | 76.8 | 80.5 | 84.4 | 74.5 | 71.6 | 58.3 |
The load-bearing comparison is the base column. SWE-1.7 vs its own `Kimi K2.7 Code` base is **+12.2** on FrontierCode Main, **+8.8** on Terminal-Bench, **+4.3** on SWE-Bench Multilingual — the largest lift on the hardest eval. That gap is the whole "no post-training ceiling" argument: it's pure RL on top of a base that was already RL-post-trained.
So: SWE-1.7 edges GPT-5.5 on SWE-Bench Multilingual (77.8 vs 76.8), trails it a hair on FrontierCode Main (42.3 vs 43.0) and Terminal-Bench (81.5 vs 84.2), and sits behind both Opus checkpoints everywhere. The jump from the previous **SWE-1.6** (9.4 / 39.7 / 58.3) is huge, but read it honestly: most of that comes from the far stronger Kimi K2.7 base, not only the new RL recipe. The rest of this piece is the recipe, which is where the interesting engineering lives. Four ideas stand out.
## 1. Asynchronous RL: don't let the trainer starve
Start with why this is hard. An agentic SE rollout is a long, **variable-length** trajectory: read files, edit, run tests, read the failure, edit again — dozens of tool-calling turns, and with self-compaction (below) some run for **six hours**. Now do RL on batches of these.
In **synchronous** RL the loop ping-pongs: the actors generate a batch of rollouts, then the learner does one optimizer step, then the next batch starts. The problem is variance. Within a batch, one trajectory finishes in minutes and another runs for hours, and the learner can only step once the **slowest** one lands. So the trainer sits idle across most of the wall-clock, and the fast actors idle too, waiting for their batch-mates. Expensive accelerators, starved.
**Asynchronous** RL breaks the ping-pong. Decouple the two fleets: actors continuously generate trajectories against the current-ish policy and push them into a buffer; the learner trains on whatever is ready. Both fleets stay busy. Flip the toggle below between the two modes and watch the learner's GPU-utilization gauge and the weight-update counter over the same wall-clock window:
The catch is honest and specific. Once the learner trains on trajectories the actors generated a few weight-versions ago, those trajectories are **off-policy** — sampled by a stale policy, not the one being updated. Cognition names the failure mode directly: a **KL-divergence mismatch between inference and training**, "since the trainer policy is usually different from the sampling policy." The fix is the standard off-policy toolkit — importance sampling, plus a bounded staleness — and a **buffer policy** after any interruption that "prevents bias from any imbalance in training-inference throughput." It's the same bias/variance trade you always pay for throughput: async keeps the GPUs full, and you spend correction terms to keep the stale gradients honest. Cognition cites [PipelineRL](https://arxiv.org/abs/2509.19128) as the async-RL lineage here. The [Agents-A1 write-up](/articles/agents-a1) is a good companion on where verified agentic trajectories come from in the first place, and the Devin side of "the environment the model is trained in" is the [agent harness](/articles/agent-harness).
## 2. Preserving entropy with top-p "sampling distribution replay"
This is the most elegant idea in the post, and it's small. Long RL runs die of **entropy collapse**: a strong policy stops exploring, the distribution sharpens to a spike, and reward plateaus within a few hundred steps. Cognition's diagnosis of *why* is worth the walk.
Take three tokens with logits $x_1 > x_2 \gg x_3$ and softmax probabilities $p_i$. Token 3 is a junk token — sampling it usually means the rollout went off the rails, so the trajectory earns low reward and its advantage is negative, $\hat{A} < 0$. The policy gradient of its log-prob on the logits is:
$$
\nabla \log p_3 = \begin{bmatrix} -p_1 \\ -p_2 \\ p_1 + p_2 \end{bmatrix}, \qquad \Delta x_i \propto \hat{A}\,\nabla \log p_3 .
$$
With $\hat{A} < 0$ this becomes
$$
\Delta x_1 \propto |\hat{A}|\,p_1, \qquad \Delta x_2 \propto |\hat{A}|\,p_2, \qquad \Delta x_3 \propto -|\hat{A}|\,(p_1+p_2).
$$
Look at what that does. Because $p_1 > p_2$, the already-dominant token's logit rises **more** than the runner-up's, and the junk token is pushed down. So *punishing* a junk sample **sharpens** the distribution — every off-track sample bleeds a little entropy. Step the widget below with replay off to watch the bars spike and the entropy gauge fall; then flip **top-p replay** on:
The fix is two moves. First, **top-p sampling**: never sample from the low-probability tail, so junk tokens never become optimization targets in the first place. But top-p naively breaks something else — the trainer computes probabilities over the *full* vocabulary while the rollout sampled from the *top-p subset*, so the two distributions diverge and you're back to a large train/inference mismatch. Second, then, **sampling distribution replay**: record the kept-set mask at rollout time and renormalize the trainer's probabilities over that **same mask**. Sampler and trainer now agree on the support, the mismatch stays bounded, and entropy holds roughly constant.
There's a free lunch hiding in the mask. A token whose probability already exceeds the top-p threshold has a **keep-set of size one** — itself — so its renormalized probability is a constant 1 and its gradient is **zeroed out**. Cognition finds a large fraction of sampled tokens sit above the threshold, so they drop out of the update entirely. The optimizer stops spending gradient on tokens the model is already sure about and focuses on the genuinely uncertain, high-learning-signal positions. Less gradient noise, for free.
This is the same idea as **Rollout Routing Replay (R3)** — [/articles/rollout-routing-replay](/articles/rollout-routing-replay) — one axis over. R3 records the MoE router's rollout-time expert choices and replays them in the trainer to align sampler and trainer on the *routing* axis; sampling distribution replay does it on the *token-sampling* axis. Both are "replay the decision the sampler actually made, so the trainer optimizes the same distribution." Cognition stacks these with importance sampling and NVFP4 low-precision rollouts, and reports gains from the [Muon optimizer](/articles/muon-optimizer) and from stripping non-deterministic trainer ops that were quietly widening the mismatch.
## 3. Multi-cluster training: ship weight deltas, not weights
Here's a structural observation that falls out of async RL: it **decomposes across clusters**. Only the trainer needs to live on a single high-bandwidth fabric — that's the one tightly-coupled, all-reduce-heavy component. The rollout inference engines are self-contained; each one needs nothing but the current weights, so it can run on whatever compute is available, anywhere.
Cognition leans all the way into that. SWE-1.7's RL spans **four datacenters across three continents**, mixing their own GPUs with third-party inference compute from **Fireworks**. The hard part is keeping every far-flung inference engine current after each optimizer step, because stale weights mean stale trajectories mean weaker gradients. Broadcasting a full ~1T-parameter model across oceans every few steps is a non-starter, so instead the trainer computes a **compressed weight delta** (XOR diff against the previous weights, then zstd) and streams it through **cloud object storage** as the single source of truth. Each engine prefetches the delta while still serving, then pauses briefly to apply it in-place with the KV cache intact.
The numbers Cognition reports for this: a delta is **>99% smaller** than the full broadcast, a cross-continental update for a 1T model lands in **1-2 minutes** end-to-end, and inference pauses only **3-4 seconds** to apply it. A **Dynamo** router fronts the inference fleet and reroutes trajectories off dead replicas; the trainer checkpoints asynchronously to local disk every step and rebuilds a dead node from peer replicas in seconds, so a hardware failure never stalls the run. The payoff loops back to idea #1: faster weight sync means less trajectory staleness, which buys room for more aggressive learning rates. This is the cost angle a sibling write-up, [why frontier RL is cheaper than you think](/articles/frontier-rl-cheaper), covers head-on — and it's [Fireworks' own post](https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) that Cognition cites, since they literally run on that compute.
## 4. Self-compaction for six-hour rollouts
The last idea handles the horizon. Two problems come with training on multi-hour tasks. First, a rollout can run far past the raw context window. Second, as DeepSeek-R1 showed, RL on reasoning tasks tends to make responses grow without bound, but I want a model that's terse on easy tasks and only elaborates on hard ones.
**Self-compaction** solves the first: when the agent approaches the context limit, it's asked to **summarize its working state**, and it resumes from its own summary. During training the model learns both halves at once — to write more informative, compact summaries, and to work well *from* them. That's what lets a single rollout stretch to **six hours** without blowing the window. An **alternating length penalty** solves the second: training alternates between *unconstrained* phases (optimize only for task success) and *budget* phases (penalize solutions that exceed a weighted cost over tokens, turns, and tool-call time). Length compresses on tasks the model can already solve, while long-horizon behavior on genuinely hard tasks is preserved.
## What the training left behind
RL this heavy leaves fingerprints on behavior, and they line up with the recipe. SWE-1.7's chain-of-thought is measurably **more condensed** than the Kimi K2.7 base — a much lower function-word ratio and nearly half the words per sentence — which Cognition attributes directly to the budget phases of the length penalty. It also **explores the codebase far more** before acting: more tool calls, file reads, and greps per run than K2.7, Opus 4.8, or GPT-5.5, and more probing of edge cases, adversarial inputs, and unstated requirements. On a bug report it chases the root cause rather than patching the one symptom, and it settles ambiguous semantics by writing a small script to test them instead of guessing. Cognition credits the data pipeline — hard verifiers that reject false positives force end-to-end solutions.
The honest cost of that thoroughness: **scope creep**. More reasoning means more doing — extra test cases, more files touched than the task strictly needs. It's an industry-wide pattern (more reasoning, wider blast radius) and Cognition flags it as an open axis, not a solved one. That's the right way to report it.
## The take
SWE-1.7 doesn't win the benchmark table, and Cognition doesn't claim it does — Opus 4.8 leads every row and SWE-1.7 sits at roughly the GPT-5.5 line. The pitch is the Pareto curve: near-frontier SE quality, served at 1000 tokens/sec through Cerebras, cheap. If that holds up in a real Devin loop rather than a `timeout=4h` harness, it's a strong practical option.
But the model is the smaller story. The training write-up is the payload, and it's unusually concrete for a launch post: async RL to stop the trainer starving, a top-p **sampling distribution replay** that kills entropy collapse *and* falls out into free gradient denoising, weight-delta streaming that makes cross-continental RL practical, and self-compaction that pushes rollouts to six hours. Each is a clean, separable idea with a plausible mechanism, and together they're a real argument that "post-training ceiling" was never a ceiling — just a recipe that hadn't been tuned yet.
The caveats are the usual ones, stated plainly. Every number is self-run and self-selected; `FrontierCode` is Cognition's own benchmark; there are no open weights to verify anything against; and the sharpest systems claims — >99% delta compression, 1-2 minute cross-continental sync, six-hour rollouts — are provider-reported, not independently measured. Take the leaderboard framing with the usual salt. Take the training ideas seriously.
---
*Built on Cognition's [SWE-1.7 launch post](https://cognition.com/blog/swe-1-7) (July 8, 2026) and its cited [FrontierCode 1.1](https://cognition.com/blog/frontier-code-1.1) eval. SWE-1.7 is served in [Devin](https://devin.ai); benchmark numbers are provider-reported. The interactive diagrams are illustrations of the mechanism, not measured traces; the entropy, mismatch, and response-length charts are reproduced from the launch post for commentary.*
---
# A6B: k-expansion, or what breaks when you force a top-8 MoE router to fire 32 experts
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/a6b-k-expansion
> date: 2026-07-08
> tags: explainer, llm, architecture, inference-optimization
**a6b-k-expansion** is a single-author research log with a sharp, testable question: a Mixture-of-Experts model already stores far more experts than it fires per token — so what happens if you just *fire more of them* at inference? The author takes **Qwen3.6-35B-A3B** (their stated base: 35B total, ~3B active, 256 routed experts per layer, native top-8 routing) and overrides the router to keep **top-32** instead of top-8. Active compute goes from ~3B to **~6.6B** parameters per token — the regime they call **"A6B"** — at **zero new weights**. Then they measure what it costs, and try to heal it.
The interesting part is that it does not work for free, and the repo says so in the title of its own findings: *the damage from naive k-expansion is smooth and monotonic — there is no free sweet spot.* This is a measurement-first log with paired A/B tests, exact McNemar significance, and the negative results kept in. That honesty is the reason it's worth reading.
Scope, up front. This is one person's **ongoing research log**, not a model release or a peer-reviewed paper. The base is **Qwen3.6-35B-A3B** as named in the repo — I take that at face value and report the author's own numbers, all of which are self-measured. "Healing" is explicitly **demonstrated, not complete**: the point estimates stay negative. Rejection fine-tuning, GRPO, and the Terminal-Bench evaluation are still in progress. Read this as a well-instrumented experiment, not a benchmark trophy.
## The one-line edit
Every MoE layer scores all 256 experts with a softmax router (in fp32), keeps the **top-k**, and — this is the load-bearing detail — **renormalizes the selected gate weights so they sum to 1**. Renormalization is baked into the architecture. So raising k is not a matter of appending a few experts on the side; it *redistributes the entire gate mass* across four times as many slots.
The experts are stored as packed 3D tensors — `gate_up_proj [256, 1024, 2048]` and `down_proj [256, 2048, 512]` per layer — so widening k adds **compute but no parameters**. If the routing here is unfamiliar, I built the top-k gate up from nothing in [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch); this article assumes that machinery and pokes at one constant inside it.
Formally, for the selected set the renormalized gate weight of expert $i$ is
$$
\tilde{g}_i = \frac{g_i}{\sum_{j \in \text{top-}k} g_j}, \qquad i \in \text{top-}k .
$$
The denominator grows with k. So every already-selected expert's $\tilde{g}_i$ *shrinks* as k rises, and the freed mass flows to the newcomers. The author profiled the router at k=32 and found the newcomers are not a rounding error: **ranks 9-32 carry 54.0% of the renormalized gate mass.** More than half the block's output now comes from experts the model never trained to fire together.
Slide k below and watch the mass move off the trained top-8 and onto the untrained tail — with the measured accuracy for each k next to it:
## No sweet spot
You might hope for a lucky k — 2-3× the native width, where the extra experts add capacity before they add noise. There isn't one. Changing **only** the inference-time top-k on the frozen base model, both a knowledge benchmark and a math benchmark decline at every single step:
| Benchmark | k = 8 | k = 16 | k = 24 | k = 32 |
| --- | --- | --- | --- | --- |
| MMLU | 0.8433 | 0.8283 | 0.8150 | 0.8067 |
| GSM8K | 0.8933 | 0.8883 | 0.8783 | 0.8650 |
The shape is the whole point: monotonic, sweet-spot-free. Every step down is paid the moment more untrained expert combinations switch on, and the noise grows with the mass those combinations carry — a mass that is real (54.0% on ranks 9-32). This is what motivates *healing* the model rather than *searching* for a lucky k. There is nothing to search for.
## Healing without touching the router
The fix is deliberately surgical, in the spirit of [ESFT — Expert-Specialized Fine-Tuning (Wang et al., 2024)](https://arxiv.org/abs/2407.01906). The recipe:
1. **Profile.** Run the target corpus through the model at k=32, count token-level routing frequency for every expert across all 40 layers.
2. **Select.** Keep experts by cumulative routing frequency up to **top-p 0.2** — the ones actually carrying the work. That is **833 of the 10,240 layer-experts** (40 × 256).
3. **Train residual deltas.** Add trainable delta tensors to the selected experts' FFN slices and train **only those deltas** — **2.62B of 35B params (7.5%)**. The router and everything else stay frozen.
4. **Toggle.** Because nothing but the deltas moved, turning them off returns the **exact stock model**. The deltas ship as one **5.2 GB** artifact, patched on or off at load time.
Toggle the deltas below to see the surgical footprint and the gap-to-baseline flip from "significant loss" to "statistical tie":
Freezing the router is not just a cost decision. The deltas are learned *relative to a specific routing distribution* — freeze it and every delta keeps meaning the same thing at inference; let it move and the coordinate system shifts underneath the deltas. Second, a router that drifts under SFT is the classic path to **routing collapse**, where a handful of experts capture all the traffic. A frozen router removes that failure mode entirely.
Each healing generation is a broader training corpus, measured as the **gap to base@k8** on its own machine (negative = still below native top-8), with the verdict from exact McNemar at p = 0.05:
| Generation | Corpus | MMLU Δ | MMLU p | GSM8K Δ | GSM8K p |
| --- | --- | --- | --- | --- | --- |
| Gen 0 | naive k32 (no training) | −3.7 pt | 0.002 · loss | −2.8 pt | 0.016 · loss |
| Gen 1 | agentic-only ESFT | −3.0 pt | 0.010 · loss | −0.8 pt | 0.487 · tie |
| Gen 2 | mixed + replay ESFT | −2.0 pt | 0.141 · **tie** | −2.2 pt | 0.263 · **tie** |
Each broader corpus removes more of the misalignment: the agentic-only patch already heals math to a tie, and adding coding, tool-calling, math and a small general/knowledge replay is what finally pulls MMLU into a tie too. But read the header honestly — a *tie* is a repair, not a gain. The MMLU point estimate is still −2.0 pt; it is no longer statistically distinguishable from base, but it is not zero.
And it does not get better with more steps. A checkpoint trajectory reads **MMLU 0.825 / 0.823 / 0.820** at steps 1200 / 2100 / 3150 — healing **saturates by ~38% of training**. The author reads this as a **capacity ceiling of the selective deltas**, not a data-volume problem: more steps on this delta set will not close the gap; a larger trainable surface or a different objective would be needed.
Where training *does* buy something beyond healing is code. On HumanEval (n = 164), the agentic patch reaches **0.902** — above both base@k8 and naive k32 — while compressing median generation to about a third of the tokens:
## The hazard: style transfers before knowledge
That fourth bar is the most useful result in the whole repo. A **coding-only** patch — same recipe, single-domain corpus — taught the model a *style* (terse code) faster than it taught it to stay correct. Median generation crashed to **186 tokens**, and accuracy collapsed with it: **HumanEval 0.762** (the worst of every arm) and **GSM8K 0.820**, which is *below even naive k32*, at p = 0.002. You can train a confident, compact, and wrong model this way.
**SFT transfers answer style before it transfers knowledge.** Corpus diversity is a safety rail, not a luxury — it is the specific guard against this failure mode, which is why the Gen-2 mix spans five domains (agentic 62%, coding 12%, tool-calling 11%, math 10%, knowledge replay 3%) plus a small replay slice.
The mixed patch is not immune to the same pressure, only more resistant. It still compressed MBPP generations (median **531** vs base's **2852** tokens) and lost MBPP by **−10.2 pt** (p < .0001) — even while HumanEval stayed at parity. Two benchmarks that both "test coding" diverged sharply, and the difference is how much each rewards long, explicit generation, which the patch has learned to suppress. Because this compression grows with training length, an *earlier* checkpoint may beat the final one on generation-heavy tasks; that evaluation is still running.
## Why I trust the numbers
The measurement discipline is where this log earns its credibility, and it's the part most self-reported results skip:
- **Paired, same-condition A/B.** Every comparison runs an arm against its own baseline on the **same machine**, same prompts, same decoding. Significance is **exact McNemar** on the paired per-item correctness vectors — not an unpaired accuracy-difference test.
- **n = 600 per benchmark** (164 for HumanEval), fixed shuffle seed so every arm sees the same items in the same order.
- **Choice-logprob MMLU** — scored by summed log-prob of each answer choice, not by parsing free-form text, so truncation and format quirks can't masquerade as a knowledge gap.
- **No-think GSM8K** — the measured quantity is the arithmetic answer, not the length or style of the scratch reasoning.
- **Re-measure on every machine.** The author observed **−0.3 to −0.6 pt cross-machine drift** on identical weights and config, so a base arm and a patched arm are always re-run together on the same box before their delta is trusted. Gen 0-1 ran on a 2× RTX PRO 6000 workstation, Gen 2 on an 8× RTX PRO 6000 Blackwell server, each with its own re-measured base@k8.
Training data also passed a hard decontamination gate against every benchmark used — exact-match, word-13-gram, short-question containment, HumanEval signature purge, Terminal-Bench instruction match — with the knowledge-replay slice further screened by embedding similarity against the full MMLU test set. Reported result: **0 residual hits**. You can disagree with the conclusions, but the instrument is honest about what it measured.
## The take
k-expansion is a clean idea with a clean negative result. Firing 4× the experts at inference is free in weights but not in accuracy, because renormalization is not a bystander — it hands 54% of every MoE block's output to 24 experts per layer that never learned to work together, and the model degrades monotonically for it. There is no lucky k. Router-frozen selective deltas — 2.62B trainable params, toggleable, 833 of 10,240 experts — heal the knowledge loss back to a **statistical tie**, which is a genuine result and an honest one: a tie, saturating at ~38% of training, with the point estimate still negative. The coding axis actually improves (HumanEval 0.902), but generation-length compression is a live hazard, and a single-domain corpus is a fast way to train a compact, confident, wrong model.
What I'd take from it, beyond A6B: the failure modes generalize. **Renormalized top-k is not free to widen.** **SFT teaches style before substance.** **Freeze the router or watch the deltas lose their coordinate system.** And measure in pairs on one machine, because a −0.5 pt result that is really cross-machine drift will fool you. Whether A6B ever nets out positive after rejection-FT and GRPO is unsettled — the repo says so plainly — but the instrumentation is the part I'd copy tomorrow.
---
*Built on the [a6b-k-expansion](https://github.com/hikarioyama/a6b-k-expansion) research log (hikarioyama, 2026) — README, `METHOD.md`, and the two HTML reports under `docs/`, MIT-licensed. All numbers are the author's own paired, McNemar-tested measurements on Qwen3.6-35B-A3B; I report them as stated and have not independently reproduced them. The interactive diagrams are my illustration of the mechanism, not measured traces; the architecture figure is reproduced from the repo for commentary.*
---
# Agent harnesses: engineering the loop around the model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/agent-harness
> date: 2026-07-08
> tags: explainer, agents, llm, systems
A base model does one thing: given tokens, predict the next ones. Everything an agent actually *does* — read a repo, run a test, spawn a subagent, decide it is finished — happens in the code wrapped around that model. Lilian Weng calls that wrapper the **harness**, and her post [Harness Engineering for Self-Improvement](https://lilianweng.github.io/posts/2026-07-04-harness/) makes a sharp claim about it: the harness is not glue you can ignore. It is "the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results." Her thesis in one line: "the layer between the raw model and the real-world context seems to be as important as the model's raw intelligence."
That reframes a lot of agent engineering. The interesting design surface is not only the model — it is the loop, the tool set, the context policy, and the evaluator you build around it. This is a walk through her framing.
Weng's post is framed around **recursive self-improvement (RSI)** — the idea, dating to I. J. Good (1965) and his "ultraintelligent machine," that a capable system can improve the machinery that produces it. In modern terms the model doesn't rewrite its own weights; it improves the *training pipeline* and the *deployment system* around itself. The harness is that deployment system, which is why it sits at the center of the story. This article follows her two-part structure: first what a harness *is* (design patterns), then how you *optimize* it (the self-improving outer loop).
## The loop
Strip an agent down and you get a loop. The model emits an action, the harness executes it, the result comes back, the model emits the next action. Weng's simplest picture of it: user input feeds model inference, which either returns a response or calls a tool; tool results feed straight back into the next inference step.
For a coding agent that loop takes a concrete, goal-oriented shape: **plan, execute, observe/test, improve, and execute again until the goal is achieved.** Weng draws it as a pipeline — observe the repo, plan, search and read files, edit and write patches, run tests, inspect errors — with a `Done` exit and a repeat arrow when the goal isn't met yet.
The static figure shows the shape; stepping through it shows the mechanic. Here is that loop walked over one real task — a failing test — where the first patch is too narrow and the harness has to loop back before the suite goes green. Watch what crosses the boundary at each phase: the model emits a **tool call**, the harness runs it, and the **observation** returns into context.
Two things are worth pulling out of that trace. First, the model touches nothing directly — every action is a tool call the harness executes, and every result is an observation the harness chooses to feed back. Second, the harness owns the **control flow**: it decides when the loop repeats and when the goal is met and the loop exits. That decision — keep going or stop — is not the model's to make unilaterally, and getting it wrong is a failure mode we'll come back to.
## The tools are the harness
If the loop is the skeleton, the tool set is the muscle. Weng's tour of a modern coding harness is essentially a table of tool categories, and the range is the point — this is a lot more than "call a function":
| Category | Representative tools |
|---|---|
| **File system** | discovery: `glob`, `grep`, `ls` · read: `read`, `read_many` · modify: `write`, `edit`, `multi_edit`, `apply_patch` |
| **Shell execution** | `bash`, `PowerShell` |
| **IO / repo** | `lsp`, `git_status`, `git_diff`, `git_commit` |
| **External context** | MCP tools, Skills |
| **Web** | `web_search`, `web_fetch`, browser tools |
| **Artifacts** | read docs/images; generate HTML/images |
| **Backend processes** | `CronCreate`, `CronDelete`, `CronList` |
| **Agent delegation** | `spawn_agent`, `resume_agent`, `wait_agent`, `list_agents`, `close_agent` |
A design note runs through the whole list: the tools are "deliberately simple and generic to enable generalization." They lean on primitives a developer already knows — a file system, a shell, git — rather than bespoke abstractions. That matters because "learning how to read, write, and edit the file system (commonly via `bash` commands) is a foundation skill for LLMs." The model has seen a million shell sessions in training; give it a shell and it already knows the idiom. The last two rows — backend processes and agent delegation — are where a harness stops being a single loop and becomes a small operating system: it can schedule work and fork subagents, which is what makes long-horizon and parallel tasks tractable.
## Context is the scarce resource
Here is the constraint that shapes everything else. A long task produces "experiment logs, code diffs, paper summaries, error traces, and past rollout trajectories" that "often grow much longer than the context window that the model has trained for." You cannot keep the whole trajectory in the prompt. So Weng's rule is blunt: "a harness should not carry the entire workflow and all logs in context; instead, it should keep durable state in files."
That single decision — context as a bounded working set, disk as the durable store — is what keeps a long task from strangling on its own history. Step through a run and watch the two strategies diverge:
The naive strategy appends everything and eventually overflows the window; past that point the model is quietly losing the early details. The file-backed harness keeps context flat and spills the history to disk, where it stays retrievable with `grep` and `read` — the same file tools from the table above, now doing double duty as memory. This is why file-system fluency is the load-bearing skill: the file system *is* the agent's long-term memory.
The same principle governs parallelism. Weng's guidance is to make it "explicit and inspectable" — store subagent outputs as "files, logs, and status records" rather than transient chat contexts, so the system can "recover after interruptions and reason over its own execution history." Durable-state-in-files isn't only a context trick; it's what makes an agent restartable.
There's a research lineage here worth naming. **Agentic Context Engineering (ACE)** maintains a "context playbook" of itemized bullet points — each with an identifier and description — updated by a Generator / Reflector / Curator trio that appends *structured entries* instead of rewriting the whole prompt. **Meta Context Engineering (MCE)** goes one level up, separating "mechanism (how to manage context) from artifact content (what is in context)." Both are the same instinct as keeping state in files: treat context as a managed, structured store, not an ever-growing transcript.
## Guardrails live outside the loop
Give an agent `bash`, `edit`, and the ability to spawn more agents and you've handed it a lot of reach. Weng is direct that this breaks abstraction boundaries: when programs can edit the systems they run on, you need a "proper design of editable surface" with "permission control and security layers outside this loop." The guardrail is deliberately *not* another prompt instruction inside the model's context — it's an enforcement layer the model cannot talk its way past. That placement is the whole point: a permission check the agent can edit is not a permission check.
## Optimizing the harness
The second half of Weng's post asks the recursive question: if the harness matters this much, can the agent improve *its own* harness? That turns the inner task loop into an **outer loop** over harness designs — run the current harness, mine where it failed, propose edits, keep the ones that survive a regression test.
This is one instance of a broader family the post surveys — **ADAS**, **AFlow**, **STOP**, **AlphaEvolve**, the **Darwin Gödel Machine** — all variations on "search over the scaffolding, not the weights." The headline result is that it works: the Darwin Gödel Machine's discovered agents went from **20% → 50%** on SWE-bench Verified and **14.2% → 30.7%** on Polyglot, matching or beating handcrafted agents.
But the honest caveat is the more instructive part. **STOP** improved performance when the base model was GPT-4 and *degraded* it with weaker models (GPT-3.5, Mixtral). Weng's reading: "recursive structure alone is not enough. The base model must be capable enough to improve the mechanism." Self-improvement is not free lift from the loop; it's a gain that depends on a model already good enough to reason about its own scaffolding. Below that bar, the outer loop makes things worse.
## Failure modes
The post catalogs where autonomous agents actually break, drawing on an analysis (Trehan & Chopra, 2026) of auto-research attempts. Six recur:
1. **Training-data defaults.** The model reaches for old libraries, stale commands, and standard formats instead of what the actual repo uses.
2. **Implementation drift.** When the proposed method gets complex, the model quietly slides toward a simpler solution than the one it was asked for.
3. **Memory degradation.** Long-horizon projects lose critical details — unless the logs were written out as persistent artifacts. (The file-system point again, stated as a failure when you skip it.)
4. **Over-optimism.** The model declares success on noisy or failed experiments — a pattern Weng names "p-hacking and eureka-ing."
5. **Insufficient domain intelligence.** It lacks the tacit craft knowledge to judge whether a result is even plausible.
6. **Weak scientific taste.** The experiments run fine but fail to answer the right question.
Notice how many of these the *harness* is supposed to catch rather than the model: memory degradation is a context-policy failure, over-optimism is an evaluator failure, drift is a control-flow failure. The whole post is an argument that these are engineering problems in the wrapper, not just intelligence gaps in the core.
## What's still hard
Weng closes with the bottlenecks between here and genuine self-improvement, and they read as an honest problem list rather than a roadmap:
- **Weak fuzzy evaluators.** "Many research claims don't have a fast/precise verifier." Taste and novelty are far harder to score than a passing test suite, and an outer loop is only as good as its evaluator.
- **Context and memory lifecycle.** Managing context growth over long autonomous runs is becoming "a core part of intelligence," not a plumbing detail.
- **Negative results.** LLMs are biased toward success and struggle to abandon a hypothesis, because their training data is skewed toward things that worked.
- **Diversity collapse.** Evolutionary and RL loops exploit known high-reward patterns; without pressure for diversity the population collapses into variants of one solution.
- **Reward hacking.** A self-improvement loop optimizes the signal it's given — including benchmark artifacts and vulnerabilities in a judge model.
- **Long-term success.** Short-horizon optimization ignores maintainability, ownership boundaries, migration cost, backwards compatibility, and debugging burden — the things that decide whether real systems survive.
- **The human role.** Her framing is that "humans should move up the stack, not be removed from the loop" — providing oversight "at the right time, the right abstraction level," not disappearing from it.
## The take
The useful shift in Weng's framing is where it puts the design surface. Agent quality is not only a function of the model you call; it's a function of the loop you wrap it in, the tools you expose, the context policy you enforce, and the evaluator you trust. Those are engineering decisions, and most of them are decisions about *state* — what stays in context, what spills to disk, what the permission layer refuses, what the regression test has to pass before a change ships. The honest bounds are stated plainly too: self-improving harnesses only help above a base-model capability threshold, the outer loop is only as good as a verifier we mostly don't have for fuzzy goals, and reward hacking and diversity collapse are unsolved. Which lands on a pragmatic note — the harness is where a lot of the near-term gains are, and it's ordinary systems engineering: files, permissions, control flow, and tests, applied to a model instead of a service.
---
*Built on Lilian Weng's [Harness Engineering for Self-Improvement](https://lilianweng.github.io/posts/2026-07-04-harness/) (2026). Quotations and the three figures are reproduced from that post for commentary; the interactive loop and context-budget diagrams are my own illustrations of the mechanism, not measured traces. Benchmark numbers (Darwin Gödel Machine, STOP) are as reported in her post.*
---
# Antidoom: breaking doom loops with Final Token Preference Optimization
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/antidoom
> date: 2026-07-08
> tags: explainer, llm, reinforcement-learning, training, inference-optimization
A **doom loop** is when a model emits a span — usually something like `Wait, let me reconsider…` — and then repeats that same span again, and again, until the context window is exhausted. Liquid AI's **Antidoom** post is about why this happens and how to train it away. Their fix is **Final Token Preference Optimization (FTPO)**: a DPO-like method that retrains a single position, the one where the loop would restart, so the model has somewhere else to go. On an early **LFM2.5-2.6B** checkpoint it takes the doom-loop rate from **10.2% to 1.4%**; on **Qwen3.5-4B** from **22.9% to 1%**. The interesting part is *why* eval scores go up when they do it: the training teaches the model nothing new about math or code, it just removes the failure mode that was keeping it from finishing.
Every number here is **provider-reported** by Liquid AI (blog + the `LiquidAI/antidoom-mix-v1.0` dataset). The two interactive widgets are my illustrations of the mechanism — schematic distributions and logits, not measured traces. The four embedded charts are the post's own figures, reproduced for commentary.
## What a loop looks like
The detector is blunt and effective: a completion is flagged as looping if a section **repeats at least four times, over at least 60 characters**. Small reasoning models hit this most on long thinking traces for hard math and coding — exactly the prompts where the model is uncertain for a long stretch.
The tokens that *start* a loop are not random. On the early LFM2.5-2.6B checkpoint, Liquid counted which token opens the repeating span:
```text
count share token
2277 11.39% ' the'
902 4.51% ' So'
644 3.22% 'Alternatively'
511 2.56% 'Wait'
493 2.46% ' But'
```
These are discourse markers and self-reflection tokens. They are not bad tokens — `Wait` or `Alternatively` can mark a genuine change of strategy. The problem is what happens when the model reaches for one *under uncertainty* and then can't get back out.
## Why the loop tightens
Three things stack up, and it's worth keeping them separate because FTPO only attacks the third.
**High priors.** Some tokens carry artificially high prior probability. Liquid points at synthetic training data inflating certain words above their natural human-text frequency — the same effect that made `delve` and `testament` model tells. In reasoning traces, the inflated tokens are the discourse markers above. When the model is unsure of the next real step, these dominate the next-token distribution, and it restarts the same local reasoning pattern instead of making progress.
**Self-reinforcing context.** This is the one that turns a stumble into a loop. Once a span is in the context, that span becomes *more* likely to appear again — and with each repetition the probability of every token inside the looping span climbs toward 1. The distribution collapses:
**Greedy decoding has no exit.** At low temperature — and especially at temp 0 — the model takes the argmax. Once self-reinforcement has pushed the loop token's probability close to 1, there is almost no mass left on anything else, so there is nothing to sample instead. Turning up temperature only helps a little: Liquid reports significant looping even at **temp=0.67**, because there just isn't enough probability left on the alternatives to escape.
The widget below is the whole story in one place. In **base model** mode, step the repeat count and watch the loop token `Wait` climb from ~41% to ~98% while the distribution collapses onto it — greedy decoding restarts the span every time. Flip to **after FTPO** once you've read the next section to see the fix.
## Why the easy fixes don't hold
The usual inference-time patch is `repetition_penalty`, which reweights the output distribution to discourage repeats. It's a band-aid: it fights the symptom at decode time and can degrade quality, and it doesn't touch the priors that caused the collapse. Reinforcement learning *can* target looping, but it needs carefully calibrated rewards and costly online rollouts. FTPO's pitch is to fix the distribution once, offline, at the exact position where the loop begins.
## Final Token Preference Optimization
FTPO is preference optimization in the DPO family. The reference form of DPO trains a policy $\pi_\theta$ against a frozen reference $\pi_{\text{ref}}$ to prefer a chosen response $y_w$ over a rejected one $y_l$:
$$
\mathcal{L}_{\text{DPO}} = -\,\mathbb{E}_{(x,\,y_w,\,y_l)}\!\left[\log \sigma\!\left(\beta \log \frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)} - \beta \log \frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)\right]
$$
FTPO keeps that skeleton and changes four things, each aimed at *not* over-correcting:
1. **Final token only.** It trains the *trailing* token of a sequence that is midway through generation — the single position where the loop would restart — not a whole response. $y_w$ and $y_l$ are one token each, at one place.
2. **Multiple chosen tokens per sample.** Instead of one $y_w$, it uses a *set* of plausible chosen tokens. This spreads the freed-up probability across several alternatives, so you aren't just replacing one overtrained token with a new overtrained token.
3. **A KL-like term in logit space.** The reference-anchoring divergence is computed on **logits**, omitting the softmax. That avoids gradient pressure leaking onto unrelated tokens through the normalization.
4. **Two-part regularization.** The tokens it means to move — chosen and rejected — are allowed to travel freely relative to the reference, while the rest of the vocabulary is held tightly near it. Loosen the ones you're fixing, pin everything else.
### Building the training row
A training row is a `[prompt prefix, one rejected token, one or more chosen tokens]` tuple, and Liquid mines it straight from the model's own failures. They generate completions on a loop-eliciting prompt mix (`LiquidAI/antidoom-mix-v1.0`) at low temperature, detect a loop with the ≥4-repeats / ≥60-chars rule, and target the **first token of the first repeat** as the rejected token. At that position they take the base model's top-k log-prob alternatives, filter out short and non-alphanumeric noise, and keep up to **20** plausible substitutes as the chosen set. Before training they regularize the two distributions, because a small set of culprits (`Wait`, `So`, `the`) would otherwise dominate — and over-suppressing them degrades reasoning.
The next widget is the same idea in logit space — the mechanism the figure above doesn't draw. Toggle **reference → after FTPO** to watch the one trained position: the rejected `Wait` logit driven down, several chosen logits lifted up, and the entire rest of the vocabulary pinned near where it started.
That last property is the reason this works without collateral damage. FTPO isn't teaching the model anything about Davy Jones or about calculus; it's redistributing probability at the exact positions where the model was getting stuck, and leaving the rest of the distribution alone.
## Results
The headline is the doom-loop rate under greedy decoding. On the early LFM2.5-2.6B checkpoint it drops from 10.2% to 1.4%; on Qwen3.5-4B, from 22.9% to 1%.
Eval scores go up across the board — but Liquid is careful about the causal story, and so am I. The training set teaches the model nothing new about math or code; it removes the failure mode that was preventing the model from reaching answers it could already produce. A completion that used to spiral into `Wait, let me reconsider…` until it ran out of tokens now finishes and gets scored. The gain is recovered credit, not new capability.
### The temperature tradeoff
This is the honest catch, and it's a genuinely interesting one. Break the LFM2.5-2.6B evals out by decoding temperature:
The average score panel tells it: Antidoom leads by roughly 8 points at temp 0 (≈47 vs ≈38), peaks around temp 0.33, and then falls back to meet the base curve near temp 1.0 (≈45). The base model, meanwhile, climbs steadily with temperature — because sampling was its only escape from the loops. So FTPO effectively **shifts the model's best operating temperature downward**: it makes low-temperature decoding safe, which is where you'd want to run a small reasoning model anyway, but it gives up the high-temperature regime, where the extra randomness now mostly adds noise instead of buying an exit. That cuts against the usual intuition that reasoning models like a bit of temperature.
### Cost and recipe
The whole thing is cheap, which is the other reason to care. For the early LFM2.5-2.6B checkpoint, generating the training set took about **1 hour on 8× MI325** GPUs (bounded by the model's own loop rate, since generation stops when it catches loops), and training took about **1–2 hours on a single MI325**. The recipe:
```yaml
method: FTPO (DPO-family, final-token)
epochs: 1
adapter: LoRA # rank 128–256 — higher learnability, less degradation
train_modules: [attention_proj, mlp_proj, lm_head]
learning_rate: 4e-6 – 2e-5
early_stop: chosen_win = 0.35 # fraction of samples where chosen beat rejected
```
Two guardrails matter. **Over-training happens easily** — training past the `chosen_win=0.35` stopping point tended to degrade the model and, ironically, spawn *new* doom loops. Stopping at that threshold typically pulled loop rates from 20–30% down to 1–2% with minimal degradation. And FTPO is **iterative by design**: after one round the loop-causing tokens are rejected and probability is reweighted toward the chosen alternatives, but that can expose new failure points where *other* tokens start looping, so a second round targets the newly surfaced loops.
## The take
Doom loops are a small-model, low-temperature, hard-problem failure, and the diagnosis here is clean: overtrained discourse-marker priors plus self-reinforcing context plus greedy decoding equals a distribution that collapses onto one token with no exit. FTPO is a tidy fix because it matches the shape of the problem — retrain the one position that restarts the loop, spread the escape probability across several tokens instead of minting a new favourite, and pin the rest of the vocabulary in logit space so you don't disturb what the model already does well. The results are strong (10.2% → 1.4%, 22.9% → 1%) and, importantly, honestly framed: the eval gains are recovered credit for answers the model could already reach, the win is concentrated at low temperature and fades by temp 1.0, over-training is a real risk with a specific stopping rule, and it can take more than one round. For anyone shipping a small reasoning model that greedy-decodes in production, a 1–2 hour LoRA pass that removes a 10–20% failure mode is an easy trade.
---
*Built on Liquid AI's [Antidoom: Reducing Doom Loops with Final Token Preference Optimization](https://www.liquid.ai/blog/antidoom) (2026) and the [`LiquidAI/antidoom-mix-v1.0`](https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0) dataset. All benchmark numbers are provider-reported; the four charts are reproduced from the blog for commentary, and the two interactive widgets are illustrations of the mechanism, not measured traces.*
---
# A field guide to attention mechanisms
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/attention-mechanisms
> date: 2026-07-08
> tags: explainer, attention, transformers, long-context, kv-cache
Attention is a single operation, and almost everything else in a transformer is plumbing around it. The zoo of named variants — MQA, GQA, MLA, sliding-window, BigBird, NSA, FlashAttention — can read like a pile of unrelated tricks. It isn't. Nearly every one is a targeted answer to a **specific bill** that plain attention runs up, and once you know which bill a mechanism is paying down, the whole field organizes itself.
This is a map, not an encyclopedia. For the fundamentals — what Q, K and V are and why the dot product means "relevance" — start with [how transformers attention works](/articles/how-transformers-attention-works). Here I assume that and go wide: the families, the mechanics, the exact costs, and the honest thing each one gives up.
## The one operation, and its two bills
For one query vector, attention scores every key by dot product, turns the scores into weights with softmax, and returns the weight-blended values:
$$
\mathrm{Attention}(Q,K,V) = \mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V
$$
$Q \in \mathbb{R}^{N\times d_k}$ are the queries, $K \in \mathbb{R}^{N\times d_k}$ the keys, $V \in \mathbb{R}^{N\times d_v}$ the values, $N$ the sequence length, and $d_k$ the per-head key dimension. The $1/\sqrt{d_k}$ is not cosmetic: if $q$ and $k$ have unit-variance entries, $q\cdot k$ has variance $\approx d_k$, so without the scale the logits grow with head width and push softmax into a near–one-hot corner where its gradient vanishes. Multi-head attention runs $h$ of these in parallel on different learned projections and concatenates — different heads settle on different relationships over the same tokens.
Now the bills. That one line hides two very different costs, and they are the two axes this whole guide is organized around.
**Bill one — the O(N²) score matrix (the compute / *mask* axis).** The product $QK^{\top}$ is an $N\times N$ matrix: every query dotted with every key. Time and memory are $O(N^2 d)$. Double the context and the work quadruples. Every *sparse* method is an answer to this bill — a rule for **which (query, key) pairs to actually compute**.
**Bill two — the KV cache (the memory axis).** At inference a decoder generates one token at a time and re-reads the whole past, so it caches the keys and values it already computed. That cache is $2 \cdot n_h \cdot d_h$ elements **per token, per layer** (one key and one value per head) — it grows linearly with context *and* with depth, and at long context it, not the FLOPs, is what pins you to the hardware. Every *KV-sharing* and *compression* method is an answer to this bill.
A third, quieter cost sits underneath both: attention is **memory-bandwidth bound** on real GPUs — moving the $N\times N$ scores in and out of HBM dominates. That is not a math problem, and it gets its own family (FlashAttention) that changes *how* the exact same numbers are computed.
Keep the two bills in mind and the families below stop looking like a zoo.
## Bill one: which pairs do we compute? (the mask axis)
The cleanest way to cut the $O(N^2)$ matrix is to not compute most of it. A **mask** says, for each query, which keys it may read; the rest are dropped. The single diagram below is the shared language for this entire axis — a query cursor lighting exactly the keys it is allowed to see, under seven different patterns. Every later sparse diagram reuses these colors.
**Bidirectional (encoder) attention** is the no-mask case: every query reads the whole sequence, both directions. It is what BERT-style encoders use, and it is the full $O(N^2)$ bill — appropriate when the input is short and you want maximum mixing.
**Causal (decoder) attention** masks the strict upper triangle: a query at position $q$ may read positions $0\ldots q$ only, never the future. This is what every autoregressive LLM uses. It halves the constant but the asymptotics are still $O(N^2)$ — the triangle is half of a square.
**Cross-attention** is a different picture entirely, because the queries and the keys come from *different sequences*. The decoder's query reads the encoder's keys and values, with no causal mask, because the whole source is already known. It is the join between two sequences — French to English in translation, image patches to caption in a vision-language model.
### Structured sparsity: a fixed, position-based pattern
The first real savings come from a *fixed* rule that depends only on position.
**Sliding-window attention** (Mistral 7B) lets each query read only the previous $w$ keys, dropping cost to $O(N\cdot w)$ and — crucially — capping the KV cache at $w$ per layer instead of $N$. Mistral uses $w=4096$ over 32 layers. The catch is obvious: one window layer cannot see past $w$. The rescue is depth. Stacked window layers **compound** their reach — after $k$ layers information can travel $k\cdot w$ tokens, so Mistral's last layer has a theoretical span of $4096\times 32 \approx 131{,}000$ tokens. That "theoretical" matters: a window model needs either stacking, or global/sink tokens, or interleaved global layers, or it genuinely loses long-range information.
**Attention sinks** (StreamingLLM) explain *why* a naive sliding window degrades, and fix it cheaply. Softmax weights must sum to 1, so a query with nothing important to attend to still has to put its mass *somewhere* — and models learn to dump that excess onto the first few tokens. Evict those tokens (as a rolling window does) and the whole distribution destabilizes; perplexity explodes. Keeping just **four initial tokens** as always-visible "sinks" plus a recent window restores stable streaming to millions of tokens, with a reported up-to-22× speedup over recomputation. The sinks carry almost no information — they are a pressure-release valve for softmax.
**Block-sparse attention** (BigBird) combines three fixed pieces at block granularity — a local **window**, a few **global** tokens every query sees, and a few **random** blocks for mixing — for $O(N)$ cost. The theory is the reassuring part: with the global and random pieces, BigBird is still a universal approximator of sequence functions and Turing complete, so the sparsity does not cost you expressive power in principle. Longformer is the same local-plus-global idea for long documents.
**Dilated / strided attention** attacks distance instead of density. The Sparse Transformer factorizes attention into **strided** and **fixed** patterns for $O(N\sqrt{N})$; LongNet's **dilated attention** grows the stride exponentially with distance so any two tokens connect in a logarithmic number of hops, reaching $O(N)$. A single dilated head skips across the sequence; combine a few at different strides and every position stays reachable.
These structured patterns are cheap and predictable, and they are the workhorses of production long-context models — usually **interleaved** with occasional full-attention layers so long-range information still has a path. Gemma 2 alternates local:global 1:1 (window 4096); Gemma 3 shifts to 5:1 with a 1024 window, so only about one layer in six caches the full 128K context and KV-cache overhead drops from roughly 60% to under 15% with little quality loss. MiMo-V2-Flash uses one global layer in six; Character.AI reports a similar 5:1, 1024-window design with over 20× KV-cache reduction. Two of these get full treatments here: [Gemma 4's interleaving and KV budget](/articles/gemma-4) and [MiMo-V2-Flash's 5:1 hybrid](/articles/mimo-v2-flash).
### Content-based sparsity: let the query choose
Fixed patterns are blind to content — they read the same positions whether or not those positions matter. The 2025–26 frontier is **learned, content-based selection**: score the past cheaply, then read only the blocks that actually matter *for this query*.
**Native Sparse Attention** (DeepSeek, 2025) is the cleanest example. For each query it runs three branches in parallel over the same KV — a **compression** branch that squashes the past into coarse block summaries (global gist), a **selection** branch that scores blocks and keeps the top-$n$ at full resolution (the important detail), and a **sliding-window** branch (local coherence) — then a learned gate blends them:
$$
o_t = \sum_{c\,\in\,\{\text{cmp},\,\text{slc},\,\text{win}\}} g_t^{\,c}\;\mathrm{Attn}\!\left(q_t,\,\tilde K_t^{\,c},\,\tilde V_t^{\,c}\right),
\qquad g_t^{\,c}\in[0,1]
$$
The gate scores $g_t^{\,c}$ come from a small MLP-plus-sigmoid on the query. The reason this matters is in the name: NSA is **natively trainable** — all three read paths are differentiable, so the sparsity is learned end-to-end rather than bolted onto a dense checkpoint at inference time, and it is designed to be hardware-aligned (Tensor-Core-friendly block sizes).
Two siblings ship the same idea in production models, and I have covered both in depth: [MiniMax Sparse Attention](/articles/minimax-sparse-attention) scores the past in 128-token blocks and keeps the top-$k$ whole blocks; [LongCat Sparse Attention](/articles/longcat-2) goes finer with a hierarchical coarse-recall-then-token-select index, shares the index across layers, and reshapes the reads for coalesced memory access. This is the least-settled family in the guide — the shape of "trained-in sparsity" is still moving — but it is where the interesting long-context work is happening.
## Bill two: how do we share or shrink K/V? (the memory axis)
The mask axis leaves attention *exact within what it reads*. The memory axis is orthogonal: keep full attention, but pay less to cache K and V. This is the difference between sharing heads and compressing them, and the two are genuinely different axes — you can combine them.
**Multi-Query Attention** (MQA) is the blunt version: keep all $h$ query heads but share a **single** key head and value head across them. The cache drops from $2 \cdot n_h \cdot d_h$ to $2 \cdot d_h$ per token — a factor of $n_h$. It targets exactly the decode-time memory-bandwidth bottleneck, and it costs some quality and can destabilize training.
**Grouped-Query Attention** (GQA) is the middle ground almost everyone now uses. Split the query heads into $G$ **groups**; each group shares one key head and one value head, so the cache is $2 \cdot G \cdot d_h$. The important thing to get right: GQA interpolates by **KV-head groups**, not by reducing query heads — you keep all $h$ query heads, they just fan into $G$ shared KV heads. $G=1$ is exactly MQA; $G=h$ is exactly MHA. Llama 2 70B adopted it, and the quality is essentially MHA's at a fraction of the cache.
**Multi-head Latent Attention** (MLA, DeepSeek-V2) sits on a *different axis*: it does not share heads, it **compresses** them. K and V are jointly projected down to a shared low-rank **latent** vector $c_t^{KV}$, cached in place of the per-head keys and values, then up-projected back to all heads at compute time. Because RoPE is incompatible with folding the up-projection into the query, MLA carries positional information on a small **decoupled** key dimension. DeepSeek-V2 caches $\tfrac{9}{2}\,d_h$ per token — about what GQA with 2.25 groups would cost — while reporting quality at or above full MHA. (Its headline "93.3% smaller KV cache" is measured against DeepSeek 67B, itself a GQA model, not against full MHA; the clean comparison is the $\tfrac{9}{2}\,d_h$ figure.) Sharing versus compressing is the real distinction between GQA and MLA.
## Drop the matrix entirely: linear and kernelized attention
Both axes above still compute a softmax. **Linear attention** asks whether we need it at all. Softmax puts a non-linearity between $Q$ and $K$, which is exactly what forces the $N\times N$ matrix to exist. Replace it with a kernel feature map $\phi$ and the product reassociates:
$$
\mathrm{softmax}(QK^{\top})V \;\longrightarrow\; \phi(Q)\big(\phi(K)^{\top}V\big)
$$
Computing $\phi(K)^{\top}V$ first gives a $d\times d$ matrix, never an $N\times N$ one. For autoregressive decoding this becomes a running state updated once per token, exactly like an RNN:
$$
S_t = S_{t-1} + \phi(k_t)\,v_t^{\top},
\qquad \mathrm{out}_t = \frac{\phi(q_t)^{\top} S_t}{\phi(q_t)^{\top} z_t},
\qquad z_t = z_{t-1} + \phi(k_t)
$$
$S_t \in \mathbb{R}^{d\times d}$ is the fixed-size state, $z_t$ the normalizer. Time is $O(N d^2)$, memory is **constant** in $N$, and context is unbounded. The diagram below is the whole pitch: the softmax triangle grows quadratically as tokens stream in, while the linear state just updates in place.
The honest cost is real, and I want to state it plainly: linear attention is an **approximation** of softmax, and pure linear models are usually **weaker on recall** — pulling one exact fact out of a long context is precisely where a compressed fixed-size state struggles. Performer's FAVOR+ makes the approximation principled (random features that provably, unbiasedly approximate the softmax kernel); "lightning" and gated-linear variants add a decay so old state fades. In practice the winning form today is **hybrid** — MiniMax-01 interleaves seven lightning (linear) blocks per one softmax block, buying linear-time bulk with periodic exact attention to restore recall.
## Exact, but IO-aware: the systems layer
This family is different in kind, and it is worth being explicit: **FlashAttention and PagedAttention are not new attention functions.** They compute the exact same softmax attention, bit for bit. They change *where the bytes move*. I include them because "attention is slow" is usually a memory-traffic statement, not a FLOP statement, and these are the fix.
**FlashAttention** attacks the fact that materializing the $N\times N$ scores in slow HBM is the real bottleneck. It **tiles** the computation: a block of queries stays in fast on-chip SRAM while blocks of K and V stream past it, and it maintains an **online softmax** — a running max $m$ and running sum $\ell$ — so it can produce the exact softmax result without ever writing the full matrix to HBM.
The payoff is an IO complexity of $\Theta(N^2 d^2 / M)$ HBM accesses, where $M$ is the SRAM size, versus $\Theta(Nd + N^2)$ for the standard implementation — many-fold fewer round-trips for typical $d$ and $M$, and the reason a long-context forward pass stopped being memory-bound. FlashAttention-2 and -3 push the same idea with better GPU work-partitioning and FP8 on Hopper. It is exact; the only thing it "gives up" is the naive implementation's simplicity.
**PagedAttention** (vLLM) does for the KV cache what FlashAttention does for the scores — a systems fix, not a math one. Before it, a serving engine reserved one *contiguous* buffer per request sized to the maximum output length, so a short generation left most of its reservation wasted (internal fragmentation), and two requests could share nothing. PagedAttention stores the cache in fixed-size **blocks** with a per-sequence **block table**, exactly like OS virtual memory: blocks are handed out on demand (near-zero fragmentation), and a shared prompt prefix maps to the **same physical blocks** via copy-on-write.
The result is many more concurrent sequences per GPU with identical model outputs — which is why paged KV is now table stakes for serving.
## A quality move, not an efficiency one: differential attention
Not every variant is about cost. **Differential attention** (Microsoft, 2024) targets a *quality* failure: softmax spends attention mass on irrelevant tokens because the weights are forced to sum to 1 — the same pressure that creates attention sinks also creates broadband "attention noise." The fix borrows from differential amplifiers: compute two softmax maps and return their difference.
$$
\mathrm{DiffAttn}(X) = \Big(\mathrm{softmax}\!\big(\tfrac{Q_1 K_1^{\top}}{\sqrt{d}}\big) - \lambda\,\mathrm{softmax}\!\big(\tfrac{Q_2 K_2^{\top}}{\sqrt{d}}\big)\Big)V
$$
The irrelevant mass is roughly common to both maps, so it subtracts away; the genuine peaks, which differ between the maps, survive. $\lambda$ is learned per head (reparameterized, initialized around 0.8), and the reported effect is sparser attention and better long-context retrieval and in-context recall.
It costs roughly double the attention compute and cache (two maps), and it is still $O(N^2)$ — this buys accuracy, not efficiency. Worth it when the failure mode is a model that gets "distracted" in long context.
## The whole map, in one table
Complexities are per attention layer; $N$ is sequence length, $d$ the model width, $n_h$ heads, $d_h$ head dim, $w$ window, $k$ selected blocks, $M$ SRAM size. "KV cache" is per token per layer.
| Mechanism | Bill it pays | Time | KV cache | Exact? | Quality / recall | Reach for it when |
|---|---|---|---|---|---|---|
| MHA (full) | baseline | $O(N^2 d)$ | $2\,n_h d_h$ | exact | reference | short context; training |
| MQA | memory | $O(N^2 d)$ | $2\,d_h$ | exact | small drop | decode memory-bound |
| GQA | memory | $O(N^2 d)$ | $2\,G\,d_h$ | exact | ≈ MHA | default for large models |
| MLA | memory | $O(N^2 d)$ | $\tfrac{9}{2}\,d_h$ | exact | ≥ MHA (reported) | long context + quality |
| Sliding-window | compute | $O(N w d)$ | capped at $w$ | exact in window | loses distance unless stacked/interleaved | cheap long context |
| Sink / StreamingLLM | compute + memory | $O(N w)$ | sinks + window | exact in kept set | stable, not true long-range recall | unbounded streaming |
| Dilated (LongNet) | compute | $O(N)$ | bounded | exact in pattern | pattern-limited | extreme length |
| Block-sparse (BigBird) | compute | $O(N)$ | bounded | exact in pattern | near-full w/ global+random | long documents |
| Content-based (NSA / MSA / LSA) | compute | $O(N k)$ | blockwise | exact in selection | near-full if selection is good | trained-in long context |
| Linear / Performer | compute + memory | $O(N d^2)$ | $d\times d$ state | approximate | weaker recall | very long, recall-tolerant |
| FlashAttention | systems (IO) | $O(N^2 d)$, $\Theta(N^2 d^2/M)$ HBM | same as base | exact | none (identical) | always — default kernel |
| PagedAttention | systems (memory) | same as base | block-paged | exact | none (identical) | serving many sequences |
| Differential | quality | $O(N^2 d)$ (≈2×) | ≈2× | exact | better retrieval | reduce attention noise |
## What's settled, what's still moving
Some of this is infrastructure now. **GQA** is the default attention for large models; **FlashAttention** is the default kernel; **PagedAttention** is the default cache manager; **sliding-window interleaved with periodic global (or sink) layers** is the standard recipe for cheap long context. If you are building a model today, those four are choices you make without much agonizing.
The frontier is the content-based sparse family — **NSA**, **MiniMax Sparse Attention**, **LongCat Sparse Attention** — where the model *learns what to read*. The promise is compelling (near-full quality at a fraction of the reads, trained end-to-end) but the designs are still diverging on granularity, how the index is shared across layers, and how to make the reads hardware-friendly; there is no settled winner yet. **Linear and kernelized attention** remains the most tantalizing and the most caveated: constant-memory unbounded context is exactly what you want, and the recall gap is exactly why pure-linear models have not displaced softmax — hybrids are the pragmatic answer for now. And attention lives inside a larger design space: whether to specialize behavior at the **head** level rather than the layer level is its own question, which I dig into in [HydraHead](/articles/hydrahead).
The map is stable even as the territory shifts. Every new mechanism you meet is answering one of the same two questions: *which (query, key) pairs do we compute*, and *how do we pay for the K/V we keep*. Place it on those axes and you already understand most of what it does — and what it gives up.
---
*The interactive diagrams are illustrations of each mechanism, not measured traces; real windows, blocks, and head counts are far larger than what fits on screen. Primary sources: Attention Is All You Need (Vaswani et al., 2017, arXiv 1706.03762); Fast Transformer Decoding / MQA (Shazeer, 2019, arXiv 1911.02150); GQA (Ainslie et al., 2023, arXiv 2305.13245); DeepSeek-V2 / MLA (DeepSeek-AI, 2024, arXiv 2405.04434); Mistral 7B (Jiang et al., 2023, arXiv 2310.06825); StreamingLLM (Xiao et al., 2023, arXiv 2309.17453); Longformer (Beltagy et al., 2020, arXiv 2004.05150); BigBird (Zaheer et al., 2020, arXiv 2007.14062); Sparse Transformer (Child et al., 2019, arXiv 1904.10509); LongNet (Ding et al., 2023, arXiv 2307.02486); Native Sparse Attention (Yuan et al., 2025, arXiv 2502.11089); Differential Transformer (Ye et al., 2024, arXiv 2410.05258); Transformers are RNNs / linear attention (Katharopoulos et al., 2020, arXiv 2006.16236); Performer / FAVOR+ (Choromanski et al., 2020, arXiv 2009.14794); FlashAttention (Dao et al., 2022, arXiv 2205.14135) and FlashAttention-2/-3 (arXiv 2307.08691, 2407.08608); PagedAttention / vLLM (Kwon et al., 2023, arXiv 2309.06180); Gemma 2 and 3 (Gemma Team, 2024/2025, arXiv 2408.00118, 2503.19786); MiniMax-01 (2025, arXiv 2501.08313).*
---
# Frontier RL is cheaper than you think: ship deltas, not mega-clusters
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/frontier-rl-cheaper
> date: 2026-07-08
> tags: explainer, reinforcement-learning, systems, inference-optimization
Fireworks' argument in [*Frontier RL Is Cheaper Than You Think*](https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) is narrow and load-bearing: the belief that reinforcement-learning post-training needs one giant co-located cluster rests on a single assumption — that **every policy update ships the full ~1 TB checkpoint to the rollout fleet**. It does not. Between adjacent RL checkpoints most weights do not change at all, so you can ship a compressed **delta** — a couple of percent of the model — and keep a rollout fleet fresh over ordinary cross-region links. That is the whole claim, and everything else follows from it.
This is a **vendor blog**. Fireworks sells the training and rollout-serving platform this argument recommends, so read the framing accordingly. The numbers below are **Fireworks-reported** from one sample setup, not an independent benchmark. What is *not* vendor-specific is the underlying physics — weight-update sparsity in RL — which a separate paper ([arXiv 2602.03839](https://arxiv.org/abs/2602.03839)) reports independently. I have kept the two apart throughout.
## RL has two jobs, not one
The mega-cluster instinct comes from pretraining, where the systems problem is keeping **one** huge synchronous job saturated. RL is a different shape. An RL run has two coupled jobs:
- The **trainer** runs forward, reward computation, backward, and the optimizer step. It wants dense, tightly-coupled hardware — the pretraining kind.
- The **rollout fleet** samples trajectories from the *current* policy — that is, it runs inference on the latest weights, across many parallel requests. It wants inference throughput, and it can live anywhere.
Pretraining only has the first job. RL has both, and the awkward part is the seam between them: how do you keep a large rollout fleet generating from a *fresh enough* policy without stalling on checkpoint transfers every step? That coupling is the whole systems problem, and it is why RL cost lives somewhere non-obvious. The same two-job structure shows up whenever RL trains agentic models — the trajectory-generation half is exactly the rollout fleet here (I wrote about the trajectory side in [Agents-A1](/articles/agents-a1)).
## The 1 TB problem
A frontier checkpoint is around 1 TB. If every policy refresh really required shipping that whole tensor to the rollout fleet, the conclusion writes itself: keep trainer and inference on the same RDMA-class fabric, avoid long-distance transfers, treat remote capacity as second class. That is the mega-cluster story, and its side effect is economic — frontier RL looks like a market only a handful of companies with a co-located supercluster can enter.
The premise is the full-checkpoint transfer. Break it and the conclusion goes with it.
## The key insight: exploiting 98% sparsity
Between nearby RL checkpoints, most weights barely move. Fireworks reports that **more than 98% of weights in `bf16` remain bit-equivalent between consecutive checkpoints**, and the unchanged fraction is higher still at lower precision. Their explanation is mechanical: RL delivers a very sparse learning signal — a few bits of reward per rollout — so training runs a small learning rate, and most parameters shift so little in `fp32` that they never cross the threshold to change their 16-bit representation. The independent paper *Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL* ([arXiv 2602.03839](https://arxiv.org/abs/2602.03839)) reports the same phenomenon, often **around 99%** sparsity in practical RL settings.
If only ~2% of the bits change, you should move ~2% of the bits. The figure's illustrative per-tensor sample makes the sparsity concrete — the change is spread thinly, and even the busiest tensor moves under 1%:
The mechanism is a periodic **full base checkpoint every N steps** to reset the chain, then a **compressed delta** in between. Only changed weight slices survive into the payload; each carries reconstruction metadata (`prev_snapshot_id`, `tensor_checksum`, `shape + dtype`) so every rollout cluster rebuilds the exact next checkpoint **losslessly** from shared object storage, then verifies the checksum before swapping.
In Fireworks' sample setup a full checkpoint is **1024 GiB**, the average delta between adjacent checkpoints is **20.3 GiB — 1.98% of the model**, and over a 50-step window that cuts cross-region transfer volume by **about 94%** versus moving the full model every time. The arithmetic is worth writing down. With window $W$ steps, a full checkpoint every $N$ steps ($f=\lceil W/N\rceil$ fulls), delta fraction $\delta$, and $R$ regions, the total cross-region volume for a checkpoint of size $C$ is:
$$
V_\text{full} = W \cdot C \cdot R, \qquad V_\text{delta} = \bigl(f\,C + (W-f)\,\delta C\bigr)\cdot R
$$
so the fraction saved is $1 - \bigl[f + (W-f)\,\delta\bigr]/W$, independent of both $C$ and $R$. Plug in $W=50$, $N=25$, $\delta=0.0198$: $f=2$, and you move $[2 + 48\cdot0.0198]/50 \approx 5.9\%$ of the naive volume — the reported ~94% cut. Drag the delta size and the cadence and watch where the crossover sits:
Two things the model makes obvious. The **percentage** saved does not depend on how many regions you feed — but the **absolute** bytes you stop moving scale with every region, which is the entire point of going distributed. And push the delta slider toward 100% and the two curves converge: at full-checkpoint-every-step you are back to the mega-cluster premise, where a co-located RDMA fabric is the only thing that can absorb the traffic.
## Async RL, and why the delta size decides it
Small deltas are necessary but not sufficient. The other half is **asynchronous RL** (also called Pipeline RL): deliberately let the rollout fleet run a little **off-policy** so that generation and training overlap instead of taking turns. Idle samplers are the expensive failure mode; a few steps of staleness is usually an acceptable price to keep them busy.
That trade only works if the handoff is fast. Delta-compressed updates keep it small: Fireworks reports distributing a new checkpoint across globally-distributed rollout clusters takes **a few minutes end-to-end**, and the actual in-GPU-memory **weight swap stays well under a minute** because download and decompression are pipelined ahead of the swap. The trainer side is pipelined too — every step uploads to shared object storage, each rank caches its previous upload and transmits only the diff, upload is sharded across training GPUs, download across inference replicas, and compression plus transfer plus signaling run in the background so training never blocks.
The payoff is where the wall-clock goes: less time waiting on checkpoint movement, more time generating rollouts on fresh weights. But the async win depends on the delta being small enough to hide behind a generation window. Drag the payload up and watch the "warm" fleet start to stall and fall off-policy:
This is the honest coupling: async RL alone does not save you, and delta compression alone does not save you. It is the two together — a handoff small enough to overlap with generation — that turns a distributed fleet into usable capacity.
## A note on staleness
Running trainer and fleet asynchronously means the fleet always serves a policy a few steps behind the trainer. That gap is **staleness**, and it is a real tradeoff, not a free lunch. The systems layer does not remove it — the *algorithm* still has to tolerate off-policy data. What delta compression buys is a staleness that is **bounded and predictable**: policy movement becomes a routine background operation instead of a stop-the-world full-checkpoint transfer. If your RL algorithm cannot stomach any off-policy data, none of this applies.
## Multi-region rollout capacity
Here the systems point turns strategic. Most teams do not have one contiguous idle supercluster for rollouts; they have GPUs scattered across regions, clouds, and availability zones, and stitching them into one co-located sampler fleet is painful even when the aggregate count exists. Once weight updates are small, that fragmented capacity becomes usable: each rollout cluster independently pulls and reconstructs weights from the same shared delta chain, with **no direct connection back to the trainer**. Add, remove, or rebalance clusters while they all track the same stream of policy updates.
Fireworks cites this in production: they say they ran Cursor's **Composer 2** RL training this way, with Federico Cassano describing the run as ["distributed across 3 (sometimes 4) different clusters around the world"](https://x.com/ellev3n11/status/2034778708163404102). Treat that as a vendor-provided data point — a single external quote, not a controlled measurement — but it is at least a concrete one, and it is a coding-agent model, the same agentic-RL regime as [Agents-A1](/articles/agents-a1).
The approach is not new territory Fireworks invented alone: it names [AReaL](https://arxiv.org/abs/2505.24298) for async RL and rollout-training disaggregation, and engineering notes from [Kimi](https://moonshotai.github.io/checkpoint-engine/) and [MiniMax](https://www.minimax.io/news/forge-scalable-agent-rl-framework-and-algorithm) on RL parameter updates and async scheduling. The contribution is running the delta-compressed, multi-region version in production.
## When this argument stops working
The blog is unusually clear about its own boundaries, and the caveats matter:
- **Small models.** If trainer and rollout inference fit on one node or a compact cluster, bandwidth was never the bottleneck and the simpler co-located setup wins. The whole argument is about the ~1 TB regime.
- **Very frequent checkpoints.** If the trainer emits updates faster than the delta pipeline can distribute and apply one, staleness becomes the limiting factor and tighter co-location can make more sense again.
- **Entangled rollout stacks.** If your rollout workers don't cleanly separate inference from training, treating them as a standard inference fleet is a poor fit and the disaggregated design loses its appeal.
There is also an assumption baked into the headline sparsity number: it is the fraction of `bf16` weights that stay **bit-identical**. That is exactly the right metric for a lossless delta, but it is a property of the *representation*, not just the math — a run at higher precision, a larger learning rate, or a more aggressive RL objective could move more bits and shrink the win. The ~2% is a measured sample, not a guarantee.
## The take
Strip the vendor framing and the load-bearing claim is a clean systems observation: RL post-training updates are sparse enough that the weight-sync between trainer and rollout fleet — the thing everyone assumed forced co-location — is ~2% of what you feared. That is corroborated independently ([arXiv 2602.03839](https://arxiv.org/abs/2602.03839)), and the engineering that follows (delta compression + checksummed reconstruction + async overlap + sharded pipelined transfer) is the ordinary, correct way to cash it in. The reproducible headline — ~1 TB checkpoint, 20.3 GiB average delta, ~94% less cross-region traffic over 50 steps — comes from one Fireworks sample setup, and the Composer 2 case is a single external quote, so treat the specific figures as illustrative rather than benchmarked. But the shape of the argument survives that discount: if the weights barely change, the mega-cluster was never load-bearing for RL. That is a genuinely useful thing to know before you go shopping for a co-located supercluster.
---
*Built on Fireworks AI's [Frontier RL Is Cheaper Than You Think](https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) (published 2026-03-23). All quantitative figures are Fireworks-reported from a sample setup unless attributed to [arXiv 2602.03839](https://arxiv.org/abs/2602.03839); the `DeltaCost` and `RolloutTimeline` widgets are my own cost models of their argument (relative/illustrative units), and the two reproduced diagrams are from the source post for commentary. The per-tensor bar chart uses the figure's own illustrative sample values.*
---
# Gemma 4: an open multimodal family, tuned to the KV-cache budget
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/gemma-4
> date: 2026-07-08
> tags: explainer, llm, architecture, kv-cache, long-context
**Gemma 4** is Google DeepMind's next open-weight family: natively multimodal (text, image, audio), released under **Apache 2.0**, with sizes from **2.3B to 31B**. Four dense models — effective **2.3B (E2B)**, **4.5B (E4B)**, **12B**, **31B** — plus a Mixture-of-Experts variant, **26B-A4B**, with **~3.8B activated** of 26B total. The headline features read like a checklist: a **thinking mode**, an **encoder-free 12B**, and a raft of efficiency work aimed at long context. I care about the last part most, because it's the part that decides whether you can actually serve these on the hardware you have.
Every benchmark here is **first-party** — Google's own report, thinking mode on unless noted, and the Gemma 3 27B column it compares against is **non-thinking**, so some of the generational jump is thinking-mode-vs-not rather than raw capability. The Arena ranking is real human eval but self-cited. Read it as a strong *open* family for its size, not an all-around SOTA claim: on the open leaderboard, several larger MoE models still sit above it.
## Where the sizes land
The family splits by how much of the model runs per token, and by how "effective" and "total" parameters differ:
| Model | Type | Total | Runs / token | Note |
|---|---|---|---|---|
| E2B | dense | 5B | 2.3B effective | per-layer embeddings (Gemma 3n trick) |
| E4B | dense | 8B | 4.5B effective | per-layer embeddings |
| 12B | dense | 12B | 12B | the encoder-free unified model |
| 31B | dense | 31B | 31B | leading dense open model on Arena |
| 26B-A4B | MoE | 26B | ~3.8B activated | sparse routing |
"Effective" is the Gemma 3n move: **E2B** and **E4B** keep 5B and 8B parameters on disk but stream **per-layer embeddings** so only 2.3B / 4.5B are resident in the compute path. It's a memory-placement trick for on-device, not a routing one — orthogonal to the MoE sparsity in 26B-A4B, where a router activates ~3.8B of 26B per token.
## Long context is a memory problem
The expensive part of a long prompt isn't the matmul — it's the **KV cache**. Every past token leaves a key and a value in every attention layer, and under full attention that cache grows linearly with context on every layer:
$$
\text{KV}_{\text{full}} \;\propto\; 2 \cdot n_{\text{layers}} \cdot L \cdot d_{\text{kv}}
$$
The factor of 2 is K *and* V; $L$ is the context length. At 128k tokens this is what fills your accelerator's memory. Gemma 4 attacks it structurally. Most layers are **local** — sliding-window attention that only looks back a fixed window $W$ — and only **1 in 6** is **global**, attending over the whole context. The ratio is **5:1** for most models, **4:1** for E2B:
That interleave changes the scaling. Local layers stop growing once the context passes $W$; only the global layers keep paying for length. Then two more moves shrink the global term itself. On global layers Gemma 4 **reuses keys as values** ($V = K$), storing one tensor where full attention stores two — except on E2B/E4B. Position uses **pp-RoPE** ($p = 0.25$) on global layers and ordinary RoPE (frequency 10k) on local ones, with global frequency 1M. Combined with KV-cache sharing, the report puts the **global KV-cache cut at 37.5%**. So the split KV cost looks like:
$$
\text{KV}_{\text{g4}} \;\propto\; \underbrace{2\,n_{\text{loc}}\,\min(L,W)\,d_{\text{kv}}}_{\text{local: bounded by }W} \;+\; \underbrace{n_{\text{glb}}\,L\,d_{\text{kv}}}_{\text{global: }V{=}K\text{, no factor of 2}}
$$
Slide the context up and watch which term dominates — the sliding window is the structural win, `values=keys` shaves what's left:
The payoff shows up in the real memory table. At 32k context the int8 KV cache adds only **+0.05 GB** (E2B), **+0.28 GB** (12B / 26B-A4B), **+1.10 GB** (31B) on top of the quantized weights — small enough that the weights, not the cache, stay the budget:
| Model | Weights (bf16) | Quantized | + int8 KV @ 32k |
|---|---|---|---|
| E2B | 4.6 GB | 0.8 GB | +0.05 GB |
| E4B | 9.0 GB | 2.3 GB | +0.14 GB |
| 12B | 24.0 GB | 7.65 GB | +0.28 GB |
| 26B-A4B | 52.0 / 7.6 GB | 16.2 / 2.8 GB | +0.28 GB |
| 31B | 64.0 GB | 19.2 GB | +1.10 GB |
The quantized column is **quantization-aware training** (QAT), not a post-hoc round: the model trains with fake-quant in the loop so int4/int8 weights land with minimal quality loss. That's how 31B fits in ~19 GB.
## The encoder-free 12B
The most unusual model is the **12B**, trained from scratch with **no vision or audio encoder at all**. Normally a multimodal LLM bolts a frozen ViT and a speech encoder onto the token stream. Gemma 4 12B replaces them with projections:
- **Vision:** it takes raw **48×48×3 RGB patches** and projects them with a **single 35M-parameter matmul** — standing in for the **550M** ViT the larger models use. 2D coordinate embeddings and a LayerNorm carry spatial position.
- **Audio:** the **305M USM conformer encoder is discarded entirely**. Raw audio is cut into **40ms chunks at 16kHz** (640-dim vectors) and projected straight into the embedding space. Audio is already temporal, so no extra positional encoding.
The bet: give a large enough LLM the raw patches and it learns the encoder's job internally, for less memory and no separate frozen tower to serve. The report's Table 8 argues the 12B stays competitive on audio-text tasks without the dedicated encoder — a genuinely different design point from the encoder-plus-LLM norm, and the honest caveat is that it only did this for one size.
For the models that *keep* a vision encoder, the input pipeline is worth a look too. Gemma 4 supports **variable aspect ratios** rather than square-cropping, resizing an image to fit a token budget $N_{\max} \in \{70, 140, 280, 560, 1120\}$ while mostly preserving shape:
## Thinking mode
Gemma 4 adds a **thinking mode**: before answering, the model can emit a reasoning trace, which lifts math and coding. It's a post-training addition on top of a Gemma 3-style recipe, toggled by a control token in a leading system turn:
```text
<|think|> # activates the reasoning trace for this turn
...user turn...
# IT models close a turn with ; base (PT) models emit
```
Because the trace is optional, you pay for it only on the prompts that need it. The flip side, for anyone reading the tables: the Gemma 3 27B baseline is non-thinking, so a row like AIME (89.2 vs 20.8) is partly measuring the mode, not only the model.
## MTP drafter: speculative decoding without prefill
Decoding is memory-bandwidth-bound — one token per forward pass, weights re-read each step. The usual fix is **speculative decoding**: a small draft model proposes several tokens, the big model verifies them in one pass, and accepted tokens are free. Gemma 4 ships a **multi-token-prediction (MTP) drafter head** for exactly this.
The drafter is a **4-layer Transformer block** (model dim 256 for E2B/E4B, 1024 for 26B-A4B/31B; three local and one global attention layers). It reuses the main model's last-layer activations and **cross-attends to the main model's KV cache**, so it needs **no prefill of its own** and supports any draft length. On E2B/E4B there's a further trick: instead of projecting the draft over the full **262k** vocabulary, it does a top-k over token clusters, cutting the final matmul from $d \times 262\text{,}000$ to $d \times 4096$ at a similar acceptance rate. The report doesn't publish an end-to-end speedup, so I won't invent one — but the design (no prefill, cross-attention to the live cache) is the part worth copying.
If speculative decoding and MTP are new, I built the idea up from scratch in [Multi-Token Prediction](/articles/multi-token-prediction).
## The numbers
Start with human eval, since it's the least gameable. On **Arena Text** (blind side-by-side, Elo, as of June 2026), Gemma 4 31B is the **top dense open model** — but the leaderboard around it is mostly much larger MoE systems, and a closed model tops it:
| Rank | Model | Elo | Open | Params / active |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 1508 | no | – |
| 15 | GLM 5.1 | 1475 | yes | 744B / 40B |
| 38 | DeepSeek V4 Pro | 1456 | yes | 1.6T / 49B |
| **43** | **Gemma 4 31B** | **1451** | **yes** | **31B dense** |
| 61 | Gemma 4 26B-A4B | 1438 | yes | 26B / 4B |
| 157 | Gemma 3 27B | 1366 | yes | 27B dense |
An Elo of 1451 at **31B dense** against 744B–1.6T MoE models a dozen ranks up is the real story: it's punching well above its parameter count, not topping the board. On static reasoning benchmarks the family scales cleanly, and the jump over Gemma 3 27B is large (thinking-mode caveat noted):
The rest of the text suite tells the same story — the 31B leads its own family, the 26B-A4B tracks close behind at a seventh of the active params, and both clear Gemma 3 27B by a wide margin:
| Benchmark | 31B | 26B-A4B | 12B | E4B | E2B | Gemma 3 27B |
|---|---|---|---|---|---|---|
| MMLU Pro | 85.2 | 82.6 | 77.2 | 69.4 | 60.0 | 67.6 |
| LiveCodeBench v6 | 80.0 | 77.1 | 72.0 | 52.0 | 44.0 | 29.1 |
| Codeforces Elo | 2150 | 1718 | 1659 | 940 | 633 | 110 |
| SciCode | 43.0 | 40.0 | 38.0 | 24.0 | 21.0 | 21.0 |
| IFEval | 98.9 | 98.5 | 97.2 | 96.7 | 94.6 | 90.4 |
| MMMLU | 88.4 | 86.3 | 83.4 | 76.6 | 67.4 | 70.7 |
Vision holds up (MMMU Pro 76.9 / MATH-Vision 85.6 / InfographicVQA 92.0 for 31B at full resolution), and long context does what the KV work promises — **RULER at 128k** stays high where Gemma 3 27B falls off:
| Long-context @ 128k | 31B | 26B-A4B | 12B | E4B | Gemma 3 27B |
|---|---|---|---|---|---|
| RULER (accuracy) | 96.4 | 89.8 | 91.2 | 86.6 | 66.0 |
| LOFT (recall@k) | 79.5 | 66.3 | 66.4 | 58.5 | 8.6 |
## The take
Gemma 4's contribution isn't a benchmark crown — it's **efficiency engineering shipped in the open**. The pieces compose: a 5:1 local:global ratio and `values=keys` on global layers keep the KV cache flat enough that a 128k prompt costs single-digit gigabytes of cache; QAT puts 31B in ~19 GB; the encoder-free 12B deletes two frozen towers; the MTP drafter does speculative decoding without a second prefill. None of these is individually novel, but the combination is a serious on-device and single-accelerator story, released under **Apache 2.0** with a **31B dense model that is the top open dense entry on Arena**.
The honest caveats are the first-party ones. The benchmarks are Google's own, thinking mode is on for Gemma 4 and off for the Gemma 3 baseline it's measured against, and "leading open model" is true only in the **dense** category — larger open MoE models (GLM 5.1, DeepSeek V4) rank above it on the same board. The encoder-free design is proven at exactly one size (12B), and there's no published MTP speedup to hold them to. For a team that wants an open, multimodal, long-context model that fits the hardware it already owns, the 12B and 31B are the ones I'd reach for — and the KV-cache design is the part I'd study regardless of which model I ended up serving.
---
*Built on the [Gemma 4 Technical Report](https://arxiv.org/abs/2607.02770) (Gemma Team, Google DeepMind, 2026), Apache 2.0 model license. All benchmark numbers are first-party from the report (thinking mode unless noted; the Gemma 3 27B baseline is non-thinking). The interactive diagrams are schematic illustrations of the mechanism, not measured traces; the two paper figures are reproduced for commentary.*
---
# Hunyuan Hy3: Tencent's 295B-A21B MoE, and the community 1M GGUF
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/hunyuan-hy3
> date: 2026-07-08
> tags: explainer, llm, mixture-of-experts, long-context, quantization, inference-optimization
**Hy3** is Tencent Hunyuan's newest open model — released as a *preview* checkpoint, `Hy3 preview (295B A21B)`. It is a Mixture-of-Experts language model with **295 billion total parameters** that activates **21 billion per token**, ships **open weights** on Hugging Face and ModelScope, and serves a **256K native context**. Tencent frames it as a *hybrid fast-and-slow-thinking* model: one set of weights, a `reasoning_effort` knob that goes from `no_think` to `high`. The pitch is efficiency — "intelligence comparable to flagship models with two to five times its parameter scale."
I read all three primary sources for this: the [research page](https://hy.tencent.com/research/hy3), the [`Tencent-Hunyuan/Hy3-preview` repo](https://github.com/Tencent-Hunyuan/Hy3-preview), and a community [`Hy3-1M-GGUF`](https://huggingface.co/satgeze/Hy3-1M-GGUF) build that YaRN-extends the context to 1,048,576 tokens for local inference. Two things are worth pinning down before any of the benchmark charts: the "1M" context is a **community artifact**, not a Tencent release, and the whole benchmark suite is **provider-reported**. Both matter, and I keep them separate below.
Everything numeric here is **provider-reported** — Tencent's own harness for the model charts, the community author's own samples for the GGUF. Hy3 tops **no** frontier benchmark: on the hardest STEM and agentic tasks **GPT-5.4**, **Gemini-3.1-Pro**, and **Claude Opus 4.6** all lead. Read it as a strong model *for its 21B active size*, competitive with **GLM-5** and **Kimi-K2.5**. The "comparable to 2–5× larger models" line and the "270-expert blind eval (2.67/4 vs GLM-5.1's 2.51/4)" are provider claims I cannot independently verify. And the **1M context is a community YaRN extension** of a model trained to **256K** — explicitly experimental and not needle-certified.
## Where the 295B lives
The parameter accounting is the whole economic argument, so start there. Hy3 is **80 decoder layers** of MoE, plus **1 extra multi-token-prediction (MTP) layer** (3.8B params). Each MoE layer holds **192 experts**; the router keeps the **top-8** per token. So of 295B total, only ~**21B is active** on any given token — roughly a **7% activation rate**. That sparsity is what lets a 295B model serve at the cost of a ~21B dense one.
Attention is **grouped-query attention (GQA)**: **64 query heads** but only **8 KV heads**, `head_dim` 128, hidden size 4096. The 8-way sharing is not a detail — it is an 8× cut in KV-cache memory versus full multi-head attention, and it is the reason a 256K window is affordable at all (more on that below). The MTP layer drafts the next token so the main model can **verify several tokens per step** — speculative decoding, built into the weights rather than bolted on.
Walk one token through a block. Flip stages to see the KV grouping, the top-8 route, and the draft head:
If the MoE routing here is new, I built it from nothing in [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch) — the router, the top-k gate, and why activating a sparse subset is the entire reason a model this large is cheap to run. The MTP head is the same idea I unpacked in [Multi-Token Prediction](/articles/multi-token-prediction): predict more than one token so a cheap draft can be verified in a single forward pass.
One honest architecture note Tencent does not hide but does not emphasize: this is the **preview** release. The card labels it `295B A21B`, and the community GGUF card rounds the same weights to `~299B / ~17B active`. I use Tencent's official figures (295B / 21B) throughout; the ~2% discrepancy is a community rounding, not a second model.
## Training: one model, two speeds
Tencent is thin on the training story, and I will not pad it. What the sources actually say: Hy3 was built with "strengthened reinforcement learning and enhanced data quality and diversity," and refined "through use by global developers and across Tencent's large-scale real-world business scenarios." No token count, no data-mix percentages, no stage breakdown. Treat the training narrative as a claim.
The concrete, testable part is the **hybrid thinking** interface. A single `reasoning_effort` parameter switches inference mode:
- `no_think` — direct answer, no chain-of-thought (the default, and the cheapest).
- `low` — moderate CoT.
- `high` — deep CoT for hard reasoning.
That is exposed at serve time, so you pay for reasoning only when the task needs it. On an OpenAI-compatible endpoint you set it per request:
```python
# after deploying Hy3 behind vLLM / SGLang (OpenAI-compatible)
resp = client.chat.completions.create(
model="hy3-preview",
messages=[{"role": "user", "content": "prove there are infinitely many primes"}],
temperature=0.9, # Tencent's recommended default
top_p=1.0,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
```
The pretrained base model is where the numbers are cleanest, because base-model evals are the least harness-sensitive. On general knowledge Hy3-Base sits a hair under the larger open models — **MMLU 87.42** (Kimi-K2 88.24, GLM-4.5 87.73, DeepSeek-V3 87.68) — but it *leads* its comparison set on several reasoning and math rows:
| Benchmark | Hy3-Base | Kimi-K2 | DeepSeek-V3 |
|---|---|---|---|
| MMLU | 87.42 | **88.24** | 87.68 |
| MMLU-Pro | 65.76 | **65.98** | 63.98 |
| MATH | **76.28** | 71.20 | 59.37 |
| GSM8K | **95.37** | — | — |
| LiveCodeBench-v6 | **34.86** | — | — |
| CRUXEval-I | **71.19** | — | — |
| SuperGPQA | **51.60** | — | — |
| MMMLU (multilingual) | **80.15** | 77.63 | 79.54 |
The pattern holds across the suite: Hy3-Base trails the biggest models slightly on broad knowledge, then pulls ahead on math (MATH 76.28 vs DeepSeek-V3's 59.37 is not close) and code reasoning. For a model activating 21B, that is the interesting shape.
## Long context: 256K trained, 1M borrowed
Here is where the honesty matters most. Hy3's **trained** context is **256K**. The headline "1M" in this article's third source is a **community** build (`satgeze/Hy3-1M-GGUF`) that applies **YaRN** to stretch positional encoding out to **1,048,576 tokens**. YaRN rescales RoPE frequencies so the model can *index* positions it never saw in training — it extends reach, it does not re-train competence. The GGUF card says so plainly: the 1M window is "unverified at full length," experimental, and **not yet needle-certified**.
So the useful mental model is two numbers: 256K you can lean on, and a 1M ceiling you should measure before trusting. Drag the window below and watch the KV-cache cost — and the point where you cross from *trained* into *extrapolated*:
The KV-cache arithmetic is worth doing by hand, because it explains both the GQA choice and why 1M is expensive regardless of quality. Per token, the cache stores K and V for every layer:
$$
\text{bytes/token} = 2 \times n_{kv} \times d_{head} \times L \times b
= 2 \times 8 \times 128 \times 80 \times b
$$
With $b = 2$ bytes (fp16) that is **320 KiB/token**. Multiply by context:
- **256K** tokens → ~**80 GiB** of KV cache (fp16), ~40 GiB at fp8.
- **1M** tokens → ~**320 GiB** (fp16), ~160 GiB at fp8 — *on top of* the weights.
Now the GQA payoff is obvious. Full MHA would use **64 KV heads**, not 8, so every one of those numbers would be **8× larger** — 640 GiB of cache at 256K. GQA is not a quality trick here; it is what makes the window fit in memory at all. If you want to push the cache down further, that is exactly the territory of [TurboQuant KV-cache quantization](/articles/turboquant-kv-cache) and why [how LLM inference works](/articles/how-llm-inference-works) spends so long on the cache.
Tencent's own long-context numbers land where you'd expect for a 256K-trained model: on **LongBench v2** Hy3 scores **65.4** (up from Hy2's 56.4), matching Kimi-K2.5 (65.6) and edging GLM-5 (62.5), while GPT-5.4 (67.4) and Gemini-3.1-Pro (67.1) lead. On **AA-LCR** it's **66.3**. Competitive at its size, not a long-context leader.
## Quantization and running it locally
The base weights are ~**590 GB** in BF16 — multi-GPU territory. Three paths bring that down.
**FP8, at serve time.** vLLM quantizes the loaded BF16 weights online with `--quantization fp8`, roughly **halving the footprint to ~295 GB** (sources conflict on whether a standalone `Hy3-FP8` checkpoint also ships — the runtime path is the one I'd rely on). Tencent's `AngelSlim` toolkit adds low-bit quantization and speculative-sampling support on top.
**GGUF, for CPU/Mac.** This is what the community `Hy3-1M-GGUF` build is for: `llama.cpp`-style quantization that runs on a single machine with lots of RAM instead of a rack of GPUs. The quant ladder trades size for quality. Pick a RAM budget and see what fits:
The sizes on that chart are the exact ones the card reports; the quality pip is an *illustrative* ordering from bits-per-weight, not a measured score — the card is explicit that even the good quants prove "coherence and basic instruction-following, not reasoning, long-context retrieval, or factual accuracy." The practical reads:
- **IQ1_M (62 GB)** fits a 128 GB Mac but is visibly weaker (it dropped list formatting in the author's samples).
- **IQ2_M (~92 GB)** is the recommended baseline; **MTP-IQ2_M (~100 GB)** bakes in a `q8_0` draft head for speculative decode.
- **Q4_K_M (183 GB)** is the highest-quality GGUF and needs a **192 GB+** box.
Running it is a `llama.cpp` server with the model's chat template and the extended context:
```bash
# 256K context; needs a hy_v3-capable llama.cpp build
llama-server -m hy3-1M-IQ2_M.gguf -c 262144 -np 1 --jinja \
--chat-template-file chat_template_llamacpp.jinja
# MTP speculative decode (draft head)
llama-server -m hy3-1M-MTP-IQ2_M.gguf -c 262144 --jinja \
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.75
```
Reported throughput: **24–25 tok/s** generation on a **MacBook Pro M3 Max (128 GB)**, dropping to 17–19 tok/s in long conversations. The MTP draft head gives **+26–37%** on CUDA (an RTX 5090) but is roughly **neutral on Apple Silicon** — speculative decode helps when verification is compute-bound, which it is on the GPU and mostly isn't on the Mac. That is an honest, useful asymmetry: don't expect the draft head to speed up your laptop.
For the datacenter path, the official serving stacks are vLLM and SGLang, both with the MTP/EAGLE draft and Hunyuan's tool + reasoning parsers:
```bash
vllm serve tencent/Hy3-preview \
--tensor-parallel-size 8 \
--speculative-config.method mtp --speculative-config.num_speculative_tokens 1 \
--tool-call-parser hy_v3 --reasoning-parser hy_v3 \
--enable-auto-tool-choice --served-model-name hy3-preview
```
## The benchmarks, in full
The clearest way to read Hy3 is as a **trajectory**: what changed from Hy2 (the previous generation, Nov 2025) to Hy3 preview (Apr 2026). On the agentic suite the jump is large, and it lands Hy3 in the pack with GLM-5 and Kimi-K2.5 — below Claude Opus 4.6.
The four agent numbers, provider-reported:
- **SWE-bench Verified**: Hy2 53.0 → **Hy3 74.4**. Field: GLM-4.7 73.8, Kimi-K2.5 76.8, GLM-5 77.8, Claude Opus 4.6 80.8.
- **Terminal-Bench 2.0**: Hy2 23.2 → **Hy3 54.4**. Kimi-K2.5 50.8, GLM-5 56.2, Claude Opus 4.6 65.4.
- **BrowseComp**: Hy2 28.7 → **Hy3 67.1**. GLM-4.7 67.5, Kimi-K2.5 74.9, GLM-5 75.9, Claude Opus 4.6 84.0.
- **WideSearch**: Hy2 53.9 → **Hy3 70.2**. GLM-5 69.8, Kimi-K2.5 72.7, Claude Opus 4.6 77.2.
The two coding-agent panels tell the "competitive, not leading" story cleanly. Hy3 tracks GLM-5 and Kimi-K2.5, and trails the Claude Opus 4.6 line:
On raw reasoning and STEM the shape splits. Hy3 **matches or leads** its size class on GPQA-Diamond (87.2, vs GLM-5 86.0, Kimi-K2.5 87.6) and the Chinese-curriculum exams (China High School Biology Olympiad 87.8 — the top bar), but the frontier models pull away on the very hardest sets: on **HLE** (Humanity's Last Exam) Hy3 is **30.0** against GPT-5.4's 39.8 and Gemini-3.1-Pro's 44.4, and on the IMO Answer Bench it's 84.3 vs 89–92 for the closed pair. Note the asterisks in the figure: domestic models are scored on a text-only subset, so cross-vendor HLE/CHSBO numbers are not strictly like-for-like.
The agentic **Claw** benchmarks (tool-use) are the honest ceiling: Hy3 improves hugely over Hy2 (ClawEval 32.4 → **55.0**, WildClawBench 33.7 → **45.3**) and edges Kimi-K2.5, but Claude Opus 4.6 is clearly ahead (ClawEval 66.3, WildClawBench 60.4).
## The take
Hy3's real claim isn't a leaderboard crown — it's **efficiency at 21B active**. It roughly matches GLM-5 and Kimi-K2.5 across coding, search, and STEM while activating a fraction of their parameters, and it packages the pieces that make a MoE cheap to serve: **GQA** (8 KV heads → 8× smaller cache), an **MTP** draft head for speculative decode, a **256K** trained window, open weights, and a `reasoning_effort` knob so you pay for chain-of-thought only when you need it. That is a coherent systems story, and the base-model math scores (MATH 76.28) back the reasoning pitch.
The caveats are the usual open-weights ones, stated plainly. Every number is Tencent's own harness; independent reproduction reports treat them as upper bounds. This is a **preview** checkpoint. Hy3 trails **Claude Opus 4.6**, **GPT-5.4**, and **Gemini-3.1-Pro** on the hardest agentic and STEM tasks — sometimes by a wide margin (HLE 30.0 vs 44.4). And the eye-catching **1M context is a community YaRN extension**, not a trained window: 256K is what I'd trust, 1M is what I'd measure. The community GGUF is a genuinely useful gift for anyone with a 128 GB Mac and patience — but the card is right to call it experimental. For a team that wants an open, agent-capable model at a real inference discount, and can either run 8×GPU tensor-parallel or a fat single box, Hy3 preview earns a look. As "flagship intelligence at 2–5× smaller" — that part is Tencent's claim, and worth checking yourself.
---
*Built from the [Hy3 research page](https://hy.tencent.com/research/hy3), the [`Tencent-Hunyuan/Hy3-preview` repo](https://github.com/Tencent-Hunyuan/Hy3-preview) (295B-A21B, 256K context), and the community [`satgeze/Hy3-1M-GGUF`](https://huggingface.co/satgeze/Hy3-1M-GGUF) build (YaRN 1M, experimental). All benchmark numbers are provider-reported; the four figures are reproduced from Tencent's model card for commentary. The interactive diagrams are illustrations of the mechanism, not measured traces — the architecture walk-through, the GGUF size ladder, and the KV-cache calculator all use the published configs and reported sizes, but the quality ordering in the quant explorer is illustrative, not benchmarked. The community 1M window is a third-party artifact, not a Tencent release.*
---
# The Jacobian lens: reading the residual stream with a derivative
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/jacobian-lens
> date: 2026-07-08
> tags: explainer, interpretability, transformers, llm
The **Jacobian lens** answers one question: *given a model's internal state at some layer, what is that state disposed to make the model say?* It is a small piece of code — Anthropic's [`jacobian-lens`](https://github.com/anthropics/jacobian-lens) repo, Apache-2.0, "reference implementation, not maintained" — that fits on any open-weights decoder transformer and reads intermediate activations out as ranked vocabulary tokens. The whole method is one line:
$$
\operatorname{lens}_\ell(h) = \operatorname{unembed}\!\big(J_\ell \, h\big), \qquad J_\ell = \mathbb{E}\!\left[\frac{\partial h_{\text{final}}}{\partial h_\ell}\right]
$$
Two moves. $J_\ell$ **transports** a residual-stream vector $h$ at layer $\ell$ into the final-layer basis. Then the model's own **unembedding** decodes it into logits over the vocabulary. The transport matrix is a *Jacobian* — the derivative of the final-layer residual with respect to layer $\ell$ — averaged over a corpus. That is the entire idea, and it is worth being precise about, because the derivative is doing something the [logit lens](#a-lens-is-a-choice-of-transport) never could.
This is the companion tool for the paper [*Verbalizable Representations Form a Global Workspace in Language Models*](https://transformer-circuits.pub/2026/workspace/index.html) (Anthropic, 2026). Everything below is drawn from the repo README, the `jlens.fitting` source, and that paper. The causal numbers I quote are the paper's own, measured on Anthropic's models (Haiku/Sonnet/Opus 4.5). I have not reproduced them; I mark them as paper-reported.
## Why a derivative
A residual stream is not written in the output vocabulary. The activation $h_\ell$ at some middle layer lives in a basis the model has been rotating and rewriting layer by layer; the unembedding $W_U$ only makes sense against the *final* layer. So you cannot just apply $W_U$ to a middle activation and expect a sensible token — the coordinates don't line up. That is the flaw in the logit lens, and the Jacobian lens is the fix: it first maps $h_\ell$ **through** the layers above it, then unembeds.
Mapping through the layers exactly would mean running the rest of the network — nonlinear, and no longer a "lens." The Jacobian is the linear stand-in. It is the best linear approximation to "run layers $\ell{+}1 \dots L$" around an operating point, and the repo takes that operating point to be the *average* over a text corpus. One matrix per layer, fit once, applied everywhere.
## How the matrix is estimated
The estimator has one subtlety worth getting right. For a source position $p$ at layer $\ell$, the influence on the final layer is not a single vector — it is spread over the current position and every *future* position (a decoder is causal, so $h_\ell[p]$ can affect the output at $p, p{+}1, \dots$). The repo sums that influence over all target positions, then averages over source positions and over prompts. In the paper's own pseudocode:
```python
# Compute J_ℓ for all layers ℓ.
# h_ℓ[t] : residual stream at layer ℓ, position t
# z[t] : residual stream at the target layer L (final by default)
for each prompt p in corpus:
run forward pass; cache h_ℓ[t] for all ℓ, t
for i in 1..d_model: # one backward pass per output dim (batched)
grad_z = e_i ⊗ 1_T # inject ∂/∂z_i = 1 at every position
for each layer ℓ:
G_ℓ = ∂(Σ_t z[t]) / ∂h_ℓ # autodiff → shape [T, d_model]
J_ℓ^(p)[i, :] = mean_t G_ℓ[t, :] # mean over source positions
for each layer ℓ:
J_ℓ = mean over prompts of J_ℓ^(p) # aggregate the average Jacobian
# apply:
lens(h_ℓ) = softmax( W_U · norm( J_ℓ · h_ℓ ) )
```
Because $\sum_t z[t]$ is differentiated, a one-hot cotangent lands at every target position at once; causality makes $\partial z[p']/\partial h_\ell[p]$ vanish for $p' < p$, so what survives is exactly the sum over current-and-future targets. The paper's lenses use **1000 sequences of 128 tokens**; the README notes quality "saturates quickly" and ~100 prompts is already usable. Cost is dominated by the model's own backward pass — one per output dimension, batched — which is why fitting is embarrassingly parallel across corpus slices (`JacobianLens.merge()`).
The rows of $W_U J_\ell$ are the interesting object: the paper calls them the **J-lens vectors**, one per vocabulary token, each a direction in residual-stream space "associated with a single token."
## A Jacobian is a tangent
Here is the mechanism I want to build intuition for, because it is also the method's main limitation. A Jacobian is a *local linear map*. Collapse the stack to one scalar activation coordinate and one token's logit, and the true function is a curve; the lens is its tangent line. Slide the probe below.
In `local` mode the tangent is re-taken at the probe, so it always touches — that is the textbook Jacobian, exact at the point and good nearby. But the real lens is `corpus-average` mode: **one** tangent, taken once at the corpus operating point, then used for every activation you feed it. Near that point the linear readout is faithful; far from it the error grows. This is the honest cost of turning "run the rest of the network" into a single matrix. The lens tells you what an activation is disposed to say *to first order, on average*; it is not a faithful simulation of the model from layer $\ell$ onward.
## Then it is just a dot product
The second half, `unembed(·)`, is the easy half. Unembedding is a matrix of one row per token; a logit is that row dotted with the transported vector. Softmax ranks them, and the lens *readout* is the top of the list — the token whose direction $J_\ell h$ points most toward. Rotate the transported vector and watch the readout hand off between neighbouring concepts:
The superscript rank you see on a real slice page — `nose³`, `smile¹⁰⁴` — is exactly a token's position in this sorted vocabulary. Rank, not just top-1, is what makes the lens useful: a concept can be climbing toward the top for several layers before it ever wins.
## What it surfaces: the ASCII-face
The repo ships one example that makes the point in a single screenshot. Give the model an ASCII-art face and ask what it depicts. The `^` character is the nose. Select that position and the lens, at *middle* layers, reads out **nose** — a word that never appears in the prompt.
Reading the page: each cell is the lens top-1 word at that `(position, layer)`; the bottom row is the model's actual output; the heatmap and line charts track a pinned concept's rank across the grid. The signal for "nose" is not at the output layer — it is a hotspot in the *middle*, then it fades as the model resolves what to actually say. That is the whole pitch: the lens shows intermediate content the output distribution has already moved past.
## A lens is a choice of transport
Every "lens" is the same unembedding applied to a transported activation; they differ only in the transport $J_\ell$.
| lens | transport $J_\ell$ | how it's obtained | early-layer behaviour |
|---|---|---|---|
| **logit lens** | identity $I$ | none | assumes one basis for all layers; recovers little early content |
| **tuned lens** | learned linear map | trained to match the output distribution (correlational) | tends to "skip ahead" to the output |
| **Jacobian lens** | $\mathbb{E}[\partial h_{\text{final}}/\partial h_\ell]$ | fit by autodiff over a corpus | corrects for cross-layer basis change by construction |
The logit lens is the $J_\ell = I$ special case — it works only where the residual basis already matches the final layer, i.e. the last few layers, and the paper notes the J-lens "agrees closely" there and diverges earlier. The tuned lens also fits per-layer linear maps, but on a *correlational* objective (match the output), which the paper finds "skips ahead" and buries exactly the unverbalized intermediates you wanted to see. The Jacobian's objective is the derivative itself, which is why it recovers interpretable content at depths where the logit lens does not.
## Does the readout mean anything causally
A readout is a correlation until you intervene on it. The paper's stronger claim is that the J-lens directions are *causally* privileged, and it tests this by swapping coordinates in the space those vectors span ("J-space", panel C above). The headline: split a concept vector into its J-space component and the orthogonal remainder, then swap along each.
The striking part is the budget: that J-space component carries a **median of only 6–7%** of the concept vector's variance, with ~93% in the orthogonal remainder — yet the small component drives the swap (59%) and the remainder barely moves it (5%). A few percent of the variance is doing almost all of the *reportable* work. Intervening at intermediate layers also propagates: swapping a concept mid-stack flips the model's top-1 output on a majority of trials, scaling with model size.
The companion result is ablation: project the top J-space contents out of the residual stream and multi-hop reasoning collapses toward zero, while shallow tasks — classification, comparison, factual recall — are essentially unaffected. Read together, the lens is not just labelling activations; the directions it finds carry content the model actually uses for the harder, chained computations.
## What it can and cannot tell you
Honest boundaries, several of them the authors' own words:
- **It is a linear approximation, and an average one.** One Jacobian per layer, taken at the corpus mean, stands in for a nonlinear stack. The `LensTangent` widget is the whole caveat — faithful near the operating point, progressively wrong away from it.
- **Single tokens only.** Each J-lens vector is tied to one vocabulary token. Concepts that span multiple tokens are not directly captured (the appendices discuss extensions). The lens sees "Mars," not "the fourth planet."
- **Approximate and incomplete.** The paper states plainly that the J-lens "only approximately and incompletely captures the model's underlying workspace structure," and that a "true workspace" may operate in layers the lens misses.
- **A readout is not a mechanism.** The lens says what an activation is *disposed* to output; it does not tell you the circuit that put it there. The causal swaps above are what upgrade a readout from suggestive to load-bearing — do the intervention before you trust the picture.
- **It is a reference implementation.** Not optimized, not maintained; fitting is dominated by the model's backward pass. Fine for research, not a production probe.
## Running it
The API is two calls — fit (or download) a lens, then apply it at chosen positions:
```python
import transformers, jlens
hf = transformers.AutoModelForCausalLM.from_pretrained("org/model").cuda()
tok = transformers.AutoTokenizer.from_pretrained("org/model")
model = jlens.from_hf(hf, tok)
lens = jlens.JacobianLens.from_pretrained("org/lens-repo", filename="model/lens.pt")
lens_logits, model_logits, _ = lens.apply(
model, "Fact: The currency used in the country shaped like a boot is",
positions=[-2])
for layer, logits in sorted(lens_logits.items()):
print(layer, [tok.decode([t]) for t in logits[0].topk(5).indices])
```
Fitting your own is `jlens.fit(model, prompts=...)`; the `walkthrough.ipynb` notebook goes end to end and renders a slice page like the ASCII-face one.
## The take
The Jacobian lens is a clean idea executed narrowly. Swap the logit lens's implicit identity transport for the real averaged derivative, and you get a readout that works in the middle of the network, where the interesting, not-yet-verbalized computation lives. The mechanism is a first-order Taylor term — a tangent — and the honest framing is exactly that: a good local, on-average picture of what a layer is disposed to say, not a faithful replay of the layers above it. What earns it more than "nice visualization" is the causal follow-through: a component holding ~6–7% of a concept's variance drives most of the reportable behavior, and ablating the J-space directions specifically breaks multi-hop reasoning while leaving shallow tasks intact. That is a real, testable claim about which internal directions the model uses to talk — reached with a derivative and an unembedding, and not much else.
---
*Built on the [`jacobian-lens`](https://github.com/anthropics/jacobian-lens) reference implementation (Anthropic, Apache-2.0) and the paper [Verbalizable Representations Form a Global Workspace in Language Models](https://transformer-circuits.pub/2026/workspace/index.html). Figures reproduced from the paper (Figure 4) and the repo (`assets/slice_vis.png`) for commentary. The `LensTangent` and `UnembedReadout` widgets are my own illustrations of the mechanism, not measured traces; all quantitative results are paper-reported on Anthropic's own models and I have not independently reproduced them.*
---
# J-space in the open: a CKA map of workspace geometry across 38 models
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/jspace-open
> date: 2026-07-08
> tags: explainer, interpretability, transformers, llm
[**jspace-open**](https://eliebak.com/viz/jspace-open) is an interactive built by **Elie Bakouch**. It takes one idea from Anthropic's [*Verbalizable Representations Form a Global Workspace in Language Models*](https://transformer-circuits.pub/2026/workspace/) — the **Jacobian lens**, which reads out the concepts a layer is disposed to say — and runs it across **38 open-weights models**, **1,411 layers** in total, from GPT-2 and Pythia up to Qwen3-32B and Llama-3.3-70B. Then it asks a single geometric question, cell by cell: *do two layers arrange the vocabulary the same way?* The map that falls out is oddly regular. Every model grows the same three-block layout — a sensory front, a workspace middle, a motor tail — and unrelated families put those blocks at nearly the same **fraction of depth**.
Two honest scoping notes up front. First, this visualizes **open models**; the causal story — that these directions are *verbalizable*, that steering them changes what the model reports, that ablating them breaks multi-hop reasoning — was established on **Claude** in the paper, not re-proven here. The open map shows the *geometry* echoes across families; it does not re-run the interventions. Second, the lens weights are pre-fitted checkpoints from [`neuronpedia/jacobian-lens`](https://huggingface.co/neuronpedia/jacobian-lens) (Anthropic companion code, fit on ~1,000 WikiText prompts) — and `qwen3-32b`'s public fit is an **80-prompt checkpoint**, so read that block with more caution than the rest.
## The J-lens vector: one steering direction per token, per layer
Start with what sits in each cell. The Jacobian lens asks: at layer $\ell$, which vocabulary tokens is this activation pushing the model toward *eventually* saying? It answers with a linearized map from the layer's hidden state to the final residual stream, averaged over contexts:
$$
J_\ell \;=\; \mathbb{E}_{t}\!\left[\frac{\partial h_{\text{final},\,t'}}{\partial h_{\ell,\,t}}\right]
$$
Push a hidden state $h_\ell$ through it and read it out on the vocabulary:
$$
\operatorname{lens}(h_\ell) \;=\; \operatorname{softmax}\!\big(W_U\,\operatorname{norm}(J_\ell\, h_\ell)\big)
$$
where $W_U$ is the unembedding and $\operatorname{norm}$ is the final normalization. Turn that around and every token gets a **direction**. For token $t$, its J-lens vector is the row of $(W_U\,\gamma)\,J_\ell$ — the direction in activation space that, added to $h_\ell$, most raises the model's disposition to verbalize $t$ downstream ($\gamma$ is the final-norm gain). It is a steering direction, indexed by *(token, layer)*. The explorer probes each layer with the same **4,096 token strings** shared by all 38 tokenizers, so the stack of directions at layer $\ell$ is a matrix
$$
V_\ell \;=\; (W_U[\text{ids}]\,\gamma)\,J_\ell \;\in\; \mathbb{R}^{4096\times d}.
$$
One row per probe token, $d$ the model width. Different models have different $d$ and different layer counts — which is exactly the problem CKA is built to sidestep.
## What each cell measures: linear CKA on the token geometry
You cannot compare $V_i$ and $V_j$ coordinate-by-coordinate — different layers (and certainly different models) live in different bases and scales. So the explorer never compares coordinates. It compares the **pairwise geometry** of the 4,096 tokens. Center each layer's directions, then form its token-by-token Gram matrix:
$$
\bar V_\ell = V_\ell - \operatorname{mean\ row}, \qquad K_\ell = \bar V_\ell\,\bar V_\ell^{\top} \in \mathbb{R}^{4096\times 4096}.
$$
$K_\ell$ tabulates *which tokens' steering directions align with which* at layer $\ell$ — a pure relational fingerprint, independent of the basis. The cell value is the cosine between two such fingerprints — **linear Centered Kernel Alignment**:
$$
\operatorname{CKA}(i,j) \;=\; \frac{\langle K_i, K_j\rangle_F}{\lVert K_i\rVert_F\,\lVert K_j\rVert_F}, \qquad 1 = \text{identical geometry},\; 0 = \text{unrelated}.
$$
Because it only ever touches Gram matrices, CKA is invariant to rotation and isotropic scaling of each activation space. That invariance is the whole reason a 12-layer GPT-2 and a 64-layer Qwen — different widths, different depths, different training data — can share one axis at all.
## Reading the big matrix
The headline view stacks all 38 models into one grid. Each block is a model against itself or against another; each cell is the CKA above.
Three things stand out. On each model's own diagonal, bright squares mark **stretches of layers that hold one geometry** — the paper's sensory / workspace / motor regions. In the cross blocks, a bright **45° band** means two models line up at matched depth. And the red outlines separate within-family from across-family, so you can see that the band survives even when you leave a family — Llama next to OLMo, Gemma next to Qwen.
The interactive below rebuilds a single model's diagonal block so you can see the structure directly. Scrub the depth marker; switch the model. The values are a deterministic reconstruction of the pattern, not the measured matrix — but the geometry it encodes is the point:
The layer count jumps from 32 to 64 as you switch models, yet the two block boundaries barely move in *relative* terms. That is the first surprise: the layout is a function of fractional depth, not layer index.
## The reindex trick: same layout at the same relative depth
If the structure lives at relative depth, then to compare two models you have to put them on a common depth axis. The explorer's **reindexed** mode does exactly that — it resamples every block onto a shared 0–100% grid (bilinear), so *matched relative depth becomes the 45° diagonal of every block*. Raw mode keeps true layer counts; reindexed mode makes the alignment legible.
The widget below is the intuition without the heatmap. Six models, wildly different depths, each split into the three stages at the relative boundaries the explorer reports. Flip between raw and reindexed:
Raw, the boundaries scatter — a 12-layer model finishes its sensory phase in a handful of layers, a 64-layer model takes dozens. Reindexed, they snap onto the same two guides. That shared relative layout is precisely what shows up in the cross blocks as a diagonal band.
The explorer also summarizes each pair with a single number: the mean CKA along its **matched-depth diagonal**, $j(i) = \operatorname{round}\!\big(i\,\tfrac{L_B-1}{L_A-1}\big)$, so "does layer 30% of A do the job of layer 30% of B?" collapses to one scalar per pair. Laid out as a model-by-model matrix, it is the same story at lower resolution:
## How universal is it, really?
"Weirdly universal" is Elie's phrase, and the map earns it — but the honest version needs the numbers, because a bright block can hide a modest effect. For the full 38-model selection the explorer reports these cross-model means, averaged over all $\binom{38}{2} = 703$ pairs:
| stat | value | what it is |
|---|---|---|
| off-diagonal block CKA | **0.548** | $\operatorname{BLK}=\tfrac{1}{L_AL_B}\sum_{i,j}C(a_i,b_j)$ — the depth-independent floor two models share |
| matched-depth CKA | **0.588** | $\operatorname{MD}=\tfrac{1}{L_A}\sum_i C(a_i,b_{j(i)})$ — the 45° diagonal only |
| depth-alignment gain | **+0.040** | $\operatorname{MD}-\operatorname{BLK}$ — similarity that is *specifically* at the right depth |
| block separation | **+0.209** | a model resembles itself more than it resembles others |
| depth order $\rho$ | **0.83** | rank correlation of each layer's best-match depth (1.0 = order perfectly preserved) |
Two of those numbers deserve a hard look. Most of the cross-model similarity is the **lexical backbone**: a floor of $0.548$ that every model shares simply because every model puts "dog" near "dogs." The part that is *specifically* about matched depth — the diagonal band over that floor — is only **+0.040**.
So the strong claim — "layer 30% of Llama and layer 30% of OLMo compute the *same thing*" — is not what the number supports. What the map actually shows is subtler and, I think, more interesting: the **ordering** is shared. Depth-order $\rho = 0.83$ says that as you walk down one model, the layer in another model that best matches you almost always walks down in step. Block separation $+0.209$ says the three-stage structure is a real, self-similar object, not an artifact of the lexical floor. The vocabulary geometry reorganizes in the same sequence, at the same relative pace, across families that never saw each other's data. That is the finding — a shared *itinerary*, more than a shared computation.
One more caveat the map makes visible: the exact widths do not match the paper. On Claude the workspace runs roughly layers 38–92% of depth; the explorer's block-finder puts the open-model sensory end at ~**46.5%** and motor start at ~**64.1%** — a much narrower workspace. The *three-block shape and its order* replicate across open models; the specific fractions do not transfer from the Claude measurement. Universality of structure, not of numbers.
## Where the pattern bends
The interesting parts of a "universal" map are the exceptions, and Elie flags a few. **Base vs instruct** checkpoints (Gemma-4 is the clearest) diverge most in the **early, sensory** layers — instruction tuning rewrites low-level parsing more than it touches the workspace middle, which stays put. **Qwen3-32B** reads as architecturally odd, looking like it skips or compresses an early phase — though that block is also the 80-prompt fit, so I would not over-read it. And on tokenizers: the shared 4,096 probe strings could in principle bias the comparison, but Elie checked with random token sampling and the pattern held, which is the right control to run.
## The take
`jspace-open` is a good piece of interpretability tooling: it takes an Anthropic method that only Anthropic could run on Claude, points it at weights anyone can download, and lets you check the geometry yourself instead of taking it on faith. The honest read is a shared *structure* — three blocks, in order, at matched relative depth, with $\rho = 0.83$ and clean block separation — rather than a shared *function*, since the matched-depth lift over the lexical floor is only +0.040 and the block widths drift from the paper's Claude numbers. That is still a real result: open models from unrelated families grow the same coarse workspace itinerary. What the map does *not* do is re-establish that these directions are causally verbalizable — that remains a claim about Claude, and the mechanism behind the lens is its own story, covered in [the Jacobian-lens explainer](/articles/jacobian-lens). Here, the contribution is the map: a way to *see* that the structure travels.
---
*Primary source: [jspace-open](https://eliebak.com/viz/jspace-open) (Elie Bakouch, 2026), the J-lens CKA explorer. Built on Anthropic's [*Verbalizable Representations Form a Global Workspace*](https://transformer-circuits.pub/2026/workspace/) and the [`neuronpedia/jacobian-lens`](https://huggingface.co/neuronpedia/jacobian-lens) lens weights. The two figures are screenshots of the explorer, reproduced for commentary; all stats are read directly from the tool's 38-model summary. The interactive diagrams are deterministic reconstructions of the pattern, not the measured CKA matrix.*
---
# Muon: orthogonalizing the update for hidden layers
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/muon-optimizer
> date: 2026-07-08
> tags: explainer, training, deep-learning, llm
**Muon** is an optimizer for the **hidden layers** of a neural network. The idea is small and specific: take the momentum update you would feed to SGD, and — before applying it to a 2D weight matrix — **orthogonalize** it. Replace the raw update direction with its nearest orthogonal matrix, so every singular direction of the step carries the same weight and no single direction dominates. The name spells out the recipe: **M**oment**U**m **O**rthogonalized by **N**ewton-schulz. Everything else — the embeddings, the final head, the norms and biases — stays on AdamW.
That one change buys two things that are worth taking seriously. Keller Jordan, who introduced Muon, used it to set a string of NanoGPT training-speed records at essentially the same wall-clock cost per step as Adam. And Moonshot AI's *Muon is Scalable for LLM Training* reports **~2× the compute efficiency of AdamW** at LLM scale — once you add two fixes that only matter once the matrices get big. This is a walk through the mechanism from first principles, then the numbers, honestly labelled.
Muon only touches **2D hidden weight matrices** — attention and MLP projections. Scalars, vectors, embeddings and the output head are optimized by AdamW, and empirically that split matters (more on why below). Every speed number here is **author- or provider-reported**: Keller's on NanoGPT/CIFAR speedruns, Moonshot's on their own scaling-law sweep and the Moonlight model. I have not re-run them.
## The idea: orthogonalize the update
Start with the problem. A momentum buffer for a weight matrix $M \in \mathbb{R}^{A \times B}$ has a spectrum — a set of singular values. Write its SVD:
$$
M = U \Sigma V^{\top}, \qquad \Sigma = \mathrm{diag}(\sigma_1, \sigma_2, \dots, \sigma_r), \quad r = \min(A, B).
$$
Here $U$ and $V$ hold the left/right **singular directions** and $\sigma_i$ are the **singular values** — how far the update reaches along each direction. Gradient updates on real transformers are badly conditioned: a few singular values are huge and the rest are tiny. So a plain momentum step lurches the weights along the top singular direction and barely moves the others. The step is *anisotropic*.
Muon's fix is to keep the directions and flatten the spectrum. Set every singular value to 1:
$$
O = U V^{\top} \quad\text{(the } \Sigma \text{ in the middle is replaced by the identity).}
$$
$O$ is the closest **semi-orthogonal** matrix to $M$. It points the same way as $M$ in every singular direction, but now each direction gets the same magnitude. This is the "spectrally normalized" step — no direction dominates. Toggle between the raw momentum step and its orthogonalized version:
That is the whole conceptual move. The rest of Muon is about doing $M \mapsto U V^{\top}$ **cheaply** — an SVD every step, on every weight matrix, would be far too slow.
## The update rule
Per step, for each 2D hidden weight $W$:
```python
# Muon, one step, per 2D hidden weight matrix W
M = mu * M + G # momentum buffer; G = this step's gradient, mu ~ 0.95
U = G + mu * M # Nesterov variant (works a bit better in practice)
O = newton_schulz5(U, steps=5) # orthogonalize: O ~ U V^T of the update
W = W - lr * O # apply the spectrally-normalized step
```
$G$ is the gradient, $M$ the momentum buffer, $\mu$ the momentum coefficient (~0.95), $\text{lr}$ the learning rate. Muon puts the momentum **before** the orthogonalization — you orthogonalize the accumulated direction, not the raw gradient — and Keller reports Nesterov-style momentum beats plain SGD-momentum in every case he tested. The only unusual line is `newton_schulz5`.
## Newton-Schulz: orthogonalize without an SVD
The trick is that $U V^{\top}$ is what you get if you take the SVD and set every singular value to 1. So you never need $U$, $V$, or $\Sigma$ explicitly — you just need a function that drives every singular value to 1 while leaving the singular *directions* alone. A matrix polynomial in $M$ does exactly that: applying $p(M) = M \, q(M^{\top}M)$ acts on the SVD as $U \, p(\Sigma) \, V^{\top}$, so it touches only the singular values. Pick the polynomial so that iterating it sends every $\sigma \in [0, 1]$ toward 1.
Muon uses a fixed **quintic**, applied $T = 5$ times. Here is the exact function, coefficients and all:
```python
def newtonschulz5(G, steps=5, eps=1e-7):
assert G.ndim == 2
a, b, c = (3.4445, -4.7750, 2.0315) # the tuned quintic coefficients
X = G.bfloat16()
X /= (X.norm() + eps) # normalize so every singular value is in [0, 1]
if G.size(0) > G.size(1):
X = X.T
for _ in range(steps): # FIXED count -- no "until converged"
A = X @ X.T
B = b * A + c * A @ A
X = a * X + B @ X # X <- a*X + b (XX^T)X + c (XX^T)^2 X
if G.size(0) > G.size(1):
X = X.T
return X
```
On the singular values this is the scalar map $\varphi(\sigma) = a\sigma + b\sigma^3 + c\sigma^5$ applied five times. Two things make it work. First, dividing by the Frobenius norm up front guarantees $\sigma_{\max} \le 1$, so every value lands in $[0, 1]$ where the iteration is designed to converge. Second, the coefficients are tuned so the slope at zero is steep — $\varphi'(0) = a = 3.4445 > 1$ — which yanks the *smallest* singular values up fast. Watch the spectrum flatten:
The honest nuance: those coefficients do **not** drive each singular value to exactly 1. $\varphi(1) = 3.4445 - 4.7750 + 2.0315 = 0.70$, and the iteration settles values into a band roughly $[0.7, 1.2]$ rather than a point. That is deliberate — Keller trades exact convergence for the steep slope at 0, which makes five steps enough. It does not matter: the goal is to *equalize* the singular values so no direction dominates, and the condition number $\kappa = \sigma_{\max}/\sigma_{\min}$ collapsing from ~33 to ~1.5 does that. Approximate orthogonalization is all Muon needs. (Moonshot report $T = 10$ gives a cleaner orthogonalization but no downstream gain, so $T = 5$ it is.)
The reason this is cheap: the iteration is just matrix multiplies in `bfloat16`, no inverses or eigendecompositions. Keller bounds the overhead at $T m / B$ FLOPs relative to the forward/backward pass, where $m$ is the model dimension and $B$ the batch size in tokens. Concretely: NanoGPT ($m=768$, $B{=}524{,}288$) pays $5 \times 768 / 524288 = 0.7\%$; a Llama-405B-shaped run ($m{=}16384$, $B{=}16\text{M}$) pays $5 \times 16384 / 16\text{M} = 0.5\%$. Orthogonalization is almost free, and it gets *cheaper* as models grow.
Muon is close kin to Shampoo/SOAP. Strip the preconditioner accumulation out of Shampoo and its update collapses to the same orthogonalized gradient $U V^{\top}$ — but Shampoo computes it with inverse-fourth-roots (an eigendecomposition), where Muon uses the Newton-Schulz iteration. Same target, far lower wall-clock and FLOP overhead.
## Why only 2D hidden layers
Newton-Schulz needs a matrix — orthogonalization is defined for 2D. So scalars and vectors (LayerNorm gains, biases) have no spectrum to flatten and go to AdamW by construction. Convolutional filters get flattened to 2D and can be included.
The less obvious rule is that the **embedding** and the **final classifier head** are 2D but should *still* use AdamW — empirically that split improves results. The intuition: those two layers are indexed per-token. Each row is one token's vector, updated only when that token appears, so the gradient is sparse and row-wise, and there is no shared "direction" across the vocabulary worth equalizing. Orthogonalizing across the vocab dimension mixes unrelated tokens. The hidden layers are the opposite — dense, shared, badly conditioned — which is exactly where flattening the spectrum pays off. So the working split is: **hidden 2D weights on Muon, everything else on AdamW.**
## The speedruns
Keller's headline result is the NanoGPT speedrun. Swapping AdamW for Muon set a new training-speed record on 2024-10-15, improving speed by **35%**, and Muon has held as the optimizer of choice through the twelve NanoGPT records set since, by seven different researchers. On the training-to-target curve it dominates: at roughly Adam's cost per step it reaches a lower validation loss than Adam, DistributedShampoo, and SOAP — in less wall-clock time.
The per-step overhead is the point. From the legend: Adam runs at **139 ms/step**, Muon at **142 ms/step** — a ~2% tax — while matching-or-beating DistributedShampoo (154–179 ms/step) and SOAP (301 ms/step) on loss. Orthogonalization is nearly free per step and the sample efficiency is better, so wall-clock wins:
The wins hold beyond NanoGPT, in Keller's own runs: a **1.5B-parameter** transformer to GPT-2-XL-level HellaSwag in **10 × 8×H100-hours** where AdamW needs **13.3** (a 1.33× wall-clock speedup); the FineWeb validation-loss record improved by **1.35×**; and the CIFAR-10-to-94% speed record cut from **3.3 to 2.6 A100-seconds**. Small models, but a consistent shape.
## Muon is Scalable — the two fixes
Muon out of the box works at NanoGPT scale. Moonshot AI's *Muon is Scalable for LLM Training* (Liu et al., 2025) is about what breaks when you push it to real LLM training, and the two fixes that close the gap. Both are one-liners once you see them.
**Fix 1 — weight decay.** Base Muon has no weight decay, and over a long run the weights (and with them the logits and activation RMS) drift upward until quality suffers. The fix is decoupled weight decay, exactly as in AdamW:
$$
W_t = W_{t-1} - \eta_t\,\big(O_t + \lambda\, W_{t-1}\big), \qquad \lambda = 0.1.
$$
$O_t$ is the orthogonalized update, $\eta_t$ the learning rate, $\lambda$ the weight-decay coefficient. Without it Muon eventually crosses *above* AdamW late in training; with it Muon stays ahead throughout.
**Fix 2 — match the update RMS.** This one is about magnitude. AdamW's per-element update has a roughly constant RMS (~0.2–0.4) regardless of a matrix's shape, so a single learning rate works everywhere. Muon's orthogonalized update does not. Moonshot's Lemma 1: for a full-rank $[A, B]$ matrix, the RMS of $O = U V^{\top}$ is
$$
\mathrm{RMS}(O) = \frac{1}{\sqrt{\max(A, B)}}.
$$
That **shrinks as matrices get wider** — a $4096 \times 11008$ MLP weight has RMS ≈ 0.01, so its effective step is ~20× smaller than a square layer's under the same learning rate. At scale the wide matrices barely move. The fix scales each Muon update to a fixed target RMS of ~0.2:
$$
W_t = W_{t-1} - \eta_t\,\Big(0.2 \cdot O_t \cdot \sqrt{\max(A, B)} + \lambda\, W_{t-1}\Big).
$$
The $\sqrt{\max(A,B)}$ cancels the shape dependence and the constant 0.2 pins the RMS to AdamW's range, so an AdamW-tuned learning rate transfers to Muon directly — no per-layer retuning.
With both fixes in, Moonshot's scaling-law sweep (dense models, 0.4B–1.5B, compute-optimal token counts) puts Muon's loss curve cleanly below AdamW's:
The fitted curves are $L_{\text{Muon}} = 2.506\,C^{-0.052}$ and $L_{\text{AdamW}} = 2.608\,C^{-0.054}$, with $C$ the compute budget. Read horizontally, matching AdamW's loss takes Muon about **52% of the training FLOPs** — the "~2× more compute-efficient" headline. Worth stating plainly: this is a fit over their own sub-2B sweep, extrapolated; it is a provider result, not an independent one.
## Moonlight
To show it holds past the toy scale, Moonshot trained **Moonlight** with Muon: a **15.3B-parameter** Mixture-of-Experts model (a DeepSeek-V3-Small-style architecture) that activates **2.24B** parameters per token, on **5.7T** tokens. The training was smooth — no loss or gradient-norm spikes. Placed on a compute-vs-MMLU frontier against open models, Moonlight sits *on* the Pareto front, matching models trained with far more compute:
Against comparable open models, Moonlight leads most of the standard suite — with one honest exception:
On **MATH** Moonlight reaches **45.3** vs Qwen2.5-3B's 42.6, Llama3.2-3B's 8.5 and DeepSeek-V2-Lite's 17.1; on **GSM8K** it is **77.4**, ahead of Llama and DeepSeek-V2-Lite but a hair *behind* Qwen2.5-3B's 79.1. And in Moonshot's own head-to-head — Moonlight (Muon) vs an identical model trained with AdamW at 1.2T tokens — Muon wins across the board, most visibly on code: HumanEval 37.2 vs 29.3, MBPP 52.9 vs 49.2. These are all provider numbers on their own harness, and Moonlight is an MoE while the scaling-law fits were on dense models — so the ~2× claim and the model result are related but not the same experiment.
## The take
Muon is a genuinely small idea with a clean mechanism: orthogonalize the momentum update for 2D hidden weights so the step is spectrally flat, and do it with a five-step Newton-Schulz iteration that costs well under 1% overhead. The Newton-Schulz coefficients are tuned for speed, not exactness — they collapse the update's condition number toward 1 rather than nailing every singular value to it, and that is enough. The scope is deliberately narrow: hidden 2D matrices only, everything else on AdamW.
The wins are real but each carries a caveat. Keller's NanoGPT/CIFAR speedruns are small-scale and self-reported, but they are reproducible and the per-step overhead (142 vs 139 ms) is visible and tiny. Moonshot's ~2× compute-efficiency is a fit over their own sub-2B dense sweep, extrapolated, and it *only* holds with the two fixes — weight decay and update-RMS matching — that Muon out of the box lacks. Moonlight is an MoE and a provider-reported result. Read together, the honest summary is: orthogonalizing the update is a cheap, well-motivated change that clearly helps hidden layers, is nearly free per step, and — with the scale fixes — looks like a real efficiency gain that others can now check for themselves.
---
*Built on Keller Jordan's [Muon: An optimizer for hidden layers in neural networks](https://kellerjordan.github.io/posts/muon) (2024) and J. Liu et al., [Muon is Scalable for LLM Training](https://arxiv.org/abs/2502.16982) (arXiv 2502.16982, 2025). Newton-Schulz code and coefficients are quoted from Keller's post; the update-RMS and weight-decay formulas from Liu et al. The interactive widgets are illustrations of the mechanism, not measured traces. Figures are reproduced from the sources for commentary; all speed and benchmark numbers are author- or provider-reported.*
---
# Rollout Routing Replay: stabilizing MoE reinforcement learning
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/rollout-routing-replay
> date: 2026-07-08
> tags: explainer, mixture-of-experts, reinforcement-learning, llm, training
Reinforcement learning is now the standard last stage of training a reasoning model: sample answers, score them, push up the probability of the good ones. It works well on dense models. On **Mixture-of-Experts** models it has a habit of blowing up — the validation curve climbs for 100 steps, then falls off a cliff. **Rollout Routing Replay (R3)**, from a Peking University and Xiaomi team, names the cause and fixes it with one idea: the router chooses different experts when you *generate* a rollout than when you *score* it for the update, so make training reuse the routing decisions the rollout already made.
All numbers here are the paper's own experiments, on one model family: **Qwen3-30B-A3B** (30B total, ~3B active) for MoE and **Qwen3-8B** for the dense baseline, trained with `VeRL` — **SGLang** for rollout, **Megatron** for the update. The tasks are math RLVR (AIME/AMC/MATH) and one multi-turn SWE-agent task. The interactive diagrams below are illustrations of the mechanism with deterministic toy numbers; the benchmark and KL values are measured and cited to the figure/table they come from.
## Why RL training runs two different models of the same policy
Modern RL frameworks split the work across two engines. An **inference engine** (SGLang, vLLM) generates rollouts fast, with its own fused kernels and quantization. A **training engine** (Megatron, FSDP) recomputes probabilities and applies gradients. They implement the same math but not the same arithmetic, so the probability a rollout token gets from each engine is slightly different.
That difference matters because PPO-style objectives are off-policy. You sample from the inference policy $\pi_\text{infer}$ but compute the loss with the training policy $\pi_\text{train}$:
$$
J(\theta) = \mathbb{E}_{x,\,y\sim\pi_\text{infer}(\theta_\text{old})}\left[\frac{1}{|y|}\sum_{t=1}^{|y|}\min\!\big(w_t(\theta)\,\hat{A}_t,\ \text{clip}(w_t(\theta),\,1-\varepsilon,\,1+\varepsilon)\,\hat{A}_t\big)\right]
$$
The whole update is scaled by the **importance-sampling ratio**
$$
w_t(\theta) = \frac{\pi_\text{train}(\theta)(y_t\mid x,y_{
The extreme tail is what actually breaks training. The paper measures it with $F(\tau)$, the fraction of tokens whose train/infer probability ratio exceeds $\tau$. For $\tau>2$ the MoE model has an order of magnitude more such tokens than the dense model — and R3 removes that excess:
## R3: replay the rollout routing mask
The fix is almost anticlimactic once the diagnosis is clear. During rollout, record the router's top-$K$ selection mask $I_\text{infer}$ for every token and every layer. During the training forward pass, **use that mask instead of recomputing one** — but still run the softmax over the *training* logits, so the router's weights keep receiving gradient.
Start from a normal MoE layer on the training side. The router scores experts and keeps the top $K$ as a binary mask:
$$
s_\text{train} = x_\text{train} W_r, \qquad I_\text{train} = \text{TopKMask}(s_\text{train}, K), \quad I_\text{train}\in\{0,1\}^M
$$
Gating weights are a softmax over the *selected* experts' logits, and the output is their weighted sum:
$$
g_{\text{train},i} = \frac{I_{\text{train},i}\,\exp(s_{\text{train},i})}{\sum_j I_{\text{train},j}\,\exp(s_{\text{train},j})}, \qquad y_\text{train} = \sum_{i=1}^{M} g_{\text{train},i}\,E_i(x_\text{train})
$$
R3 changes exactly one term: replace the training mask with the **inference** mask captured during rollout, $I_\text{infer} = \text{TopKMask}(s_\text{infer}, K)$, while keeping the softmax on the training logits:
$$
g_{\text{replay},i} = \frac{I_{\text{infer},i}\,\exp(s_{\text{train},i})}{\sum_j I_{\text{infer},j}\,\exp(s_{\text{train},j})}, \qquad y_\text{replay} = \sum_{i=1}^{M} g_{\text{replay},i}\,E_i(x_\text{train})
$$
Two properties fall out of that single substitution. **Alignment:** the training pass now activates exactly the experts the rollout used, so the layer output matches and $w_t$ returns to ~1. **Gradient survives:** only the discrete mask $I_\text{infer}$ is borrowed; the softmax still runs over $s_\text{train}$, so $\partial/\partial W_r$ keeps flowing and the router keeps training. You borrow the *decision*, not the *weights*.
Below is the same mechanism at one layer. The rollout router picks its top-2; the training router, on a different engine, recomputes logits and can land on a different top-2 — and when it does, the importance ratio blows up. Toggle **R3 on** to replay the rollout mask and watch every token snap back to $w \approx 1$:
The paper's own schematic makes the data flow explicit: the rollout selection `select (1, 4)` is captured once and fed into both later forward passes (the recompute of the old policy and the update of the new one), overriding whatever the training routers would have chosen on their own:
If the router and top-$K$ gate are unfamiliar, I built them up from one MLP in [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch) — R3 is a small surgery on exactly that dispatch step.
### What it costs, and the caching trick that makes it free at rollout
Storing a mask per token per layer sounds expensive. It is not: a top-$K$ mask is a handful of small integers, and the paper reports **under 3% latency overhead** during rollout. The neat part is multi-turn. Inference engines already cache the KV of a prefix so repeated turns don't re-prefill; R3 caches the **routing masks alongside that KV**. Same prefix, same masks, no recomputation. That is what keeps R3 cheap on agent tasks — [software-engineering and browsing agents](/articles/agents-a1) that interleave many generation and tool-call turns — where re-prefilling to regenerate masks would otherwise dominate.
## Does it work
Three things to check: does it kill the extreme tokens, does it stop the collapse, does it score better.
**Alignment.** Replaying the masks drops the train–inference KL from **1.54×10⁻³ to 0.75×10⁻³**, essentially the dense model's 0.64×10⁻³, and cuts the large-ratio tail by an order of magnitude — the green curve in Figure 2 above.
**Stability.** In the single-mini-step setting, all three runs *without* R3 collapsed. The tell was mechanical: KL and $F(\tau{=}2)$ climbed together, and once $F(\tau{=}2)$ passed 0.1 — 10% of tokens differing by more than 2× between engines — the run fell over (SFT + GRPO collapsed at step 60). With R3, $F(\tau{=}2)$ stayed below 10⁻⁴ for most of training and nothing collapsed. The clearest picture is the validation curve: without R3 it climbs to ~0.62 and then craters; with R3 it climbs smoothly past 0.70.
**Performance.** On the single-mini-step SFT model, R3 beats TIS by **5.58 points** of average math score — and unlike GRPO and GRPO+TIS, it never crashes:
In the multi-mini-step setting the story repeats against GSPO: GRPO+R3 edges GSPO by 1.29, and stacking R3 on GSPO adds another 0.95 — while plain GRPO collapsed at step 120:
And it generalizes past math. On a multi-turn SWE-agent task (R2E-Gym train, SWE-bench Verified eval), GRPO collapses at step 90; GRPO+R3 stays stable and finishes **6.8 points higher** on Pass@1:
## How it differs from GSPO's routing replay
GSPO already proposed a "routing replay," so it is worth being precise about what R3 changes. An RL step has three forward passes: **rollout** (generate), **recompute** (score the old policy), **update** (score the new policy). GSPO's *Recompute* Routing Replay caches the mask from the recompute pass and replays it in the update pass — it fixes routing drift *caused by the weight update*, but does nothing about the rollout-vs-training **framework gap**. R3 caches from the **rollout** pass and replays it in both recompute and update, so it fixes the framework gap *and*, because both training passes now share one mask, the update drift too.
The distinction bites at `mini_step=1`, where the old and new policies are identical and GSPO's recompute-based replay has nothing to correct — yet the framework gap is still there, and only R3 closes it. One honest caveat from the same experiments: R3 already removes most of the discrepancy, so **stacking TIS on top does not help and can hurt** — TIS+R3 scored 1.69 below R3 alone on the single-mini-step SFT model. If you run R3, drop the importance-sampling patch.
## The take
R3 is the kind of fix that reads as obvious only after someone isolates the cause. The instability everyone attributed vaguely to "MoE being finicky" turns out to be a specific, measurable thing — the router selecting different experts in the two engines — and the remedy is to stop letting the training pass re-decide something the rollout already decided. It aligns the two policies at the source rather than clipping the symptom downstream, it costs under 3% at rollout, it caches cleanly for multi-turn agents, and it is orthogonal to GRPO/GSPO/DAPO so you can bolt it on.
The caveats are the usual single-paper ones: everything is one model family (Qwen3-30B-A3B / Qwen3-8B) on math plus one SWE task, and it needs you to reach into the inference engine to capture and store routing masks — free in principle, real integration work in practice. But the mechanism is clean, the diagnosis is well-measured, and the result — MoE RL that trains as stably as a dense model — is worth the plumbing.
---
*Built on [Rollout Routing Replay](https://arxiv.org/abs/2510.11370) (Ma et al., 2025; arXiv:2510.11370). Figures are reproduced from the paper for commentary. The interactive diagrams use deterministic toy numbers to illustrate the mechanism; all KL, $F(\tau)$, and benchmark figures are the paper's measured values, cited to their source figure or table.*
---
# Reading a torch.profiler trace: overhead-bound vs compute-bound
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/torch-profiler
> date: 2026-07-08
> tags: explainer, systems, inference-optimization, training
Every "my GPU is slow" bug is one of two things: the GPU is doing too much work, or it is doing nothing while the CPU flails. You cannot tell which by staring at the code. You attach a profiler and read the trace. Hugging Face's [torch.profiler guide](https://huggingface.co/blog/torch-profiler) teaches this with the smallest possible workload — `y = matmul(x, w) + b`, bf16, on an **NVIDIA A100-SXM4-80GB** — and shows the same three lines of code land on opposite ends of that spectrum depending only on the matrix size. This is a walk through what the profiler prints, how to read it, and the two bottleneck regimes it exposes.
All numbers here are from the post's runs on one A100 in bf16. Kernel timings drift a few percent run to run (GPU clocks, thermals, power caps), so treat them as representative, not exact constants. The interactive diagrams are redrawn from the post's traces to explain the mechanism — they are not live captures.
## The 20-line setup
The workload is deliberately trivial so the profiler output is the whole story, not the model:
```python
# 01_matmul_add.py — profile y = matmul(x, w) + b on an A100, bf16
import torch
N = 64 # start small; bump to 4096 later
x = torch.randn(N, N, dtype=torch.bfloat16, device="cuda")
w = torch.randn(N, N, dtype=torch.bfloat16, device="cuda")
b = torch.randn(N, N, dtype=torch.bfloat16, device="cuda")
def fn(x, w, b):
return torch.add(torch.matmul(x, w), b)
def step():
with torch.profiler.record_function("matmul_add"): # a named region in the trace
fn(x, w, b)
```
`record_function("matmul_add")` is the one line people skip and then regret: it draws a labelled box around your code in the trace so you can find it among thousands of `aten::*` ops. The harness wraps `step()` in the profiler:
```python
schedule = torch.profiler.schedule(wait=1, warmup=1, active=3, repeat=1)
with torch.profiler.profile(
activities=[
torch.profiler.ProfilerActivity.CPU,
torch.profiler.ProfilerActivity.CUDA,
],
schedule=schedule,
) as prof:
for _ in range(5): # 1 wait + 1 warmup + 3 active = 5 steps
step()
prof.step() # advance the schedule one step
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=15))
prof.export_chrome_trace("trace.json") # open in https://ui.perfetto.dev
```
Two knobs matter here. `activities` records both CPU-side dispatch and CUDA kernels — you want both, because the whole point is comparing them. `schedule` decides which steps count.
## schedule: which steps land in the trace
`prof.step()` advances a little state machine. `wait` steps are skipped entirely, `warmup` steps run but get discarded, `active` steps run and get recorded. The warmup exists to throw away the first-step cold start — on step zero the CPU sits idle for a couple hundred microseconds before it issues its first launch, and you do not want that artifact averaged into your numbers.
With `wait=1, warmup=1, active=3` over `range(5)`, exactly **3 steps** are recorded — which is why every row in the table below reads `# of Calls = 3`, and why the first recorded box in the trace is labelled `ProfilerStep#2`.
## Reading the table: Self CPU vs Self CUDA
`key_averages().table()` is the first thing to read, before any timeline. Here is the 64x64 run, sorted by CUDA time:
The two columns that decide everything are **Self CPU** and **Self CUDA**. Self time is a row's own time, excluding its children — so `aten::matmul` shows big CPU total but ~0 self time (its work lives in the child `aten::mm` and the kernel it launches). Read the self columns and the picture is stark:
- **Self CUDA time total: 23.104 us.** The GPU does 23 microseconds of real work. That splits into two kernels — the GEMM (`ampere_bf16_s16816gemm_bf16_64x64_...`, 14.272 us, ~62%) and a `vectorized_elementwise_kernel` for the add (8.832 us, ~38%).
- **Self CPU time total: 2.314 ms** — a hundred times larger. And 1.786 ms of it (77.20%) is a single row: `cudaDeviceSynchronize`, the CPU blocking to wait for the GPU.
When Self CPU dwarfs Self CUDA like this, the GPU is starved. The kernel isn't slow; there is barely any kernel. This is **overhead-bound**.
## Reading the trace: two lanes, one dependency
The table tells you *what* is expensive; the timeline tells you *when* and *why there are gaps*. Export the trace, open it in [Perfetto](https://ui.perfetto.dev), and you get two lanes that matter — the CPU main thread and the GPU stream:
The mental model: **the CPU launches, the GPU executes, and the two lanes are offset in time.** The CPU calls `cudaLaunchKernel`, which returns almost immediately — the kernel is queued, not run. The GPU picks it up a moment later on its own stream. So a fast CPU op and a slow GPU kernel show up as *staggered* boxes, not stacked ones. The dashed dependency in the diagram below is that hand-off.
In the 64x64 trace the GPU lane is nearly empty: 23 us of kernels scattered in a wall that is milliseconds wide. The GPU is idle roughly **98%** of the time. Flip the interactive to see the same lanes fill up when the matrices grow:
## Same code, two regimes
Change one number — `N = 64` to `N = 4096` — and rerun. Nothing else moves. The table inverts:
| run | Self CPU total | Self CUDA total | dominant cost | verdict |
|---|---|---|---|---|
| `64 x 64` | 2.314 ms | 23.104 us | `cudaDeviceSynchronize` 1.786 ms (77%) | overhead-bound |
| `4096 x 4096` | 4.908 ms | 4.495 ms | gemm kernel 4.285 ms (95% of GPU) | compute-bound |
At 4096, Self CUDA (4.495 ms) finally rivals Self CPU (4.908 ms). One kernel — `ampere_bf16_s16816gemm_bf16_128x256_...`, 4.285 ms, 95.33% of GPU time — is now the entire budget. Note the tile even changed: cuBLAS picks a `128x256` GEMM tile for the big matrices where it used `64x64` for the small ones. `cudaDeviceSynchronize` is still 94% of the CPU wall, but the reading is different: here the CPU is legitimately blocked on real GPU work, not spinning on launch overhead. Same row, opposite meaning — which is exactly why you read both lanes.
The verdict changes what you do next:
- **Overhead-bound** (64x64): stop launching so many tiny kernels. Batch more work per launch, fuse ops, or hand it to `torch.compile`. Making the kernel faster buys you nothing — it is already 23 us.
- **Compute-bound** (4096x4096): the gemm *is* the job. Optimize the kernel — lower precision, better tiling, a fused epilogue — or reduce FLOPs. Cutting launch overhead buys you nothing here.
The one-line diagnostic: compare **Self CUDA time total** against **Self CPU time total**. GPU much smaller than CPU means overhead-bound — you are launch- and sync-limited. GPU comparable to or larger than CPU means compute-bound — go optimize kernels. Everything else is detail.
## What torch.compile actually does here
The obvious fix for the overhead-bound case is to stop dispatching `matmul` and `add` as two separate ops. `torch.compile` does that:
```python
cfn = torch.compile(fn)
def step():
with torch.profiler.record_function("matmul_add"):
cfn(x, w, b)
```
In the trace, the two ops collapse into a single `aten::addmm` dispatch — a GEMM with the bias folded into its epilogue instead of a separate elementwise kernel:
Two things are worth seeing honestly. First, there is still a **Device-to-Device `cudaMemcpyAsync`** in the region — the bias has to be staged/broadcast before it folds into the GEMM, so "fused" does not mean "zero extra work". Second, the compiled path adds its own CPU cost: a `CompiledFxGraph` call plus Dynamo's guard and cache lookup roughly **double** the per-step CPU overhead versus eager. Underneath, it is still the same `ampere` cuBLAS GEMM kernel doing the math.
So `torch.compile` is a real win when the fusion removes launches across *many* ops or feeds a big epilogue — but on a single `matmul + add` over small inputs, its fixed dispatch overhead is not amortized and can cost more CPU than it saves. The profiler is how you tell the difference instead of guessing. Measure both, keep the faster one.
## The workflow, condensed
- **Wrap regions** with `record_function("name")` so you can find your code in the trace.
- **Use a `schedule`** with at least one `warmup` step; discard the cold start.
- **Read `key_averages().table()` first.** Compare Self CUDA total vs Self CPU total to get the regime.
- **Then open the trace in Perfetto.** CPU lane launches, GPU lane executes, offset in time. Empty GPU lane = overhead-bound; a fat kernel filling the GPU lane = compute-bound.
- **Fix the regime you actually have.** Fuse/batch for overhead; optimize the kernel for compute. Re-profile to confirm the win is real and not a torch.compile tax.
None of this needs a big model. A 64x64 matmul on an A100 is enough to show the difference between a GPU that is busy and a GPU that is waiting — and that difference is most of GPU performance work.
---
*Built on Hugging Face's [Understanding the torch.profiler](https://huggingface.co/blog/torch-profiler) (2026). All timings are the post's A100 / bf16 runs, reproduced from its `key_averages()` tables and Perfetto traces for commentary; the `TraceLanes` and `ScheduleStrip` widgets are redrawn illustrations of the mechanism, not live captures.*
---
# zvec: an in-process vector database, and the ANN search inside it
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/zvec
> date: 2026-07-08
> tags: explainer, vector-search, systems, quantization, information-retrieval
**zvec** is Alibaba's open-source (**Apache 2.0**, Tongyi Lab) **in-process vector database** — it links into your app as a library instead of running as a server. The pitch is "SQLite for vectors": no cluster, no network hop, one embedded engine that does approximate nearest-neighbour (ANN) search over embeddings. The core is C++ (built on **Proxima**, Alibaba's older vector-search engine), with SDKs for **Python, Node.js, Go, Rust, and Dart/Flutter** and builds for Linux (x86_64/ARM64), macOS (ARM64), Windows, and Android.
The headline number is a throughput claim: on VectorDBBench's **Cohere 10M** set (10M vectors, 768-d) a 16-vCPU / 64-GiB instance serves **8,475 QPS**, roughly 2× the next-fastest entry on that board. That is the top bar in the cover figure. Two mechanisms do the work, and neither is unique to zvec — they are the two levers every fast ANN index pulls. This post explains both from first principles, then puts the self-reported numbers up in full.
Every number here is **self-reported** by zvec, measured with [VectorDBBench](https://github.com/zilliztech/VectorDBBench). And the comparison is not apples-to-apples: zvec runs **embedded on local hardware**, so its number has no network in it, while several of the competitors on the same chart (ZillizCloud, Pinecone, Qdrant Cloud) are **managed cloud services** whose QPS includes round-trip latency. The hardware also differs row to row (core counts, node counts, index versions are baked into each label). Read it as "fast for an in-process engine," not as a clean head-to-head. No independent third-party run has been published yet.
## Why nearest-neighbour search is hard
The naive query is brute force: score the query vector against all $N$ stored vectors, keep the top-$k$. That is $O(N)$ distance computations per query. At $N = 10^7$ and 768 dimensions, each query touches ten million dot products — correct, and far too slow to serve at thousands of QPS.
The standard fix is a **graph index**. zvec's default is **HNSW** (Hierarchical Navigable Small World): wire every vector to its nearest neighbours, then answer a query by *walking* the graph — start somewhere, repeatedly hop to whichever neighbour is closer to the query, stop at a local minimum. You compute distances only to the nodes you actually visit, and that count grows roughly *logarithmically* with $N$, not linearly. Step through one descent:
That is the first lever: **visit fewer vectors**. The graph decides *which* candidates to score. It trades a small, bounded loss in recall (you can land on an approximate neighbour, not the exact one) for skipping the overwhelming majority of the dataset. `ef-search` is the knob — a larger search frontier means more nodes visited, higher recall, lower QPS. The benchmark run uses `--ef-search 118` with `--m 50` (`m` = neighbours per node in the graph).
## Quantize to make each score cheap
The graph decides how *many* distances you compute. Quantization decides how *expensive* each one is. A 768-d `fp32` vector is 3072 bytes; scoring it is a 768-wide float dot product. Compress it and both the memory footprint and the per-distance cost drop:
- **int8** (scalar quantization) — 768 bytes/vector, a 4× shrink. Distances become int8 dot products with a tiny quantization error. This is what the headline run uses.
- **int4 / fp16** — zvec also exposes 4-bit and half-precision codes for tighter memory/accuracy tradeoffs.
- **RaBitQ** (1-bit, added in v0.3.0 via [the SIGMOD 2024 method](https://github.com/gaoj0017/RaBitQ)) — one bit per dimension, 96 bytes for 768-d, a 32× shrink. A distance collapses to a `popcount`. RaBitQ's selling point is a theoretical error bound that, its authors argue, keeps recall high *without* a re-ranking pass.
The catch: aggressive codes distort distances, so the *ranking* off the compressed vectors is wrong. zvec's answer is the **refiner** (the `--is-using-refiner` flag). Retrieve a broad shortlist using the cheap quantized distances, then **re-score just that shortlist with the original full-precision vectors** and return the exact-scored top-$k$. Coarse pass to go fast, fine pass to stay accurate. Toggle refinement off and watch a true neighbour fall out of the result:
The refiner is why the benchmark can run int8 codes and still report high recall: the int8 pass is only a *filter*, and the returned order is decided by full-precision math on a handful of survivors. RaBitQ is the more aggressive bet — it aims to skip that refine step entirely, which is a stronger claim and the one I'd want independent numbers on before trusting.
## What zvec actually ships
The two levers above sit inside a fuller engine. The index and quantization menu:
| Layer | Options |
|---|---|
| Index | HNSW (dense + sparse), IVF, Flat (brute-force), HNSW-RaBitQ, Vamana / DiskANN (on-disk) |
| Quantization | fp16, int8, int4, RaBitQ (1-bit) |
| Distance | full-precision refiner pass over any quantized index |
| Retrieval | dense + sparse vectors, multi-vector queries, full-text search with hybrid fusion |
Systems details that matter for the throughput story:
- **CPU auto-dispatch.** zvec detects `AVX2`, `AVX512`, and `NEON` at runtime (via its `ailego` kernel library) and dispatches the SIMD distance kernels accordingly — so the same binary uses Ice Lake AVX512 on the benchmark box and NEON on ARM. int8 L2 distance is computed in batches.
- **Persistence.** A write-ahead log (WAL) for crash recovery; **RocksDB** holds metadata and the scalar index; vectors live in auto-scaling mmap'd segment files.
- **Concurrency.** Many concurrent readers; writes are single-process exclusive — the embedded, single-writer model, same as SQLite.
- **Filtered search.** Scalar predicates are pushed *into* the HNSW traversal instead of filtering after the fact, so a filtered query doesn't first retrieve then discard.
The exact config behind the headline run, as published:
```bash
# VectorDBBench · Cohere 10M (10M × 768-d) · Alibaba Cloud g9i.4xlarge (16 vCPU / 64 GiB)
zvec-bench \
--index hnsw \
--quantize-type int8 \
--m 50 \
--ef-search 118 \
--is-using-refiner \
--threads 12-20
# → 8,475 QPS, index build ≈ 1 hour
```
The Python surface is the usual embedded-DB shape — a `CollectionSchema` to declare fields and the vector index, `Doc` objects to insert, and a `VectorQuery` to search — so wiring it into a RAG loop is a library import, not a service to stand up.
## The numbers, in full
The cover chart is the VectorDBBench QPS ranking, with zvec highlighted at the top:
Only the top of the field, redrawn so the gap is legible. The label suffixes (`16c64g`, `8cu-perf`, `p2.x8-1node`) are each entry's own hardware and version — this is a leaderboard of different setups, not one controlled sweep:
QPS is the throughput axis; VectorDBBench measures it at a matched recall level per entry, and that recall panel isn't shown on this chart, so treat the ranking as "throughput at comparable accuracy" rather than raw speed at any accuracy. The `*` rows (OpenSearch, ElasticCloud) use `force_merge`, a build-time optimisation that trades index time for query speed. zvec's own build is the ~1-hour figure above.
## The take
The genuinely interesting thing about zvec is not the top bar — it's the *form factor*. An embedded, Apache-2.0, single-file-ish vector engine with real SDKs across five languages, that runs on-device down to Android, is a useful thing to have for local RAG where standing up Milvus or a managed cloud is overkill. "SQLite for vectors" is the right mental model, and the single-writer/many-reader concurrency model matches it exactly.
The 8,475 QPS is real but oversold by the framing. It is an in-process number — no network — sitting on a chart next to managed services that pay round-trip latency, on hardware that varies row to row. The mechanisms getting it there are the standard two: a graph index that visits a logarithmic slice of the data, and int8 quantization with a full-precision refiner so the cheap codes only *filter* while exact math decides the final order. Both are well-understood; zvec's contribution is a clean, SIMD-dispatched, embeddable implementation of them, plus a newer **RaBitQ** 1-bit path whose "high recall without re-ranking" claim is the one I'd hold out for independent verification on. For a team that wants vector search *inside* the application binary, it's worth a real evaluation — just run VectorDBBench yourself, on your hardware, against your recall target, before believing any single bar.
---
*Built on the [zvec release](https://github.com/alibaba/zvec) (Alibaba Tongyi Lab, Apache 2.0) — GitHub README, [zvec.org docs](https://zvec.org/en/docs/db/benchmarks/), and the [v0.3.0 notes](https://github.com/alibaba/zvec/releases/tag/v0.3.0). All QPS numbers are self-reported via [VectorDBBench](https://github.com/zilliztech/VectorDBBench) on Cohere 10M; the two interactive diagrams are illustrations of greedy graph descent and quantize-then-refine, not measured traces. The benchmark figure is reproduced from the project's published chart for commentary.*
---
# Laguna's Model Factory: treating model development as an industrial process
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/laguna-model-factory
> date: 2026-07-03
> tags: llm, agents, mixture-of-experts, systems, explainer
Most model reports are about a model. Poolside's **Laguna** report is unusual: its real subject
is the *factory*. Yes, it ships two Mixture-of-Experts models for agentic software engineering —
**Laguna M.1** (225.8B total, 23.4B active) and the open **Laguna XS.2** (33.4B total, 3B active) —
and they're solid, competitive-in-class coding models. But the argument the report actually makes
is that the models are downstream of something more valuable: a **Model Factory** that treats
foundation-model development as an *industrial process* rather than a craft. The evidence they
offer is a number — they built XS.2 from inception to release in **five weeks**.
## The two models, briefly
The models are conventional-but-careful MoE. Both are pre-norm Transformers with RMSNorm and
token-choice routing with a shared expert. The report spells out XS.2's config — **8 of 256
routed experts** per token plus one shared expert, the routed output scaled by 2.5 before it
rejoins the residual (a DeepSeek-V3 / Nemotron-style modulation), a sigmoid router normalized
after top-k, and a dense first layer for stability — but gives M.1's totals without its routing
width, so I won't guess it. Both use GQA (not MLA), the **Muon** optimizer, and a served context
of **256K** — XS.2's attention runs 8 KV heads with softplus per-head gating.
The one architectural wrinkle worth noting is XS.2's attention: where M.1 runs **global attention
on every layer**, XS.2 interleaves **sliding-window and global at a 3:1 ratio** with a 512-token
window — the same [hybrid-attention idea as MiMo-V2-Flash](/articles/mimo-v2-flash), tuned for a
smaller model that has to be cheap to serve. And the data is unambiguously code-first: XS.2's
pretraining mix is ~30% raw code plus a large synthetic-code share, on top of >30T tokens (from a
pool of ~27T unique). But by the report's own framing, none of this is the point. The point is how
it was all *made*.
XS.2 is really M.1 with four ablated deltas flipped on: the hybrid attention above, a
Warmup-Stable-Decay LR schedule instead of cosine, the routed-expert modulation, and one dense
layer instead of three. Each was a config change against the M.1 baseline, chosen on a small MoE
proxy — cheap precisely because the factory made the data pipeline, trainer, and eval harness
transfer for free. The peak LR came from a fitted **WSD scaling law**, $\text{lr}^\star \propto
N^{-0.46}\,D^{-0.27}$ in active params and tokens, rather than a re-tuned sweep. The stability
lessons transferred too: M.1's pre-training surfaced expert collapse around **450B tokens** (fixed
by Moonlight-style LR scaling, so Muon runs at AdamW-scale weight decay) and a logit-drift blowup
traced to a BF16 all-reduce on the LM-head input gradient (fixed by forcing that one reduction to
FP32). XS.2 hit none of them — the factory's job is to make the second model boring.
## What "industrial process" actually means
The Model Factory is defined as "a tightly-integrated stack of versioned data, training, evaluation,
and inference." Three principles hold it together:
- **Experiments as code.** Every run's inputs and config are committed to one repo and get a unique
ID; a Dagster DAG is the control plane for what runs and what depends on what. The payoff is
*end-to-end lineage* — a single token in a packed training shard traces back through dedup,
filtering, and synthesis to its source document, and every checkpoint, eval, and deployment traces
to the exact run that produced it. Nothing is a mystery artifact.
- **One code base for research and production.** A promising research idea isn't re-implemented to
ship — it's "promoted into production by flipping a configuration flag." The inference library
(Atlas, on vLLM) consumes the trainer's (Titan) model definitions *bit-accurately*, so what you
evaluate is what you serve is what generates your RL rollouts.
- **Reserve human attention for novel decisions.** A custom Kubernetes scheduler places jobs in
under a minute, auto-recovers from hardware failure, and only pages a human when recovery itself
fails. Cross-replica bit-identical weight-hash checks catch the silent data corruption a defective
GPU would otherwise smear through a run.
The components have names, and the flywheel is literal: **Titan** trains, **Blender** streams the
data mix, **Hive** generates synthetic data, **Atlas** serves inference *and* RL rollouts, and a
containerized **Code Execution** platform spanning ~1M repositories provides synthetic tasks,
evaluation, *and* RL execution rewards — one component doing all three. The RL harness they train
against is the same harness they ship to customers. That's the loop the interactive above is
gesturing at: data → train → eval → infer, where inference feeds the next round of data, and the
whole thing is versioned tightly enough that a five-week model is possible.
That pipeline also flipped a habit. For M.1 they ran a high-precision filter that aggressively
dropped noisy documents; for XS.2 they went **high-recall** instead — the composite score fully
rejects only ~25.8% of web samples as pure noise and *recovers ~34%* of documents the old static
rules had thrown away, then treats quality as a ranking signal and samples from score buckets
rather than hard-filtering. Under a >30T-token budget, controlling repetition and diversity beats
maximizing average quality; over-filtering starves the long tail.
## Choosing the data mix by optimization, not taste
The web pipeline decides *which documents survive*. A separate problem is *how much of each
source to train on* — the mixture weights. Done by hand, that's a few knobs set by taste and a
handful of ablations. The report's quietest radical move, **AutoMixer**, turns it into an
optimization loop, and it's the cleanest single instance of the factory thesis: a decision that
used to be craft becomes a versioned, automated search.
The setup is a surrogate-model sweep. For each data ablation they train a **swarm of ~60 proxy
models** — each a ~0.5B-parameter MoE on ~60B tokens — from **different mixtures** sampled over
**50+ heterogeneous dataset groups** (web, curated edu, academic, raw / grounded / synthetic code,
math web, conversational and knowledge sets). Every proxy is one labeled example of "mixture in,
capabilities out."
Formally they learn a surrogate $\mathcal{M}: x \to y$ where $x \in \Delta^{d}$ is a mixture over
$d$ dataset groups and $y \in \mathbb{R}^{k}$ is a vector of downstream metrics across $k$
capability groups — coding, math reasoning, STEM knowledge, commonsense, general knowledge.
Candidate mixtures are drawn near a hand-designed prior $x_0$ as $x \sim \text{Dirichlet}(\alpha
x_0)$ subject to $\lVert x - x_0 \rVert_1 < \epsilon$, so the search stays in realistic regions.
For each capability $j$ they fit a regressor $f_j(x) \approx y_j$ — linear in the simplified
picture, $\hat{y}_j = \beta_j^{\top} x + b_j$, non-linear in practice. The mixture is then chosen
by maximizing a weighted sum of the surrogates over the simplex:
$$
\max_{x}\ \sum_{j=1}^{k} w_j\, f_j(x)
\quad \text{s.t.}\quad \sum_i x_i = 1,\ \ x_i \ge 0,\ \ \lVert x - x_0 \rVert_1 < \epsilon
$$
with a $\lambda\, D_{\mathrm{KL}}(x \Vert x_0)$ penalty keeping the answer from collapsing onto a
few dominant sources. The knobs $w_j$ are where intent enters: weight coding and math and the
optimizer allocates data toward them.
The learned surrogate recovers relationships you'd expect — synthetic and curated code lift coding
evals; conversational and knowledge corpora lift commonsense — plus finer cross-effects. On a
3B-param / 1.5T-token check, the optimized mix posts large gains on the targets (HumanEval+ **+43%**,
CRUX-I **+54%**, GSM8K **+41%**, MultiPL-E **+27%**) and, encouragingly, **generalizes to held-out
benchmarks** it wasn't optimized against (MATH **+25%**, LiveCodeBench **+39%**, BigCodeBench +16%).
The cost is stated honestly and it's small: a few commonsense tasks regress (ARC-C **−6.8%**, the
rest under 1.5%), which is exactly what you sign up for when the objective down-weights them. XS.2's
final mixture — **30.6% raw code, 25.4% synthetic/code-text, 25.2% web, 9% math**, the rest
knowledge / instruction / academic / books (Table 4) — shifted toward web, synthetic, and math
relative to M.1's while keeping the code-heavy spine. That's the thesis in one artifact: a data-mix
decision made by an optimizer over a learned model, logged and re-runnable, instead of argued in a
meeting.
## Agentic training, from real commits
The coding ability comes from training on the actual job. Poolside turns **real git commits from
public repos into verifiable tasks** — a problem statement, a repo checkout, and a hidden test
patch, with the gold answer being the commit's own diff. A two-sided filter keeps only commits
where the gold diff passes the tests *and* an empty patch fails them (discarding trivial or
non-exercising tests), yielding **30–60k tasks from a ~236k-commit pool**. Those tasks feed both
SFT (as teacher-generated trajectories, sometimes wrapped in synthetic system messages for
instruction-following pressure) and the RL pool (where the repo's own test suite is the verifier).
The RL stage is online: the policy itself drives the **production agent harness** across several
thousand live containers at a time, and each rollout's reward comes from that container's verifier.
The objective is **CISPO** — the importance-ratio-clipping surrogate from MiniMax-M1, not a Poolside
invention — paired with a length-weighted leave-one-out group baseline; clipping is asymmetric,
an effective $[0, 5]$ on the ratio, so it only bites on heavily off-policy tokens. Reward is a
deterministic chain of checks: a malformed tool call or template violation is $-0.1$, giving up
before a minimum number of tool calls is $-0.1$, a timeout is $0.0$, and the **only positive reward
is the binary task verifier** ($1.0$) — unit tests for SWE tasks, bash assertions for terminal
tasks, exact-match for tool-integrated math. A small $-0.05$ per-token penalty lands on exactly the
tokens of a failing tool step to sharpen credit assignment; everything else is carried by the
terminal 1/0. It's the same "make the process the trainable target" instinct as [Agents-A1's
verifier-graded trajectories](/articles/agents-a1), wired straight into the factory — the execution
environment that grades RL is the one that generates data and runs evals.
## The numbers, honestly
Here's where the modest framing matters. The report claims the models are "competitive with
state-of-the-art open models in their respective weight classes," and that's accurate — *competitive*,
not leading. M.1 lands mid-pack among the ~200B-class open models on SWE-bench Verified:
XS.2 is the more interesting result, because it's competitive in the ~30B class while activating only
**3B** parameters per token — against dense 24–31B rivals:
XS.2 beats Devstral Small 2 and Gemma 4 and edges Qwen3.5, though Qwen3.6 leads the class. Two honesty
notes the report itself makes: the baseline numbers are Poolside-selected published scores (not re-run
in their harness, so there's provider-config bias), and they patched all four benchmarks to remove
git-history leaks before scoring — so these differ slightly from public leaderboards by construction.
It's the rare case where the *methodology* disclosure is more reassuring than the raw scores.
One finding from the quantization work is worth carrying away regardless of the leaderboard: **bad
quantization hurts agentic benchmarks far more than single-turn ones.** A small per-token error
that's invisible on a one-shot question compounds across a hundred-step trajectory, and the report
saw exactly that — intermediate schemes that barely moved single-turn scores cratered the agentic
ones. The fix started by looking at *where* the error comes from. Naive INT4 (AWQ, `W4A16`) lost
quality because outlier activations pile up in the residual stream starting around **layer 30 of
the 40-layer network**:
So they went mixed-precision: **`INT4` for the first 30 layers, `INT8` for the last 10** (with a
SpinQuant rotation as a pre-pass). `NVFP4` needed more — direct post-training quantization lost too
much, so they recovered it with **quantization-aware distillation**, training the quantized student
to match a higher-precision teacher on a fixed dataset. The KV cache goes to `FP8` across the full
131K context, roughly doubling how many trajectories a replica can hold. The through-line is the
same as everything else here: the diverse eval harness is what caught the agentic-only regression
that a single-turn benchmark would have waved through.
## The take
I went in expecting an architecture paper and came out thinking about CI. The Laguna models are
genuinely fine — a competent ~200B flagship and a genuinely efficient 3B-active open model that holds
its own in a crowded class — but they're not what the report is selling. It's selling the claim that
*iteration speed* is the frontier lever: if every run is reproducible code, every win ships by flipping
a flag, and the same execution environment grades your RL, runs your evals, and generates your data,
then you can turn out a from-scratch model in five weeks and keep pace with model complexity instead of
drowning in it. Whether the "factory is the moat" thesis holds as everyone industrializes is the open
question — but as a piece of honest infrastructure writing, with the models presented as *outputs of a
process* rather than heroic artifacts, it's a refreshing shape for a technical report. XS.2's weights
are open (Apache-2.0); the factory, of course, is not.
---
*Built on the [Laguna M.1/XS.2 Technical Report](https://arxiv.org/abs/2605.27605) (Poolside, 2026) and
the [XS.2 model release](https://huggingface.co/collections/poolside/laguna-xs2) (Apache-2.0). Benchmark
figures are from the report's tables (baselines are Poolside-selected published scores on leak-patched
benchmark images); SWE-bench Verified figures use the report/model-card values (74.6 / 69.9), higher
than the earlier launch checkpoint.*
---
# LongCat 2.0: a 1.6T open-weights MoE, and the sparse attention behind it
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/longcat-2
> date: 2026-07-05
> tags: llm, mixture-of-experts, attention, inference-optimization, explainer
**LongCat 2.0** is Meituan's largest open model: a Mixture-of-Experts language model with **1.6 trillion total parameters** that activates only **~48 billion per token**, released with **MIT-licensed weights** on Hugging Face and ModelScope (plus a separate `LongCat-2.0-FP8` artifact for deployment). Two things make it worth a close read. First, the whole training run and serving stack were built on **AI ASIC superpods** rather than GPUs — a claim about frontier-scale training on alternative silicon. Second, its headline architectural bet is **LongCat Sparse Attention (LSA)**, a redesigned sparse-attention indexer aimed squarely at long-context serving.
Every benchmark here is **provider-reported** and, unless marked otherwise, run in-house under Meituan's own harness. On the published suite LongCat 2.0 does **not** top a single benchmark against the full frontier set: it is competitive with — and on a few coding/agentic tasks beats — **Gemini 3.1 Pro** and **GPT-5.5**, while the strongest closed model (usually **Claude Opus 4.8**) leads every row. Read it as a strong *open-weights* model, not overall SOTA. The ASIC training story and "millions of accelerator-days" are provider details we cannot independently verify.
## Where the parameters live
1.6T total but ~48B active is a **~3% activation rate** — the sparsity that makes a model this large affordable to run. But LongCat 2.0 has a second, less common parameter store: an **N-gram Embedding** of **135B parameters**, inherited from LongCat-Flash-Lite. Crucially, these are *not* extra experts — they expand parameters along **sparse dimensions orthogonal to the MoE**, adding capacity through a token-n-gram lookup rather than a wider expert pool. Meituan frames it as a scaling principle: once MoE sparsity has "crossed the sweet spot," a bounded slice of N-gram parameters beats simply adding equivalent MoE capacity.
If the MoE routing here is unfamiliar, we built it up from nothing in [Mixture of Experts, from scratch](/articles/mixture-of-experts-from-scratch) — the router, the top-k gate, and why activating a sparse subset is the whole economic argument for a trillion-parameter model.
## LongCat Sparse Attention
The expensive part of long context is attention: under full attention every query reads every past token, so a 1M-token context is brutal to serve. The field's fix is to make attention **sparse** — read only a chosen subset of the past. LongCat's starting point is [DeepSeek's Sparse Attention (DSA)](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp) and its "Lightning Indexer," whose weaknesses Meituan names directly: **output discontinuity** and a **quadratic scoring bottleneck**. LSA answers with three *orthogonal* improvements.
The core one is **Hierarchical Indexing (HI)**, and it is best understood against the sparse-attention design we covered earlier — [MiniMax Sparse Attention (MSA)](/articles/minimax-sparse-attention), which scores the past in **128-token blocks** and keeps the top-k *whole blocks*. LSA treats that block selection as only a **coarse recall**: a cheap block-level pass proposes candidate blocks, then a **fine pass selects the most relevant individual tokens** inside them. Same idea of scoring first and reading less — but token-precise, over a smaller candidate set the indexer has to score. Flip the mode below to watch the budget move from whole blocks to individual tokens:
The second improvement, **Cross-Layer Indexing (CLI)**, cuts cost on a different axis. Deciding what to read costs an *index pass* per layer; but attention saliency is empirically stable across adjacent layers, so LongCat **shares one index every 2 layers** — half the layers reuse a neighbour's selection instead of recomputing it, taught by cross-layer distillation during training. The same trick collapses the model's 3-step [Multi-Token-Prediction](/articles/multi-token-prediction) draft into a single shared pass for speculative decoding:
The third, **Streaming-aware Indexing (SI)**, is a memory-systems move: it reshapes the token-selection budget to combine hardware-aligned **contiguous** access with dynamic random selection, turning fragmented reads into predictable sequential ones for coalesced HBM bandwidth. Together the three attack the same target from different sides — HI shrinks *what* the indexer scores, CLI shrinks *how often* it runs, SI makes the reads it does issue cheaper.
One honest deployment note: the public SGLang integration **drops the hierarchical stage for simplicity** and serves LongCat 2.0 on 16× H20 with tensor + expert parallelism. So the shipping inference path today is the CLI/SI part of LSA, not the full HI pipeline.
## Training, and what "1M context" means
Pretraining spans **35T+ tokens** with, Meituan reports, **no rollbacks or irrecoverable loss spikes** — a stability claim about frontier-scale training on ASICs. Long-context ability comes from training on **hundreds of billions of tokens of 1M-context data**. That "1M" is a **training-data** figure: the sources describe the data the model saw, and the README does not publish a separate usable-context window or a long-context retrieval eval (e.g. RULER/HELMET) to pin down how far that quality actually holds at inference. Worth keeping the two apart.
## The numbers, in full
LongCat 2.0 is evaluated against Gemini 3.1 Pro, GPT-5.5, and three Claude Opus checkpoints (4.6 / 4.7 / 4.8). The official chart:
On **SWE-bench Pro** LongCat's 59.5 edges past GPT-5.5 (58.6), Gemini 3.1 Pro (54.2) and Opus 4.6 (57.3) — but the newer Opus checkpoints pull ahead (4.7 = 64.3, 4.8 = 69.2):
**IFEval** shows the opposite shape: LongCat (90.0) trails Gemini 3.1 Pro (96.1) and GPT-5.5 (95.0), but the newest Claude checkpoints regressed here, so LongCat sits *above* Opus 4.8 (86.0):
The rest of the suite tells the same "competitive, not leading" story. LongCat edges Gemini 3.1 Pro on **Terminal-Bench 2.1** (70.8 vs 70.7) and **FORTE** (73.2 vs 70.3, matching Opus 4.6), leads it on **RWSearch** (78.8 vs 76.3) and **IMO-AnswerBench** (81.8 vs 79.5 for GPT-5.5) — yet on each of those a closed model still tops the row (GPT-5.5 = 77.8 on FORTE; 85.3 on RWSearch; Opus 4.8 = 78.9 on Terminal-Bench; Gemini = 90.0 on IMO-AnswerBench, 96.1 on IFEval, 94.3 on GPQA-diamond, where LongCat is 88.9). On **BrowseComp** it clearly trails (79.9 vs Gemini's 85.9). No single number here is a headline win over the field.
## The take
LongCat 2.0's real contribution isn't a benchmark crown — it's the *combination*: a genuinely large (1.6T) MoE, trained end-to-end on non-NVIDIA silicon, shipped under **MIT** weights, with a sparse-attention design that is a real step past the DSA/MSA lineage. LSA's three moves are cleanly separated — **HI** goes finer than MSA's whole-block selection (token-precise, over a coarsely-recalled candidate set), **CLI** amortizes the indexer across layers, **SI** makes the memory access regular — and each targets a distinct cost, which is the kind of engineering that turns theoretical sparsity into wall-clock savings. The honest caveats are the usual open-weights ones: the benchmarks are self-run and self-selected; the model is competitive with Gemini 3.1 Pro and GPT-5.5 on coding/agentic tasks but trails Claude Opus 4.8 on most; the striking "1M context" and "millions of accelerator-days" are provider claims, not independently measured; and the part of LSA that actually ships today (in SGLang) is CLI/SI, with the hierarchical stage dropped for simplicity. For a team that wants frontier-scale, open, and long-context — and can run 16× H20 — it's a serious option. As a claim that you don't need NVIDIA to train at this scale, it's the more interesting story.
---
*Built on the [LongCat 2.0 release](https://github.com/meituan-longcat/LongCat-2.0) (Meituan, 2026) — GitHub README and [Hugging Face model card](https://huggingface.co/meituan-longcat/LongCat-2.0), MIT license. All benchmark numbers are provider-reported (in-house harness unless marked `*` = official model report); the interactive diagrams are illustrations of the mechanism, not measured traces. The benchmark figure is reproduced from the model card for commentary.*
---
# MiniMax Sparse Attention: let each query pick its own blocks
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/minimax-sparse-attention
> date: 2026-07-04
> tags: llm, attention, inference-optimization, mixture-of-experts, explainer
The thing that makes long context expensive is that, under full attention, **every query token reads
every past token**. The [KV cache](/articles/how-llm-inference-works) that stores those keys and values
grows linearly with sequence length, and the attention itself scales with it — so a 1M-token context is
brutal to serve. The field's answer has been to make attention *cheaper*: linear-attention hybrids
[like HydraHead](/articles/hydrahead), sliding windows [like MiMo-V2-Flash](/articles/mimo-v2-flash), or
KV compression [like TurboQuant](/articles/turboquant-kv-cache). **MiniMax Sparse Attention (MSA)** takes
a different route: keep attention **exact**, but let each query attend to only a **small, learned subset**
of the context.
The trick is to think in **blocks**. Partition the past KV into fixed blocks of 128 tokens; for each
query, a cheap *index branch* scores every block, keeps only the highest-scoring few, and the main
attention runs — exactly — over just those. Scrub the query and watch which blocks it actually reads:
Two design choices make this work. First, the selection happens **per GQA group**: MiniMax uses grouped-
query attention with 64 query heads tied to 4 KV heads, so there are 4 groups, and *each group picks its
own blocks* off the same keys and values — flip the group above and the arrows swing. Second, the index
is **trained, not heuristic**: an auxiliary KL loss aligns the index branch's block scores with the main
attention's true distribution over blocks, with a stop-gradient so it only trains the tiny index
projections and never disturbs the backbone.
## The mechanism, precisely
Each attention layer splits into two branches:
- **Index branch (selection).** It adds one lightweight index-query head per group and a single shared
index-key head. For query `i` it computes `Q_idx · K_idxᵀ`, pools those token scores to the **block**
level by max-pooling, and takes the **top-k blocks** — deployed as **block size 128, k = 16**, so a
fixed budget of **2,048 tokens** per query. The block the query sits in (the *local* block) is always
kept.
- **Main branch (compute).** Given the selected block indices, it runs **exact** attention over only
those blocks. Because every head in a group reuses the same block set, KV reads stay block-contiguous —
which is what lets a custom kernel turn the FLOP savings into real wall-clock speedup.
## Why the fixed budget is the whole point
Because every query attends to the same **2,048 tokens** no matter how long the context is, the *fraction*
of the context each query reads collapses as context grows — but, honestly, the measured speedup is much
smaller than that fraction would suggest, because the index branch still scans every block. Slide the
context length and watch both numbers:
## The numbers
MSA is trained two ways — from scratch (**MSA-PT**) and by converting a full-attention checkpoint
(**MSA-CPT**) — and compared against MiniMax's *own* GQA full-attention model at a matched 3T-token
budget. The headline is parity, not a free lunch: MSA holds or slightly beats full attention on most of a
28-benchmark suite, with real regressions on a few. The efficiency, meanwhile, is decisive at long
context:
Against MiniMax's own full-attention model (values in parentheses), MSA-PT holds or edges ahead —
RULER-8K 84.2 (79.8), GSM8K 77.7 (76.2), MMLU 67.2 (67.0), HumanEval 64.0 (61.0), VisualWebBench 68.4
(55.6):
The efficiency wins that motivate all of this, at 1M context on H800:
A few honesty notes. The baseline is **provider-selected and internal** — MSA vs MiniMax's own GQA model,
not against other sparse-attention methods (NSA, MoBA) or external frontier models. Quality is "on par,"
and the converted **MSA-CPT** variant does trail full attention on some tasks (GSM8K 73.7 vs 76.2,
HumanEval 57.9 vs 61.0, HELMET-128K −0.60) — the fixed budget shows up as small losses on
retrieval-heavy long-context tasks. And the striking efficiency figures are the **1M** extreme with a
fixed 2,048-token budget on a specific head config; at 32k the advantage is barely 1.6×. The **28.4×** is
a theoretical FLOP count read off a chart, not measured throughput — the honest wall-clock numbers are
14.2× and 7.6×.
## The take
MSA's contribution is that it makes *selection* a first-class, trainable part of attention rather than a
bolt-on. The block granularity is the quiet key: picking whole 128-token blocks (not individual tokens)
keeps memory access regular enough that a kernel can actually realize the savings, and sharing the
selection across a GQA group keeps it cheap. Set against the other long-context playbooks — linear
attention trades exactness for O(1) state; sliding windows drop the far past; KV quantization shrinks
each entry — MSA keeps attention **exact and full-range** and simply reads *less of it*, chosen per query
and per group. Whether a fixed 2,048-token budget holds up as tasks demand genuinely global reasoning is
the open question the retrieval regressions hint at; but as a way to serve a 1M context at a fraction of
the cost while staying on the full-attention quality curve, it's a clean, well-engineered bet. The
production model, MiniMax-M3, ships with it (open weights, minimax-community license).
---
*Built on [MiniMax Sparse Attention](https://arxiv.org/abs/2606.13392) (Lai, Xu, Yang et al.; MiniMax,
2026) and the [MiniMax-M3 release](https://huggingface.co/MiniMaxAI/MiniMax-M3). Benchmark and efficiency
figures are quoted from the paper (the 109B-total / 6B-active experimental model; block size 128, top-16);
the interactive diagrams are illustrations of the mechanism. Speedups are measured on H800 at 1M context.*
---
# MrFlow: climb the resolution in pixel space, not diffusion steps
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mrflow-diffusion-acceleration
> date: 2026-07-04
> tags: diffusion, image-generation, inference-optimization, flow-matching, explainer
A modern text-to-image model — a flow-matching or [diffusion](/articles/set-diffusion) network like
FLUX or Qwen-Image — makes a picture by running a stack of denoising steps. The expensive fact about
that stack is where the compute lands: **every step runs at the output resolution**. A 20-step sample
at 1024² pays the full-resolution price twenty times over, and most of those steps are spent on detail
the early ones can't even see yet. That's the cost MrFlow goes after — and it does so **without any
fine-tuning**, purely by rearranging *when* the model works at full size.
The idea is a budget reshuffle. Structure is decided early and is cheap to compute at small sizes;
sharp high-frequency detail is what full resolution actually buys you. So MrFlow spends the pricey
diffusion steps at **low resolution**, does the climb to full resolution in **pixel space** with an
off-the-shelf super-resolution network, and then spends just **one** diffusion step at high resolution
to clean up. Walk the four stages:
## The pipeline, precisely
MrFlow is four stages, and only the outer two touch the full-resolution latent:
1. **Low-resolution generate.** Run the flow/diffusion model for ~12 of its 20 steps at a low
resolution (e.g. 512²). Steps are cheap at this size, and this is where composition and layout are
fixed. Decode to a low-resolution image.
2. **Pixel-space super-resolution.** Upsample that image to full resolution with a **pixel-space SR
network** — Real-ESRGAN — in a single forward pass. This is the resolution jump, and crucially it is
*not* more latent diffusion steps.
3. **Low-strength noise re-injection.** Encode the upscaled image back to latent and add a small amount
of noise, `σ_t ∈ [0.1, 0.15]`, computed **closed-form** from the flow schedule. No extra network,
no training — just enough noise to give the model something to denoise.
4. **One HR refine step.** A single high-resolution diffusion step blends the upscaled detail back into
the model's own distribution and cleans up the SR network's artifacts.
The contrast with the usual "latent upsampling" trick (middle row of the figure) is the whole point.
Those schemes also start small, but they climb by running **more latent diffusion steps** at the higher
resolution — you pay full-resolution diffusion twice. MrFlow does the climb with a cheap pixel-space
network and buys back fidelity with a *single* diffusion step.
## Where the budget actually goes
The reason this is faster isn't subtle: high-resolution diffusion steps are the expensive line item,
and MrFlow runs almost none of them. If you count an HR step as roughly 4× an LR step (2× the linear
resolution is 4× the pixels), a native 20-step HR sample and a MrFlow `(12, 1)` config are worlds apart
in compute. Drag the config and watch the budget — and the honest quality tradeoff — move:
That schematic is where the **config dependence** lives, and it's the first honest caveat. The speedup
is the reliable part — fewer, cheaper steps is unambiguously less compute. Quality is the part that
swings: strip the refine steps and Real-ESRGAN's artifacts survive to the final image; starve the
low-resolution stage and the composition never locks in.
The advertised **"within 1% of native quality" is a best-case figure**, and it lives at the conservative
end of that slider. Push the config harder and the gap widens fast — FLUX at the aggressive `(12, 1)`
setting drops roughly **18% on OneIG**. Degradation is real, and it is both **config- and
model-dependent**.
## Why pixel-space SR instead of more diffusion
The subtle design question is: once you have a low-resolution image, how do you get to full resolution?
The latent-upsampling answer is "more diffusion." MrFlow's answer is "a super-resolution network, then
one diffusion step to fix it." The paper compares SR backbones — bilinear interpolation, SwinIR,
OSEDiff, Real-ESRGAN — with and without the final HR refine step:
Two things read off that grid. First, the SR network alone is **not enough** — plain interpolation
stays soft, and even a strong SR net can hallucinate wrong detail (look at the "9am" text). Second, the
single refine step is doing real work: the bottom row is visibly cleaner than the top. Real-ESRGAN can
introduce its own artifacts, and the HR refine step is there **precisely to fix them** — the two stages
are a pair, not alternatives.
## The speed–quality frontier
The payoff is best seen as a Pareto plot: quality against speedup, versus the obvious baseline of just
running the model with fewer native steps. Toggle the model and walk the refine-step configs:
The shape is the honest summary. On the good configs, MrFlow's frontier dominates both "fewer native
steps" and prior accelerators like TeaCache and RALU — you get several× speedup while staying close to
the native star. But the frontier still slopes downward: more speed costs some quality, and the `+1`
config (fastest) sits measurably below `+3`. Qwen-Image holds its quality far better than FLUX across
the same speedups, which is exactly the point about **model dependence** — some base models tolerate the
reshuffle much better than others.
## The honest headline
The clean numbers to keep are the training-free ones: roughly **4–9× faster** on the frontier configs at
**near-native** quality, no fine-tuning required. The bigger figures you may see quoted come with strings:
The **10×/25× speedups are config-dependent**, and the top end is not training-free. The **25× figure is
reached only with distillation stacked on top** of MrFlow's staged sampling — not by the training-free
pipeline alone. Quote the ~4–9× training-free range if you want the number MrFlow earns on its own.
A few smaller honesty notes. "Within 1%" is a best case measured on forgiving configs and models; the
same pipeline pushed to `(12, 1)` on FLUX loses ~18% on OneIG. The compute-unit accounting in the
diagrams above is a schematic (an HR step is *roughly* 4× an LR step) — the paper's speedups are
measured wall-clock, and they depend on the SR network's own cost, which the unit count glosses over.
And Real-ESRGAN, being a GAN, can invent detail that isn't in the low-resolution image; the refine step
mitigates but doesn't fully erase this.
## The take
MrFlow's contribution is a reframing more than a new network: the resolution climb doesn't have to be
paid for in diffusion steps. Set against the usual acceleration playbooks — caching redundant
computation ([TeaCache-style](/articles/how-llm-inference-works)), or simply taking fewer steps — it
makes a sharper bet: **structure is cheap and settles early; resolution is expensive and can be borrowed
from a pixel-space SR net; artifacts are cleanable in one diffusion step**. Because every piece is
training-free and closed-form, it drops onto an existing checkpoint with no retraining, which is the
practical reason to care. The catch is the one every honest acceleration paper carries — the best-case
headline is best-case. On a forgiving model at a conservative config it really is near-free; push the
config or pick a brittle model and you pay for the speed in quality. As a way to make an existing
image model several times cheaper to sample without touching its weights, though, it's a clean idea,
cleanly executed.
For the diffusion background this builds on, see [Set Diffusion](/articles/set-diffusion) and, for how
denoising models differ from autoregressive ones, [the diffusion *language* model
walkthrough](/articles/illada-diffusion-language-model).
---
*Built on MrFlow (arXiv 2607.01642), a training-free staged-sampling accelerator for flow and diffusion
image models. Benchmark and speedup figures (GenEval, OneIG; FLUX.1-dev and Qwen-Image) are quoted from
the paper; the interactive diagrams are illustrations of the mechanism, and the per-stage compute units
are a schematic, not measured FLOPs. Speedups are config- and model-dependent; the top-end figures
require distillation on top of the training-free pipeline.*
---
# Nemotron in NVFP4: training a frontier model natively in 4-bit
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nemotron-nvfp4
> date: 2026-07-04
> tags: llm, quantization, training, mixture-of-experts, explainer
Almost every "4-bit model" you have heard of is 4-bit *after the fact*: the network is trained in BF16,
and then a [post-training quantizer](/articles/turboquant-kv-cache) compresses the finished weights so
they are cheaper to serve. Training itself stays in 16-bit, because the gradients are delicate and 4 bits
is a very small number. NVIDIA's Nemotron does something more aggressive: it runs the actual training
GEMMs — the big matrix multiplies in the forward *and* backward pass — natively in **NVFP4**, a 4-bit
floating-point format, while the model is still learning. The headline is a training-loss gap under
**0.4%** against a BF16 reference. The interesting part is everything that had to be true to get there.
## What NVFP4 actually is
Four bits cannot, on their own, represent the range of values a weight tensor spans. NVFP4's answer is to
split the job: store each value in a tiny 4-bit **element**, and recover dynamic range from a **shared
scale** that a whole block of elements multiplies through. The element is **E2M1** — 1 sign bit, 2
exponent bits, 1 mantissa bit. The scale is where NVFP4 differs from the MXFP4 format you may have seen:
it shares one **FP8 (E4M3)** scale across every **16** elements, plus a single FP32 scale for the whole
tensor. Flip between the formats:
That two-level scaling is the whole trick. A 4-bit element with an 8-bit block scale over 16 values
behaves, in effective dynamic range, more like a ~10-bit ("E6M4"-ish) number than a raw 4-bit one — while
costing **4.5 bits per element** to store (4 for the element, 8/16 = 0.5 for the scale). MXFP4 uses a
coarser power-of-2 (E8M0) scale over a wider 32-element block: cheaper at 4.25 bits, but blunter, because
one outlier drags a bigger block and a power-of-2 scale can only snap to coarse steps. NVFP4's finer block
and FP8 scale are what make it stable enough to *train* in, not just serve in.
## Why 4-bit training normally falls apart — and the three fixes
Post-training quantization only has to preserve the *forward* pass of a frozen network. Training in 4-bit
is harder on two counts: the **gradients** flow through the same low-precision GEMMs, and any systematic
error compounds over trillions of tokens. Nemotron leans on three stabilizers, each aimed at a specific
failure:
1. **2D block quantization on weights.** Instead of quantizing weights in 1D strips, quantize them in 2D
blocks, so a block's shared scale better fits the local structure of the weight matrix and fewer values
get clipped.
2. **Random Hadamard transform on the wgrad inputs.** The weight-gradient (wgrad) matmul is the one most
poisoned by outliers. Multiplying its inputs by a random Hadamard matrix *spreads* those outliers
across the block before quantization (and is undone analytically), so no single large value blows out
the block scale.
3. **Stochastic rounding on gradients.** Deterministic rounding biases small gradients toward zero — over
trillions of steps that lost signal is fatal. Rounding gradients *stochastically* is unbiased in
expectation, so the gradient direction survives quantization even when individual values don't.
Where each of these lives is easier to see than to say. Here is one path through the stack, colored by
precision, with the stabilizers attached to the FP4 GEMMs — flip to the backward pass:
## The honest part: this is not end-to-end FP4
The precision map makes the biggest caveat visual: **most of the network is not FP4.** Native FP4 training
means the heavy expert/MLP weight-GEMMs run in NVFP4 — that is the bulk of the FLOPs — but a meaningful
fraction of the model is deliberately kept at higher precision, because FP4 is most fragile exactly there:
**Kept at higher precision:** the **final ~16 layers**, the **Mamba-2 projection layers**, the **QKV
projections**, the **MTP (multi-token-prediction) module**, and the **embeddings**. So "native FP4
training" is really *mixed-precision* training with FP4 doing the heavy lifting on the compute-bound
GEMMs — not a network where every tensor is 4-bit. Read the quality claim the same way: it's a
**training-loss gap under 0.4%**, not a downstream task-benchmark parity result. A small loss gap is
necessary for parity but does not prove it.
And it was not a smooth ride. The run **diverged twice** — around ~8T and ~16T tokens — and each time the
team had to roll back to an earlier checkpoint and restart the segment (with an FP32-rounding fix and a
re-annealed learning rate) to recover. Four-bit training is not a free lunch you turn on and forget:
## The model underneath
The precision story rides on a specific architecture: a **550B-total / 55B-active** hybrid that interleaves
**Mamba-2** state-space blocks, periodic **attention**, and **[LatentMoE](/articles/mixture-of-experts-from-scratch)**
expert layers, with a **[multi-token-prediction](/articles/multi-token-prediction)** head on top. The
Mamba-2 blocks carry most of the sequence mixing cheaply; attention appears sparingly for exact long-range
recall; the MoE layers are where the parameters (and the FP4 GEMMs) live. The repeating layer pattern:
Storing that many parameters in 4.5-bit elements is a real memory win, but — honestly — not the clean 4×
the "4-bit" label implies, once you count the scale overhead. Slide the parameter count:
For the effective storage cost per element, the formats line up cleanly — and NVFP4's 4.5 bits sits
between MXFP4's leaner-but-blunter 4.25 and BF16's 16:
## One more thing not to conflate
There is a *second* quantization result in this work that is easy to mix up with the training story:
**inference PTQ** — post-training-quantizing the finished model down for cheaper serving. Those serving
numbers are a separate experiment about deployment, measured on the trained checkpoint; they are not
evidence about training precision. The claim on the table here is narrower and more interesting: that you
can run the *training* GEMMs in 4-bit and land within a fraction of a percent of a BF16 loss curve. Keep
the two apart.
## The take
The quiet lesson is that "native 4-bit training" is an engineering result about *where* you dare to put
4 bits, not a claim that the whole network is 4-bit. NVFP4's two-level scale (4-bit element + FP8 block
scale + FP32 tensor scale) buys back enough dynamic range to make the compute-heavy expert GEMMs
trainable in 4 bits; the three stabilizers — 2D block quantization, Hadamard-smeared wgrad inputs, and
stochastic gradient rounding — keep the gradients honest; and mixed precision quietly protects the fragile
edges (embeddings, final layers, projections, MTP). The payoff is a real reduction in training compute and
memory bandwidth on hardware built for FP4, at a training-loss gap under 0.4%. The caveats are equally
real: it is not end-to-end FP4, a loss gap is not benchmark parity, the run diverged twice, and the
inference-PTQ numbers are a different story. As a demonstration that frontier-scale training can leave the
16-bit comfort zone, though, it is a genuinely aggressive, well-instrumented bet — and it stuck.
---
*Built on NVIDIA's Nemotron NVFP4 training report. NVFP4 is a 4-bit E2M1 element with a per-16-element
FP8 (E4M3) block scale and an FP32 per-tensor scale; figures are reproduced from the paper for commentary,
and the interactive diagrams are illustrations of the mechanism. Related: [why quantization is
hard](/articles/turboquant-kv-cache), [mixture-of-experts from
scratch](/articles/mixture-of-experts-from-scratch), [multi-token
prediction](/articles/multi-token-prediction), and [large-scale
training](/articles/megatrain-single-gpu-training).*
---
# Program-as-Weights: compiling a natural-language spec into a tiny local model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/program-as-weights
> date: 2026-07-04
> tags: llm, inference-optimization, fine-tuning, systems, explainer
Some everyday programming tasks refuse to be written as rules. Alerting a human on only the log lines
that matter, repairing malformed JSON, ranking snippets by intent, deciding whether a message is
urgent — even a regex for "parse this messy text" shatters on edge cases. The modern move is to punt:
call a large language model API and let it decide, per input. That works, but you pay for it on
**every call**, and you give up locality, reproducibility, and price.
**Program-as-Weights (PAW)** — from Waterloo, Cornell, and Harvard — proposes a different deal. Treat
that fuzzy function like source code: **compile it once** from a natural-language specification into a
compact, locally-executable neural artifact, and then *run* that artifact cheaply for every subsequent
input. The compiler is a 4B model; the artifact is a small **LoRA adapter**; the thing that executes it
is a **frozen 0.6B interpreter**. Play with the loop first — pick a spec, compile it, then feed inputs
through the interpreter:
The headline result is the payoff of that split: a **0.6B Qwen3 interpreter** running PAW programs
**matches direct prompting of Qwen3-32B**, while using **roughly one-fiftieth of the inference memory**
and running at **30 tokens/second on a MacBook M3**.
## The compiler–interpreter system
PAW borrows the oldest idea in programming — separate the compiler from the runtime — and instantiates
it with neural parts. The **compiler** reads a function specification (plus an optional pseudo-program
sketch) and emits a parameter-efficient adapter. The **interpreter** is a small, *frozen* model that,
loaded with that adapter, executes the function on real inputs.
The load-bearing piece is the **LoRA mapper**. Rather than have the compiler regress raw adapter
weights (a huge, unstructured output), it mean-pools the compiler's features and passes them through an
MLP that produces mixing coefficients over a **set of learned LoRA basis matrices** — the emitted
adapter is a combination of reusable low-rank factors. That keeps the compiler's output small and
well-conditioned, and it's what makes "spec → weights" learnable at all. (An earlier prefix-tuning
variant works too; Text-to-LoRA is just the current best.)
To train the compiler, the authors built **FuzzyBench**, a **10-million-example** dataset of fuzzy
functions spanning text processing, search and matching, custom classification, code and NL commands,
safety and verification, agentic tool use, and format repair — with clean and noisy specification
variants so the compiler learns to be robust to sloppy prompts:
## Why compile at all? The economics
The reason this framing matters is cost structure. Calling a big model's API is a cost you pay on
**every input**. PAW pays a **one-time compile** — a single pass of the 4B compiler — and then each
application runs locally for almost nothing. So cumulative cost crosses over fast, and the gap only
widens. Drag the number of calls:
That's the thesis in one picture: PAW **reframes the foundation model from a per-input problem solver
into a tool builder**. You invoke the big model once per *function definition*, get back a small
reusable artifact you own, and every *function application* after that is cheap, offline, and
reproducible. Compile a library of them and each becomes a tiny local endpoint:
Two practical wins reinforce it: the paper reports the adapters **quantize with no measurable accuracy
loss**, and the whole thing runs at interactive speed (30 tok/s) on a laptop — so the compiled function
is genuinely local, not a cloud dependency in disguise.
## The honest caveats
PAW is a genuinely new framing, but it buys its efficiency with real constraints, and the paper is
candid about them:
- **The compiler and interpreter are a coupled pair.** An adapter compiled for one frozen interpreter
isn't portable to a different one — you commit to an interpreter.
- **The compiled program isn't interpretable.** Unlike source code, you can't read a LoRA adapter to
see what the function *does*; you can only run it. "Program" is an analogy for the compile-once
workflow, not a claim of legibility.
- **Single-step fuzzy functions.** PAW targets one-shot transformations (classify, repair, rank), not
long multi-step agentic control flow.
- **Trained on synthetic data.** FuzzyBench is model-generated; the compiler's competence is bounded by
the distribution of tasks it was synthesized from, and real specs can fall outside it.
- **The best adapter type is task-dependent.** Text-to-LoRA wins overall, but the paper finds no single
parameter-efficient method dominates every task — there's still a choice to make.
- **"Matches 32B" is on FuzzyBench-style tasks.** The parity is measured on the fuzzy-function
distribution PAW is built for, not a claim that a 0.6B model equals a 32B model in general.
## The take
What I like about PAW is that it's a *systems* idea wearing an ML paper's clothes. The interesting move
isn't a new architecture — it's noticing that we've quietly turned every fuzzy function into a recurring
API bill, and that the compile/runtime split we use for ordinary code applies here too: pay the big
model once to *build the tool*, then run the tool locally forever. The Text-to-LoRA mapper is the clever
engineering that makes "spec → weights" tractable, but the reframing is the contribution — a foundation
model as a **compiler for behaviors** rather than an always-on oracle. Whether it generalizes past
single-step functions is the open question; as a way to make a hundred small fuzzy tasks local,
private, and nearly free, it's one of the freshest ideas of the season. Code, the 4B compiler, and the
10M-example FuzzyBench are released (CC BY 4.0).
---
*Built on [Program-as-Weights: A Programming Paradigm for Fuzzy Functions](https://arxiv.org/abs/2607.02512)
(Zhang, Hotsko, Kim, Nie, Shieber, Deng; University of Waterloo, Cornell, Harvard; 2026, CC BY 4.0).
Figures are reproduced from the paper; benchmark and efficiency figures (0.6B ≈ 32B, ~1/50 memory, 30
tok/s on an M3) are quoted from it. The interactive playground and cost model are illustrations of the
mechanism, not runs of the released model.*
---
# TabFM: a foundation model that learns tables in-context
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/tabular-foundation-model
> date: 2026-07-04
> tags: tabular, foundation-models, in-context-learning, explainer, transformers
For a decade the honest answer to "what should I use on this spreadsheet?" has been a
gradient-boosted tree — XGBoost or LightGBM — fit fresh on each dataset. Deep learning kept
losing to it on tabular data. **TabFM**, Google's tabular foundation model, is a bet that the
transformer recipe finally transfers, but not by training a network on *your* table. Instead it
does **in-context learning**: you hand it your labeled rows as *context*, and it predicts new rows
in a single forward pass — no gradient descent on your data at all. The knowledge to do this was
baked in ahead of time, by pre-training on **hundreds of millions of synthetic tables**.
The cleanest way to feel what that means is to watch it. Feed the model labeled example rows and
it answers a test row; add rows and the prediction firms up; take them all away and it can only
guess. Crucially, **no training happens** at any point below — the "training set" is just context:
If that shape looks familiar, it's because TabFM descends from **TabPFN** and **TabICL** — the
line of *prior-fitted* / in-context tabular transformers that reframed prediction as inference over
a prompt rather than a fit loop. TabFM scales that idea up: a bigger model, a richer synthetic
prior, and an architecture built specifically for the two things tables have that text doesn't —
columns with no natural order, and rows that are exchangeable examples.
## The synthetic prior: learning to predict, from causal make-believe
TabFM never sees your data in training. It is trained on tables **generated from structural causal
models (SCMs)** — random causal graphs that emit synthetic features and a label with a known
generative structure. Sample hundreds of millions of such tables, each a fresh little prediction
problem, and train one transformer to solve all of them *in-context*. What the model learns is not
any particular dataset but the **algorithm** of tabular prediction: given some labeled rows, infer
the rule and apply it to a new row. Meeting your real table at inference time is then just another
draw from a distribution it has already seen a hundred million cousins of. (If synthetic-data
priors interest you, the same generate-to-learn logic shows up in [set diffusion](/articles/set-diffusion).)
## The architecture, in three stages
TabFM's body is a pipeline that respects tabular structure in three moves. Step through them:
1. **Column attention.** A **Set Transformer** attends *across the features* of a row. Because a
set has no order, the model is **permutation-invariant** over columns — shuffle your feature
order and nothing changes. Numeric values, which transformers otherwise handle poorly, are
embedded with **Fourier features** (sin/cos of the value at several frequencies) so the network
can represent magnitude and periodicity.
2. **Row compression.** Each row, now a set of attended feature embeddings, is pooled down to a
single **CLS token** — one vector per row. **RoPE** positional encoding gives the rows an order
to work with, so the sequence of rows is something the next stage can index into.
3. **In-context transformer.** A **24-block** transformer reads the sequence of row tokens — the
labeled context rows *and* the test row — and does the actual in-context learning, emitting the
test row's predicted label. This is the stage that "learns" your dataset, and it does so with
attention, in one pass, weights frozen.
The paper's own diagram lays out the same three stages — the alternating row/column attention that
feeds row compression, then the in-context stack that predicts the missing label:
## In-context vs fit-a-tree: the shape of the work
It helps to set TabFM beside the thing it wants to replace. A boosted tree has to *fit* your table
first — a sequential loop of boosting rounds — before it can predict. TabFM has no fit loop for
your data at all; the table enters as context and the answer comes out of one forward pass. Drag
the rounds and watch the sequential work pile up on one side and vanish on the other:
This is a claim about **workflow**, not accuracy. "No training loop" is a real ergonomic win — you
stand up a predictor in one pass — but whether it *predicts better* than a tuned tree is a separate
question, and one worth being careful about.
## The numbers, and what they actually compare
TabFM is evaluated on **TabArena**, a leaderboard that scores tabular methods with a **relative Elo**
across a suite of datasets. On classification, TabFM lands at **1727** Elo, and a small ensemble of
TabFM runs (**TabFM-Ensemble**) reaches **1815** — ahead of the strongest AutoGluon configuration
and the TabPFN/TabICL baselines:
The regression picture is wider: TabFM at **1940**, the ensemble at **2125**, clear of the field.
Here is the load-bearing caveat. The banner "beats GBDTs" rests on beating **AutoGluon** — an
*AutoML ensemble* that stacks and tunes many models — not a single tuned XGBoost or LightGBM. In
the headline chart there is **no standalone GBDT bar**; the strongest tree-based entries are the
AutoGluon configurations. So the honest framing is **ensemble-vs-ensemble**: TabFM-Ensemble edges
AutoGluon, and single TabFM is competitive with it. That is a genuinely strong result — but it is
not "TabFM beats your tuned XGBoost," which the chart does not measure. The paper's full TabArena
chart shows exactly this — AutoGluon as the tree-side comparison, TabFM and its ensemble on top:
**Read the fine print before you reach for it.** (1) **License:** the released weights are
**non-commercial** — fine for research, not for shipping a product. (2) **Ensemble-vs-ensemble:**
the "beats GBDTs" comparison is against AutoGluon's AutoML ensemble; there is no standalone
XGBoost/LightGBM bar in the headline chart. (3) **~10-class ceiling:** the classifier handles up to
about 10 classes. (4) **~500-feature ceiling:** it does not scale to very wide tables. (5) **Elo is
relative** — a ranking on TabArena's specific dataset suite, not an absolute accuracy guarantee on
*your* data. Validate on a held-out slice of your own table before trusting it.
## The take
TabFM is the tabular field finally getting its "just add context" moment. The mechanism is worth
internalizing even if the weights' license keeps you from deploying them: prediction reframed as
**in-context inference** over a synthetic prior, with an architecture that takes tables seriously —
order-invariant column attention, Fourier-embedded numerics, row-to-CLS compression, and a
24-block stack that does the learning at inference time. It inherits this from
[TabPFN and TabICL](/articles/how-transformers-attention-works) and pushes the scale, and the
TabArena numbers say the bet largely pays off: state-of-the-art *among ensembles*, and a real
challenger to the AutoML pipelines that have owned tabular ML.
What it is not, yet, is a drop-in XGBoost replacement for production. The non-commercial license,
the ~10-class and ~500-feature ceilings, and the ensemble-framed comparison all matter, and the
Elo is a leaderboard ranking, not a promise about your dataset. But as a research artifact it moves
the frontier: the fit-a-tree-per-dataset era now has a serious foundation-model rival, and the
interesting question is no longer "can transformers do tables" but "how far does one forward pass
go before you still need to train."
---
*Built on Google's **TabFM** — the [Hugging Face model card](https://huggingface.co/google/tabfm)
and Google's accompanying [blog post](https://research.google/blog/). There is no primary arXiv
paper; figures are from the model card, and the interactive diagrams are our own illustrations of
the mechanism. TabFM descends from [TabPFN](https://arxiv.org/abs/2207.01848) and TabICL; Elo scores
are quoted from TabArena.*
---
# Leanstral: proving theorems by being a code agent, not a prover
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/leanstral-formal-proofs
> date: 2026-07-03
> tags: llm, agents, reinforcement-learning, formal-methods, explainer
Formal theorem proving is the rare corner of ML where correctness is not a vibe: a proof either
type-checks against Lean's kernel or it doesn't. The catch is that the strongest Lean systems tend to
win with *machinery around the model* — Seed-Prover runs conjecture proposing and lemma pools at a
budget of ~10 H20-days **per problem**; Goedel-Architect generates a blueprint and proves lemmas in
parallel. Powerful, but none of it is the interface a person actually uses to write Lean.
Mistral's **Leanstral 1.5** makes the opposite bet. It's a **119B-parameter Mixture-of-Experts** that
activates just **6B** per token, and it proves theorems by doing exactly what a developer does in
[Mistral Vibe](https://mistral.ai): edit files, run shell commands, and query the Lean language server
for goals, types, and errors — then revise, and keep going. No prover scaffold. Its test-time scaling
is nothing more exotic than *running that loop longer*. And it works: it **saturates miniF2F** (100% on
validation and test), solves **587/672 PutnamBench**, and sets a new open state-of-the-art of **87 on
FATE-H** and **34 on FATE-X** — while being Apache-2.0.
## The loop is the whole method
There's no separate "prover mode." A theorem — or a repository with a missing proof — enters the same
agent harness used for ordinary software engineering, and the model works it turn by turn:
Mistral's own diagram of the loop shows the same thing at the token level: the system prompt hands the
model `lean-lsp-mcp` tools, and from there it's assistant turns interleaving ``, tool calls, and
tool results — the agent editing a Lean file and checking it exactly as it would edit and test any other
codebase, through Vibe's raw filesystem / bash / MCP surface.
Why insist on the ordinary loop? Two reasons. It makes the model **usable** — you point it at a Lean repo
the way you'd point a coding assistant at any repo — and, more importantly, it means the model is
**trained in the same long-horizon interaction pattern it uses at inference**. There's no train/test
interface gap to paper over.
## Test-time scaling, without a trick
Because the model just spends tokens inside that loop, its accuracy is a smooth function of the
**per-attempt token budget**. This is the headline result, and it's worth playing with — drag the budget
and watch PutnamBench solved climb:
That monotonic climb — 44 solved at 50k tokens, 244 at 200k, 493 at 1M, 587 at 4M — is Leanstral's whole
argument in one curve. Rather than giving up when a proof runs long, it keeps reasoning, editing, and
compacting context, turning budget directly into solved theorems. Here's Mistral's version of the same
figure:
The economics are the striking part. On PutnamBench, Leanstral edges Seed-Prover 1.5's high setting by 7
problems (587 vs 580) at roughly **$4 per problem** against an estimated **$300+** for Seed-Prover, whose
high setting budgets ~10 H20-days per problem. The only systems above it run under different rules —
natural-language proof guidance, or a much larger cost like Aleph Prover at $54–68 per problem.
## Why you can't fake a proof
Turning compiler feedback into an RL reward is dangerous: if the objective is just "make Lean stop
complaining," the cheapest policies are to *cheat* — leave a `sorry`, call the unsound `native_decide`,
assume an extra `axiom`, or loosen the checker with `set_option`. Leanstral is graded by a fork of
**SafeVerify**, which is built to reject every one of those. Full reward requires a proof that compiles,
uses **only standard Lean axioms** (checked via `#print axioms`), and took no shortcut. Toggle the cheats
and watch the verdict flip:
This adversarial verifier is what makes the reinforcement learning honest. Leanstral trains on two RL
environments through a **CISPO** objective: a **multiturn** environment where it must prove *or disprove*
a theorem, getting Lean compiler feedback between attempts until it succeeds or runs out of budget; and a
**code-agent** environment where it acts as a developer across a whole repository. Mistral's diagram of
the multiturn loop is the picture of a verifier-gated reward — the same instinct as
[Agents-A1's verifier-graded RL](/articles/agents-a1), specialized to Lean:
## The numbers
On competition math, Leanstral is the best *open* result across the board — the caveat being that a
couple of systems above it on PutnamBench either use natural-language guidance or cost 10–75× more to run:
The result that best captures the "keep working the problem" behaviour is **FLTEval** — proof-engineering
tasks drawn from real pull requests to the Fermat's Last Theorem repository. Here Leanstral 1.5 tops even
frontier general models, at a fraction of the cost:
Mistral's own charts show the full comparison set — the three-benchmark bar chart (PutnamBench / FATE-H /
FATE-X) and the FLTEval scaling across pass@1/2/4, where 1.5 pulls clearly ahead of much larger
open models:
## It generalizes past math
Because the skill it learned is *proof engineering in a repository*, not competition-problem pattern
matching, it transfers to code verification:
- **AVL trees.** Leanstral proved the `O(log n)` time-complexity guarantees for a real self-balancing-tree
implementation — structural induction mirroring the recursion, unfolding a `TimeM` monad to expose the
step counts, exhaustive rebalancing-case analysis. It ran **over 2.7 million tokens across 22 context
compactions** to close it: the test-time-scaling curve in practice.
- **Finding real bugs.** In an automated pipeline — Aeneas translates Rust to Lean, Leanstral infers the
intended properties and tries to prove each (or disprove it) — across **57 repositories** it flagged 47
violated properties, **11 genuine bugs, 5 previously unreported**. One was an integer overflow in
`datrs/varinteger`'s zigzag decode: on `U64.MAX`, `value + 1` overflows — a crash in debug, silent
corruption in release, exactly the edge case fuzzing tends to miss.
## The take
Leanstral's contribution isn't a clever proof-search algorithm; it's a stance. The formal-proving
leaderboard has been climbing by wrapping models in ever-more-specialized test-time scaffolds, and
Leanstral shows that a plain code agent — trained in the loop it's evaluated in, graded by a verifier it
can't cheat — matches or beats that machinery at a fraction of the cost, while staying usable by anyone
who can drive a coding assistant. The MoE is efficient (6B active), the license is open (Apache-2.0), and
the scaling story is the honest kind: no plateau trick, just compute converted into checked proofs. The
open question is how far "just run the agent longer" goes as problems get genuinely harder than Putnam —
but as a demonstration that formal verification can be *practical* infrastructure rather than a
prover-lab specialty, it's the most encouraging Lean release in a while.
---
*Built on the [Leanstral technical report](https://github.com/mistralai/LeanstralSafeVerify/blob/main/LeanstralReport.pdf)
and [Mistral's Leanstral 1.5 announcement](https://mistral.ai/news/leanstral-1-5/) (Mistral AI, 2026;
Apache-2.0, weights on Hugging Face). Figures are reproduced from Mistral's report and blog post; benchmark
and cost figures are quoted from those sources.*
---
# MiMo-V2-Flash: a 128-token window and one global layer in six
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mimo-v2-flash
> date: 2026-07-03
> tags: llm, mixture-of-experts, attention, agents, explainer
Xiaomi's **MiMo-V2-Flash** is, on the numbers, the **strongest open-source model for software
engineering** — 73.4 on SWE-Bench Verified, and the top score in its comparison table on
SWE-Bench Multilingual. It's a **309B-parameter Mixture-of-Experts** that activates just **15B**
per token. But the part worth an article isn't the leaderboard; it's the **attention design** that
lets a 309B model serve a 256K context quickly. MiMo doesn't use full attention, and it doesn't use
[linear attention like HydraHead](/articles/hydrahead) either. It interleaves **sliding-window** and
**global** layers — with a window so small it's almost provocative.
## The 5:1 hybrid, and a 128-token window
The backbone is 48 layers arranged into **8 hybrid blocks**, each block being **5 sliding-window
attention (SWA) layers followed by 1 global attention (GA) layer** — a 5:1 ratio. Each SWA layer
attends only to the previous **128 tokens**. That window is tiny on purpose; the report frames it as
an *inductive bias*: "smaller windows force the model to focus on local context... mitigat[ing]
overfitting." Two things keep 128 from being crippling:
- **Stacking compounds reach.** Like a stack of small convolutions, five SWA layers see far more than
128 tokens — each layer's window sits on top of the last, so information propagates several windows
back before you ever hit a global layer.
- **One global layer per block.** Every sixth layer attends to the *entire* sequence, mixing whatever
the local layers couldn't reach directly. Full context is restored periodically, at a sixth of the
cost of making every layer global.
A few architecture details worth having right: attention is **grouped-query (GQA), not MLA** — SWA
layers use 64 query / 8 KV heads, GA layers 64 / 4 — with **partial RoPE** on the first 64 dims and a
**learnable attention-sink bias** that the report credits with much of the hybrid's stability. The
MoE is **256 experts, top-8, no shared expert**; only the first block runs a dense FFN.
## Why the window is the point
The reason to accept a 128-token window is the **KV cache**. In a full-attention model, every layer's
cache grows with the context length. In MiMo, the five SWA layers per block stop growing once the
context passes 128 tokens — only the 1-in-6 global layers keep scaling. So at long context the cache
is dominated by that one-sixth:
The report claims "nearly a **6× reduction** in KV-cache storage and attention computation for long
contexts," and the interactive shows exactly where it comes from — the SWA layers flatten, the global
layers carry the linear term. This is the same economics as the
[KV-cache story in inference](/articles/how-llm-inference-works): the cache is what caps concurrency
and context, so shrinking it 6× is what makes a 309B model cheap to serve at 256K. And crucially,
long-context *quality* holds — MiMo posts the best open-source LongBench V2 score and near-perfect
needle retrieval (96.7% at 256K), evidence the hybrid isn't paying for its speed with reach.
## Multi-token prediction, for speed
MiMo also ships with **multi-token prediction** built in. It trains one MTP head during pretraining,
then replicates it into a **3-layer MTP module** for inference — a built-in draft model for
**self-speculative decoding**. The reported acceptance length reaches ~3.6 tokens and the measured
speedup is up to **2.6× decoding** (2.70× at batch size 96). If you want the mechanism, the
[multi-token prediction write-up](/articles/multi-token-prediction) covers exactly this
predict-several-verify-in-one-pass lineage; MiMo is a clean production instance of it.
## How it was trained
Pretraining is **27T tokens** in FP8 with the MTP objective, native 32K context, over three stages
(general → code-heavy with synthetic reasoning → context extension to 256K via RoPE base rescaling,
not YaRN). The post-training is the notable bit: it uses the **MOPD recipe** — multi-teacher on-policy
distillation, the same [MOPD from the recent arXiv digest](/arxiv/2026-06-30) that
[Agents-A1](/articles/agents-a1) also builds on — in three stages: SFT, domain-specialized RL, then
distilling several domain teachers into the student on its own rollouts. It's another data point that
on-policy multi-teacher distillation is becoming the default way to fuse agentic capabilities into one
deployable model.
## The numbers
The headline is agentic coding. On **SWE-Bench Multilingual** — resolving real GitHub issues across
languages — MiMo tops its entire comparison table, closed models included:
On **LiveCodeBench-v6** it's the best open model and edges GPT-5-High, trailing only Gemini-3.0-Pro:
And the one that validates the whole attention bet — **LongBench V2**, where the sliding-window hybrid
could have hurt but instead lands best-open, a whisker behind Claude:
It's not uniformly frontier — MiMo trails the reasoning leaders on HMMT (84.4), long-context MRCR
(45.7), and Humanity's Last Exam (22.1), and closed models still edge single-language SWE-Bench
Verified. The fair summary is: **the best open model for agentic software engineering and
long-context**, competitive with closed frontier systems on those axes, from a design that's cheaper
to serve. (The "150 tokens/second" and per-token pricing you'll see quoted are from Xiaomi's product
page, not the peer-reported tech report — treat them as marketing, not measurements.)
## The take
MiMo-V2-Flash is a satisfying piece of systems thinking more than a scaling flex. The bet is that most
of what attention does is *local* — so make most layers local and cheap (a 128-token window is an
aggressive way to commit to that), and buy back global reach with one full-attention layer per block
and a periodic reset. Stack MTP on top for decode speed and MOPD for capability, and you get a 309B
model that serves 256K context at a fraction of a full-attention model's KV cost while topping the
open-source agentic-coding board.
The through-line with [HydraHead](/articles/hydrahead) is worth noticing: both argue you shouldn't pay
full-attention cost on every layer, they just cut the budget differently — HydraHead per *head* by
interpretability, MiMo per *layer* on a fixed 5:1 schedule. The schedule is blunter, but it's simple,
it's proven at 309B, and the KV-cache math is undeniable. Open weights (Apache-2.0), MTP included — a
strong, honest release.
---
*Built on the [MiMo-V2-Flash Technical Report](https://arxiv.org/abs/2601.02780) (Xiaomi LLM-Core,
2025) and the [model release](https://github.com/XiaomiMiMo/MiMo-V2-Flash) (Apache-2.0, open weights +
3-layer MTP). Benchmark figures are quoted from the report's tables; throughput and pricing figures
are from Xiaomi's product page.*
---
# Set Diffusion: one knob from autoregression to diffusion
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/set-diffusion
> date: 2026-07-03
> tags: llm, diffusion, language-models, inference-optimization, explainer
There's a spectrum hiding behind two families of language model that usually get treated as
opposites. **Autoregressive** models generate strictly left-to-right, one token per step —
sequential, but they support the [KV cache](/articles/how-llm-inference-works) that makes serving
cheap. **Diffusion** LMs ([like iLLaDA](/articles/illada-diffusion-language-model)) denoise many
tokens in parallel and in any order — flexible, but fixed-length, and they *can't* cache (every step
needs full bidirectional context). **Set Diffusion** (Arriola & Kuleshov, Cornell — the block
diffusion authors) makes the spectrum explicit and shows you can slide along it with a single idea:
change *which token sets you decode together, and in what order.*
## The spectrum, precisely
The prior bridge between these worlds was **block diffusion** (BD3-LM): generate fixed-size
contiguous blocks left-to-right, with diffusion inside each block. That buys variable length and a
per-block KV cache, but the block is rigid — it can only extend left-to-right, so no infilling, and
the cache can only update once a whole block finishes decoding (the within-block denoising needs
bidirectional context until then).
Set Diffusion's move is to stop thinking in blocks and think in **sets**. A *set* is an
arbitrary-position, arbitrary-length subset of token positions you decode together. Factorize the
likelihood over a sequence of disjoint sets that cover the whole sequence, and the familiar models
fall out as special cases:
- **Autoregression** — every set is a single token, in left-to-right order.
- **Order-agnostic diffusion** — one set of the whole sequence, decoded in random order.
- **Block diffusion** — fixed contiguous blocks, left-to-right.
They're not different architectures; they're different **set schedules** for the same object. That's
the whole conceptual payoff — and once you see it, the interesting question is how to pick a schedule
*between* the corners.
## Two knobs, one window
Set diffusion exposes two knobs the block-size parameter conflates: **set size** (how many tokens you
commit per step — parallelism) and **ordering bias** (how left-to-right you stay — quality). In
practice both are controlled by one schedule parameter, a window width `w`: each position gets an
active generation window of width `w`, and the widths determine how much decoding overlaps.
The paper makes the endpoints rigorous. As `w → 1/L` the windows stop overlapping, tokens generate
one at a time in order, and the training objective *becomes the tight autoregressive ELBO* — best
perplexity, no parallelism. As `w → 1` every position shares one schedule and you recover
order-agnostic diffusion — maximally parallel and any-order. Set diffusion lives in between: a sliding
window that decodes a few tokens per step, mostly in order but flexible enough to fill gaps. Smaller
`w` buys perplexity; larger `w` buys parallelism and any-order decoding. One dial, the whole spectrum.
## Why sets get to keep the KV cache
The systems win is that generation is **set-causal**: each set attends to itself and to all
*previously decoded* sets, but not to future ones. Because the ordering across sets is causal,
finished sets never need reprocessing — their keys and values are cached and reused, and the cache
**updates after every inference step**. That's the thing pure diffusion can't do (it needs full
bidirectional context, so nothing is ever "final" enough to cache) and the thing block diffusion does
only *per block* (bidirectional context *within* the block blocks earlier caching). The ablation is
stark: turn the causal mask and KV caching off and GSM8K accuracy collapses from **26.6 to 6.4** while
throughput drops too — the causal set structure is buying both.
The flexibility also gives **infilling** for free. Because sets are flexible-position, the schedule can
select gap tokens and condition them on the clean tokens on *both* sides — something block diffusion's
strict left-to-right blocks structurally cannot do.
## The numbers
At GPT-2-small scale (110M params), the headline is that set diffusion beats block diffusion on *both*
axes at once — accuracy and speed. On GSM8K it tops the whole diffusion field:
— and it does so at higher throughput than any of those block-diffusion settings (60.4 vs 55.4
tok/s). It won't be lost on you that an ordinary AR transformer scores higher still (75.7) and, at this
small scale with full diffusion sampling steps, is even a bit faster; the point of a diffusion LM isn't
to beat AR on a left-to-right benchmark, it's to keep AR's efficiency *while* offering any-order
generation. Which is exactly where the second result lands — **infilling**, filling a gap given text on
both sides, where set diffusion clearly beats block diffusion:
At ~25% faster decoding, no less. Across the rest of the suite it's the same shape: on OpenWebText
it matches block diffusion's perplexity at **22% higher throughput** (and runs ~13× faster than
cacheless MDLM), on LM1B it posts the best diffusion perplexity *and* the highest diffusion throughput,
and on CNN/DailyMail it's competitive on ROUGE at up to 10% faster. A strictly better speed-quality
frontier than block diffusion, plus the infilling block diffusion gives up.
## The take
What I like about Set Diffusion is that it's a *reframing* that pays off, not a new mechanism bolted
on. "Interpolate between AR and diffusion by varying block size" was already a good idea (that's block
diffusion); the insight here is that block size was the wrong knob — the right one is the **order in
which token sets are generated**, and once you factorize over flexible sets instead of rigid blocks you
get a strictly larger design space that still contains AR, still contains diffusion, and adds a
KV-cacheable, any-order, infilling-capable middle that block diffusion couldn't reach.
The honest caveats: everything is at 110M parameters, an AR model still wins the straight
left-to-right benchmarks, and the ideal window schedule is currently hand-tuned (learning it is future
work). But as a clean statement of *what the AR↔diffusion spectrum actually is*, and a practical model
that sits usefully in its middle, it's the most satisfying diffusion-LM paper I've read since block
diffusion itself — which makes sense, given it's the same group closing the loop on their own idea.
---
*Built on [Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion](https://arxiv.org/abs/2607.01775)
(Marianne Arriola, Volodymyr Kuleshov; Cornell, ICML 2026), which generalizes the same authors'
[Block Diffusion](https://arxiv.org/abs/2503.09573). Code and weights are at
[kuleshov-group/setdlms](https://github.com/kuleshov-group/setdlms). Benchmark figures are quoted from
the paper's tables (110M-parameter models).*
---
# BM25: the ranking function that refuses to die
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/bm25
> date: 2026-07-02
> tags: information-retrieval, search, ranking, algorithms, explainer
Here's a fact that should be more embarrassing for machine learning than it is: a ranking
function designed in the early 1990s, with **no learned parameters** and no notion of meaning,
is still the default text scorer in Lucene, Elasticsearch, and OpenSearch — and still a
baseline that dense neural retrievers regularly *fail to beat* out of domain. It's called
**BM25**, and if you do search, retrieval, or RAG, it's worth understanding exactly, because
you're almost certainly running it.
"BM25" is *Best Matching 25* — roughly the 25th matching function tried in the Okapi
information-retrieval project at City University London (Robertson, Spärck Jones, and
colleagues), grounded in the probabilistic relevance framework. Strip away the theory and it's
a small, honest idea: score a document by summing, over the query terms it contains, how
*rare* each term is times how *emphatically* the document uses it — with two corrections that
make it work in practice. Let's build it.
## Start with TF-IDF, and its two flaws
The classic bag-of-words score is **TF-IDF**: for each query term, multiply its frequency in
the document (**term frequency**, TF — "this document uses the word a lot") by its rarity
across the corpus (**inverse document frequency**, IDF — "and the word is distinctive"). Sum
over query terms. It's a reasonable instinct, and it has two clear problems:
1. **TF grows linearly and forever.** A document that says "jaguar" 40 times scores 40× one
that says it once. But the difference between 1 and 2 mentions is meaningful; the difference
between 39 and 40 is noise. Relevance *saturates*.
2. **Length is unaccounted for.** A 10,000-word document will rack up term counts just by being
long, drowning out a tight 100-word document that's genuinely more on-topic.
BM25 is, almost exactly, **TF-IDF with those two flaws fixed.** Here's the whole thing:
$$
\text{score}(D, Q) = \sum_{q \in Q} \text{IDF}(q)\cdot\frac{f(q, D)\,(k_1 + 1)}{f(q, D) + k_1\left(1 - b + b\,\dfrac{|D|}{\text{avgdl}}\right)}
$$
Three pieces do all the work: the **IDF** weight, the **saturating** term-frequency factor
(the $k_1$ part), and the **length normalization** (the $b$ part). Here they are, color-coded —
the map for the rest of the piece:
Take them one at a time.
## Rarity: the IDF term
$f(q,D)$ is just the count of term $q$ in document $D$. The multiplier in front is inverse
document frequency — how much a term's presence should count, based on how rare it is:
$$
\text{IDF}(q) = \ln\!\left(1 + \frac{N - n(q) + 0.5}{n(q) + 0.5}\right)
$$
$N$ is the number of documents, $n(q)$ the number containing $q$. A term in every document
(like "the") gets an IDF near zero — matching it tells you nothing. A term in one document out
of a million gets a large IDF — matching it is almost the whole story. This is why BM25 needs
no stopword list: common words are down-weighted *automatically* because their IDF collapses.
(The $\ln(1 + \cdots)$ form is Lucene's; it keeps IDF non-negative. The original
Robertson–Spärck-Jones IDF drops the $+1$ and can go slightly negative for terms in more than
half the corpus.)
## Saturation: the k1 term
Now the first fix. Instead of using the raw count $f(q,D)$, BM25 passes it through a saturating
function $\frac{f\,(k_1+1)}{f + k_1\cdot(\dots)}$ that rises fast for the first few occurrences
and then flattens toward an asymptote. The parameter $k_1$ controls how fast:
The first mention of a term is strong evidence; each additional mention adds less. That curve —
not the straight line of TF-IDF — is how relevance actually behaves. Set $k_1 = 0$ and it
becomes binary (any occurrence counts the same); crank $k_1$ up and it straightens back toward
linear TF. Lucene's default is **$k_1 = 1.2$**.
## Length: the b term
The second fix lives in the denominator: the $k_1$ term is scaled by
$\left(1 - b + b\,\frac{|D|}{\text{avgdl}}\right)$, where $|D|$ is the document's length and
$\text{avgdl}$ the average across the corpus. A longer-than-average document gets a bigger
denominator, so its term-frequency factor is discounted — the same two mentions count for less
when they're diluted across more text:
The knob $b \in [0,1]$ sets how aggressively. At $b = 0$ length is ignored entirely; at $b = 1$
it's fully normalized. The default **$b = 0.75$** is a compromise that's proven hard to beat.
## Put it together: a live BM25 engine
That's the entire algorithm. Here it is running on a tiny corpus — edit the query, drag $k_1$
and $b$, and every document is re-scored with the exact formula above. Expand a document to see
each query term's contribution (its IDF times the saturated, length-normalized factor):
Notice the behaviors fall out on their own: rare query terms dominate the ranking, repeating a
common word barely moves the score, and a long document doesn't win just for being long. No
training, no embeddings — just term statistics arranged sensibly.
## Thirty lines of Python
There's no magic hiding in a library. The whole thing is a couple of counters and the formula:
```python
import math
from collections import Counter
class BM25:
def __init__(self, corpus, k1=1.2, b=0.75):
self.k1, self.b = k1, b
self.docs = [doc.lower().split() for doc in corpus]
self.N = len(self.docs)
self.avgdl = sum(len(d) for d in self.docs) / self.N
self.tf = [Counter(d) for d in self.docs] # term counts per doc
self.df = Counter() # docs containing each term
for d in self.docs:
for term in set(d):
self.df[term] += 1
def idf(self, term):
n = self.df.get(term, 0)
return math.log(1 + (self.N - n + 0.5) / (n + 0.5))
def score(self, query, i):
d, tf = self.docs[i], self.tf[i]
norm = self.k1 * (1 - self.b + self.b * len(d) / self.avgdl)
s = 0.0
for term in query.lower().split():
f = tf.get(term, 0)
if f:
s += self.idf(term) * f * (self.k1 + 1) / (f + norm)
return s
def rank(self, query):
return sorted(((i, self.score(query, i)) for i in range(self.N)),
key=lambda x: -x[1])
```
In production you don't loop over every document — you keep an **inverted index** (term →
postings list of documents that contain it) and only score documents that share a term with the
query. That's what makes BM25 fast enough to serve web-scale corpora on commodity hardware, and
it's the same index that's been powering Lucene since 2011.
## The variants you'll meet
The core formula spawned a small family, mostly patching edge cases:
- **BM25+** adds a small constant $\delta$ (default 1.0) to the term-frequency factor, fixing a
subtle bug where very long documents can be over-penalized to the point that a document
*containing* a rare term scores below one that doesn't.
- **BM25F** ("fielded") scores structured documents — title, body, anchor text — by combining
per-field term frequencies *before* saturation, with a weight per field, so a title match
counts more than a body match. It's what real search engines actually run.
- **BM25L** re-weights to stop long documents from being unfairly buried.
- Lucene's implementation is BM25 with the non-negative IDF above and per-field length norms
quantized into a single byte — the pragmatic engineering version of the equation.
## Why it won't die
Neural retrieval was supposed to make this obsolete years ago. It hasn't, for reasons worth
naming:
- **It's a brutal baseline.** On out-of-domain benchmarks (the BEIR suite made this famous),
BM25 beats or ties many dense retrievers — because it can match *any* term, including names,
codes, and jargon a fixed-vocabulary embedding never saw in training. It never has an
"out-of-distribution" moment.
- **It's the sparse half of hybrid search.** The current default in serious systems is to run
BM25 *and* a dense retriever and fuse the results (often with reciprocal-rank fusion). Lexical
precision plus semantic recall beats either alone, which is why "BM25 is obsolete" quietly
became "BM25 is one of your two retrievers."
- **It's cheap and interpretable.** No GPU, no training, no embedding drift. When it ranks a
document highly you can point at exactly which rare terms did it — which matters when a
ranking has to be debugged or defended.
The lesson I take from BM25 is that a model doesn't have to learn anything to encode real
knowledge about a problem. Every piece of it is a hypothesis about relevance — rarity matters,
repetition saturates, length dilutes — written as arithmetic instead of learned from data. Three
good hypotheses, two tunable knobs, and thirty-five years later it's still the thing your search
bar is probably running.
---
*BM25 originates in the Okapi project (Stephen Robertson, Karen Spärck Jones, et al.) and the
probabilistic relevance framework; the formulation and non-negative IDF here follow
[Lucene's `BM25Similarity`](https://lucene.apache.org/core/9_9_1/core/org/apache/lucene/search/similarities/BM25Similarity.html)
(defaults $k_1 = 1.2$, $b = 0.75$). BM25+ / BM25L are from Lv & Zhai (2011).*
---
# TwoTower: giving a diffusion LM a frozen autoregressive memory
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nemotron-twotower
> date: 2026-07-02
> tags: llm, diffusion, inference-optimization, architecture, explainer
An autoregressive (AR) language model emits one token per forward pass — the sequential axis
is the entire sequence, which is exactly why decode is the slow, [memory-bound half of
inference](/articles/how-llm-inference-works). **Diffusion language models** ([like
iLLaDA](/articles/illada-diffusion-language-model)) offer the escape: denoise many tokens per
step and refine iteratively, so generation can be *parallel*. The catch is that every diffusion
LM so far has made one network do two jobs at once — and those jobs pull in opposite directions.
NVIDIA's **Nemotron-Labs-TwoTower** fixes that by refusing to share. It splits the model into
two towers: a **frozen autoregressive context tower** and a **trainable diffusion denoiser
tower** that reads from it. Built on a 30B Mamba-Transformer MoE, it retains **98.7%** of the
autoregressive baseline's quality while generating **2.42× faster**.
## One network, two jobs that fight
Here's the tension existing diffusion LMs live with. At every denoising step, the same decoder
has to (1) *represent the clean tokens* already committed — which wants strong **causal**
processing, the thing AR pretraining is great at — and (2) *denoise the corrupted block* —
which wants **bidirectional** attention over the noisy tokens. As the paper puts it, this
"entanglement pulls the same set of weights in different directions, limiting their capacity to
excel at either." A single set of weights forced to be both a causal reader and a bidirectional
denoiser ends up mediocre at both.
TwoTower's move is to stop asking one network to be both:
- The **context tower** is the *frozen* pretrained AR model. It causally processes clean tokens
and never gets a gradient — so it keeps every bit of the 25T-token backbone's context ability
intact. It carries the persistent left-context (KV cache and Mamba states) across blocks.
- The **denoiser tower** is trained from the diffusion objective and does nothing but refine the
current noisy block with bidirectional attention. It reads the context through **layer-aligned
cross-attention**: denoiser layer *i* attends to context layer *i*, over both the frozen
tower's committed blocks and its own in-block tokens.
The base is `Nemotron-3-Nano-30B-A3B`, an open hybrid **Mamba-Transformer MoE** — 30B total,
~3B active, 52 layers (23 Mamba-2, 6 attention, 23 MoE). Cross-attending *into* a Mamba-hybrid
sounds awkward (Mamba is a recurrence, not a KV cache), and the trick is neat: the **Mamba chunk
size is matched to the diffusion block size**, so the existing kernel exposes clean recurrent
states exactly at block boundaries — right where the denoiser needs them.
## Block-wise autoregressive diffusion
TwoTower isn't fully parallel and isn't fully sequential — it's **autoregressive across blocks,
diffusion within a block**. Text is chunked into blocks (size **16** by default); blocks are
generated left-to-right, each conditioned on the finished ones, but *within* a block all tokens
are denoised together over a few steps:
The diffusion is masked/absorbing-state — the same LLaDA-style "replace tokens with `[MASK]`
and predict them back" as [iLLaDA](/articles/illada-diffusion-language-model), with a linear
noise schedule. The number of denoising steps is *adaptive*: a confidence sampler commits any
token whose prediction clears a threshold (γ = 0.8) immediately and lets the uncertain ones wait
another step. In practice most tokens of a block resolve in the **first** step, so a block costs
far fewer forward passes than its token count — which is the whole source of the speedup. The
sequential axis is now the number of *blocks*, not the number of *tokens*.
One sharp constraint falls out of this: you have to **sample with the same block size you
trained on**. Sample with blocks *larger* than training and generation collapses — GSM8K drops
from 89.8 to 2.2 at a sampling block of 64. The block size isn't a free inference knob; it's
baked in at training time.
## Why decoupling is the whole point
The paper's central experiment is an ablation that isolates the decoupling. Build the model
three ways from the same backbone and measure how much quality survives versus the AR baseline:
The entangled single tower — one network trained jointly for both roles — loses 21–26% across
general, code, and math. Continued AR training does better. But freezing the context tower and
training a *separate* denoiser keeps the most, losing only 6–11%. That gap is the argument:
neither role compromises the other when they don't share weights. It's the same instinct as
[HydraHead](/articles/hydrahead) — match the mechanism to the job — applied to whole towers
instead of individual heads.
## The numbers
The released checkpoint is a genuinely strong model in absolute terms — this isn't a toy that
trades away quality for speed:
Aggregated, that's **98.7%** of the autoregressive baseline's quality — the headline claim.
And the speed lever is the block size: bigger blocks mean more tokens denoised in parallel per
step, so higher throughput (the released checkpoint reaches **2.42×**):
A few honest caveats, because this is a fresh preprint and the framing invites them. The
quality "retention" is an aggregate; the per-category drops aren't uniform — **code (−10.5%) and
math (−11.3%)** take the biggest hits, exactly the tasks where a single wrong token derails the
answer. The paper reports **no comparison to other diffusion LMs** (LLaDA, Dream, external
block-diffusion) — only its own AR baseline and internal ablations — so "best diffusion LM"
is not a claim it makes or supports. Throughput is reported only as a relative speedup (no
tokens/sec), the 2.42× is the released checkpoint while the ablation tables show 2.02× at the
same block size under a different recipe, and running two towers means the frozen context
tower's weights sit resident in memory on top of the denoiser.
## The take
The appeal here is architectural honesty. Diffusion LMs have quietly been asking one network to
be a causal historian and a bidirectional editor simultaneously, and TwoTower's contribution is
mostly the observation that you shouldn't — plus the engineering to make cross-attention into a
frozen Mamba-hybrid actually work (the chunk-size-equals-block-size trick is the load-bearing
detail). Keeping the pretrained AR tower *frozen* is the elegant part: you inherit a 25T-token
backbone's context ability for free and spend all your training budget teaching the one thing
that's genuinely new, bidirectional block refinement.
Whether this is the design that finally makes diffusion decoding a default is still open — the
code/math gap is real, and 2.42× on two H100s with extra resident weights is a solid but not
seismic win. But "give the diffusion model a frozen autoregressive memory instead of making it
grow its own" is the kind of clean decomposition that tends to stick, and the weights are out
(CC BY 4.0) if you want to poke at it.
---
*Built on [Nemotron-Labs-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive
Context](https://arxiv.org/abs/2606.26493) (Reda, Kamalu, Waleffe, Patwary, Shoeybi, Catanzaro;
NVIDIA, 2026). Benchmark and throughput figures are quoted from the paper's tables; category-level
AR-vs-TwoTower comparisons are from its Figure 2.*
---
# TurboQuant: rotate first, then quantize the KV cache
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/turboquant-kv-cache
> date: 2026-07-02
> tags: llm, inference-optimization, quantization, kv-cache, explainer
The [KV cache runs the economics of LLM serving](/articles/how-llm-inference-works): it grows
linearly with context length, per layer, and it's what decides how many requests fit on a GPU.
The obvious lever is to **quantize** it — store the cached keys and values in 2–4 bits instead
of 16. The catch is that key and value vectors have **outlier coordinates**: a few dimensions
carry most of the magnitude, so a fixed set of quantization levels either clips the outliers or
wastes precision on the empty middle. The standard fix — calibrate a per-channel quantizer on
sample data — doesn't work *online*, where new tokens stream in and you have no calibration set.
**TurboQuant** (Zandieh, Daliri, Hadian, Mirrokni — Google Research, [arXiv 2504.19874](https://arxiv.org/abs/2504.19874))
solves this with one idea: **rotate the vector before you quantize it.**
## Why rotation is the whole trick
Multiply a vector by a random orthogonal matrix and you don't change its length or its inner
products with other (equally rotated) vectors — but you *do* spread its energy evenly across all
coordinates. After the rotation, the coordinates of a unit vector provably follow a known
**Beta distribution**, $f(x) \propto (1 - x^2)^{(d-3)/2}$ on $[-1, 1]$, which for large $d$
tightens to a Gaussian $\mathcal{N}(0, 1/d)$. Every coordinate of every rotated vector looks the
same, and there are no outliers left to clip.
That's what makes the quantizer **data-free**. Because you know the post-rotation distribution
analytically, you can solve for the optimal scalar quantizer (Lloyd–Max: the levels that minimize
expected squared error against that Beta density) *once, ahead of time*, and it's optimal for
every vector you'll ever see — no calibration pass, which is exactly what an online KV cache
needs. The paper proves this is **near-optimal**: TurboQuant's distortion is within a small
constant factor (about **2.7×**) of the information-theoretic lower bound for *any* vector
quantizer, at every bit-width and dimension.
## The bias nobody mentions
There's a subtlety the paper is unusually careful about. An MSE-optimal quantizer minimizes
reconstruction error — but attention doesn't reconstruct vectors, it takes **inner products**
(query · key) and softmaxes them. And an MSE-optimal quantizer is *biased* as an inner-product
estimator: on average its scores are systematically off, which quietly warps the attention
weights. TurboQuant fixes this with a two-stage scheme — quantize with the MSE quantizer, then
store the **sign of the residual** under a random projection (a 1-bit *Quantized Johnson–
Lindenstrauss* transform, from the same authors' earlier QJL work). That 1-bit correction is
exactly what cancels the bias, giving an **unbiased** inner-product estimate:
Unbiasedness is the part that lets quality hold at aggressive bit-widths: the paper reports
**absolute quality neutrality at 3.5 bits per channel**, and only marginal degradation at 2.5.
## What it costs, what it buys
An implementation ([`0xsero/turboquant`](https://github.com/0xsero/turboquant)) wires this into
vLLM and makes the tradeoff concrete. It allocates bits **asymmetrically** — keys get **3 bits**
with the unbiased inner-product quantizer (attention scores are precision-sensitive), values get
**2 or 4 bits** with simpler group quantization (value aggregation is more forgiving). The
reconstruction quality splits cleanly along that line:
Keys reconstruct essentially perfectly; 4-bit values are near-lossless; 2-bit values are where
the quality actually gives (0.94), which is why the implementation recommends 4-bit values for
anything sensitive. Net, it compresses the full-attention KV cache about **4.4×**, and that turns
straight into context length:
On the reported runs, that's **30 GB of KV cache freed** on a 4-GPU RTX 5090 box and a **2.0×**
jump in max context (457K → 914K tokens) for a dense model; a MoE model with linear-attention
layers gets less (**1.45×**), because those layers keep a recurrent state that doesn't compress.
Throughput barely moves (**+5.7%** prefill, **+3.1%** decode) — this is a *memory* win, not a
speed one, and the honest read is that the value comes from fitting longer contexts and bigger
batches, not faster tokens. The current build also still allocates a full cache during prefill
and only frees it afterward, and its hybrid decode path dequantizes history to fp32 each step —
real limitations the repo names outright.
## The same algorithm, three places
What makes TurboQuant worth an article isn't just the KV-cache result — it's that "rotate, then
quantize with a data-free optimal quantizer" is a **general** vector-quantization primitive, and
it's showing up in very different systems:
- **KV cache** (`0xsero/turboquant`, above): compress the attention cache in vLLM for longer
contexts.
- **Model weights** (`turbo-tan/llama.cpp-tq3`): a `TQ3` quantization type that applies the same
rotate-then-quantize idea to *weights* in llama.cpp (see below).
- **Vector search** ([`turbovec`](/articles/turbovec)): the same rotation + Lloyd–Max +
bit-packing, in Rust with SIMD kernels, as a FAISS-competitive similarity index — a 10M-vector
corpus in 4 GB instead of 31. (Its own write-up is [here](/articles/turbovec).)
One paper, one primitive — a quantizer whose optimality comes from *reshaping the data into a
known distribution first* rather than learning a codebook from samples — and it drops into
inference caches, weight files, and ANN indexes alike.
### TQ3 in llama.cpp: the same idea, on weights
The `llama.cpp-tq3` fork adds `TQ3_1S` / `TQ3_4S` — **3-bit weight** quantization types that run the
TurboQuant pipeline (a Walsh–Hadamard rotation, then Lloyd–Max scalar quantization per block) on
model weights. Worth clearing up a name collision: llama.cpp already ships `TQ1_0` and `TQ2_0`,
but those are *ternary* formats unrelated to this paper — the "TQ" match is coincidental. TQ3 is
genuinely TurboQuant-based, and at ~3.5 bits per weight it hits **Q4-class quality about 10%
smaller** (on Qwen3.5-27B, `TQ3_4S` measures a hair *better* perplexity than `Q3_K_S` at ~12.9 GiB),
which is what lets a 27B model run on a 16 GB GPU.
The speedup is the interesting engineering twist. Because TQ3 blocks are 3-bit-after-rotation, they
map cleanly onto **FP4 tensor cores** on Blackwell-class GPUs — the fork fuses the rotation into an
FP4 activation quantizer and runs the matmul in FP4. Turning that path on roughly **doubles
prompt-processing throughput**: on an RTX 3090, Gemma-12B goes 737 → **1,819 tok/s** (+147%) and
SuperGemma-26B 983 → **2,005** (+104%); on a DGX Spark (GB10), a 27B MTP model goes 360 → **920
tok/s** (+155%). That's a rotate-then-quantize weight format turning a hardware FP4 unit into free
speed — the same primitive, paying off a third way. (The exact `llama-quantize` invocation isn't
documented in the repo yet; the types ship as pre-quantized models on Hugging Face.)
## The take
The elegant part of TurboQuant is that the hard problem (outliers, calibration) is dissolved
rather than fought. Instead of detecting and special-casing outlier channels, you rotate them
away; instead of calibrating on data, you compute the optimal quantizer against the distribution
the rotation guarantees. The QJL residual is the tasteful finish — a one-bit patch that turns a
good reconstruction quantizer into an unbiased *inner-product* quantizer, which is the thing
attention actually needs.
It's not magic: the KV-cache implementation is a memory win, not a throughput one, 2-bit values
visibly degrade, and MoE/linear-attention models compress less. But the underlying result —
near-optimal, data-free vector quantization with a formal distortion bound — is the kind of solid
primitive that ends up everywhere, which is exactly what's happening.
---
*Built on [TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate](https://arxiv.org/abs/2504.19874)
(Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni; 2025), building on the authors' earlier
QJL work. Implementation details and benchmarks are from [`0xsero/turboquant`](https://github.com/0xsero/turboquant)
(GPLv3) and [`turbo-tan/llama.cpp-tq3`](https://github.com/turbo-tan/llama.cpp-tq3).*
---
# TurboVec: FAISS-competitive vector search with no training phase
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/turbovec
> date: 2026-07-02
> tags: vector-search, information-retrieval, quantization, rust, explainer
Vector search is memory-bound. A few million embeddings at float32 are gigabytes of RAM, and
the standard fix — **product quantization** (PQ), the workhorse inside FAISS — compresses them by
learning a codebook: run k-means over a sample of your data, replace each sub-vector with the
nearest centroid's index. It works well, but it has a cost that's easy to forget: a **training
phase**. You need representative data up front, you call `train()` before you can add anything,
and as your corpus drifts the learned codebook goes stale and wants a rebuild.
**TurboVec** ([`RyanCodrai/turbovec`](https://github.com/RyanCodrai/turbovec)) removes the training
phase entirely. It's a Rust vector index built on **TurboQuant** — the same
[rotate-then-quantize algorithm](/articles/turboquant-kv-cache) that compresses LLM KV caches —
whose quantizer is *data-free*, so there's nothing to learn. The headline: a 10-million-vector
corpus that costs ~31 GB as float32 fits in ~4 GB, and searches faster than FAISS.
## Why there's nothing to train
The full derivation is in the [TurboQuant write-up](/articles/turboquant-kv-cache), but the short
version is the whole reason there's no training step. TurboVec encodes each vector in six moves:
1. **Normalize** to the unit sphere (the length is stored separately for scoring).
2. **Rotate** by a random orthogonal matrix — built once from a seeded Gaussian via QR, so it's
deterministic and, crucially, *data-independent*.
3. **TQ+ calibration** — a small per-coordinate shift/scale, fit once on the first batch, to snap
real embeddings onto the ideal post-rotation marginal.
4. **Lloyd–Max scalar quantization** — the optimal quantization levels, precomputed against the
*known* Beta distribution the rotation produces. This is the data-free part: because you know
the distribution analytically, the optimal quantizer is a fixed table, not something learned
from your vectors.
5. **Bit-pack** to 2, 3, or 4 bits per coordinate — 16× smaller than float32 at 2-bit.
6. **Store a per-vector correction scalar** so the quantized dot product is an *unbiased* estimate
of the true inner product (a RaBitQ-style length renormalization).
PQ learns its codebook from your data; TurboQuant reshapes your data into a distribution whose
optimal codebook is already known. That's the trade — and it's why TurboVec can ingest online with
no `train()`, no parameter tuning, and no rebuilds as the corpus grows.
## Does it actually beat FAISS?
Mostly yes, and the repo is honest about where it doesn't. Across its benchmark configs (100K
vectors, 1K queries, k=64), here's TurboVec against FAISS `IndexPQ` on recall, latency, and
compression — flip through the datasets and bit-widths:
On the OpenAI embedding sets it wins recall@1 outright (up to +1.9 points at 2-bit) *and* runs
12–19% faster on ARM, at 8–16× compression with no training pass. The exceptions are real: on x86
it trails a few percent at 2-bit (it wins the 4-bit configs), and on low-dimensional GloVe vectors
at 2-bit FAISS edges it by 0.06 of a point — the rotation has less room to spread energy in only
200 dimensions. Net, it's genuinely competitive with a mature, heavily-optimized library, which is
a high bar for a quantizer with no learned codebook.
## The systems half
A near-optimal quantizer only matters if the search is fast, and TurboVec is a real systems
project, not a reference implementation:
- **Hand-written SIMD kernels** — NEON on ARM, AVX-512BW on x86, with an AVX2 fallback and runtime
feature detection. Scoring runs on the *packed codes directly* via table lookups; there's no
decompression step.
- **32-vector blocks**, FAISS FastScan-style, with the query LUT built per search. Filtering is a
bitmask checked at block granularity — whole blocks with no allowed vectors are skipped, and the
filter is applied *inside* the kernel so a restricted search returns the true top-k among allowed
items with **no recall penalty and no over-fetch**.
- **`IdMapIndex`** gives stable `uint64` external IDs with **O(1) removal** (swap-remove, no
tombstones) — the vector you delete is replaced by the last one and both ID maps update in
constant time.
- **Online ingest and plain persistence** (`.tv` / `.tvim` files), plus drop-in adapters for
LangChain, LlamaIndex, Haystack, and Agno.
The Python API is what you'd hope for — no training call anywhere:
```python
from turbovec import TurboQuantIndex
index = TurboQuantIndex(dim=1536, bit_width=4) # 2, 3, or 4 bits
index.add(vectors) # float32 (n, dim) — indexed immediately
scores, ids = index.search(query, k=10) # searches the packed codes
index.write("corpus.tv")
```
## Where it fits (and where it doesn't)
The honest scope: TurboVec is a **flat, exhaustive-scan** index — it scores every (packed) vector
per query. That's exactly the regime where it competes with FAISS's flat PQ scan, and at a hundred
thousand to a few million vectors it's excellent: no training, tiny memory, strong recall. It is
*not* a billion-scale graph index — if you need sub-linear search over hundreds of millions of
vectors you still want an IVF or HNSW structure (and you could quantize *those* with TurboQuant
too). Think of it as the compression-and-scoring core done unusually well, not a replacement for
every ANN system.
## The take
What I like about TurboVec is that it makes the [TurboQuant](/articles/turboquant-kv-cache) thesis
concrete in a second domain: the same "rotate into a known distribution, then quantize with a
data-free optimal quantizer" that shrinks KV caches also shrinks embedding indexes — and here it
buys something PQ structurally can't, the elimination of the training phase. Pair it with a lexical
scorer like [BM25](/articles/bm25) and you've got both halves of hybrid retrieval, each running on
commodity hardware with no GPU and no learned index. It won't dethrone HNSW at billion scale, but
for the very common case of a few million embeddings that need to fit in RAM and update live, "as
good as FAISS, with nothing to train" is a genuinely nice place to land.
---
*Built on [`RyanCodrai/turbovec`](https://github.com/RyanCodrai/turbovec) (Rust + Python, MIT),
which implements [TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate](https://arxiv.org/abs/2504.19874)
(Zandieh, Daliri, Hadian, Mirrokni; ICLR 2026). Benchmark figures are from the repo's published
results (100K vectors, k=64; ARM = Apple M3 Max, x86 = Xeon Sapphire Rapids).*
---
# HydraHead: hybrid attention at the head, not the layer
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/hydrahead
> date: 2026-07-01
> tags: llm, attention, long-context, interpretability, explainer
The attention that made transformers work is also what makes them expensive: every token
attends to every other, so cost grows with the **square** of the context length. That
quadratic term is fine at 4K tokens and ruinous at 512K — exactly the regime long-context
models are pushing into. **Linear attention** (LA) fixes the scaling by keeping a
fixed-size recurrent state instead of a full attention matrix, so cost grows linearly. But
that fixed state is lossy: it can't do exact, long-range token retrieval — the "find the
one sentence 400K tokens ago" trick that full attention nails.
So the field mixes them — **hybrid attention**. Almost everyone does it *per layer*:
interleave whole full-attention (FA) layers with whole linear-attention layers at some
fixed ratio (3:1, 7:1), sometimes searched with NAS (as in [GLM-5](/articles/glm-5-2)).
**HydraHead**, from Alibaba, argues that the layer is the wrong unit — and backs it with an
interpretability result.
## The finding: layers are smooth, heads are not
HydraHead's authors probe a pretrained dense model (Qwen3-1.7B) and look at two things:
- **Across layers**, the outputs vary *smoothly* — the layer-to-layer output-similarity
matrix is one gradual block, with no crisp boundary that says "put full attention here,
linear attention there." Layer-wise hybridization is cutting where there's no clean seam.
- **Within a layer**, the heads are sharply *heterogeneous*. Reading the same input, they do
different jobs — and only a few do the long-range retrieval that genuinely needs FA. The
per-layer Gini coefficient of head importance averages **0.62**: importance is concentrated
in a handful of heads. Across all 448 query heads, only about **6.5% are essential** for
retrieval; ~91% can be swapped to linear attention with negligible loss. And the critical
ones are *scattered* — almost every layer mixes a couple of retrieval heads with a dozen
replaceable ones.
You can see the heterogeneity directly — heads in one layer specialize, and only some need
to reach far back:
That's the whole argument in one observation. If retrieval capability lives in a sparse,
scattered set of *heads*, then the head — not the layer — is the natural granularity for
deciding where to spend full attention.
## Picking the retrieval-critical heads
HydraHead keeps FA for about **25% of heads** by default and runs the rest as **Gated
DeltaNet** (GDN, the linear-attention variant it builds on). The question is *which* 25%.
The selection is a causal interpretability procedure, not a guess:
- Build **counterfactual pairs** from RULER needle-in-a-haystack probes — swap the needle's
value for a same-length distractor while holding the rest of the context fixed, so
activations stay in-distribution.
- Run **activation patching** (for heads that *receive* the retrieved information) and
**path patching** (for heads that *send* it), scoring each head by how much restoring it
recovers the correct-answer logit.
- Fuse the per-capability scores, rank all heads, and keep the top-K as FA.
It's cheap — the ranking stabilizes from roughly **six calibration samples** — and it's
*faithful*: knock out just the top ~1% of heads by this score and needle-retrieval accuracy
collapses, while ablating random heads barely moves it. Crucially, the ablation confirms the
selection **beats fixed or random head assignment** — the interpretability signal is doing
real work.
## Reconciling two kinds of output
You can't just concatenate FA and GDN heads and project them — their outputs live on
different scales. Softmax attention is **query-magnitude-modulated**: it produces sharp,
low-entropy distributions peaked on a few tokens. Linear attention cancels that magnitude
out, giving smoother, higher-entropy, more uniform outputs. Splice the two naively and the
model degrades badly.
HydraHead's **scale-normalized fusion** handles it with two moves: an **independent per-head
RMSNorm** on every head's output, then a **learnable per-head scalar** $\gamma_h$ that
re-weights each head before the shared output projection. It's a small module, but it's
load-bearing — remove the normalization and RULER's extended-context score drops from
**87.5 to 71.4**. A learnable *scale* also beats a learnable *gate* by ~20 points, so the
model keeps every head's contribution and just rescales it, rather than gating heads off.
## Building it cheaply
You don't train HydraHead from scratch — you *convert* a pretrained FA model in a three-stage
transfer pipeline that reuses as much as possible:
1. **Parameter migration + alignment** — the FA heads keep their pretrained weights; the new
GDN heads reuse the base model's Q/K/V projections (repeated channel-wise to bridge the
GQA→multi-head shape gap), so nothing starts from random. A per-layer MSE loss aligns the
hybrid's hidden states to the original.
2. **Logit distillation** — unfreeze the whole model and match its output distribution to the
original FA teacher with a KL objective.
3. **Long-context fine-tuning** — ordinary next-token prediction at 16K context.
The controlled conversion runs on **~2.3B tokens**; the scaled-up model in the paper uses
**~15B**. Either way it's a rounding error next to pretraining — the capability is inherited,
not learned fresh.
## The payoff: long context that doesn't collapse
The headline result is retention. On RULER single-needle retrieval, most hybrid models — and
the base model itself — fall to **near zero** by 256K. HydraHead holds:
The harder multi-key retrieval shows the same shape — everything else craters, HydraHead
degrades gracefully:
Across the full context sweep the retention gap only widens: HydraHead tracks the baselines
up to 64K, then holds as they collapse — reaching **+86.6** points on single-needle and
**+69.2** on multi-key at 512K, closing on Qwen3.5-2B-Base (which ships native 256K support):
And it buys that without wrecking short-context ability — the usual tax on linear-attention
conversions. On general reasoning it lands within ~3.4 points of the full-attention base
model, and on MMLU it essentially matches it:
The efficiency claim is the one to internalize: at a **7:1** GDN-to-FA head ratio — only one
head in eight keeping full attention — HydraHead matches a **3:1 layer-wise** hybrid's
long-context average, while doing *better* on hard reasoning. Same quality, far less full
attention, which means a smaller KV cache (HydraHead's is ~0.35× a full-attention model's).
If you've read the [inference write-up](/articles/how-llm-inference-works), that cache number
is the whole game at long context — it's what decides how many requests fit on the GPU.
## The take
What I like here is that the architecture change *follows from* an interpretability result
instead of being reverse-justified by one. The claim "retrieval lives in a sparse, scattered
set of heads" is measured with causal patching, the ablations show fixed/random selection is
worse, and the fix — hybridize per head, keep FA where the retrieval heads are — falls
straight out of the measurement. It's a clean example of interpretability paying rent.
Two honest caveats. The flashiest numbers — *"69% improvement at 512K"* and *"approaching
Qwen3.5"* — come only from a figure, with no supporting table; the tabulated results stop at
256K, so treat the 512K story as directional. And this is all at the **1.7B** scale on
retrieval-style benchmarks; whether head-level hybridization holds its edge at 30B+ and on
messier long-context reasoning is the open question. But the core idea — that the *head* is
the right unit for spending your quadratic-attention budget — is the kind of insight that
tends to generalize.
---
*Built on [HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention
Hybridization](https://arxiv.org/abs/2606.20097) (Tan, Chen, Shen, Liu, Shen, Wu, Ye;
Alibaba Group, 2026). Benchmark figures are quoted from the paper's tables; the 512K and
Qwen3.5 comparisons are from its Figure 1.*
---
# Tapered Language Models: spend your width where the work is
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/tapered-language-models
> date: 2026-07-01
> tags: llm, transformers, architecture, scaling, explainer
Here's an assumption baked into every transformer, recurrent, and memory-based language
model since 2017: the layers are **identical**. Same residual dimension, same attention
shape, same MLP width, stacked N deep. It's a default inherited from the original
transformer and almost never questioned — parameters are spread *uniformly* across depth.
**Tapered Language Models** (TLMs), from Bayat, Behrouz, and Courville, ask what happens if
you don't. Their answer is a one-line change with a free-lunch flavor: under a *fixed*
parameter budget, give the early layers more MLP width and the late layers less, on a smooth
cosine schedule. Same parameters, same FLOPs — better model.
## The default no one questions
Why would non-uniform be better? Because a growing pile of evidence says layers *don't*
contribute equally. The paper points at three findings:
- **Early-exit** methods show the residual stream often converges to its final prediction
well before the last layer.
- **Layer-skipping** shows later layers can be bypassed at inference with minimal damage.
- Interpretability work finds lower layers capture shallow syntactic patterns while upper
layers encode semantics — different jobs, not equal ones.
The unifying picture: **later layers refine the residual stream rather than transform it.**
An early layer rewrites the representation; a late layer nudges it. If you measure how much
each layer actually changes things, the work is front-loaded:
And if late layers are only refining, giving them the *same* width as the layers doing real
transformation is wasted capacity — capacity you could have spent where it matters.
## The taper, precisely
TLMs taper exactly one thing: the **MLP intermediate width** $d_{ff}$. Attention — residual
dimension, head count, key/value dims — is left identical to the baseline. That's a
deliberate choice: MLPs dominate the parameter count in every modern LM family, and width is
a single clean axis to vary. The schedule is a cosine over depth:
$$
d_{ff}(l) = d_{end} + \frac{d_{start} - d_{end}}{2}\left(1 + \cos\frac{\pi l}{L-1}\right)
$$
Layer 0 is widest ($d_{ff} = d_{start}$), the last layer narrowest ($d_{ff} = d_{end}$),
monotonically decreasing in between. The endpoints are multipliers of the baseline width, and
the best config is **$1.5\times$ at the front, $0.5\times$ at the back** — a 3:1 taper. The
crucial property: the per-layer widths **average to the baseline**, so
$$
\frac{1}{L}\sum_l d_{ff}(l) = d_{ff}^{\text{baseline}}
$$
Early layers get wider than baseline, late layers narrower, and the integral is preserved.
Total parameters don't change. Total FLOPs don't change. Tapering only shifts *where* the
compute is spent, not how much. That's what makes it a redistribution rather than a bigger
model — and why the comparison to the uniform baseline is honest.
## Direction is everything
The foundational experiment splits the stack into three blocks and moves the extra capacity
around — early, middle, or late — at fixed budget. The result is unambiguous:
Front-loading helps. Centering the capacity is *worse* than uniform. And back-loading —
wider late layers — is the worst option of all, **+1.01 perplexity** over just doing nothing.
This is the control that makes the whole paper: the gain isn't from "more capacity somewhere,"
it's specifically from putting it where the representational work happens.
The *shape* matters too. Sweeping cosine against linear and sigmoid schedules, cosine wins in
every setting; sigmoid is often worse than the uniform baseline. And taper strength is
U-shaped — too gentle leaves gains on the table, too aggressive (a 7:1 ratio) starves the late
layers and regresses. The 1.5→0.5 cosine is the bottom of that U.
## How consistently it holds
The tuned config — cosine 1.5→0.5, found once on a 440M Transformer — is then transferred
*unchanged* to three scales (440M/30B tokens, 760M/50B, 1.3B/100B) and four architectures:
plain Transformer, Gated Attention, and Behrouz's own **HOPE** and **Titans** memory models.
It keeps improving almost everywhere. At 760M, the perplexity reduction from the exact same
parameter budget:
And it carries through to downstream accuracy — the average over eight commonsense benchmarks
(LAMBADA, PIQA, HellaSwag, WinoGrande, ARC-easy/challenge, SIQA, BoolQ) rises for every
architecture at 760M:
The full picture, uniform → tapered, is small-but-consistent rather than dramatic:
| scale | architecture | WikiText ppl | LAMBADA ppl | commonsense avg |
|---|---|---|---|---|
| 760M | Transformer | 21.86 → **21.42** | 22.29 → **21.25** | 52.25 → **52.84** |
| 760M | Gated Attention | 20.74 → **19.98** | 21.85 → **21.44** | 52.61 → **52.88** |
| 760M | Titans | 21.58 → **20.77** | 23.09 → **22.92** | 52.30 → **53.29** |
| 1.3B | Transformer | 17.39 → **17.17** | 17.62 → **16.93** | 56.05 → **56.38** |
| 1.3B | Titans | 16.05 → **15.76** | 14.19 → **14.04** | 56.73 → **57.08** |
The honest read: at scale the gains are typically 0.1–1.0 perplexity and a few tenths of a
point of accuracy — improving in ~15 of 16 measured cells (the lone regression is 1.3B
HOPE's WikiText, off by 0.03). It's not a step change. It's a **free, universal nudge in the
right direction** from a config that was never even tuned at these scales.
## Why it works
The mechanism check makes the story tight. Measure the cosine similarity between each MLP's
output and the residual stream it writes into, and it *rises* with depth — later MLPs produce
updates increasingly *aligned* with the residual (Pearson r ≈ 0.49–0.71 vs. layer index). An
update aligned with the residual is a refinement; an orthogonal one is a transformation. So
late MLPs, writing residual-aligned updates, aren't using their extra width — the hidden
dimension is spent nudging in a direction the stream already points. Tapering removes width
exactly where that alignment is highest and moves it to the early layers, which write the
orthogonal, representation-defining updates that actually need the capacity. The
[MLP is where the parameters live](/articles/mixture-of-experts-from-scratch); this is just
allocating them by how hard each layer is working.
## The take
I like this paper for the same reason I liked [HydraHead](/articles/hydrahead): the
architectural change is *derived from* a measurement, not reverse-justified. "Later layers
refine, not transform" is an old observation; TLMs turn it into a concrete lever — cosine-taper
the MLP width — and then run the control (reverse it, and it hurts) that proves the direction
is what's doing the work.
The caveats are real and the authors state them. The gains at scale are modest, the single
1.5→0.5 cosine schedule was tuned only on a 440M Transformer and transferred without
re-tuning (so there's likely more on the table), and there's no code release. But the appeal
is that it costs *nothing* — same parameters, same FLOPs, one function applied to the MLP
widths. For a default that's gone unquestioned for eight years, "uniform depth was leaving
free perplexity on the floor" is a satisfying result, and an easy one to try.
---
*Built on [Tapered Language Models](https://arxiv.org/abs/2606.23670) (Reza Bayat, Ali
Behrouz, Aaron Courville, 2026). Perplexity, accuracy, and ablation figures are quoted from
the paper's tables; the residual-alignment correlation is from its Figure 4.*
---
# Agents-A1: scaling the agent horizon, not the parameter count
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/agents-a1
> date: 2026-06-30
> tags: llm, agents, reinforcement-learning, distillation, explainer
Most of the frontier's agentic gains over the last year came from one move: make the model
bigger. Kimi-K2.6 and DeepSeek-V4-pro are *trillion*-parameter systems, and they top the
hard agent benchmarks. **Agents-A1**, from InternScience, makes the opposite bet. It's a
**35B Mixture-of-Experts** model with only **~3B active parameters per token** (it's
initialized from `Qwen3.5-35B-A3B`), and it lands in the same benchmark band as those 1T
models. Its thesis, verbatim from the tech report: *scaling the horizon, not the
parameters.*
The chart is the whole argument. Agents-A1 and `Qwen3.5-35B-A3B` are the **same model** —
same 35B size, same ~3B active params. The ~18-point vertical gap between them is *entirely*
the training recipe, and it lifts the 35B student into the band held by models 30× its size.
The question this article answers is: what is "the horizon," and how do you scale it instead
of the parameter count?
## Two axes of horizon
"Agent horizon" splits into two things you can grow independently, and Agents-A1 pushes both:
1. **Long-horizon trajectories** — how *far* a single agent run goes. Agents-A1 is trained on
agentic trajectories averaging **~45K tokens** (deep-research runs average 44K, coding 48K,
scientific reasoning 37K, general agentic 39K; only short instruction-following tasks pull
the mean down at ~3K). That's hundreds of think→act→observe steps per task, not a single
question-answer turn.
2. **Heterogeneous abilities** — how *many kinds* of agent the one model can be. Agents-A1
unifies **six domains**: long-horizon search, engineering, scientific research, instruction
following, general agentic tasks, and scientific agentic tasks.
A single long-horizon run looks like this — a chain of tool calls, each with an observation
and a **verifier outcome** that says whether the step actually worked:
The verifier is the load-bearing part. A raw transcript of an agent flailing is not training
data; a transcript where every step is *checked* — did the code converge, does the answer
match the cited evidence — is. That check is what turns a 45K-token trace into a sequence of
trainable targets. Which is the next problem: where do verified 45K-token trajectories at
scale come from?
## The knowledge-action graph
You can't hand-write 100K agent trajectories. Agents-A1's answer is a **knowledge-action
graph** (KAG) per domain — a typed four-tuple
$$
\mathcal{G}_d = (\mathcal{C}_d,\ \mathcal{A}_d,\ \mathcal{O}_d,\ \mathcal{V}_d)
$$
| symbol | what it holds |
|---|---|
| $\mathcal{C}_d$ — **corpus** | evidence chunks, entities, facts, constraints — the domain's grounded knowledge |
| $\mathcal{A}_d$ — **actions** | tool calls, retrieval queries, code edits and executions, reasoning steps |
| $\mathcal{O}_d$ — **observations** | tool returns, retrieved evidence, execution states, intermediate artifacts |
| $\mathcal{V}_d$ — **verifiers** | automatic checks over correctness, evidence support, constraint satisfaction, goal completion |
The graph is populated by a **proposer–solver–verifier self-play game**: a proposer policy
$\pi_P$ samples regions of the graph to pose constrained tasks, a solver $\pi_S$ attacks them
with retrieval and tools, and a verifier $\pi_V$ checks the answer, the evidence, the
execution result, and the trajectory for shortcut-taking. A candidate task is kept **only if**
it's verifiable, valid, process-informative, evidence-covering, and unambiguously specified.
Each accepted step is logged as a record $(s_t, a_t, o_t, v_t)$ — prior state, action,
observation, verifier outcome — and *that* tuple is the trainable target. The data engine and
the agent are the same machinery: the graph that grounds the agent's actions is the graph that
generates its training data.
## The three-stage recipe
With verified trajectories in hand, the model is built in three stages — broaden, specialize,
then re-unify:
Stages 1 and 2 are familiar: a **full-domain SFT** pass aligns the base model with broad agent
behavior across all domains (~100K trajectories, response-token cross-entropy, one epoch at up
to 131K sequence length), then a set of **domain-level teachers** is trained, each with its own
recipe — the search teacher with SFT then GRPO over web-search/read/code tools; the science
teacher with reasoning-enhanced then tool-augmented SFT; the instruction-following and
tool-calling teachers with their own GRPO setups and reward shaping. Each teacher goes deep
where a single generalist would be pulled thin.
Stage 3 is the interesting one, and it's the same idea three separate papers landed on this
same week ([MOPD and DOPD](/arxiv/2026-06-30)): **on-policy distillation** as the way to fuse
capabilities. The full name is a mouthful — *multi-teacher domain-routed on-policy distillation
with salient vocabulary alignment* — so here's what each piece means:
- **On-policy.** The *student* generates the rollout, and the teacher supervises the student's
own tokens — not a fixed teacher transcript. This kills exposure bias: the student learns to
recover from the states it actually visits, not the ones a teacher would have.
- **Domain-routed.** Routing is hard and per-sample: each trajectory carries a domain label, and
it's supervised *only* by that domain's teacher ($\theta_{t,i} = \theta_t^{d_i}$). No learned
per-token gate — the task picks the specialist.
- **Salient vocabulary alignment (SVA).** At each position, the distillation loss is computed
*only over the teacher's top-$k$ vocabulary* — the handful of tokens the teacher actually puts
probability on. Both distributions are renormalized onto that support and matched with a
forward-KL term. The long low-probability tail, which carries no decision information, is
dropped. You align where the capability lives.
- **Heterogeneity-aware.** Losses are averaged *within* a domain first, then *across* domains, so
a high-volume domain can't drown out a small one — each active domain gets comparable influence
on the update.
The result is one deployable 35B student that inherits all six specialists, with **no teacher
shipped at inference**. If you've read the [inference write-up](/articles/how-llm-inference-works),
this is the training-time mirror of the serving-time story: the whole game is getting frontier
behavior out of a model small enough to actually run.
## The numbers
The payoff is parity with — and on several benchmarks, victory over — models ~30× larger. The
sharpest case is **FrontierScience-Research**, where the gap to the trillion-parameter field is
not subtle:
The base model scores **2.5**; the trained 35B scores **40.0** — above GPT-5.5's 26.7 and more
than double the trillion-parameter Kimi and DeepSeek. On long-horizon search, it takes overall
SOTA on **Seal-0**, edging out frontier systems that are far larger:
And on instruction following it leads outright, which matters because it's the capability most
likely to *degrade* when you fuse many domains into one model:
Across the full suite, Agents-A1 takes overall SOTA on **six** benchmarks (Seal-0, HiPhO 46.4,
FrontierScience-Olympiad 79.0, FrontierScience-Research 40.0, IFBench 80.6, IFEval 94.8) and is
the best ~35B-class model on most of the rest (BrowseComp 75.5, GAIA 96.0, XBench-DS 86.0,
SciCode 44.3, HLE-w/-tools 47.6, MolBench-Bind 56.8). It does **not** win everywhere — GPT-5.5
still leads pure web search (BrowseComp 84.4) and engineering (SciCode 56.1, MLE-Lite 72.7), and
DeepSeek-V4-pro tops GAIA (98.1) and the general-agentic τ²-Bench. The honest summary is parity
in the frontier band, with clear leads in science and instruction following — from a model you
can serve on a fraction of the hardware.
## Running it
Agents-A1 is **Apache-2.0** and runs on the standard stack — Hugging Face Transformers, vLLM, or
SGLang with OpenAI-compatible endpoints, at a served context of **262K tokens**. The release
recommends specific sampling for the long-horizon behavior to hold up:
```python
# vLLM / SGLang OpenAI-compatible call
sampling = dict(
temperature=0.85,
top_p=0.95,
top_k=20,
min_p=0.0,
presence_penalty=1.1, # discourages the repetitive loops long agents fall into
)
```
The `presence_penalty` is the non-obvious one: long agent rollouts are prone to getting stuck
repeating a failing action, and a mild penalty keeps the trajectory exploring.
## The take
What I like about Agents-A1 is that it's an honest systems argument, not a parameter flex. The
recipe is the contribution: a knowledge-action graph that makes verified long-horizon data a
renewable resource, and an on-policy distillation stage that folds many specialists into one
small model without the capability erosion you'd expect. It converges with a clear 2026 theme —
[on-policy distillation](/arxiv/2026-06-30) is becoming the default way to *integrate*
capabilities rather than trade them off, and [horizon, not size](/articles/how-llm-inference-works),
is where the agentic gains are now coming from.
The caveats are the usual ones for a benchmark-led release: the expert count and MoE routing
aren't disclosed, the "trillion-parameter performance" framing rests on benchmark parity rather
than a fitted scaling law, and benchmark SOTA is not the same as robustness in a messy
production loop. But the direction is the point. If a 35B model with 3B active parameters can be
trained to sit in the frontier's agentic band, the interesting frontier stops being *how big*
and becomes *how far* — how long the horizon, how many the domains, how good the verifiers.
---
*Built on [Agents-A1: Reaching Trillion-Parameter Performance with a 35B Agent](https://arxiv.org/abs/2606.30616)
(InternScience, 2026), the [project page](https://internscience.github.io/Agents-A1/), and the
[model release](https://huggingface.co/InternScience/Agents-A1) (Apache-2.0). Benchmark figures
are quoted from the tech report and model card.*
---
# FAST-LIO2 from scratch: LiDAR-inertial odometry you can actually reproduce
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/fast-lio2-lidar-inertial-odometry
> date: 2026-06-28
> tags: slam, lidar, state-estimation, point-cloud, explainer
FAST-LIO2 is the LiDAR-inertial odometry I keep coming back to: it's accurate, it runs at
100 Hz on a laptop, it survives 1000 deg/s rotations, and the whole thing is one tight loop
around a Kalman filter. But the official code is a maze of templates, and the cleanest
annotated fork is in Chinese. So this is the article I wanted: what FAST-LIO2 actually
*does*, derived from first principles, with real code — and a path to rebuild the core
without ROS, because once you can do that, you understand SLAM.
If the Kalman filter isn't fresh, read [my Kalman piece](/articles/kalman-filter) first —
FAST-LIO2 is exactly the "iterated, error-state, on-manifold" filter that article ends on,
fed by a high-rate IMU and corrected by thousands of LiDAR points per scan.
## The whole system in one loop
The problem: a LiDAR gives you ~100k 3D points per scan at 10 Hz, but a scan takes ~100 ms
during which the sensor moves, so the points are distorted; and LiDAR alone is slow to
register and fragile under fast motion. An IMU gives you 200–1000 Hz acceleration and
angular velocity — great for short-term motion, but it drifts. Fuse them tightly and each
fixes the other's weakness.
Before the math, here's the cycle as a sequence — what each stage does and, just as
important, *why it has to be there*. It runs once per LiDAR scan; the IMU drives the
prediction in between:
Two contributions set FAST-LIO2 apart from its predecessor:
1. **Direct registration.** No edge/plane feature extraction — it registers *raw* points to
the map by point-to-plane residuals. Less to tune, and it uses all the geometry.
2. **The ikd-Tree.** An incremental k-d tree that inserts, deletes, downsamples, and
re-balances in place, so the map updates in real time instead of being rebuilt.
Underneath both is a **tightly-coupled iterated error-state Kalman filter on a manifold**.
Let's build it piece by piece. I'll quote C++ from
[`zlwang7/S-FAST_LIO`](https://github.com/zlwang7/S-FAST_LIO) — a clean reimplementation
that writes the filter out explicitly instead of hiding it in template magic — and give
simplified NumPy alongside.
## The state lives on a manifold
You can't store an orientation as a 3-vector and add to it — rotations live on the manifold
$SO(3)$. FAST-LIO2 tracks a **24-dimensional nominal state** but a **23-dimensional error
state** (and covariance), because each $SO(3)$ rotation needs 3 tangent dimensions, not the
4 of a quaternion, and gravity lives on the sphere $S^2$ (2 dimensions). The state is:
$$
\mathbf{x} = \big[\ \mathbf{p},\ \mathbf{R},\ \mathbf{R}_{L}^{I},\ \mathbf{t}_{L}^{I},\
\mathbf{v},\ \mathbf{b}_g,\ \mathbf{b}_a,\ \mathbf{g}\ \big]
$$
position, attitude, LiDAR→IMU extrinsic rotation and translation, velocity, gyro bias,
accel bias, and gravity. In the clean code that's one manifold declaration:
```cpp
// include/use-ikfom.hpp — 24-D nominal, 23-D tangent
MTK_BUILD_MANIFOLD(state_ikfom,
((vect3, pos)) ((SO3, rot))
((SO3, offset_R_L_I)) ((vect3, offset_T_L_I))
((vect3, vel)) ((vect3, bg)) ((vect3, ba))
((S2, grav)));
```
We update on the manifold with $\boxplus$ (retraction) and measure differences with
$\boxminus$ — for the $SO(3)$ part, $\mathbf{R}\boxplus\delta = \mathbf{R}\,\mathrm{Exp}(\delta)$
and $\mathbf{R}_1\boxminus\mathbf{R}_2 = \mathrm{Log}(\mathbf{R}_2^{\!\top}\mathbf{R}_1)$.
Everything else is ordinary vector $+/-$.
## Forward propagation: ride the IMU
Between LiDAR scans, the IMU drives the state forward. The continuous kinematics are the
standard strapdown model — position integrates velocity, attitude integrates de-biased
angular velocity, velocity integrates de-biased, gravity-corrected acceleration:
$$
\dot{\mathbf{p}} = \mathbf{v}, \qquad
\dot{\mathbf{R}} = \mathbf{R}\,(\boldsymbol{\omega}_m - \mathbf{b}_g)^\wedge, \qquad
\dot{\mathbf{v}} = \mathbf{R}(\mathbf{a}_m - \mathbf{b}_a) + \mathbf{g},
$$
with the biases doing a slow random walk. That's `get_f` verbatim:
```cpp
// f(x,u): continuous-time kinematics
vect3 omega = in.gyro - s.bg; // ω = ω_m − b_g
vect3 a_inertial = s.rot * (in.acc - s.ba); // R(a_m − b_a)
res(i) = s.vel[i]; // ṗ = v
res(i + 3) = omega[i]; // Ṙ ← ω
res(i + 12) = a_inertial[i] + s.grav[i]; // v̇ = R(a_m−b_a) + g
```
The predict step pushes the *mean* forward by $\mathbf{x} \boxplus (\Delta t\,\mathbf{f})$ and
the *covariance* forward with the error-state Jacobians $\mathbf{F_x},\mathbf{F_w}$:
$$
\hat{\mathbf{x}} = \mathbf{x}\boxplus(\Delta t\,\mathbf{f}), \qquad
\hat{\mathbf{P}} = \mathbf{F_x}\,\mathbf{P}\,\mathbf{F_x}^{\!\top} +
(\Delta t\,\mathbf{F_w})\,\mathbf{Q}\,(\Delta t\,\mathbf{F_w})^{\!\top}.
$$
```cpp
void predict(double &dt, Matrix &Q, const input_ikfom &i_in) {
Matrix f_ = get_f(x_, i_in); // 24×1
Matrix df_dx_ = df_dx(x_, i_in); // ∂f/∂x
Matrix df_dw_ = df_dw(x_, i_in); // ∂f/∂w
x_ = x_.plus(f_, dt); // x ⊞ (dt·f)
// F_x = I + dt·A·df_dx , F_w = dt·df_dw (assembled via the boxplus Jacobian)
P_ = F_x1 * P_ * F_x1.transpose() + (dt*F_w1) * Q * (dt*F_w1).transpose();
}
```
In NumPy the shape of it is just:
```python
def predict(x, P, imu, dt, Q):
f = get_f(x, imu) # kinematics above
F_x, F_w = jacobians(x, imu, dt) # error-state transition + noise maps
x = boxplus(x, f * dt) # advance the mean on the manifold
P = F_x @ P @ F_x.T + F_w @ Q @ F_w.T
return x, P
```
This runs once per IMU sample, and the per-sample poses are cached — we need them next.
## Backward propagation: deskew the scan
Because the scan sweeps over time while the platform moves, every point was measured from a
slightly different pose. Stack them naively and a flat wall comes out sheared:
FAST-LIO2 fixes this with **backward propagation**: walk the cached IMU poses from the
scan-end time backward, and transform each point from the pose it was *actually* sampled at
into the scan-end frame. For a point sampled at time $\rho_j$ with the IMU pose
$(\mathbf{R}_j,\mathbf{t}_j)$ relative to the scan-end pose $(\mathbf{R}_e,\mathbf{t}_e)$:
$$
\mathbf{p}^{\text{end}}_j = \mathbf{R}_{L}^{I\,\top}\!\Big(\mathbf{R}_e^{\!\top}\big(\mathbf{R}_j(\mathbf{R}_{L}^{I}\mathbf{p}_j + \mathbf{t}_{L}^{I}) + (\mathbf{t}_j - \mathbf{t}_e)\big) - \mathbf{t}_{L}^{I}\Big)
$$
which is exactly the compensation in `UndistortPcl`:
```cpp
M3D R_i(R_imu * Exp(angvel_avr, dt)); // attitude at this point's sample time
V3D T_ei(pos_imu + vel_imu*dt + 0.5*acc_imu*dt*dt - imu_state.pos);
V3D P_compensate = imu_state.offset_R_L_I.conjugate() *
(imu_state.rot.conjugate() * (R_i * (imu_state.offset_R_L_I * P_i
+ imu_state.offset_T_L_I) + T_ei) - imu_state.offset_T_L_I);
```
Now every point lives in one consistent frame and it's safe to register against the map.
## The measurement: point-to-plane
FAST-LIO2 doesn't extract features. For each deskewed point it transforms it into the world
with the current state, finds its 5 nearest map points via the ikd-Tree, fits a plane to
them, and the **residual is the point-to-plane distance** — zero when the point sits exactly
on the surface:
For a point $\mathbf{p}$ in the body frame, transformed to world
$\mathbf{p}^W = \mathbf{R}(\mathbf{R}_L^I\mathbf{p}+\mathbf{t}_L^I)+\mathbf{p}$, a plane with
unit normal $\mathbf{u}$ and offset $d$ gives residual $z = \mathbf{u}^{\!\top}\mathbf{p}^W + d$.
That's `h_share_model`:
```cpp
V3D p_global(s.rot * (s.offset_R_L_I * p_body + s.offset_T_L_I) + s.pos); // to world
ikdtree.Nearest_Search(point_world, NUM_MATCH_POINTS, points_near, sqDis); // 5 nearest
if (esti_plane(pabcd, points_near, 0.1f)) { // fit plane (a,b,c,d)
float pd2 = pabcd(0)*x + pabcd(1)*y + pabcd(2)*z + pabcd(3); // point-to-plane dist
...
}
// Jacobian row (w.r.t. attitude θ and extrinsic), residual = −distance
V3D C(s.rot.conjugate() * norm_vec); // Rᵀu
V3D A(point_I_crossmat * C); // (R_L^I p + t_L^I)^∧ Rᵀu
ekfom_data.h_x.block<1,12>(i,0) << norm_p.x, norm_p.y, norm_p.z, A, ...; // [ u | A | … ]
ekfom_data.h(i) = -norm_p.intensity; // the residual
```
The crucial detail: the Jacobian $\mathbf{H}$ is $m\times 12$ — $m$ thousands of points,
but only **12 columns** (6 for pose, 6 for the extrinsic), because a single LiDAR scan can't
observe velocity, biases, or gravity directly. Hold that thought; it's why the next step is
fast. In NumPy:
```python
def build_H_z(points_body, x, ikdtree, map_pts, R_LI, t_LI):
H, z = [], []
for p in points_body:
pw = x.R @ (R_LI @ p + t_LI) + x.p # body → world
nn = ikdtree.nearest(pw, k=5) # 5 nearest map points
n, d = fit_plane(map_pts[nn]) # unit normal, offset
r = n @ pw + d # point-to-plane distance
if abs(r) < 0.1: # keep confident matches
pI = R_LI @ p + t_LI
A = skew(pI) @ (x.R.T @ n) # ∂r/∂θ (attitude block)
H.append(np.concatenate([n, A])) # [ normal | attitude ]
z.append(-r)
return np.array(H), np.array(z) # H: m×6 (here), z: m
```
## The iterated update, and the gain that makes it cheap
A single EKF update would linearize the very-nonlinear point-to-plane fit once, at a
possibly-wrong pose, and be off. So FAST-LIO2 **iterates**: re-associate, rebuild $\mathbf{H}$
at the latest estimate, take one Kalman step, repeat until the correction is tiny. Watch the
scan snap onto the map:
Each iteration is
$$
\delta\mathbf{x} = \mathbf{K}\,\mathbf{z} + (\mathbf{I}-\mathbf{K}\mathbf{H})(\mathbf{x}^\kappa \boxminus \hat{\mathbf{x}}),
\qquad \mathbf{x}^{\kappa+1} = \mathbf{x}^\kappa \boxplus \delta\mathbf{x},
$$
iterating $\kappa$ until every component of $\delta\mathbf{x}$ drops below $10^{-3}$. The
piece that makes FAST-LIO *fast* is the **reformulated Kalman gain**. The textbook form,
$$
\mathbf{K} = \hat{\mathbf{P}}\mathbf{H}^{\!\top}(\mathbf{H}\hat{\mathbf{P}}\mathbf{H}^{\!\top}+\mathbf{R})^{-1},
$$
inverts an $m\times m$ matrix — and $m$ is *thousands* of points. FAST-LIO uses the
information-form identity to rewrite it as
$$
\mathbf{K} = (\mathbf{H}^{\!\top}\mathbf{R}^{-1}\mathbf{H} + \hat{\mathbf{P}}^{-1})^{-1}\mathbf{H}^{\!\top}\mathbf{R}^{-1},
$$
which inverts a $23\times 23$ matrix — the **state** dimension — no matter how many points
there are. That's the whole trick, and in clean code it's one block:
```cpp
// R is a scalar (LASER_POINT_COV = 0.001), so R⁻¹ = 1/R
auto K_front = (HTH / R + P_.inverse()).inverse(); // (HᵀR⁻¹H + P⁻¹)⁻¹ — 23×23
K = K_front.block<23,12>(0,0) * H.transpose() / R; // … Hᵀ R⁻¹
Matrix dx_ = K * dyn_share.h // K z
+ (Matrix::Identity() - K*H) * dx_new; // (I−KH)(x ⊟ x̂)
x_ = x_.boxplus(dx_);
// convergence: every |dx_[j]| < epsi (0.001); then update covariance
P_ = (Matrix::Identity() - K*H) * P_;
```
The same loop in NumPy, with the cheap gain spelled out:
```python
def update_iterated(x, P, points_body, ikdtree, map_pts, R=1e-3, max_iter=4, eps=1e-3):
x_prior = x.copy()
n = P.shape[0] # 23 (error-state dim)
for _ in range(max_iter):
H, z = build_H_z(points_body, x, ikdtree, map_pts, R_LI, t_LI) # relinearize
# information-form gain: invert (state × state), independent of len(z)
S = H.T @ H / R + np.linalg.inv(P) # n×n
K = np.linalg.solve(S, H.T) / R # K = S⁻¹ Hᵀ R⁻¹
dx = K @ z + (np.eye(n) - K @ H) @ (-boxminus(x, x_prior))
x = boxplus(x, dx)
if np.max(np.abs(dx)) < eps:
break
P = (np.eye(n) - K @ H) @ P
return x, P
```
That's the engine. The converged $\mathbf{x}$ is your odometry output, published at LiDAR
rate.
## The map: an incremental k-d tree
The nearest-neighbor search in the measurement step is the hot path, and the map is
*growing and moving*. A static k-d tree would be rebuilt every scan — fatal. The **ikd-Tree**
instead inserts points in place, downsamples on the tree, deletes whole regions with one
box-wise delete as the local map window slides with the sensor, and lazily re-balances only
the subtrees that get lopsided:
In code it's a handful of calls:
```cpp
ikdtree.Build(feats_down_world->points); // first scan
ikdtree.Add_Points(PointToAdd, true); // incremental insert + on-tree downsample
ikdtree.Delete_Point_Boxes(cub_needrm); // box-wise delete (window slid)
ikdtree.Nearest_Search(point_world, 5, near, d); // kNN, inside the measurement step
```
The payoff is real: on the authors' benchmarks FAST-LIO2 spends *less* time per scan than
FAST-LIO while holding a *larger* map, on both Intel and Arm.
## Putting it together — and dropping ROS
Here's the entire main loop, which is shorter than you'd expect:
```cpp
while (running) {
if (sync_packages(Measures)) { // group IMU + one LiDAR scan by time
p_imu->Process(Measures, kf, feats_undistort); // forward-propagate + deskew
downSizeFilterSurf.filter(*feats_down_body); // voxel-downsample the scan
kf.update_iterated_dyn_share_modified( // the iterated point-to-plane EKF
LASER_POINT_COV, feats_down_body, ikdtree, Nearest_Points,
NUM_MAX_ITERATIONS, extrinsic_est_en);
state_point = kf.get_x(); // odometry output
map_incremental(); // ikdtree.Add_Points(...)
}
}
```
Notice what's *not* algorithm here: `sync_packages` is just time-aligning two streams,
`publish_odometry`/`publish_frame_world` are ROS topics, and `tf` is bookkeeping. **None of
that is the filter.** To reproduce FAST-LIO2 without ROS you only need:
| You need | You don't need |
|---|---|
| read IMU samples (t, ω, a) from a file/array | ROS subscribers / message types |
| read LiDAR points (x, y, z, per-point time) | rosbag, nodelets |
| forward-propagate + deskew (the IMU code) | tf tree |
| a k-d tree over the map (ikd-Tree, or even scipy `cKDTree` rebuilt per scan to start) | rviz, publishers |
| the iterated point-to-plane update | the IKFoM template layer |
A no-ROS skeleton is just the loop, fed from arrays:
```python
x, P = init_state(), init_cov()
ikdtree = KDMap(voxel=0.5) # or scipy cKDTree to begin with
for scan in lidar_scans: # each: points + per-point timestamps
imu_batch = imu_between(prev_t, scan.t_end)
for imu in imu_batch: # 1) forward propagation
x, P = predict(x, P, imu, imu.dt, Q)
pts = deskew(scan.points, imu_poses, x) # 2) backward propagation
pts = voxel_downsample(pts, 0.5) # 3) downsample
x, P = update_iterated(x, P, pts, ikdtree, ikdtree.points) # 4) iterated EKF
ikdtree.add(transform_to_world(pts, x)) # 5) grow the map
yield x.p, x.R # pose = odometry
```
Start with a `cKDTree` rebuilt each scan to get the algorithm working end to end, then swap
in a true incremental tree once you care about speed. That ordering — correctness first,
then the ikd-Tree — is exactly how to learn it.
To anchor yourself in the real repo, here's the file → concept map for the clean version:
| File | What it owns |
|---|---|
| `use-ikfom.hpp` | the state manifold, `get_f`, `df_dx`, `df_dw` |
| `esekfom.hpp` | the explicit ESEKF: `predict`, `h_share_model`, `update_iterated_dyn_share_modified`, the reformulated gain |
| `IMU_Processing.hpp` | IMU init, forward propagation, `UndistortPcl` (deskew) |
| `ikd_Tree.cpp` | `Build`, `Add_Points`, `Delete_Point_Boxes`, `Nearest_Search` |
| `laserMapping.cpp` | the ROS glue + main loop (the part you replace) |
## A complete, runnable implementation
I wrote the whole thing as one dependency-light file —
[`fastlio2_mini.py`](/articles/fastlio2/fastlio2_mini.py) (≈390 lines, `numpy` +
`scipy` + `rosbags`). It's a faithful teaching implementation: the on-manifold state,
forward/backward propagation, the iterated point-to-plane update with the reformulated
gain — all the code blocks above, assembled and tested. It takes a real LiDAR→IMU
extrinsic and does per-scan voxel downsampling; the remaining simplifications, called out
honestly, are gravity fixed after init and a `scipy` cKDTree rebuilt per scan instead of a
true ikd-Tree.
The driver is the no-ROS loop, fed from plain arrays — read IMU, propagate (caching
poses), deskew into the IMU frame, downsample, iterated-update, grow the map:
```python
def run_offline(imu_stream, lidar_scans, voxel=0.4, scan_voxel=0.5,
T_LI=None, R_LI=None, acc_cov=1e-2, gyr_cov=1e-2,
bacc_cov=1e-4, bgyr_cov=1e-4, init_secs=0.5):
Q = np.diag([gyr_cov]*3 + [acc_cov]*3 + [bgyr_cov]*3 + [bacc_cov]*3)
R_LI = np.eye(3) if R_LI is None else np.asarray(R_LI, float)
T_LI = np.zeros(3) if T_LI is None else np.asarray(T_LI, float)
to_imu = lambda p: (R_LI @ p.T).T + T_LI # LiDAR points -> IMU frame
g, bg = imu_init([s for s in imu_stream if s[0] < imu_stream[0][0] + init_secs])
kf = ESEKF(g); kf.x.bg = bg
lmap = LocalMap(voxel); traj = []; imu_i = 0; bootstrapped = False
for scan in lidar_scans:
poses = []
while imu_i < len(imu_stream) and imu_stream[imu_i][0] <= scan['t_end']:
t, acc, gyro = imu_stream[imu_i]
dt = t - (imu_stream[imu_i-1][0] if imu_i > 0 else t)
if dt > 0: kf.predict(acc, gyro, dt, Q) # 1. forward propagation
poses.append((t, kf.x.R.copy(), kf.x.p.copy()))
imu_i += 1
body = to_imu(scan['points']) # extrinsic
pts = deskew(body, scan['dts'], poses, scan['t_end']) # 2. backward deskew
pts = voxel_downsample(pts, scan_voxel) # sparse, even set
if not bootstrapped:
lmap.add((kf.x.R @ pts.T).T + kf.x.p); bootstrapped = True # seed the map
else:
kf.update(pts, lmap) # 3. iterated point-to-plane EKF
lmap.add((kf.x.R @ pts.T).T + kf.x.p) # 4. grow the map
traj.append((scan['t_end'], kf.x.p.copy(), kf.x.R.copy()))
return traj, lmap
```
It ships with a synthetic world (a robot looping through a 10×10×3 m room) so you can run
it with **no dataset at all** — and that's how I validated it:
```
$ python fastlio2_mini.py
in-memory : ATE rmse = 0.037 m final = 0.019 m
via .bag : ATE rmse = 0.069 m final = 0.057 m
```
The second line is the important one: the file also writes the simulated data to a real
ROS1 `.bag` (`sensor_msgs/Imu` + `PointCloud2`), reads it back through `read_bag()`, and
re-runs — exercising the exact bag-parsing path you'd use on real hardware, end to end, to
**4–7 cm** of absolute trajectory error. The math is correct.
## Reading a real `.bag` — including Livox
`read_bag()` uses the pure-python `rosbags` (no ROS install) and handles both standard
`sensor_msgs/PointCloud2` (Velodyne/Ouster) and Livox's custom `CustomMsg`. Livox is the
catch with FAST-LIO data — its bags aren't PointCloud2, they're a custom message, so you
register the type definition and parse it yourself:
```python
from rosbags.typesys import Stores, get_typestore
from rosbags.typesys.msg import get_types_from_msg
ts = get_typestore(Stores.ROS1_NOETIC)
ts.register(get_types_from_msg( # the Livox point struct
"uint32 offset_time\nfloat32 x\nfloat32 y\nfloat32 z\n"
"uint8 reflectivity\nuint8 tag\nuint8 line\n", 'livox_ros_driver/msg/CustomPoint'))
ts.register(get_types_from_msg( # the Livox scan message
"std_msgs/Header header\nuint64 timebase\nuint32 point_num\nuint8 lidar_id\n"
"uint8[3] rsvd\nlivox_ros_driver/CustomPoint[] points\n", 'livox_ros_driver/msg/CustomMsg'))
# then: msg.points -> (x,y,z, offset_time); offset_time is per-point time for the deskew
```
So fetching and running an actual HKU dataset is two steps:
```bash
pip install gdown
gdown 1YqxHuDKzWUcda80QKBV61lXI86TXsGjP -O avia.bag # a Livox Avia indoor bag
python fastlio2_mini.py avia.bag
```
## What happens on real data — and the bug that taught me the most
I ran exactly that on the HKU Avia "quick-shack" bag (49 s, 9953 IMU + 491 Livox scans).
It tracks. The sensor is waved roughly in place — **47.4 rad of cumulative rotation** over
**38 m of path** in a small room — and the filter stays locked the whole way, returning to
**within ~0.6 m of its start** and reconstructing a crisp room with single, sharp walls:
But it did **not** track on the first try. Getting from "reads the bag" to the figure above
took three separate fixes, and each one is a lesson worth more than the result — because
none of them was the filter *math*. Here they are in the order I hit them.
### Bug 1 — a sensor-clock mismatch: the filter never moved
The first run collapsed exactly the way a broken LIO does: the trajectory froze within a few
centimetres of the origin while the IMU clearly showed the sensor swinging through ~1 rad/s
of rotation. I almost wrote it off as "the teaching filter isn't robust enough." It wasn't
that. I instrumented the scan timestamps and found them
landing at `t ≈ -1.6e9` *relative to the IMU* — an impossible 50-year gap. The Livox
`CustomMsg` header stamps each scan on the **sensor's own clock** (seconds since the LiDAR
booted, ~361 s into this bag), while the `/livox/imu` messages are stamped on the **bag's
record clock** (Unix time, ~1.6 billion). My propagate-up-to-scan-end loop compares the two:
```python
while imu_stream[imu_i][0] <= scan['t_end']: # IMU time vs LiDAR header time
kf.predict(...) # ...never true → never runs
```
Because every IMU timestamp (1.6e9) was vastly larger than every scan's header time (361),
that condition was *never* true. **The IMU never propagated.** The filter sat at its
initial pose, the update snapped each scan onto the origin-seeded map, and the whole thing
looked like a plausible "data-association collapse" — when really it was a unit/epoch bug
two layers down. The fix is to ignore the Livox header entirely and timestamp each scan
with the **bag record time** (which `rosbags` gives you for every message, on one
consistent clock):
```python
for conn, t, raw in reader.messages(connections=conns): # t = bag record time (ns)
...
lidar_scans.append({'t_end': t * 1e-9, 'points': pts, 'dts': dts}) # not msg.header!
```
That one change is the difference between a frozen origin and a moving trajectory. The
lesson: in sensor fusion, **check your clocks first.** A mismatched epoch or a
nanosecond-vs-second unit error masquerades perfectly as a modelling failure, and you can
waste a day tuning covariances that were never the problem. (While here, I also wired in the
calibration any real deployment needs: the **LiDAR→IMU extrinsic** from `avia.yaml`,
the **real IMU noise** `acc_cov = gyr_cov = 0.1` — my synthetic `Q` was 100× too small — and
per-scan voxel downsampling.)
### Bug 2 — a stub deskew: the map smeared
Now it moved, but the reconstructed map came out with **doubled, smeared walls** — the same
physical wall drawn twice, slightly rotated. That's within-scan distortion. My first `deskew`
was a stub: it dropped each point into the *nearest* cached IMU pose with no compensation for
the motion *across* the sweep. But a Livox sweep takes ~100 ms, and at ~1 rad/s and walking
speed the sensor rotates and translates meaningfully in that window — so every point has to
be carried from the pose it was **actually sampled at** to the scan-end frame. That's the
backward propagation from [earlier](#backward-propagation-deskew-the-scan): I interpolate the
IMU-propagated trajectory (rotation on $SO(3)$, position linearly) to each point's capture
time before registering. Single-line idea, big effect — the per-plane thickness of the
reconstructed map drops to **~4 cm** and the doubled walls collapse into one.
### Bug 3 — an outlier gate that starved the update
With the deskew fixed it tracked cleanly for ~150 scans and then **diverged** — a clean
straight ramp off into space, the unmistakable signature of the LiDAR constraint dropping out
and the IMU dead-reckoning. The cause was a gate I'd copied *too* faithfully. FAST-LIO accepts
a point-to-plane match with a range-normalized test (`s = 1 − 0.9|d|/√range`); on its dense
clouds that's fine. On my sparse, voxel-downsampled scans, the moment the prediction was
slightly off it rejected **every** correspondence, the update was skipped, and with nothing
to correct it the pose ran away. A gentler metric gate (tolerance scaled mildly with range)
keeps hundreds of inliers per scan, and the filter stays locked for the whole bag. The
lesson: **an outlier gate that's correct in a dense reference can starve a sparse
reimplementation** — watch the inlier *count*, not just the residual.
All of these are wired into `run_offline`, and the CLI uses the Avia values by default, so
`python fastlio2_mini.py avia.bag` reproduces the figure above. The honest caveats that
remain are the ones this is a *teaching* filter for: a `cKDTree` rebuilt per scan (so a full
bag is a few minutes offline, not sensor-rate), gravity fixed at init rather than estimated
on $S^2$, no loop closure, and a handful of stray points where fast rotation meets the narrow
~70° FOV. Real FAST-LIO2's ikd-Tree, in-state gravity, and tighter handling close those gaps.
But the spine — the five steps, the manifold state, the reformulated gain — is exactly what's
running here, and it's enough to track a real Livox bag and rebuild the room.
## The whole file, end to end
Everything above — the SO(3) helpers, the manifold state, forward/backward propagation,
the iterated point-to-plane update with the reformulated gain, the map, and the no-ROS
bag reader — is one self-contained file. It's deliberately unoptimized for readability
(a `cKDTree` rebuilt per scan, plain Python loops), but it runs and it tracks the real
Avia bag. Here it is in full (393 lines) — expand to read or copy the whole thing,
or [download it](/articles/fastlio2/fastlio2_mini.py):
```python
"""
fastlio2_mini.py — a minimal, ROS-free FAST-LIO2-style LiDAR-inertial odometry.
A teaching reimplementation of the FAST-LIO2 core: an iterated error-state Kalman
filter on SO(3), fed a high-rate IMU and corrected by raw point-to-plane LiDAR
residuals over an incremental k-d-tree map. It supports a LiDAR->IMU extrinsic and
per-scan voxel downsampling; the simplifications vs. the paper (called out where
they matter) are gravity fixed after init and a scipy cKDTree rebuilt per scan
instead of a true ikd-Tree. Everything else — the manifold state, forward/backward
propagation, the reformulated Kalman gain — is faithful.
pip install numpy scipy rosbags
python fastlio2_mini.py # runs a synthetic demo (no bag needed)
python fastlio2_mini.py avia.bag # runs a real Livox Avia bag (calibrated)
# or: from fastlio2_mini import read_bag, run_offline
Validated: ~4 cm ATE on an 8 s synthetic trajectory, and it tracks the real HKU
Livox Avia bag (491 scans, ~50 s) — see run_offline's calibration arguments.
"""
import numpy as np
from scipy.spatial import cKDTree
# ============================================================ SO(3) utilities
def hat(w): # vector -> skew-symmetric matrix
return np.array([[0, -w[2], w[1]], [w[2], 0, -w[0]], [-w[1], w[0], 0]])
def Exp(w): # so(3) -> SO(3) (Rodrigues)
th = np.linalg.norm(w)
if th < 1e-9:
return np.eye(3) + hat(w)
K = hat(w / th)
return np.eye(3) + np.sin(th) * K + (1 - np.cos(th)) * K @ K
def Log(R): # SO(3) -> so(3)
c = np.clip((np.trace(R) - 1) / 2, -1, 1)
th = np.arccos(c)
v = np.array([R[2, 1] - R[1, 2], R[0, 2] - R[2, 0], R[1, 0] - R[0, 1]])
return 0.5 * v if th < 1e-9 else (th / (2 * np.sin(th))) * v
# ============================================================ state on the manifold
# error-state layout (15): [ p(0:3) th(3:6) v(6:9) bg(9:12) ba(12:15) ]
class State:
def __init__(s):
s.p = np.zeros(3); s.R = np.eye(3); s.v = np.zeros(3)
s.bg = np.zeros(3); s.ba = np.zeros(3)
def copy(s):
t = State()
t.p, t.R, t.v, t.bg, t.ba = s.p.copy(), s.R.copy(), s.v.copy(), s.bg.copy(), s.ba.copy()
return t
def boxplus(x, d): # x ⊞ d (retract onto the manifold)
y = x.copy()
y.p += d[0:3]; y.R = x.R @ Exp(d[3:6]); y.v += d[6:9]
y.bg += d[9:12]; y.ba += d[12:15]
return y
def boxminus(a, b): # a ⊟ b (tangent so that a = b ⊞ d)
d = np.zeros(15)
d[0:3] = a.p - b.p; d[3:6] = Log(b.R.T @ a.R); d[6:9] = a.v - b.v
d[9:12] = a.bg - b.bg; d[12:15] = a.ba - b.ba
return d
# ============================================================ the filter
class ESEKF:
def __init__(s, g):
s.x = State(); s.P = np.eye(15) * 1e-2; s.g = g.copy()
def predict(s, am, wm, dt, Q):
"""Forward propagation: integrate one IMU sample, inflate covariance."""
x = s.x
w = wm - x.bg # de-biased angular velocity
a = x.R @ (am - x.ba) + s.g # de-biased, gravity-corrected accel (world)
# --- nominal mean ---
x.p = x.p + x.v * dt + 0.5 * a * dt * dt
Rn = x.R @ Exp(w * dt)
x.v = x.v + a * dt
x.R = Rn
# --- error-state transition F_x and noise map F_w (paper Eq. 7/8) ---
A = np.zeros((15, 15))
A[0:3, 6:9] = np.eye(3) # dp/dv
A[3:6, 3:6] = -hat(w); A[3:6, 9:12] = -np.eye(3) # dth/dth, dth/dbg
A[6:9, 3:6] = -x.R @ hat(am - x.ba); A[6:9, 12:15] = -x.R # dv/dth, dv/dba
Fx = np.eye(15) + A * dt
Fw = np.zeros((15, 12))
Fw[3:6, 0:3] = -np.eye(3); Fw[6:9, 3:6] = -x.R
Fw[9:12, 6:9] = np.eye(3); Fw[12:15, 9:12] = np.eye(3)
s.P = Fx @ s.P @ Fx.T + (Fw * dt) @ Q @ (Fw * dt).T
def update(s, pts_body, lmap, R=1e-3, max_iter=4, eps=1e-3):
"""Iterated point-to-plane update with the reformulated Kalman gain."""
x_prior = s.x.copy()
n = 15; K = None; Hfull = None
for _ in range(max_iter):
x = s.x
pw = (x.R @ pts_body.T).T + x.p # body -> world at current estimate
H_rows, z = [], []
for i in range(len(pts_body)):
nrm, off, ok = lmap.fit_plane(pw[i]) # nearest-5 plane via kd-tree
if not ok:
continue
r = nrm @ pw[i] + off # point-to-plane distance
# FAST-LIO weights acceptance by range (`s = 1 - 0.9|d|/sqrt(range)`),
# but on a sparse, voxel-downsampled scan that gate can starve a slightly-
# off prediction of *all* correspondences — the update is then skipped and
# the pose dead-reckons away. We keep a plain metric gate (correct deskew
# already removes the smear a tight gate was meant to fight) with a mild
# range allowance so far points must still fit reasonably.
rng = np.linalg.norm(pts_body[i])
if abs(r) > 0.3 + 0.05 * rng:
continue
Hr = np.zeros(15)
Hr[0:3] = nrm
Hr[3:6] = hat(pts_body[i]) @ (x.R.T @ nrm) # d(residual)/d(theta)
H_rows.append(Hr); z.append(r)
if len(H_rows) < 10:
break
H = np.array(H_rows); z = np.array(z)
dx_prior = boxminus(s.x, x_prior)
# reformulated gain: invert a 15x15 (state), NOT an mxm (measurements)
S = H.T @ H / R + np.linalg.inv(s.P)
K = np.linalg.solve(S, H.T) / R # K = (H'R^-1 H + P^-1)^-1 H' R^-1
Hfull = H
dx = -K @ z - (np.eye(n) - K @ H) @ dx_prior
s.x = boxplus(s.x, dx)
if np.max(np.abs(dx)) < eps:
break
if K is not None:
s.P = (np.eye(n) - K @ Hfull) @ s.P
# ============================================================ map (stand-in for ikd-Tree)
class LocalMap:
def __init__(s, voxel=0.4, cap=60000):
s.voxel = voxel; s.cap = cap; s.pts = None; s.tree = None
def add(s, world_pts):
s.pts = world_pts if s.pts is None else np.vstack([s.pts, world_pts])
if len(s.pts) > s.cap:
s.pts = s.pts[-s.cap:]
s.tree = cKDTree(s.pts) # a real ikd-Tree updates in place instead
def fit_plane(s, p, k=5, max_d=1.0, thick=0.1):
d, idx = s.tree.query(p, k=k)
if d[-1] > max_d:
return None, None, False
near = s.pts[idx]; c = near.mean(0)
_, _, Vt = np.linalg.svd(near - c) # smallest singular vector = normal
nrm = Vt[2]
if np.max(np.abs((near - c) @ nrm)) > thick:
return None, None, False # neighbours aren't planar enough
return nrm, -nrm @ c, True
# ============================================================ deskew (backward propagation)
def deskew(points, point_dts, imu_poses, t_end):
"""Backward propagation: transform each point from the pose it was *sampled* at
into the single scan-end frame, undoing the shear a moving sensor bakes into a
sweep. imu_poses: list of (t, R, p) propagated across the sweep; point_dts:
per-point time before scan end. We interpolate the propagated trajectory to each
point's capture time — SO(3) for rotation, linear for position — so the
within-sweep *rotation and velocity* are both compensated (using only the nearest
pose, as a naive version does, leaves fast scans warped and smears the map)."""
ts = np.array([q[0] for q in imu_poses])
R_end, p_end = imu_poses[-1][1], imu_poses[-1][2]
n = len(imu_poses)
out = np.empty_like(points)
for i, pb in enumerate(points):
t = t_end - point_dts[i]
j = min(max(np.searchsorted(ts, t) - 1, 0), n - 2) if n >= 2 else 0
if n >= 2:
t0, R0, p0 = imu_poses[j][0], imu_poses[j][1], imu_poses[j][2]
t1, R1, p1 = imu_poses[j + 1][0], imu_poses[j + 1][1], imu_poses[j + 1][2]
a = 0.0 if t1 == t0 else min(max((t - t0) / (t1 - t0), 0.0), 1.0)
R_c = R0 @ Exp(a * Log(R0.T @ R1)) # interpolate rotation on SO(3)
p_c = p0 + a * (p1 - p0) # interpolate position (carries velocity)
else:
R_c, p_c = imu_poses[0][1], imu_poses[0][2]
wpt = R_c @ pb + p_c # point in world at its capture pose
out[i] = R_end.T @ (wpt - p_end) # back into the scan-end frame
return out
# ============================================================ downsample (voxel grid)
def voxel_downsample(pts, voxel=0.5):
"""One representative point per occupied voxel — FAST-LIO's per-scan downsample.
100k raw points per scan is overkill; a sparse, even set keeps the update real-time."""
if len(pts) == 0:
return pts
keys = np.floor(pts / voxel).astype(np.int64)
_, idx = np.unique(keys, axis=0, return_index=True)
return pts[np.sort(idx)]
# ============================================================ offline driver
def imu_init(imu_samples, g_mag=9.81):
"""Estimate gravity direction and gyro bias from a short static window."""
a = np.mean([s[1] for s in imu_samples], 0) # mean specific force
w = np.mean([s[2] for s in imu_samples], 0) # mean angular velocity = gyro bias
g = -a / np.linalg.norm(a) * g_mag # gravity opposes measured accel
return g, w
def run_offline(imu_stream, lidar_scans, voxel=0.4, scan_voxel=0.5,
T_LI=None, R_LI=None, acc_cov=1e-2, gyr_cov=1e-2,
bacc_cov=1e-4, bgyr_cov=1e-4, init_secs=0.5):
"""imu_stream: list of (t, acc[3], gyro[3]); lidar_scans: list of dict with
't_end', 'points'(N,3 body), 'dts'(N per-point time before scan end).
T_LI / R_LI: LiDAR->IMU extrinsic (point in IMU frame = R_LI @ p_lidar + T_LI).
acc_cov/gyr_cov: IMU noise densities (Avia's avia.yaml uses 0.1; the synthetic
demo is quieter). The process-noise Q is built from these — too small and the
filter trusts a stale prediction and refuses to move, too large and it's jumpy."""
Q = np.diag([gyr_cov]*3 + [acc_cov]*3 + [bgyr_cov]*3 + [bacc_cov]*3)
R_LI = np.eye(3) if R_LI is None else np.asarray(R_LI, float)
T_LI = np.zeros(3) if T_LI is None else np.asarray(T_LI, float)
to_imu = lambda p: (R_LI @ p.T).T + T_LI # LiDAR points -> IMU body frame
# --- init gravity + gyro bias from the first static window of IMU ---
t0 = imu_stream[0][0]
static = [s for s in imu_stream if s[0] < t0 + init_secs]
g, bg = imu_init(static)
kf = ESEKF(g); kf.x.bg = bg
lmap = LocalMap(voxel)
traj = []; imu_i = 0; bootstrapped = False
for scan in lidar_scans:
poses = []
# forward-propagate every IMU sample up to scan end, caching poses for deskew
while imu_i < len(imu_stream) and imu_stream[imu_i][0] <= scan['t_end']:
t, acc, gyro = imu_stream[imu_i]
dt = t - (imu_stream[imu_i-1][0] if imu_i > 0 else t)
if dt > 0:
kf.predict(acc, gyro, dt, Q)
poses.append((t, kf.x.R.copy(), kf.x.p.copy()))
imu_i += 1
if not poses:
poses = [(scan['t_end'], kf.x.R.copy(), kf.x.p.copy())]
body = to_imu(scan['points']) # into the IMU body frame
pts = deskew(body, scan['dts'], poses, scan['t_end'])
pts = voxel_downsample(pts, scan_voxel) # sparse, even set for the update
if not bootstrapped: # seed the map from the first scan
lmap.add((kf.x.R @ pts.T).T + kf.x.p); bootstrapped = True
else:
kf.update(pts, lmap) # the iterated EKF correction
lmap.add((kf.x.R @ pts.T).T + kf.x.p)
traj.append((scan['t_end'], kf.x.p.copy(), kf.x.R.copy()))
return traj, lmap
# ============================================================ read a .bag without ROS
_LIVOX_DEFS = (
"uint32 offset_time\nfloat32 x\nfloat32 y\nfloat32 z\nuint8 reflectivity\nuint8 tag\nuint8 line\n",
"std_msgs/Header header\nuint64 timebase\nuint32 point_num\nuint8 lidar_id\nuint8[3] rsvd\n"
"livox_ros_driver/CustomPoint[] points\n",
)
def read_bag(path, imu_topic='/livox/imu', lidar_topic='/livox/lidar', g_mag=9.81):
"""Read IMU + LiDAR from a ROS1 bag with the pure-python `rosbags` (no ROS install).
Handles sensor_msgs/PointCloud2 (Velodyne/Ouster) AND livox_ros_driver/CustomMsg
(Livox Avia/Horizon). Livox accel (reported in g) is auto-scaled to m/s^2."""
from pathlib import Path
from rosbags.rosbag1 import Reader
from rosbags.typesys import Stores, get_typestore
from rosbags.typesys.msg import get_types_from_msg
ts = get_typestore(Stores.ROS1_NOETIC)
ts.register(get_types_from_msg(_LIVOX_DEFS[0], 'livox_ros_driver/msg/CustomPoint'))
ts.register(get_types_from_msg(_LIVOX_DEFS[1], 'livox_ros_driver/msg/CustomMsg'))
imu_stream, lidar_scans = [], []
with Reader(Path(path)) as reader:
conns = [c for c in reader.connections if c.topic in (imu_topic, lidar_topic)]
for conn, t, raw in reader.messages(connections=conns):
msg = ts.deserialize_ros1(raw, conn.msgtype)
if conn.topic == imu_topic:
a, w = msg.linear_acceleration, msg.angular_velocity
imu_stream.append((t * 1e-9, np.array([a.x, a.y, a.z]), np.array([w.x, w.y, w.z])))
elif 'CustomMsg' in conn.msgtype: # Livox
pts, dts = parse_livox(msg)
lidar_scans.append({'t_end': t * 1e-9, 'points': pts, 'dts': dts})
else: # PointCloud2
pts, dts = parse_pointcloud2(msg)
lidar_scans.append({'t_end': t * 1e-9, 'points': pts, 'dts': dts})
if imu_stream and np.mean([np.linalg.norm(s[1]) for s in imu_stream[:50]]) < 2.0:
imu_stream = [(t, a * g_mag, w) for t, a, w in imu_stream] # g -> m/s^2
return imu_stream, lidar_scans
def parse_livox(msg):
"""Decode a livox_ros_driver/CustomMsg into (N,3) xyz and per-point dt-before-scan-end.
Note: we deliberately ignore msg.header.stamp here. On real Avia bags the Livox
header runs on the sensor's own clock (seconds-since-boot), while the IMU is stamped
with the bag's record clock (Unix time). Mixing them silently breaks IMU/LiDAR sync,
so read_bag uses the bag record time `t` for every scan's t_end and only uses the
per-point offsets here for deskew."""
P = msg.points
xyz = np.array([[p.x, p.y, p.z] for p in P], float)
off = np.array([p.offset_time for p in P], float) * 1e-9 # ns -> s from scan start
keep = np.linalg.norm(xyz, axis=1) > 0.5
xyz, off = xyz[keep], off[keep]
return xyz, (off.max() - off if len(off) else off) # dt before scan end
def parse_pointcloud2(msg):
"""Decode a sensor_msgs/PointCloud2 into (N,3) xyz + per-point time offset."""
dtype = np.dtype({'names': [f.name for f in msg.fields],
'formats': [_PF[f.datatype] for f in msg.fields],
'offsets': [f.offset for f in msg.fields],
'itemsize': msg.point_step})
arr = np.frombuffer(msg.data, dtype=dtype, count=msg.width * msg.height)
xyz = np.stack([arr['x'], arr['y'], arr['z']], -1).astype(float)
# the per-point time field is named 'time'/'t'/'offset_time' depending on driver
tcol = next((n for n in ('time', 't', 'offset_time', 'timestamp') if n in arr.dtype.names), None)
dts = (arr[tcol].astype(float) if tcol else np.zeros(len(xyz)))
if dts.max() > 1.0: # ns/us -> s heuristics
dts = dts * (1e-9 if dts.max() > 1e6 else 1e-3)
return xyz, dts
_PF = {1: 'i1', 2: 'u1', 3: 'i2', 4: 'u2', 5: 'i4', 6: 'u4', 7: 'f4', 8: 'f8'}
# ============================================================ synthetic world (no dataset)
def simulate_room(seconds=8, seed=0):
"""A robot looping through a 10x10x3 m room. Returns (imu_stream, lidar_scans,
truth) in exactly the format run_offline / a real bag would give you."""
rng = np.random.default_rng(seed)
Rz = lambda a: np.array([[np.cos(a), -np.sin(a), 0], [np.sin(a), np.cos(a), 0], [0, 0, 1]])
truth = lambda t: (np.array([2*np.sin(0.4*t), 1.5*(1-np.cos(0.4*t)), 0.0]), Rz(0.3*np.sin(0.5*t)))
acc = lambda t: np.array([-0.32*np.sin(0.4*t), 0.24*np.cos(0.4*t), 0.0])
yawrate = lambda t: 0.15*np.cos(0.5*t)
g = np.array([0, 0, -9.81])
wall = []
for _ in range(4000):
f = rng.integers(0, 5); u, v = rng.uniform(-5, 5), rng.uniform(0, 3)
wall.append([[-5, u, v], [5, u, v], [u, -5, v], [u, 5, v], [u, rng.uniform(-5, 5), 0]][f])
wall = np.array(wall, float)
bg_t, ba_t = np.array([2e-3, -1e-3, 1.5e-3]), np.array([2e-2, -3e-2, 1e-2])
imu_stream = [] # 200 Hz IMU
for k in range(int(seconds * 200)):
t = k / 200; _, R = truth(t)
am = R.T @ (acc(t) - g) + ba_t + rng.normal(0, 0.01, 3) # specific force in body
wm = np.array([0, 0, yawrate(t)]) + bg_t + rng.normal(0, 1e-3, 3)
imu_stream.append((t, am, wm))
scans = [] # 10 Hz LiDAR, skewed over the sweep
for s in range(1, int(seconds * 10)):
tc = s / 10; p, _ = truth(tc)
vis = wall[np.linalg.norm(wall - p, axis=1) < 8]
idx = rng.choice(len(vis), size=min(400, len(vis)), replace=False)
pts, dts = [], []
for j, kk in enumerate(idx):
tau = tc - 0.1 + (j / len(idx)) * 0.1; pp, RR = truth(tau)
pts.append(RR.T @ (vis[kk] - pp) + rng.normal(0, 0.01, 3)); dts.append(tc - tau)
scans.append({'t_end': tc, 'points': np.array(pts), 'dts': np.array(dts)})
return imu_stream, scans, truth
def write_demo_bag(path, seconds=6):
"""Write the simulated room to a real ROS1 .bag (sensor_msgs/Imu + PointCloud2),
so you can exercise the read_bag() path without downloading a dataset."""
import struct
from rosbags.rosbag1 import Writer
from rosbags.typesys import Stores, get_typestore
ts = get_typestore(Stores.ROS1_NOETIC)
Imu, PC2, PF = (ts.types[f'sensor_msgs/msg/{n}'] for n in ('Imu', 'PointCloud2', 'PointField'))
Header, Time = ts.types['std_msgs/msg/Header'], ts.types['builtin_interfaces/msg/Time']
Quat, Vec3 = ts.types['geometry_msgs/msg/Quaternion'], ts.types['geometry_msgs/msg/Vector3']
H = lambda t, f: Header(seq=0, stamp=Time(sec=int(t), nanosec=int((t % 1) * 1e9)), frame_id=f)
imu_stream, scans, _ = simulate_room(seconds)
with Writer(path) as w:
ci = w.add_connection('/imu', Imu.__msgtype__, typestore=ts)
cp = w.add_connection('/velodyne_points', PC2.__msgtype__, typestore=ts)
for t, a, wv in imu_stream:
m = Imu(header=H(t, 'imu'), orientation=Quat(x=0., y=0., z=0., w=1.),
orientation_covariance=np.zeros(9),
angular_velocity=Vec3(x=wv[0], y=wv[1], z=wv[2]), angular_velocity_covariance=np.zeros(9),
linear_acceleration=Vec3(x=a[0], y=a[1], z=a[2]), linear_acceleration_covariance=np.zeros(9))
w.write(ci, int(t * 1e9), ts.serialize_ros1(m, Imu.__msgtype__))
flds = [PF(name=n, offset=o, datatype=7, count=1) for n, o in (('x', 0), ('y', 4), ('z', 8), ('time', 12))]
for sc in scans:
blob = b''.join(struct.pack('ffff', *p, d) for p, d in zip(sc['points'], sc['dts']))
m = PC2(header=H(sc['t_end'], 'lidar'), height=1, width=len(sc['points']), fields=flds,
is_bigendian=False, point_step=16, row_step=16 * len(sc['points']),
data=np.frombuffer(blob, np.uint8).copy(), is_dense=True)
w.write(cp, int(sc['t_end'] * 1e9), ts.serialize_ros1(m, PC2.__msgtype__))
def _ate(traj, truth):
err = [np.linalg.norm(p - truth(t)[0]) for t, p, _ in traj]
return float(np.sqrt(np.mean(np.square(err)))), float(err[-1])
if __name__ == '__main__':
import sys
if len(sys.argv) > 1: # python fastlio2_mini.py path/to/real.bag
# defaults target the HKU Livox Avia bag: its avia.yaml extrinsic + noise
imu, scans = read_bag(sys.argv[1], imu_topic='/livox/imu', lidar_topic='/livox/lidar')
traj, lmap = run_offline(imu, scans, T_LI=[0.04165, 0.02326, -0.0284],
acc_cov=0.1, gyr_cov=0.1, scan_voxel=0.5, init_secs=1.5)
ps = np.array([p for _, p, _ in traj])
plen = float(np.sum(np.linalg.norm(np.diff(ps, axis=0), axis=1)))
print(f"{len(traj)} poses, map={len(lmap.pts)} pts, path={plen:.2f} m, "
f"final={np.round(ps[-1], 2)}")
else: # self-contained demo: in-memory + a real .bag
imu, scans, truth = simulate_room()
traj, _ = run_offline(imu, scans)
print("in-memory : ATE rmse = %.3f m final = %.3f m" % _ate(traj, truth))
write_demo_bag('/tmp/fastlio2_demo.bag')
imu_b, scans_b = read_bag('/tmp/fastlio2_demo.bag',
imu_topic='/imu', lidar_topic='/velodyne_points')
traj_b, _ = run_offline(imu_b, scans_b)
print("via .bag : ATE rmse = %.3f m final = %.3f m" % _ate(traj_b, truth))
```
## Honest notes
- **It needs a decent IMU and initialization.** The filter assumes a high-rate IMU
(200 Hz+) and a short static period at start to estimate the gravity direction and biases.
Garbage init, garbage trajectory.
- **Geometric degeneracy is the real failure mode.** Point-to-plane constraints vanish in a
long featureless tunnel or an open field — the update becomes unobservable along the
degenerate direction and the IMU drift takes over. This is fundamental to LiDAR odometry,
not a bug.
- **It's odometry, not loop-closing SLAM.** FAST-LIO2 drifts slowly but has no global loop
closure; pair it with a pose-graph backend (e.g. a FAST-LIO-SLAM setup) if you need
globally consistent maps.
- **The "100 Hz" is real but hardware-shaped.** The headline rates assume the ikd-Tree and a
reasonable CPU; the per-scan cost grows with map density and point count, which is exactly
what the downsampling and the moving window are there to bound.
The thing I'd take away: FAST-LIO2 is not a pile of heuristics — it's *one* iterated
error-state Kalman filter, fed a deskewed point cloud, corrected by point-to-plane residuals,
over an incremental map, with a gain rewritten so thousands of measurements cost the same as
a few. Understand those five steps and you can rebuild it, and you understand the spine of
modern LiDAR SLAM.
---
*Built on [FAST-LIO2: Fast Direct LiDAR-Inertial Odometry](https://arxiv.org/abs/2107.06829)
(Xu, Cai, Bai, Zhang, 2021), the original [FAST-LIO](https://arxiv.org/abs/2010.08196) and
[ikd-Tree](https://arxiv.org/abs/2102.10808) papers, and the
[HKU-MARS/FAST_LIO](https://github.com/hku-mars/FAST_LIO) source. C++ snippets are from the
clean reimplementation [zlwang7/S-FAST_LIO](https://github.com/zlwang7/S-FAST_LIO); Python is
simplified for teaching.*
---
# How LLM inference works: prefill, decode, and where the time goes
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/how-llm-inference-works
> date: 2026-06-28
> tags: llm, inference-optimization, systems, kv-cache, explainer
When a model feels slow in production, the first question I ask is *which phase* is
slow. Because a single `generate()` call isn't one workload — it's two, with opposite
bottlenecks running on the same GPU:
- **prefill** processes the prompt and is **compute-bound**,
- **decode** generates tokens one at a time and is **memory-bound**.
Almost every inference optimization you'll read about targets one of these two phases.
So before reaching for a fix, you have to know which one is hurting. Here's the whole
pipeline, and where the time actually goes.
## From text to vectors
Before either phase, the text becomes numbers. A tokenizer — usually byte-pair encoding
(BPE) — splits the string into integer IDs from a vocabulary of roughly 50,000 entries.
Each ID indexes a row of the embedding table, a learned `[vocab_size, hidden_dim]`
matrix, so for `hidden_dim = 4096` every token becomes a 4096-dimensional vector.
Position is injected here. Modern models use rotary position embeddings (RoPE), which
encode position by *rotating* each query/key vector by an angle proportional to its
index, rather than adding a separate positional vector. It's cheap and it's what lets
the same weights generalize across lengths.
## Inside a layer
The embedded sequence flows through a stack of transformer layers — 32 for a 7B, 80+ for
the big ones. Each layer is two operations:
1. **Self-attention** projects every token into a query `Q`, key `K`, and value `V`.
Each token's query is scored against every token's key; scale, softmax, and the scores
become weights that mix the values. This is the only place information moves *between*
positions.
2. **Feed-forward network (FFN)** — a two-layer MLP applied to each token independently.
Attention routes information across positions; the FFN transforms it in place.
After the last layer, the final position's hidden state is projected back to vocabulary
size, softmaxed, and sampled — that's one output token. How that projection-and-sample
gets *driven* is exactly what differs between the two phases.
## Prefill: compute-bound
Prefill processes the entire prompt at once. `Q`, `K`, `V` are computed for every prompt
token in parallel, and attention is a big **matrix-matrix multiply**. That's dense
arithmetic, and it saturates the GPU's math units — utilization runs near 100%. The
metric that captures this phase is **Time To First Token (TTFT)**: how long before the
first output appears.
Prefill also populates the **KV cache** — the `K` and `V` tensors for every layer get
written to GPU memory so they never have to be recomputed. That cache is what makes the
next phase cheap, and also what makes it expensive.
## Decode: memory-bound
Once the first token exists, generation switches to one token per step. For each new
token the model computes `Q`, `K`, `V` for *that token only*; the keys and values for
everything before it are already cached. So the attention is one query vector against a
cached key matrix — a **matrix-vector** multiply, almost no arithmetic.
And yet decode is the slow part per token, because the GPU still has to **stream every
weight matrix and the entire KV cache out of memory** to do that tiny computation. The
bottleneck flips from arithmetic to **memory bandwidth**. The metric here is **Inter-Token
Latency (ITL)** — the gap between consecutive tokens, which is what makes a stream feel
fast or sluggish. GPU utilization during decode can sit at 30% on a fully loaded server,
because the math units are starved waiting on memory.
| | prefill | decode |
|---|---|---|
| Work | whole prompt, parallel | one token at a time |
| Attention shape | matrix × matrix | matrix × vector |
| Bottleneck | compute (arithmetic) | memory bandwidth |
| Metric | TTFT | ITL |
| GPU util | ~95% | ~30% |
| Optimize by | more FLOPs, better kernels | smaller cache, faster memory, batching |
## The KV cache runs the economics
The cache is the single most important object in LLM serving. Prefill writes one entry per
prompt token in a single pass; then each decode step appends exactly *one* entry and reuses
everything already there, recomputing nothing. Watch it accumulate — bright is written this
step, faded is reused:
That reuse is the whole point. Without the cache, generating a 1000-token response would
re-attend over the whole growing sequence every step — quadratic work. With it, each step
does constant new work — linear. Toggle it and watch the per-step cost, then drag the
context length to see what the cache costs in memory:
The trade is brutal and unavoidable: the cache grows linearly with sequence length, *per
layer*. For a 13B model it's roughly **1 MB per token**, so a 4K context is ~4 GB of VRAM
spent on cache alone — before a single weight. And that memory competes directly with
batch size: every gigabyte on cache is a gigabyte not serving another request. Long
contexts are expensive not because of compute, but because they evict concurrency.
The standard mitigations all attack the cache from different angles:
- **Quantize it** to INT8 or INT4 (it's just tensors).
- **Sliding-window attention** — drop tokens outside a fixed window.
- **Grouped-query attention (GQA)** — share `K`/`V` across attention heads so there are
fewer cached tensors. (This is exactly the change [iLLaDA](/articles/illada-diffusion-language-model)
and most modern models make.)
- **PagedAttention** — the trick behind vLLM: page the cache in fixed-size blocks like an
OS pages virtual memory, killing fragmentation and packing in more concurrent requests.
## Redesigning attention around the cache
The deeper move is to make the cache structurally smaller from the start, by changing
attention itself. DeepSeek's V4 series does this with a hybrid of two compressed
mechanisms: **Compressed Sparse Attention** (compress KV ~4× with softmax-gated pooling,
then attend sparsely) and **Heavily Compressed Attention** (consolidate KV across 128
tokens into one entry, attend densely over those). At a 1M-token context, V4-Pro needs
about **27% of the single-token inference FLOPs and 10% of the KV cache** of its
predecessor — in absolute terms, ~9.62 GiB of cache per sequence in bf16 versus an
estimated ~83.9 GiB for the older design, and fp4/fp8 halves it again. (I went deeper on
V4's drafter in the [DSpark write-up](/articles/deepseek-dspark).) The cache has become
the constraint the architecture is being designed *around*.
## Quantization
Training needs FP32/BF16 for gradient stability. Inference doesn't. Dropping bit-width
saves memory linearly, and quality barely moves. Pick a size and precision:
INT4 is the reason a 7B model runs on a 4–6 GB laptop GPU at all. Methods like GPTQ and
AWQ use per-channel scaling to keep the lossy compression within 1–2 points of full
precision on standard benchmarks. And going FP16 → INT8 often roughly halves latency with
negligible quality loss — which makes quantization the highest-leverage single change for
most deployments.
## The serving layer
On top of the prefill/decode loop sits the infrastructure that makes a GPU economical:
- **Continuous batching** interleaves tokens from many requests on the same GPU step. This
is the big one: decode leaves most of the arithmetic idle, so you fill that idle
capacity with other requests' tokens. It's why one GPU serves dozens of users at once.
- **Speculative decoding** drafts several tokens with a cheap model and verifies them in
one pass of the big model — turning sequential decode steps into one parallel
verification when acceptance is high. (Two whole articles' worth:
[DSpark](/articles/deepseek-dspark) and
[multi-token prediction](/articles/multi-token-prediction).)
- **PagedAttention** for the cache memory, as above.
Frameworks like vLLM, TensorRT-LLM, and TGI combine all of this. The throughput they get
comes mostly from the fact that decode is memory-bound, so there's spare arithmetic lying
around for batching to soak up.
## The full path
1. **Tokenize** — text → integer IDs via BPE.
2. **Embed** — IDs → vectors; RoPE rotates in position.
3. **Prefill** — all prompt tokens through every layer in parallel; compute-bound; KV
cache populated; first token emitted (TTFT).
4. **Decode loop** — one token per step: project `Q`, attend over cached `K`/`V`, run FFN,
sample, append to cache; memory-bound (ITL).
5. **Detokenize** — IDs → text, streamed out.
## How to actually use this
The whole point of splitting it this way is diagnosis. When something is slow:
- **Slow to start** → you're prefill-bound. Long prompts dominate TTFT; optimize the
prompt path (caching, chunked prefill, more compute).
- **Slow to stream** → you're decode-bound. Long outputs dominate ITL; the fix is *not*
more compute — it's a smaller cache, faster memory, or better batching.
- **Context length is never free.** It bloats the KV cache and directly cuts how many
requests fit on the GPU, so it shows up as reduced throughput long before it shows up as
an out-of-memory error.
That last instinct is the one I'd internalize: during decode the arithmetic units are
mostly idle, so when a decode-bound server is slow, throwing a bigger compute budget at it
does nothing. The bottleneck is the memory bus. Optimize the thing that's actually full.
---
# iLLaDA: how far a masked-diffusion language model scales
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/illada-diffusion-language-model
> date: 2026-06-28
> tags: llm, diffusion, language-models, architecture, explainer
Almost every language model you use is **autoregressive**: it factorizes text
left-to-right, $p(x) = \prod_i p(x_i \mid x_{
### The forward process: corrupt by masking
Pick a masking ratio $t \sim \mathcal{U}[0,1]$. Each token is independently replaced by a
special `[MASK]` token with probability $t$. At $t=0$ the sequence is clean; at $t=1$
it's fully masked. Drag $t$ and watch the corruption — and the loss weighting — change:
The model $p_\theta$ sees the corrupted sequence $x_t$ and is trained to predict the
*original* tokens at every masked position at once. The objective is a masked
cross-entropy, computed only on masked positions and reweighted by $1/t$:
$$
\mathcal{L}(\theta) \;=\; -\,\mathbb{E}_{t,\,x_0,\,x_t}\!\left[\frac{1}{t}\sum_{i=1}^{L}
\mathbf{1}\!\left[x_t^{i} = \mathrm{M}\right]\,\log p_\theta\!\left(x_0^{i} \mid x_t\right)\right]
$$
The indicator $\mathbf{1}[x_t^i = \mathrm{M}]$ restricts the loss to masked positions;
the $1/t$ factor re-normalizes so heavily- and lightly-masked samples both contribute
correctly. Averaged over $t$, this is a Monte-Carlo **upper bound on the negative
log-likelihood** — a principled training objective, not a heuristic. iLLaDA keeps this
*same* objective through pre-training **and** supervised fine-tuning.
### Bidirectional attention comes for free
An autoregressive model *must* hide the future — if token $i$ could attend to token
$i{+}1$, it would just read the answer it's supposed to predict. A masked diffusion model
predicts *masked* positions anywhere in the sequence, not "the next" one, so there's
nothing to hide. Every position attends to every other, left and right:
That full context is the structural argument for diffusion LMs: on infilling and tasks
where later text disambiguates earlier text, seeing both sides at every layer should
help.
### Generation: unmask in parallel, over a few steps
Generation runs the process backward. Start from a block of all-`[MASK]` tokens. Each
denoising step, the model predicts every masked position, **commits its most confident
predictions**, and **re-masks the low-confidence ones** to try again next step. A whole
block resolves over a handful of steps, in confidence order — not reading order. Flip
between the two paradigms:
This is the crux. Autoregression spends one forward pass per output token, in series.
Diffusion spends a fixed, smaller number of denoising passes over the whole block — the
promise being fewer sequential steps, at the cost of needing enough steps for quality.
## What iLLaDA changes over LLaDA
iLLaDA is, more than anything, a careful **scale-up** of LLaDA — proof that the recipe
keeps paying off with more tokens and a better post-training pass.
### A bigger, leaner backbone
The architecture is a standard dense Transformer (RMSNorm, SwiGLU, RoPE, no biases), but
re-tuned for cheaper inference:
| | iLLaDA | LLaDA |
|---|---|---|
| Attention heads | 32 | 32 |
| Key/Value heads | **8 (GQA)** | 32 (MHA) |
| FFN dim | 14,336 | 12,288 |
| Vocabulary | 155,136 | 126,464 |
| Max sequence length | **8192** | 4096 |
| Embedding / LM head | **tied** | untied |
| Total parameters | 7.62B | 8.02B |
The load-bearing change is **grouped-query attention** (8 KV heads instead of 32),
adopted to shrink the cached key/value footprint at inference — plus a larger vocab,
doubled context, and tied embeddings.
### The headline spend: 12T tokens
- **Pre-training: 12T tokens**, up ~5.2× from LLaDA's 2.3T. AdamW, weight decay 0.1, LR
warmed to $2\times10^{-4}$, held, then cosine-decayed to $5\times10^{-6}$.
- **SFT: a 25B-token instruction corpus for 12 epochs.** The new wrinkle: SFT now applies
the *same* masking as pre-training across the entire sequence (prompt, response, EOS),
rather than keeping the prompt fully visible — a more consistent objective end to end.
And the fine-tuning clearly hadn't saturated. The SFT-epoch ablation rises monotonically
through all 12 epochs (they stopped on compute, not convergence):
### Two inference-side tricks
- **Variable-length generation.** Instead of committing to a fixed output block and
denoising all of it, iLLaDA appends a mask block, runs the sampler, commits confident
tokens, and continues until termination — so it only denoises as many positions as the
answer needs, rather than padding to a worst case. (The paper argues the efficiency,
but — notably — does not report latency or step-count numbers for it.)
- **Confidence-based multiple-choice scoring.** Rather than a likelihood estimate, they
score a candidate by revealing its tokens one at a time, each step unmasking the
highest-confidence position, and summing the log-probs:
$$
S_{\text{conf}}(y \mid p) \;=\; \sum_k \log p_\theta\!\left(y^{i_k} \mid p,\, \tilde{y}_{k-1}\right),
\quad i_k = \arg\max_{i \in \mathcal{M}_{k-1}} p_\theta\!\left(y^i \mid p,\, \tilde{y}_{k-1}\right)
$$
The authors are upfront that this is "not a likelihood estimate" but a task-specific
surrogate. Its ablation is modest: +1.3 PIQA, +0.6 ARC-C, +2.3 HellaSwag over
likelihood scoring.
## The results, honestly
Two stories live in these tables, and they point in different directions.
### Base models: genuine parity with Qwen2.5
As a base model, iLLaDA improves broadly over LLaDA and lands **even with Qwen2.5-7B** on
average — winning several benchmarks outright:
The original LLaDA paper made the same point with its headline figure — its 8B base
model tracing out a wide envelope over the LLaMA baselines across general, math, code, and
Chinese tasks:
The per-benchmark picture, with the gains over LLaDA that the abstract leads on:
| Base | iLLaDA | LLaDA | Qwen2.5 | Δ vs LLaDA |
|---|---|---|---|---|
| MMLU | 74.8 | 65.9 | 71.9 | +8.9 |
| BBH | 71.3 | 49.7 | 63.9 | **+21.6** |
| ARC-C | 60.8 | 45.9 | 51.5 | **+14.9** |
| HellaSwag | 76.6 | 70.5 | 79.0 | +6.1 |
| GSM8K | 81.9 | 70.3 | 78.9 | +11.6 |
| MATH | 38.4 | 31.4 | 41.1 | +7.0 |
| HumanEval | 50.0 | 35.4 | 56.7 | +14.6 |
| MBPP | 57.8 | 40.0 | 63.6 | +17.8 |
iLLaDA-Base beats Qwen2.5-Base on MMLU, BBH, ARC-C, and GSM8K; Qwen still wins on
HellaSwag, MATH, and code. But the average edges ahead — and *that's the real result*:
a from-scratch bidirectional diffusion model matching a strong autoregressive base.
### Instruct models: the gap that's left
This is the part the abstract's "competitive on several benchmarks" softens. After
instruction tuning, iLLaDA **trails Qwen2.5 by ~10 points on average**, with double-digit
gaps on the hard reasoning and coding tasks:
| Instruct | iLLaDA | LLaDA | Qwen2.5 | Δ vs LLaDA |
|---|---|---|---|---|
| MMLU | 71.6 | 65.5 | 76.6 | +6.1 |
| MMLU-Pro | 52.3 | 37.0 | 56.3 | +15.3 |
| GSM8K | 89.0 | 77.5 | 91.6 | +11.5 |
| MATH | 56.7 | 42.2 | 75.5 | **+14.5** |
| HumanEval | 65.9 | 49.4 | 84.8 | **+16.5** |
| MBPP | 58.0 | 41.0 | 79.2 | +17.0 |
The improvement *over LLaDA* is huge and real (+12.6 average). The gap *to Qwen* on
MATH (56.7 vs 75.5), HumanEval (65.9 vs 84.8), and MBPP (58.0 vs 79.2) is also real, and
the authors don't hide it — they point to the lack of RL alignment as part of the cause.
## What I make of it
- **The base-model result is the one that matters, and it's solid.** A bidirectional
masked diffusion LM, trained from scratch, reaching autoregressive base parity is a
genuine data point: diffusion LMs *scale* like AR LMs. The paradigm is viable, not a
curiosity.
- **"Competitive" is doing some work in the abstract.** On instruction-tuned reasoning
and code, AR still wins by 10–20 points. Read the instruct tables before repeating the
headline.
- **The efficiency case is asserted, not measured.** GQA and variable-length generation
are motivated by cost, but the paper reports no sampling-step counts, no latency, no
tokens/sec — and the number of denoising passes is *exactly* diffusion's central
liability. "More efficient" is a design argument here, not a demonstrated result.
- **Parity wasn't cheap.** 12T tokens at 8B is a frontier-scale data spend, ~5× LLaDA and
on par with what strong AR models of this size consume. iLLaDA shows diffusion can reach
AR base parity — by paying full AR-scale training cost, and still trailing after
post-training. There's also an honest failure mode noted: the sampler can fall into
repetitive reasoning loops that need inference-time mitigation.
The fair summary: bidirectional masked diffusion is now a **scalable paradigm at parity
with autoregressive base models** — no longer something you can wave off — but not yet a
proven win on post-trained quality, and not yet a proven efficiency advantage. That's a
meaningful place to have gotten to, stated without the gloss.
---
*Built on [Improved Large Language Diffusion Models](https://arxiv.org/abs/2606.25331)
(Nie et al., Renmin University & ByteDance Seed, 2026) and its predecessor
[LLaDA](https://arxiv.org/abs/2502.09992). All numbers are from the paper's Tables 1–3;
the SFT-epoch curves are redrawn from its Figure 1.*
---
# The Kalman filter from first principles
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/kalman-filter
> date: 2026-06-28
> tags: state-estimation, kalman-filter, slam, robotics, explainer
I spend a lot of time on odometry and SLAM, and underneath almost all of it is the same
idea: you have a noisy sensor, and a model of how the thing you're tracking moves, and you
want to fuse them into a single best estimate that also tells you how much to trust itself.
That's the Kalman filter. It's the optimal recursive estimator for a linear system with
Gaussian noise, and once it clicks, you see it everywhere — GPS, IMUs, radar tracking, and
the LiDAR-inertial filters I'll build on in the next article.
Let me derive it the way it actually makes sense: as fusing two Gaussians.
## Two sources of information
You're tracking a state $\mathbf{x}$ — say the position and velocity of something. You have
two things:
1. A **process model**: how the state evolves on its own. For constant velocity,
$p_{t} = p_{t-1} + v\,\Delta t$. You can *predict* forward, but the prediction drifts —
it accumulates uncertainty.
2. A **measurement model**: a sensor that observes some function of the state, noisily.
A position sensor gives you $z = p + \text{noise}$.
Neither is enough alone. The prediction drifts; the sensor is noisy and jittery. The
Kalman filter combines them, and the key fact is that **combining two Gaussian estimates
yields a third Gaussian that is sharper than either input**.
That's the whole update step. If the prediction is $\mathcal{N}(\mu_1, \sigma_1^2)$ and the
measurement is $\mathcal{N}(\mu_2, \sigma_2^2)$, their normalized product has
$$
\mu = \mu_1 + K(\mu_2 - \mu_1), \qquad \sigma^2 = (1-K)\,\sigma_1^2,
\qquad K = \frac{\sigma_1^2}{\sigma_1^2 + \sigma_2^2}.
$$
$K$ is the **Kalman gain** — the fraction of the way you move from the prediction toward
the measurement, set by which one is more certain. Trust the sensor ($\sigma_2 \to 0$) and
$K \to 1$; trust the prediction ($\sigma_1 \to 0$) and $K \to 0$. Everything else is this,
generalized to vectors.
## The predict-update cycle
In general the state is a vector $\mathbf{x}$ with covariance $\mathbf{P}$, and the models
are linear maps with Gaussian noise:
$$
\mathbf{x}_t = \mathbf{F}\mathbf{x}_{t-1} + \mathbf{B}\mathbf{u}_t + \mathbf{w},
\quad \mathbf{w}\sim\mathcal{N}(0,\mathbf{Q});
\qquad
\mathbf{z}_t = \mathbf{H}\mathbf{x}_t + \mathbf{v}, \quad \mathbf{v}\sim\mathcal{N}(0,\mathbf{R}).
$$
$\mathbf{F}$ is the state-transition, $\mathbf{H}$ maps state to measurement, $\mathbf{Q}$
is process noise (how much the model drifts), $\mathbf{R}$ is measurement noise. The filter
alternates two steps forever:
**Predict** — push the state and its uncertainty forward through the model:
$$
\hat{\mathbf{x}} = \mathbf{F}\mathbf{x} + \mathbf{B}\mathbf{u}, \qquad
\hat{\mathbf{P}} = \mathbf{F}\mathbf{P}\mathbf{F}^{\!\top} + \mathbf{Q}.
$$
**Update** — correct with the measurement, by the matrix Kalman gain:
$$
\mathbf{y} = \mathbf{z} - \mathbf{H}\hat{\mathbf{x}}
\quad(\text{innovation}), \qquad
\mathbf{S} = \mathbf{H}\hat{\mathbf{P}}\mathbf{H}^{\!\top} + \mathbf{R}
\quad(\text{innovation covariance}),
$$
$$
\mathbf{K} = \hat{\mathbf{P}}\mathbf{H}^{\!\top}\mathbf{S}^{-1}, \qquad
\mathbf{x} = \hat{\mathbf{x}} + \mathbf{K}\mathbf{y}, \qquad
\mathbf{P} = (\mathbf{I} - \mathbf{K}\mathbf{H})\hat{\mathbf{P}}.
$$
Watch it run on a constant-velocity tracker. The state is $[\text{position}, \text{velocity}]$,
the sensor sees only position, and the filter has to infer velocity and smooth the noise.
Drag $R$ and $Q$ to feel the trust trade-off the gain encodes:
## In code
The whole thing is a few lines of linear algebra. Here's the constant-velocity tracker
above, in NumPy — no framework, runnable as-is:
```python
import numpy as np
dt = 1.0
F = np.array([[1, dt], # constant-velocity transition
[0, 1]])
H = np.array([[1.0, 0.0]]) # observe position only
Q = 0.5 * np.array([[dt**3/3, dt**2/2], # process noise (drift)
[dt**2/2, dt]])
R = np.array([[49.0]]) # measurement noise (sensor variance)
x = np.array([[0.0], [0.0]]) # initial state
P = np.eye(2) * 50.0 # initial uncertainty
def step(x, P, z):
# predict
x = F @ x
P = F @ P @ F.T + Q
# update
y = z - H @ x # innovation
S = H @ P @ H.T + R # innovation covariance
K = P @ H.T @ np.linalg.inv(S) # Kalman gain
x = x + K @ y
P = (np.eye(2) - K @ H) @ P
return x, P
for z in measurements: # stream of noisy position readings
x, P = step(x, P, np.array([[z]]))
# x[0] is the smoothed position estimate, P[0,0] its variance
```
Two knobs do all the tuning. `R` says how noisy the sensor is; `Q` says how much you let
the model drift. Get their ratio right and the filter is optimal; get it wrong and it either
lags reality (`Q` too small) or chases noise (`R` too small). In practice you start from the
sensor's datasheet for `R` and tune `Q` until the innovation $\mathbf{y}$ looks like white
noise.
For numerical robustness in real systems, use the **Joseph form** of the covariance update,
$\mathbf{P} = (\mathbf{I}-\mathbf{KH})\hat{\mathbf{P}}(\mathbf{I}-\mathbf{KH})^{\!\top} +
\mathbf{KRK}^{\!\top}$, which stays symmetric positive-definite even with floating-point
error, where the compact $(\mathbf{I}-\mathbf{KH})\hat{\mathbf{P}}$ can drift and diverge.
## When the world isn't linear
The plain Kalman filter assumes $\mathbf{F}$ and $\mathbf{H}$ are linear. Robotics is full
of rotations and projections that aren't. Three extensions matter, and the last one is the
bridge to LiDAR-inertial SLAM:
- **Extended KF (EKF).** The dynamics $f$ and measurement $h$ are nonlinear, so you
*linearize* them at the current estimate: use the Jacobians
$\mathbf{F} = \left.\frac{\partial f}{\partial \mathbf{x}}\right|_{\hat{\mathbf{x}}}$ and
$\mathbf{H} = \left.\frac{\partial h}{\partial \mathbf{x}}\right|_{\hat{\mathbf{x}}}$ in
place of the matrices, run the same equations. Cheap, and it works when the nonlinearity
is mild over one step.
- **Iterated EKF (iEKF).** One linearization point can be bad if the prior is far from the
truth. So *relinearize*: after the update, recompute $\mathbf{H}$ at the new estimate and
redo the update, iterating until it converges. Each iteration is a Gauss-Newton step on
the maximum-a-posteriori objective. This is what FAST-LIO uses — it matters because the
point-to-plane LiDAR residual is very nonlinear in the pose.
- **Error-state / on-manifold KF.** You can't add a vector to a rotation
$\mathbf{R}\in SO(3)$ and stay on the manifold. So you track the state on the manifold but
the *error* (and its covariance) in the tangent space, fusing with the $\boxplus/\boxminus$
operators instead of $+/-$. This keeps rotations valid and the covariance minimal
(3 numbers for orientation, not 9).
That last combination — an **iterated, error-state Kalman filter on a manifold** — is
exactly the engine inside FAST-LIO2, fusing a high-rate IMU prediction with thousands of raw
LiDAR points per scan. The Kalman gain there gets one more clever rewrite to handle those
thousands of measurements cheaply, which is where I'll pick up next.
## The one-paragraph summary
A Kalman filter holds a Gaussian belief over a state. **Predict** moves the belief through a
motion model and inflates its uncertainty; **update** multiplies it by the measurement's
Gaussian, which sharpens it and pulls the mean toward the data by the gain $\mathbf{K}$ —
optimally weighted by which source is more certain. Linearize for nonlinear systems (EKF),
relinearize for hard ones (iEKF), and track the error in the tangent space for rotations.
That's the whole toolkit, and it runs real robots.
---
# MegaTrain: training a 120B model on one GPU by inverting where memory lives
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/megatrain-single-gpu-training
> date: 2026-06-28
> tags: llm, training, systems, cuda, explainer
The thing that stops most people from training a big model isn't FLOPs — it's GPU
memory. A 70B model's parameters, gradients, and Adam moments don't fit in an 80GB card,
so you reach for tensor/pipeline parallelism across a cluster you may not have, or for
offloading systems that thrash and OOM as the model grows.
**MegaTrain** takes the other path: invert the memory hierarchy. Host CPU memory becomes
the authoritative store for *all* persistent state — parameters, gradients, and optimizer
moments — and the GPU is demoted to a **transient compute engine** that holds only the
layer it's working on right now. On a single H200 with 1.5 TB of host RAM, that's enough
to train models up to **120B parameters at full precision** — no quantization, no second
GPU. It's a systems paper, and a good one, so let's read it as systems.
## The inversion
Start with the accounting. Mixed-precision Adam costs about **12 bytes per parameter**: 2
for the BF16 weight, 2 for the BF16 gradient, and 8 for the FP32 first and second moments.
A GPU-centric trainer keeps all of that in HBM, so the moment $12 \times \text{params}$
exceeds the card, you're done. Move that state to host memory and stream one layer at a
time, and the device footprint goes flat while the host's terabytes set the ceiling. Drag
the model size and watch where it OOMs:
That's the whole idea in one slider: a single H200's 141 GB HBM caps a GPU-centric run
around 11–12B parameters, but with the persistent state in 1.5 TB of host RAM and only a
streamed layer resident, the same card reaches past 120B. The paper's own architecture
makes the split concrete — everything lives in the CPU domain; the GPU domain is scratch:
## The bottleneck, and the step
Inverting the layout creates one obvious problem: **PCIe bandwidth**. Every layer's
weights now have to cross the bus into the GPU, and its gradients have to cross back —
~128 GB/s on H200's PCIe Gen5, versus 4.8 TB/s of on-card HBM. If you did this naively the
GPU would spend most of its life waiting on DMA.
The training step is three phases built to keep that traffic off the critical path:
1. **Streaming forward** — stream each layer's weights in (H2D), compute, checkpoint
activations every $K$ layers, release the weights immediately.
2. **Block-wise backward** — recompute activations from the nearest checkpoint, stream the
layer's weights back in, compute gradients in reverse, offload them (D2H), release.
3. **Optimizer update** — run Adam **entirely on the CPU** (AVX-512), so the freshly
computed gradients and the moments never make another round trip to the device.
Block-wise recomputation bounds activation memory at $O(N \cdot A_{\max} \cdot L/K)$ —
independent of total depth — which is what lets depth scale without the activation memory
exploding.
## The double-buffered pipeline
This is the optimization that makes it fast instead of merely possible. MegaTrain runs
**three CUDA streams** concurrently — one for compute, one for H2D weight transfer, one
for D2H gradient evacuation — and double-buffers the weights so that while the compute
stream works on layer $i$ out of one buffer, the next layer's weights prefetch into the
other. Flip between naive serialization and the double-buffered schedule:
The coordination is three events — *weights-ready*, *backward-done*, *buffer-free* — and
the payoff is a compute lane with no gaps: the GPU never stalls on PCIe. The ablation
makes the importance unambiguous. Remove double-buffering and throughput drops **31.3%**
(266 → 183 TFLOPS at 14B) — by far the largest single contributor, more than the gradient
slab pool (−3.3%) or tighter checkpointing. Here's the paper's own timeline of the overlap:
A few more systems details earn their keep:
- **Stateless layer templates.** A persistent autograd graph assumes weights stay
resident — incompatible with streaming and eviction. MegaTrain uses kernel templates
with no baked-in weight pointers and a `Bind` primitive that maps streamed buffer views
into the template's input slots, so device memory never exceeds a single layer.
- **Layer-contiguous tiling.** BF16 weights, BF16 grads, and FP32 moments for a layer are
packed into one 4 KB-aligned block, so a layer moves as a single large-burst DMA that
saturates PCIe instead of many fragmented transfers.
- **Pinned slab pool.** A fixed pool of pinned staging slabs (default 12), each sized to
the *largest layer* rather than the whole model, JIT-packed by a CPU worker — you get
pinned-memory transfer speed without pinning the entire model.
## What it delivers
The headline is a capability, and it's the most convincing part: **120B parameters on one
H200**, and a **512K-token context on a single GH200**. These are regimes where the
offload baselines simply OOM — so the comparison is binary, which is the strongest kind.
Where the baselines *can* run but are memory-starved — a 14B model on a PCIe A100 — the
margin is large:
That's 8.1× over Gemini and 12.2× over ZeRO-3 — and on a 48 GB A6000 or a 24 GB RTX 3090,
MegaTrain trains 14B at all (56.8 and 30.2 TFLOPS) while ZeRO-3 OOMs. Crucially, accuracy
doesn't move — full precision means no drift:
| MetaMathQA accuracy | MegaTrain | ZeRO-3 | ZeRO-Infinity | PyTorch |
|---|---|---|---|---|
| 7B | 88.99 | 88.93 | 88.97 | 88.91 |
| 14B | 92.52 | 92.41 | — | — |
On depth, it's the only system that keeps going: ZeRO-3 OOMs by 132 layers and FSDP by 84,
while MegaTrain runs the whole range and is **6.14× faster than FSDP at 56 layers**. On
width, both baselines OOM at 4.0× while MegaTrain alone reaches 5.0×.
## What I make of it
- **The capability claims are real and well-supported.** Training 120B at full precision
on one GPU, 512K context on one GH200, and 14B on a 3090 — in each case the baseline
*cannot run*. "Only system that works here" is the most honest result a systems paper
can have, and the accuracy-parity table backs the full-precision claim.
- **Double-buffering is the load-bearing idea, and the ablation proves it.** −31.3% without
it is a clean, isolated attribution. The rest — contiguous tiling, stateless templates,
CPU Adam — are the supporting cast that make the streaming viable.
- **Read the throughput claims with the regime attached.** This is single-GPU only — no
multi-node scaling, and the metric is TFLOPS, not MFU, which a recomputation-heavy design
inflates (you do extra FLOPs re-deriving activations). At small or unconstrained sizes
the baselines are actually *faster* (FSDP 501 vs MegaTrain 406 TFLOPS at 1.0× width);
MegaTrain wins specifically once the model is large enough that offload systems are
thrashing or out of memory. The 1.84× and 6–12× numbers live near that memory cliff, not
everywhere.
- **The best numbers lean on expensive hardware.** GH200's 900 GB/s NVLink-C2C and H200's
1.5 TB host RAM do a lot of work; on a plain PCIe Gen4 box the absolute throughput is far
lower (122 vs 266 TFLOPS). So it democratizes *what fits*, more than it democratizes
*speed*.
The honest summary: MegaTrain redefines the memory ceiling for single-GPU training —
provably, at full precision, with public code — and the double-buffered pipeline is a
genuinely nice piece of CUDA-stream engineering. Just don't read "1.84× faster" as a
general speedup; read it as "it runs, fast enough, where nothing else runs at all."
---
*Built on [MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a
Single GPU](https://arxiv.org/abs/2604.05091) (Yuan et al., 2026;
[code](https://github.com/DLYuanGod/MegaTrain)). All numbers are from the paper's tables
and figures.*
---
# DeepSeek DSpark: making speculative decoding draft better and verify smarter
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/deepseek-dspark
> date: 2026-06-27
> tags: llm, inference-optimization, speculative-decoding, deepseek, explainer
The first thing to get straight: **DSpark is not a new model.** The Hugging Face
cards say it plainly — `DeepSeek-V4-Flash-DSpark` "is not a new model. It is the
same checkpoint with an additional speculative decoding module attached." DSpark is
an *inference accelerator* for the existing DeepSeek-V4 weights, shipped alongside an
open training repo called [DeepSpec](https://github.com/deepseek-ai/DeepSpec). It
makes generation faster without changing a single output token.
That last clause is the whole reason to care. Speculative decoding is **lossless** by
construction: a cheap draft model proposes a block of tokens, the full target model
verifies the whole block in one forward pass, and the acceptance rule keeps exactly
the prefix the target would have produced anyway, plus one free "bonus" token. Output
is bit-identical to plain decoding. You only buy latency.
So the entire game is the draft-and-verify loop, and DSpark improves both halves of
it. Recall the per-token latency of speculative decoding:
$$
t_{\text{token}} \;\approx\; \frac{t_{\text{draft}} + t_{\text{verify}}}{g}
$$
where $g$ is the **accepted length** — how many real tokens one expensive target
forward bought. You win three ways: draft faster, draft better (raise $g$), or verify
smarter. Prior work chased the first two. DSpark is the first to seriously attack the
third.
## What's being accelerated: DeepSeek-V4
DSpark exists to serve [DeepSeek-V4](https://arxiv.org/abs/2606.19348) — two MoE models,
**V4-Pro** (1.6T params, 49B activated) and **V4-Flash** (284B, 13B activated), both with
**1M-token context**. V4 is already aggressively efficiency-engineered: a hybrid of
Compressed Sparse Attention and Heavily Compressed Attention, Manifold-Constrained
Hyper-Connections (mHC), and the Muon optimizer, trained on >32T tokens.
That context matters: when the model itself is this optimized, the decoding loop is where
the last big latency wins live, and speculative decoding is the lever. DSpark is the
drafter that lever needed.
## The loop, one round at a time
First, watch the loop run end to end. The drafter proposes a block (faint), the target
verifies it in one forward, the matching prefix locks in and a mismatch is corrected for
free — and the meters track the mean accepted length $g$, which *is* the speedup over
vanilla one-token-at-a-time decoding:
Now the same thing in slow motion, one round at a time, so the accept/reject/bonus rule
is unambiguous — step through and watch $g$ change round to round:
The mismatch case is the one to internalize. When the target disagrees at position
$k$, everything after $k$ is thrown away — but because the target *also* tells you its
own token at $k$, the round still nets $k+1$ accepted tokens. You never lose ground,
and you never change the answer. The only question is how big $g$ gets on average.
## Why long draft blocks were a trap
Early drafters were **autoregressive** — each draft token conditions on the previous
one (EAGLE-style). Quality is high, but drafting latency grows linearly with block
size, so you're forced into short, shallow blocks.
**Parallel drafters** (DFlash, Medusa) flipped this: produce all draft logits in one
forward pass, so drafting latency is nearly independent of block size. In principle
you can now draft long blocks cheaply. In practice two things break:
- **Quality.** Each position is predicted independently, so it can't condition on the
tokens actually sampled elsewhere in the block. Given a context with two plausible
continuations — "of course" and "no problem" — a parallel drafter happily emits
"of problem" or "no course", because each slot marginalizes over all predecessors
instead of committing to one. Acceptance decays fast down the block.
- **System efficiency.** Even when long blocks *are* good, indiscriminately verifying
all of them wastes target-model batch capacity. Under high concurrency that
capacity is the bottleneck, and verifying tokens that will be rejected is pure loss.
DSpark's two components map one-to-one onto these two failures.
## Component 1: semi-autoregressive drafting
The fix for the quality problem is to put a *little* sequentiality back, cheaply. A
heavy **parallel backbone** (DeepSeek uses DFlash here) runs one forward pass over the
whole block and emits per-position hidden states $h_1,\dots,h_\gamma$ and base logits.
Then a **lightweight sequential head** runs over those, injecting intra-block
dependencies so position $j$ can finally see the token sampled at $j-1$.
The released config keeps the head tiny: a draft network of **three MoE layers** with
mHC and sliding-window attention of 128, max block size $W=5$. The sequential head
comes in two flavors:
- **Markov head** — first-order, memoryless: position $j$ conditions only on the
immediately preceding sampled token. Cheap, and scales to large vocabularies. Once
position 1 samples "of", the Markov head boosts "course" and suppresses "problem"
at position 2 — exactly the collision the parallel drafter couldn't avoid.
- **RNN head** — carries more history than the memoryless Markov variant, at a little
more cost.
The shipped drafter, "DSpark-5", uses the Markov head. It keeps almost all of the
parallel drafter's speed — drafting latency is still nearly flat in $W$ — while
recovering the acceptance rate a fully-parallel block throws away.
The parallel backbone is [DFlash](https://arxiv.org/abs/2602.06036) (ICML 2026), which
fuses the target model's context features into the draft model's KV cache so a single
forward pass can predict the whole block:
On the offline metric that isolates draft quality — macro-average accepted length per
round, target models Qwen3-4B/8B/14B at temperature 1.0 across the DeepSpec eval suite
— DSpark's semi-autoregressive drafter beats both the autoregressive and the
fully-parallel baselines:
Roughly **+27–31% over the autoregressive EAGLE-3** and **+16–18% over the parallel
DFlash** it's built on — the semi-autoregressive head recovers most of what pure
parallelism gave up, without paying EAGLE-3's per-token drafting cost.
## Component 2: confidence-scheduled, load-aware verification
This is the genuinely new lever. Bolt a **confidence head** onto the drafter, trained
end-to-end and then *post-hoc calibrated* — the paper cares about calibration error
(ECE), not just ranking, because the scores have to mean something. The head estimates
per-position prefix-survival probabilities. Then a **hardware-aware scheduler** reads
live engine throughput and chooses, per request, how much of the draft block to bother
verifying.
The intuition: verification consumes target-model batch capacity, which is the scarce
resource under concurrency. Spending it on a low-confidence tail token the target will
reject is pure waste. So trim the block to its confident head when the system is busy,
and verify everything when it's idle.
There's a real systems subtlety underneath the slider. To avoid GPU pipeline stalls —
you'd need the *next* step's capacity estimate before the current step finishes — the
scheduler approximates upcoming capacity using confidence-head outputs from **two
steps prior**, while still sorting candidate tokens by up-to-date cumulative
confidence. The two-steps-stale signal only sets the dynamic truncation length; the
acceptance itself is always exact, so the lossless guarantee holds.
Calibration — not just ranking — is the reason this works. The paper's reliability
diagram shows the raw confidence estimator already *discriminates* well (it ranks
survivors above doomed tokens) but is poorly *calibrated* — a raw score of 0.8 doesn't
mean an 80% survival chance. A scheduler that truncates on a probability threshold needs
the second property, not just the first, so DSpark calibrates the head post-hoc and
measures ECE. Once it's calibrated, the threshold means what it says: as it tightens, the
acceptance rate among verified tokens climbs from roughly 76.9% / 67.6% / 45.7% to about
92.5% / 92.0% / 95.7% on Math / Code / Chat respectively — the scheduler is keeping the
tokens that actually survive.
DSpark also studied how deep to make the drafter and how long to draft. Deeper drafters
help up to a point (the released config uses three MoE layers), and accepted length keeps
rising with proposal length $W$ where a fully-parallel DFlash block would have decayed —
which is the whole argument for the semi-autoregressive design, and why $W=5$ is a
sensible default rather than a hard ceiling.
## What it does in production
DSpark-5 replaced the previous production setup (a static MTP-1 single-token drafter)
on DeepSeek's own V4 serving engines. MTP-1 was the incumbent precisely *because*
naively deploying a static multi-token drafter degrades aggregate throughput under
high concurrency — the exact problem the scheduler exists to solve.
The honest reading of those annotations matters. At matched, practical throughput,
DSpark accelerates per-user generation by **60–85% on V4-Flash** and **57–78% on
V4-Pro**. The eye-popping "+661% throughput" point is a *specific operating regime* —
a strict 120 tok/s/user SLA where the single-token baseline is already pinned at its
operational boundary. The paper itself flags it as evidence of "extending the feasible
interactivity frontier," not a representative multiplicative speedup. Don't quote +661%
as a generic number; quote the 57–85% per-user range.
The gains concentrate where GPUs are *under*-utilized — low batch, strict latency,
RL-style long-tail decoding. When the system is already compute-saturated, smarter
verification has less slack to recover, and the benefit shrinks. That's the tradeoff:
DSpark buys interactivity, and interactivity is worth most exactly when you have spare
compute to spend on it.
## Where DSpark sits
| | Autoregressive (EAGLE-3) | Parallel (DFlash) | DSpark |
|---|---|---|---|
| Draft cost vs block size | grows linearly | ~flat | ~flat |
| Intra-block dependency | full | none | first-order (Markov head) |
| Acceptance decay | low | rapid | low |
| Verification length | fixed | fixed | scheduled per request |
| Lossless | yes | yes | yes |
EAGLE-3 drafts well but slowly; DFlash drafts fast but loosely; DSpark keeps DFlash's
parallel speed, threads just enough sequential dependency back in to fix acceptance,
and then adds the verification scheduler nobody else had. It's built directly on
DFlash (the parallel backbone) and DeepSeek-V4 (the target), and ships open under MIT.
## What I make of it
- **The framing is right.** Speculative decoding's latency formula has three levers,
and "verify smarter" was the neglected one. A calibrated confidence head plus a
load-aware scheduler is a clean, principled way to pull it — and because acceptance
stays exact, it costs zero quality.
- **The semi-autoregressive head is the quiet win.** A first-order Markov head is
almost free and recovers most of the acceptance a parallel block throws away. That's
a better engineering trade than going back to a slow autoregressive drafter.
- **Read the numbers carefully.** The +16–31% accepted-length gains are clean
apples-to-apples and should reproduce via DeepSpec. The production speedups are real
but regime-dependent; the headline ratio is a boundary artifact, not a uniform
multiplier. DSpark shifts the Pareto frontier — it doesn't move every point on it by
6×.
---
*Built on DeepSeek's [DSpark: Confidence-Scheduled Speculative Decoding with
Semi-Autoregressive Generation](https://github.com/deepseek-ai/DeepSpec) (paper in the
DeepSpec repo), the [DFlash](https://arxiv.org/abs/2602.06036) parallel drafter it
extends (ICML 2026), and [DeepSeek-V4](https://arxiv.org/abs/2606.19348), the target
it accelerates. Weights: [V4-Flash-DSpark](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark)
and [V4-Pro-DSpark](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-DSpark), MIT-licensed.*
---
# Multi-token prediction: training a model to see further than one step
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/multi-token-prediction
> date: 2026-06-27
> tags: llm, multi-token-prediction, inference-optimization, speculative-decoding, explainer
Standard language models are trained on a deceptively narrow task: given everything so
far, predict the *single* next token. The objective is one cross-entropy term per
position,
$$
L_{\text{next}} \;=\; -\sum_t \log P_\theta\big(x_{t+1}\mid x_{\le t}\big).
$$
**Multi-token prediction (MTP)** changes one thing: from each position, predict the
next $n$ tokens at once. That small change buys two unrelated-looking wins — a model
that *trains* better, and a model that *decodes* faster — and the second one is why
it's now in shipping products from DeepSeek to Google.
A provenance note up front, because the "Google multi-token prediction" framing gets
the credit wrong. The modern MTP training objective was defined by **Meta (FAIR)** in
[Gloeckle et al., 2024](https://arxiv.org/abs/2404.19737). The *predict-several-then-
verify* decoding idea it reuses traces to **Google Brain's** 2018 [blockwise parallel
decoding](https://arxiv.org/abs/1811.03115). **DeepSeek-V3** productionized a sequential
variant at pretraining scale. And **Google's** genuine 2026 contribution is applied:
MTP-based speculative decoding in Gemma 4 and on-device Gemini Nano. I'll attribute as
I go.
## The objective: predict n futures
Keep one **shared transformer trunk** that turns the context into a latent $z_t$. Attach
$n$ output heads. The loss sums cross-entropy over the next $n$ positions:
$$
L_{\text{MTP}} \;=\; -\sum_t \sum_{i=1}^{n} \log P_\theta\big(x_{t+i}\mid z_t\big).
$$
From position $t$, head $i$ predicts $x_{t+i}$. Pick $n$ and a flavor and read the loss
off directly:
The obvious worry is memory: materializing $n$ full vocabulary-sized logit tensors at
once is brutal. The fix is mundane and important — compute each head's forward/backward
*sequentially* and accumulate gradients at the trunk, so peak memory stays flat in $n$.
You pay a little time, not a lot of VRAM.
## Two flavors: parallel heads vs sequential modules
The flavor toggle above is the real architectural fork.
- **Meta / Medusa — parallel independent heads.** All heads hang off the same trunk and
predict in parallel; head 2 does *not* see head 1's token. Cheap, and it composes with
a clean inference trick: at deployment you can **discard the extra heads** and recover
an ordinary next-token model with zero overhead, or **keep them** to self-speculate.
- **DeepSeek-V3 — sequential modules that keep the causal chain.** Module $k$ takes the
previous depth's hidden state, concatenates the embedding of the (already-known) token,
RMS-norms both, projects, and runs its own transformer block. So depth-2's prediction
is conditioned on depth-1's token — the drafts are internally coherent, at more cost
than parallel heads.
$$
h'^{\,k}_i \;=\; M_k\big[\operatorname{RMSNorm}(h^{\,k-1}_i)\,;\ \operatorname{RMSNorm}(\operatorname{Emb}(t_{i+k}))\big]
$$
The trade is exactly what you'd guess: parallel heads are cheaper and discardable;
sequential modules draft more coherent blocks because each step sees the last.
## Win one: it trains a better model
The training objective is the whole reason for the quality gain. Slide the window: from
each position the model predicts the next $n$ tokens, so $n$ loss terms fire where an
ordinary model gets one. Flip $n$ between 1 and 4 to feel the supervision get denser:
Predicting further is a denser, more demanding signal, and at scale it produces a model
that's better even when you *throw the extra heads away*. Meta's 13B model, trained with
MTP, solves materially more coding problems than the matched next-token model on the
same data and compute:
Two caveats keep this honest. The gains **scale with model size** — small models barely
benefit, and very large $n$ erodes quality ($n=4$ is the sweet spot for ~7B on code).
And whether the *quality* gain transfers beyond pretraining is genuinely contested:
["Multi-Token Prediction Needs Registers"](https://arxiv.org/abs/2505.10518) (MuToR,
NeurIPS 2025) exists precisely because the benefit hasn't consistently generalized to
fine-tuning without help. MTP is not a free quality lunch in every regime.
## Win two: it decodes ~3x faster
The extra heads are a *built-in draft model*. They cheaply propose the next $n-1$ tokens
in one pass; the model then verifies all of them in a single batched forward and accepts
the longest correct prefix — **self-speculative decoding**. This is lossless: rejected
drafts fall back to the true next-token distribution, so output is unchanged.
That verify-and-accept scheme is the part with Google's fingerprints — Stern, Shazeer &
Uszkoreit's 2018 **blockwise parallel decoding** at Google Brain is the ancestor:
predict several future positions with auxiliary heads, then accept the longest correct
prefix. MTP simply folds those auxiliary heads into the training objective.
How much you gain depends entirely on the **acceptance rate** — how often a drafted
token matches what the target would have produced. Because acceptance compounds along
the block, the marginal value of the $i$-th drafted token decays, and the whole scheme
hits a ceiling no matter how far you draft. That's why a *more coherent* drafter is worth
more than a *longer* one — and why DeepSeek's sequential modules (higher acceptance) and
DSpark's semi-autoregressive head exist at all. Drag the acceptance rate and block size:
The speedups, across the lineage:
DeepSeek-V3 is the cleanest production data point: with MTP depth $D=1$ (predict two
tokens total), the **second token is accepted 85–90%** of the time, and repurposing the
module for speculative decoding gives **~1.8× tokens/sec**. The training loss weight was
annealed — $\lambda = 0.3$ for the first 10T tokens, then $0.1$ for the remaining 4.8T —
a detail worth noting because MTP at pretraining is a *secondary* objective, not the
main one.
## Google's actual 2026 role: applied MTP
Where Google genuinely shows up is deployment, not the objective:
- **Gemma 4 MTP drafters** (April 2026): MTP-style draft heads for lossless speculative
decoding, with a runtime heuristic that adapts how many tokens to draft. Google's
release claims **up to ~3x faster inference, no quality loss** — present that as a
vendor figure; the official docs only say "significant speedups".
- **Gemini Nano frozen MTP** (June 2026): a *frozen-backbone* MTP head for on-device
speculative decoding on Pixel. It correctly predicts ~2 extra tokens per pass, gives a
**50%+ speedup on Pixel 9** over a standalone drafter, lifts token acceptance ~55% on
structured text, and — the on-device kicker — costs **−130MB per instance** by sharing
the KV cache zero-copy instead of running a separate draft model.
The on-device framing is the interesting one: a separate draft model is a non-starter
when you're counting megabytes on a phone, so folding the drafter into the main model as
a frozen head is exactly the right move.
## The open questions
MTP isn't a closed book — two recent threads are worth knowing because they bound where
the simple story breaks.
- **Does the quality gain survive fine-tuning?** Meta's gains are a *pretraining*
phenomenon, and they don't reliably transfer when you only have a fine-tuning budget.
["Multi-Token Prediction Needs Registers"](https://arxiv.org/abs/2505.10518) (MuToR,
NeurIPS 2025) addresses this by interleaving learnable **register tokens** into the
sequence, each responsible for predicting a future token — adding almost no parameters
and no architectural surgery, so MTP's benefit shows up in the fine-tuning regime where
plain MTP heads underdeliver.
- **Can you extract more drafts from a model that already exists?** Apple's ["Your LLM
Knows the Future"](https://arxiv.org/abs/2507.11851) argues a standard model already
encodes multi-token information, and unlocks it with **masked-input MTP** plus a gated
LoRA and a learnable sampler — reporting roughly **5× on code/math** and **2.5× on
general chat**, lossless. The framing is telling: MTP capability may be latent in
next-token models, waiting for the right decoding head.
Both reinforce the same lesson the speedup curve shows: the value is in *acceptance and
coherence*, and the active research is about getting more of both without paying a full
retrain.
## Who did what
| Work | Org | Contribution |
|---|---|---|
| Blockwise parallel decoding (2018) | Google Brain | predict-several-then-verify/accept — the decoding ancestor |
| Better & Faster LLMs via MTP (2024) | Meta / FAIR | the canonical MTP training objective (n parallel heads) |
| Medusa (2024) | academic | multiple decoding heads + tree attention (not Google) |
| DeepSeek-V3 MTP (2024) | DeepSeek-AI | sequential MTP modules at pretraining; ~1.8× TPS |
| MuToR — "MTP needs registers" (2025) | academic | register tokens so MTP helps in fine-tuning |
| Gemma 4 / Gemini Nano MTP (2026) | Google | applied MTP speculative decoding, incl. on-device |
## What I make of it
- **One change, two payoffs.** Predicting $n$ futures is a denser training signal *and*
a free draft model. That two-for-one is why MTP spread so fast from a 2024 paper to
2026 phones.
- **The flavors matter.** Parallel heads are cheap and discardable; sequential modules
draft coherent blocks. Pick by whether you care more about training overhead or draft
acceptance.
- **Keep the credit straight.** Meta defined the objective, Google Brain seeded the
verify/accept decoding, DeepSeek productionized it, and Google's 2026 work is applied
speculative decoding — strongest exactly where a separate draft model can't fit, like
on-device.
- **Mind the caveats.** Quality gains scale with size and don't automatically survive
fine-tuning; the headline speedups are real but partly vendor-reported. The lossless
*speed* win is the part to trust unconditionally — it's guaranteed by the acceptance
rule, not a benchmark.
---
*Built on Meta's [Better & Faster Large Language Models via Multi-token
Prediction](https://arxiv.org/abs/2404.19737), the [DeepSeek-V3 Technical
Report](https://arxiv.org/abs/2412.19437) (§2.2), Google Brain's [Blockwise Parallel
Decoding](https://arxiv.org/abs/1811.03115), [MuToR](https://arxiv.org/abs/2505.10518),
and Google's 2026 [Gemini Nano frozen-MTP
work](https://research.google/blog/accelerating-gemini-nano-models-on-pixel-with-frozen-multi-token-prediction/).*
---
# Nous Hermes and Mixture-of-Agents: when models confer before they answer
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/nous-hermes-moa
> date: 2026-06-27
> tags: llm, multi-agent, mixture-of-agents, nous-research, open-weights, explainer
"Hermes MoA" isn't one thing, so let me separate the threads before building anything,
because the difference between them is the difference between a real result and a
marketing claim.
- **Mixture-of-Agents (MoA)** is a real, well-cited technique from Together AI / Duke /
Stanford — [arXiv:2406.04692](https://arxiv.org/abs/2406.04692). It's the foundation
everyone means by "MoA".
- **Nous Hermes** is a real open-weight model family. The current
[Hermes 4](https://arxiv.org/abs/2508.18255) (14B / 70B / 405B) is a single
hybrid-reasoning model — it does *not* itself use MoA.
- The genuine bridge between them is Nous's **Forge Reasoning API** (Nov 2024), which
really did run MoA — plus Monte Carlo Tree Search and Chain of Code — on top of
Hermes 70B.
- There is also a body of **2026 "Hermes Agent MoA" claims** that I'll get to at the
end, and explicitly flag as unverified.
So the honest construction is: here's the MoA mechanism, here's the open-weight Hermes
line and why it's a natural host for it, and here's how Nous actually shipped the two
together. Let's build it.
## The core idea: proposers and aggregators
A single LLM gets one shot at your prompt. MoA's bet is that several models, allowed to
*read each other's drafts and synthesize*, beat any one of them — even when the
individual drafts are mediocre.
The structure is a stack of **L layers**, each with **n agents**. Agents play two
roles:
- **Proposers** generate candidate responses. Diversity matters more here than any
single proposer's quality — you want different mistakes, not the same answer four
times.
- **Aggregators** take the candidates and synthesize a single, better response.
The data flow is the load-bearing part: every agent in layer $i$ receives **all
outputs from layer $i-1$**, concatenated into an *Aggregate-and-Synthesize* prompt that
tells the model to critically evaluate the candidates and fuse them. The final layer's
aggregator emits the answer. Watch one round play out — four diverse proposers, each
right about a different piece, fused into an answer that beats all of them:
Stack that into layers and the synthesized answer sharpens further. Add depth and watch
the quality climb:
## Why conferring helps: collaborativeness
The empirical observation that motivates the whole thing: an LLM produces a *better*
answer when shown other models' responses — **even when those responses are
individually weaker than what it would have written alone.** The paper calls this
collaborativeness, and it's the reason MoA isn't just "best-of-n with extra steps".
The mechanism is concrete. The aggregator isn't voting; it's reading. A proposer that
nailed the units, another that caught an edge case, a third that structured the
explanation — the aggregate-and-synthesize prompt lets one model keep the part each
proposer got right and drop the rest.
## What MoA scores
The headline result is that a stack of **open-source** models, conferring, beats a
single frontier model. On AlpacaEval 2.0 (length-controlled win rate):
Open-only MoA hits **65.1%** against GPT-4 Omni's **57.5%** — a +7.6-point margin with
no closed model in the loop. On MT-Bench it scores **9.25** (9.40 with GPT-4o added)
versus GPT-4 Omni's 9.19, and on FLASK it leads on robustness, correctness, factuality,
and completeness against the strongest single proposer.
The cost is exactly what you'd expect: $n$ proposers across $L$ layers means many model
calls per answer, and latency stacks with depth. MoA buys quality with compute. Whether
that trade is worth it depends entirely on how much you value the marginal correctness.
## Design choices that actually move the needle
MoA has three knobs, and they don't all behave the way intuition says:
- **Width (proposers) over depth (layers).** The bulk of the gain comes from having
several *diverse* proposers in the first layer; stacking more layers helps less and
costs latency linearly. The toy above saturates fast in $L$ for exactly this reason.
- **Diversity beats raw strength.** Proposers that fail differently give the aggregator
more to work with than several copies of the strongest model. The paper's pool is
deliberately heterogeneous — Qwen, WizardLM, Llama-3, Mixtral, dbrx — not six clones.
- **The aggregator is a real choice.** Not every strong model is a good *synthesizer*;
the aggregator has to read several candidate answers and fuse them faithfully rather
than just re-emit its own. The paper uses Qwen1.5-110B-Chat as the final aggregator,
and the role-suitability of a model as aggregator vs proposer is measured separately.
You can see the effect dimension-by-dimension on FLASK, which scores along twelve skill
axes rather than one number:
The shape of that result is the tell: MoA helps exactly where having several independent
attempts to cross-check is valuable, and barely moves dimensions that a single competent
model already nails.
## Where Hermes comes in
Nous Research builds the **Hermes** line — open-weight models post-trained with a
deliberate *neutral alignment* philosophy: minimal gratuitous refusals, maximal user
steerability. Hermes 4 (70B and 405B on Llama-3.1 bases, 14B on a Qwen3 base) adds
**hybrid reasoning** — a single checkpoint with a toggleable `…` block,
so you get reasoning and instruct behavior from one model — plus strong function
calling and JSON-schema structured output. It was trained on ~60B tokens (~5M samples)
built with Nous's DataForge and the Atropos RL environment, rejection-sampled against
roughly a thousand task-specific verifiers, on 192× B200 GPUs.
On capability it's competitive with the open frontier (Hermes 4 405B, reasoning mode):
| Benchmark | Hermes 4 405B (reasoning) | non-reasoning |
|---|---|---|
| MATH-500 | 96.3 | 73.8 |
| AIME'24 | 81.9 | 11.4 |
| AIME'25 | 78.1 | 10.6 |
| GPQA Diamond | 70.5 | 39.4 |
| LiveCodeBench v6 | 61.3 | 28.1 |
| MMLU | 87.2 | 73.6 |
But the number that captures the *philosophy* is RefusalBench — Nous's own measure of
how often a model refuses across 32 categories of typically-refused requests (higher =
fewer refusals, except for a few inverted safety categories scored the other way):
That steerability is what makes Hermes a natural MoA citizen. Open weights mean you can
run a whole proposer pool yourself; neutral alignment means the aggregator won't refuse
to synthesize half its inputs. Hermes is built to be *driven*, which is exactly what a
multi-agent harness does to it.
## The real bridge: Forge
The genuine "Nous ran MoA on Hermes" artifact is the **Forge Reasoning API** (beta,
Nov 2024). Forge combined three inference-time techniques on top of Hermes 70B:
Mixture-of-Agents, Monte Carlo Tree Search, and Chain of Code. The MoA piece is exactly
the mechanism above — "models respond, confer, and synthesize new answers" — applied to
a Hermes-centric pool. If you want a concrete instance of MoA on the Hermes line that
actually shipped, Forge is it.
Forge stacked three inference-time techniques that compose cleanly because they attack
different failure modes:
- **Mixture-of-Agents** — breadth. Several models propose and an aggregator synthesizes,
the mechanism above.
- **Monte Carlo Tree Search** — depth. Instead of one greedy chain, explore a tree of
reasoning continuations and back up value estimates, spending more search on promising
branches. This is the "think longer on hard problems" axis.
- **Chain of Code** — grounding. Offload the steps that are better *executed* than
*reasoned about* (arithmetic, string manipulation, logic) into code that actually
runs, so the model isn't bluffing its way through a calculation.
Breadth, depth, and grounding are orthogonal, which is why bolting all three onto a
fixed Hermes backbone bought more than any one alone.
A practical MoA instantiation Nous-style also collapses the textbook diagram into
something cheap: a small **reference** model runs first *without* tool schemas (avoiding
refusals and saving tokens), its output is appended as private context, and the
**aggregator** — the real Hermes agent — does the actual tool-calling loop with the
reference draft in hand. One layer, two roles, most of the benefit. It's a reminder that
"MoA" in production rarely looks like the 3×6 textbook diagram; it's whatever
proposer/aggregator split pays for itself.
There is a wave of **June 2026 "Hermes Agent MoA 2.0"** content claiming MoA presets
that beat "Claude Opus 4.8" and "GPT-5.5" on an unpublished "HermesBench" (e.g. a
quoted 0.8202 vs 0.7607/0.7412 for the individual models). I could not verify any of
it: the cited models aren't confirmably released, the benchmark has no published
leaderboard, and the supporting sources are a crypto-news post (which hedges with
"claiming") and social posts. Treat the *mechanism* as faithful MoA, but treat the
*numbers* as marketing-stage and unverified — not established fact.
## What I make of it
- **The result is real and a little counterintuitive.** Open models that read each
other's drafts beat a single frontier model on AlpacaEval, and the lift comes from
collaborativeness — synthesis from diverse, even weaker, drafts. That's a genuine,
reproducible finding with public code.
- **Hermes is the right host, not the inventor.** MoA is Together AI's; Hermes is
Nous's open-weight, neutral-alignment line; Forge is where Nous actually combined
them. Keep the attribution straight and the story is clean.
- **The cost is the catch, as always.** $n \times L$ model calls per answer and latency
that grows with depth. MoA is for when correctness is worth real compute — agentic
pipelines, hard reasoning — not for chat you need back in 200ms.
- **Be skeptical of the 2026 leaderboard claims.** The mechanism is sound; the
benchmark numbers floating around are not yet something I'd cite.
---
*Built on Together AI's [Mixture-of-Agents Enhances Large Language Model
Capabilities](https://arxiv.org/abs/2406.04692) (Wang et al., 2024;
[code](https://github.com/togethercomputer/moa)), the [Hermes 4 Technical
Report](https://arxiv.org/abs/2508.18255) (Nous Research, 2025), and Nous's [Forge
Reasoning API](https://nousresearch.com/introducing-the-forge-reasoning-api-beta-and-nous-chat-an-evolution-in-llm-inference).*
---
# Unconventional AI's Un-0: generating images with coupled oscillators
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/unconventional-un-0
> date: 2026-06-27
> tags: generative-models, neuromorphic, physics, image-generation, explainer
Every image model you know is built from the same parts: neural-network layers, a lot
of matrix multiplies, and — for the generative step — either a diffusion schedule or an
adversary. **Un-0** throws all of that out. Its computational core is a population of
**coupled oscillators**, and the generative step is just *letting them settle*.
This is the first release from **Unconventional AI**, the company Naveen Rao (ex-Databricks
AI head, founder of Nervana and MosaicML) started with Michael Carbin and Sara Achour,
on a $475M seed. The thesis is one sentence: **physics as a computational primitive.**
Instead of simulating a dynamical system on a von Neumann machine, run the dynamical
system directly in analog silicon, and let the chip's physics *be* the computation —
chasing brain-like (~20 W) efficiency.
Un-0 is explicitly the "hello world" of that program: a proof, in software simulation,
that the math produces real images. The chip doesn't exist yet. Keep that line bright;
I'll come back to it.
Here is the thing itself: a field of oscillators, each pulled toward its neighbours.
Raise the coupling and watch incoherent speckle organise into travelling waves. That
self-organisation — not a matrix multiply — is the computation Un-0 runs.
## The primitive: Kuramoto oscillators
An oscillator is just a phase $\theta_i \in [0, 2\pi)$ turning at its own natural
frequency $\omega_i$. Couple a population of them and each one also feels a pull toward
its neighbours' phases. That's the **Kuramoto model**:
$$
\dot{\theta}_i \;=\; \omega_i \;+\; \frac{K}{N}\sum_{j=1}^{N} \sin(\theta_j - \theta_i)
$$
$K$ is the coupling strength. The behaviour has a sharp phase transition. Below a
critical $K$, everyone runs at their own frequency and the phases scatter — incoherent.
Above it, the population spontaneously **synchronizes** into one travelling cluster. The
standard measure is the order parameter
$$
r\,e^{i\psi} \;=\; \frac{1}{N}\sum_{j=1}^{N} e^{i\theta_j},
$$
where $r \to 0$ is total incoherence and $r \to 1$ is full lock. You've seen this in the
physical world — pendulum metronomes started out of step on a shared, freely-moving base
pull each other into perfect synchrony:
Each oscillator can be drawn as its own dial. Drag $K$ through the transition and watch
the hands go from smeared to locked, and $r$ climb:
The point for Un-0: that transition, and the rich partially-synchronized regime around
it, is a programmable dynamical system. If you can *shape* the coupling and the
frequencies, the settled phase pattern can encode something — like an image.
## The pipeline: condition, evolve, read out
Un-0's main class is a `ConditionalImplicitKuramotoGenerator`. There's no denoising
schedule, no adversary, no iterative refinement loop in the diffusion sense — just an
ODE you integrate forward once.
Step by step:
1. **Initialize** every oscillator's phase randomly.
2. **Condition** on the class label through a *separate* oscillator array that couples
one-directionally into the main population — the label bends the dynamics without
being bent back.
3. **Evolve** the coupled ODE forward for a fixed time $T$ with explicit Euler
integration. This is the entire "generation" — no schedule, no sampler loop.
4. **Read out** the settled phases via $\sin\theta, \cos\theta$.
5. **Decode** with a small conventional network — capped at ≤15% of total parameters —
to produce pixels.
Training learns the coupling matrix $K$, the natural frequencies $\omega$, and the
decoder weights, via a "drifting loss" that uses a frozen DINOv2 feature extractor, with
AdamW. So the *learning* is conventional gradient descent; what's unconventional is that
the thing being learned is the physics of a dynamical system, not a stack of attention
layers.
## Training: differentiating through the dynamics
The subtle part is how you get gradients into an ODE. The forward pass *is* the Euler
integration of the Kuramoto system — a long chain of $\sin$-coupled updates — and the
decoder reads the final state. Because every step is differentiable, you can
backpropagate through the unrolled trajectory and update $K$, $\omega$, and the decoder
end-to-end. The "drifting loss" supervises in a perceptual feature space (a frozen
DINOv2 encoder) rather than raw pixels, which is what lets a tiny decoder — capped at
≤15% of parameters — get away with so little work: the oscillator field is doing the
heavy lifting, and the loss only has to match high-level features, not paint exact RGB.
Two things fall out of this design that are worth stating plainly:
- **Capacity lives in the coupling.** Almost all the model's parameters are the coupling
matrix $K$ (it's $O(n^2)$ in the oscillator count $n$), which is exactly why FID
improves monotonically as you scale $n$ — you're literally adding interaction terms to
the dynamical system.
- **The natural frequencies $\omega$ are learned, not fixed.** The model gets to choose
each oscillator's intrinsic rhythm, so it can place itself wherever in the
synchronize/desynchronize landscape is most useful for a given class.
## Oscillators vs diffusion
It's tempting to file Un-0 under "another iterative generator," but the comparison is
instructive precisely because of how it *differs*:
| | Diffusion model | Un-0 (oscillators) |
|---|---|---|
| Generative step | reverse a noising schedule, T denoising passes | integrate one coupled ODE to time T |
| Core compute | matrix multiplies in NN layers | sin-coupled phase updates |
| Conditioning | cross-attention / adaLN on the class | one-directional coupling from class oscillators |
| Stochasticity | injected noise at each step | random initial phases only |
| Why it might be efficient | — | the *physics* can run in analog silicon |
A diffusion model spends its compute pushing tensors through learned layers many times.
Un-0 spends its compute letting a physical system relax. On a GPU that's a wash at best
— more on that below — but the bet is that the relaxation is free when the substrate is
the right kind of analog hardware.
## Does it actually generate images?
Yes — and the honest version is "yes, modestly". Un-0 is class-conditional and low-res,
and the company is upfront that it underperforms state-of-the-art generators like EDM.
The headline is FID **6.74** on ImageNet 64×64, which they frame as matching early
conventional generators.
And here is the generation *happening* — a row of samples resolving out of the
oscillator field as the ODE integrates forward in time. There's no denoising loop; this
is the population relaxing toward its conditioned attractor and the decoder reading it
out frame by frame:
FID scales the way you'd hope with oscillator count $n$ (more oscillators, lower FID):
| Dataset | config | params | FID (↓) |
|---|---|---|---|
| CIFAR-10 32×32 | n1024 | 1.3M | ~11.0 |
| CIFAR-10 32×32 | n2048 | 4.9M | ~9.3 |
| CIFAR-10 32×32 | n4096 | 19.4M | ~8.8 |
| ImageNet 64×64 | n6656 | 57M | ~8.4 |
| ImageNet 64×64 | n10240 | 130M | ~8.0 |
| ImageNet 64×64 | n16384 | 322M | **6.74** |
The same monotone scaling holds on CIFAR-10, where even a 1.3M-parameter field already
reaches a usable FID:
Note the FID values wobble slightly between the blog and the repo README (e.g. 8.41 vs
8.36) — these are self-reported, not third-party-reproduced, so treat them as
approximate. The compute is non-trivial too: the largest ImageNet run is reported around
640 B200-GPU-hours — *simulating* the oscillators on conventional GPUs is the expensive
part, which is exactly the cost the proposed chip is meant to erase.
## The hardware bet
This is where the whole thing either pays off or doesn't. Today's accelerators are von
Neumann machines: weights live in memory, you stream them to compute units, multiply,
and write back. That shuffle — not the arithmetic — is where most of the energy goes.
Unconventional AI's proposal is to build the oscillators in physical silicon (CMOS ring
oscillators are the usual candidate), so that the coupled dynamics *happen* rather than
being computed. There's no weight streaming because the coupling is the wiring; the
system's settling to a synchronized state is the forward pass. The aspiration is
brain-like efficiency — order-of-tens-of-watts, against data-center GPUs — and the
"1000×" figure is a projection of what that substrate could do relative to simulating
the same ODE on a GPU.
It's a real idea with real lineage — analog and neuromorphic computing has chased this
for decades — and the team (Naveen Rao, plus Michael Carbin from MIT and Sara Achour
from Stanford on the hardware/compiler side) is credible. But it is, today, a
*proposal*. The repo says chip schematics are "coming soon".
## The part to keep straight
Un-0 runs on a **software simulation of hardware that does not yet exist.** No oscillator
chip has been built, and no chip schematics had been released at launch. The headline
**"1000× lower energy" is a projection by the founders, not a measured result** — there
is no analog silicon to measure. Press lines claiming Un-0 "matches Stable Diffusion"
overstate what are class-conditional, 32×32/64×64 benchmarks behind SOTA. And there is
**no peer-reviewed or arXiv paper** — the release is a company technical blog plus an
MIT-licensed [GitHub repo](https://github.com/unconv-ai/Un-0). (Two real Kuramoto arXiv
papers surface in searches — "Artificial Kuramoto Oscillatory Neurons" and "Kuramoto
Orientation Diffusion Models" — but they are *unaffiliated* prior art, not Un-0.)
So separate two claims cleanly. **Demonstrated today, in simulation:** a coupled-oscillator
ODE, conditioned and evolved once, decodes into recognizable class-conditional images at
FID 6.74 (ImageNet-64). **Proposed, not yet built:** the analog oscillator chip whose
physics would run that ODE for ~1000× less energy. The first is a real, open, checkable
result. The second is a hardware vision — credible given the team and funding, but
unbuilt and unverified.
## What I make of it
- **The idea is genuinely different, not a reskin.** Replacing layers + a diffusion
schedule with "set up a dynamical system and let it settle" is a real departure. The
generative step is an ODE integration, and the learned object is the physics itself.
- **The demo is honest and modest.** FID 6.74 on ImageNet-64 is a proof-of-concept that
the math closes, deliberately framed as a "hello world", explicitly behind SOTA. That
honesty is worth more than a cherry-picked headline.
- **The whole bet lives in the hardware that isn't here.** On a GPU, simulating
oscillators is *slower and costlier* than just running a normal generator — the entire
payoff is conditional on the analog chip materializing and delivering the projected
efficiency. Until silicon exists, "1000×" is a hypothesis, and the right way to read
Un-0 is as a credible research demonstration of physics-based generative computing —
not a shipping efficiency win.
---
*Built on Unconventional AI's [Un-0 technical
writeup](https://unconv.ai/blog/introducing-un-0-generating-images-with-coupled-oscillators/)
and the MIT-licensed [Un-0 code](https://github.com/unconv-ai/Un-0). Benchmarks are
self-reported; the analog-hardware efficiency claim is a founder projection, not a
measured result.*
---
# GLM 5.2: long-horizon coding at a million tokens
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/glm-5-2
> date: 2026-06-23
> tags: llm, glm, long-context, agentic-coding, explainer
GLM 5.2, from Z.ai (Zhipu AI), is the flagship of the GLM-5 line: a 744-billion-
parameter mixture-of-experts with 40B active per token, MIT-licensed open weights,
and — the headline — a genuine **1-million-token context**. It is tuned for one thing
in particular: long-horizon agentic coding, the sessions that run hundreds of rounds
and thousands of tool calls without losing the thread.
There is no standalone GLM 5.2 paper. It builds on the GLM-5 technical report
([arXiv 2602.15763](https://arxiv.org/abs/2602.15763)) and, for the context trick at
its center, a method paper — IndexCache / IndexShare
([arXiv 2603.12201](https://arxiv.org/abs/2603.12201)). This pulls from both, plus the
[release blog](https://z.ai/blog/glm-5.2).
## What changed from 5.1
GLM-5 → 5.1 → 5.2 share the same 744B/40B backbone. What 5.2 adds:
- a real **1M-token context**, up from 200K;
- **IndexShare**, the architecture change that makes that context affordable;
- a shift to **critic-based PPO** for very long RL rollouts;
- faster speculative decoding (**+20% acceptance length**);
- a **thinking-effort** dial (High / Max).
The first two are the load-bearing pair: the long context, and the trick that keeps
it cheap.
## The model
744B total parameters, 40B active per token — a mixture-of-experts on an 80-layer,
256-expert backbone. Attention is **DeepSeek Sparse Attention (DSA)**: Multi-head
Latent Attention plus a lightweight *indexer* that, for each query, selects the
top-$k$ tokens worth attending to instead of the whole sequence. That sparsity is
what makes a million-token context tractable at all.
## IndexShare: making 1M context cheap
DSA has a catch. The indexer runs at *every layer*, and as the context grows toward
1M tokens, that per-query top-$k$ search becomes the dominant cost. The IndexCache
paper's observation is the whole insight: **adjacent DSA layers select almost the same
tokens** — 70–100% of their top-$k$ overlap.
So compute the indexer once per group of layers and reuse its selection for the rest.
GLM 5.2 shares one indexer across every 4 layers — skipping it in 3 of every 4:
If the indexer's cost per layer scales with selecting top-$k$ over $L$ tokens, then
sharing it across a group of $g$ layers amortizes that cost to $O(L/g)$ per layer.
With $g = 4$ and the rest of each layer unchanged, GLM 5.2 reports **2.9× lower
per-token FLOPs at a 1M-token context**, with quality essentially intact.
The honest tradeoff: push reuse too far — share across 8 layers instead of 4 — and
long-context fidelity starts to degrade. One indexer per four layers is the sweet
spot the paper settles on.
## Faster decoding: MTP and KVShare
GLM 5.2 also sharpens its multi-token-prediction layer (speculative decoding). With
IndexShare, KVShare, and end-to-end training, the average **acceptance length rises
~20% — from 4.56 to 5.47 tokens** per verification pass. More accepted tokens per
pass means faster generation, which matters most when you are streaming long agent
traces.
## Training for the long horizon
Pretraining scaled to **28.5T tokens** (up from GLM-4.5's 23T). But the interesting
change in 5.2 is the agentic post-training. It moves from group-relative RL to a
**critic-based PPO** that estimates token-level advantages from individual rollouts —
which accommodates *trajectory compaction* without capping how long a trace can get.
That is exactly what you need when a single agent run is thousands of tool calls long
and won't fit in one rollout.
It also adds an **anti-reward-hacking module**: a rule-based filter first catches
likely hacks (tuned for recall), then an LLM judge checks intent; on a detected hack
the system blocks the call and returns dummy information so the rollout continues
instead of being thrown away. All of it runs on Zhipu's open asynchronous RL
framework, **slime**.
## Benchmarks
The headline result: GLM 5.2 is the **strongest open-weights model on standard and
long-horizon coding**, closing much of the gap to Claude Opus 4.8 and GPT-5.5.
Where it stands out most is *long-horizon* coding — runs that have to stay coherent
over many rounds — where it nearly catches Opus 4.8 and leaves the rest behind:
Reasoning is strong — a near-perfect AIME — though it trails the very top closed
models on the hardest knowledge benchmarks (GPQA, HLE):
## Thinking effort, and what 1M costs to serve
GLM 5.2 exposes two reasoning-effort levels — `high` for everyday speed and `max` for
hard multi-step coding — and Z.ai positions its capability between Claude Opus 4.7 and
4.8 at similar token spend.
The 1M context is not free to serve. The bottleneck moves from raw compute to
**KV-cache capacity, long-context kernels, and CPU-side overhead**; the throughput
advantage grows with context length, but you need 8×H100-class hardware and ~1.5 TB
for the weights, and the API meters at 3× during peak hours.
## What I make of it
- **The genuinely new bit is IndexShare** — a clean, well-motivated systems trick
(reuse what's nearly identical instead of recomputing it), with a paper that shows
*why* it's almost lossless. That's what turns "1M context" from a spec-sheet number
into something you can actually serve.
- **It's the strongest open-weights model for long-horizon agentic coding**, and it's
MIT-licensed. That combination matters more than the benchmark deltas — you can run
and fine-tune it yourself.
- **It still trails the best closed frontier models** on most hard coding and
reasoning axes (SWE-Bench Pro 62.1 vs Opus 4.8's 69.2), and it is heavy to
self-host. The bet was never "beat Opus 4.8 everywhere" — it's "match the frontier on
long-horizon work, in the open, at a million tokens." On that, it largely delivers.
---
*Sources: the [GLM 5.2 release blog](https://z.ai/blog/glm-5.2), the GLM-5 technical
report ([arXiv 2602.15763](https://arxiv.org/abs/2602.15763)), and the IndexCache
method paper behind IndexShare ([arXiv 2603.12201](https://arxiv.org/abs/2603.12201)).
Benchmark figures are from Z.ai; numbers quoted as reported.*
---
# Sakana Fugu: a multi-agent system as a model
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/sakana-fugu
> date: 2026-06-23
> tags: llm, multi-agent, orchestration, reinforcement-learning, explainer
No single LLM wins everywhere. One model leads on competition math, another on
agentic coding, a third on multilingual work, and open models win on cost. The
usual response is to pick one and absorb its weak spots. Sakana AI's bet is the
other one: don't pick a model — *orchestrate a pool of them*, and make the
orchestration itself the model.
That product is **Sakana Fugu** (and a heavier tier, Fugu Ultra), shipped behind a
single API. Underneath are two ICLR 2026 papers that attack the same problem from
opposite ends: [TRINITY](https://arxiv.org/abs/2512.04695) *evolves* a tiny
coordinator over frozen models, and the
[Conductor](https://arxiv.org/abs/2512.04388) *reinforcement-learns* a 7B model to
write orchestration plans in natural language. This is a walk through both, and what
they add up to.
## A multi-agent system as a model
Fugu's framing is the whole pitch: one OpenAI-compatible endpoint. You send a
request to `model: fugu`; behind it a learned coordinator assembles a team from a
pool of frontier and open models, runs them over several turns, and returns one
answer. You never see the routing.
The pool is swappable — you can opt a model out for compliance and the coordinator
routes around it — and billing is a single top-tier rate rather than stacked
per-model fees. There's even an export-controls angle: because Fugu can hit
frontier-level quality by coordinating open and semi-open models, you get the
capability without hard dependence on any one restricted vendor.
But the API is the boring part. The interesting part is that the coordinator is
*learned*, not hand-written. There are two ways to learn it.
## TRINITY: evolve a tiny coordinator
TRINITY's constraint shapes everything: you cannot fine-tune GPT-5's weights, and
merging models with incompatible architectures doesn't work. So freeze every model
in the pool, and learn only a tiny thing on top that decides who does what.
### The coordinator is under 20,000 parameters
A small model — Qwen3-0.6B — reads the current problem state and produces a hidden
vector; a linear head turns that into a choice of *agent* and *role*. Given the
penultimate-token hidden state $h(s)\in\mathbb{R}^{d}$ from the small model, a head
$f_\theta$ of roughly 10K parameters emits logits over $L$ agents plus 3 roles, and
the coordinator samples its action $a$ from
$$
\pi_\theta(a \mid s) \;\propto\; \exp\!\big(f_\theta(h(s))_a\big),
\qquad a \in \{1,\dots,L\}\cup\{\mathrm{T},\mathrm{W},\mathrm{V}\}
$$
where $s$ is the running transcript, $\mathrm{T},\mathrm{W},\mathrm{V}$ are the three
roles below, and $\theta$ is everything that gets trained. On top of the head, TRINITY
adds *singular-value fine-tuning*: take an SVD of one or two of the small model's
weight matrices and learn only the singular-value scales, keeping the orthogonal
factors fixed. That's a few thousand more numbers. Total trainable: **under 20K
parameters.** The 0.6B backbone and all seven frontier and open models stay frozen.
### Three roles, looped until accepted
Each turn, the coordinator gives the chosen agent one of three roles:
- **Thinker** — plan, decompose, or critique; no direct work.
- **Worker** — do the work: derive, compute, write code.
- **Verifier** — check the current answer and return `ACCEPT` or `REVISE`.
It loops, accumulating a transcript, and halts the moment a Verifier accepts (or a
fixed turn budget $K$ is exhausted):
$$
\tau \;=\; \min\{\, k \le K \;:\; R_k = \mathrm{V} \ \text{and}\ u_k = \mathrm{ACCEPT} \,\}
$$
where $R_k$ is the role at turn $k$ and $u_k$ is the verifier's verdict. Step through
one problem — watch a wrong answer get caught and revised before it's accepted:
### Trained by evolution, not gradients
Why not just RL the head? Because the reward is binary — the final answer is right or
wrong — and the head is tiny, so the per-parameter gradient signal is buried in
noise. TRINITY instead optimizes the coordinator with a *derivative-free* evolution
strategy, maximizing expected terminal reward:
$$
J(\theta) \;=\; \mathbb{E}_{\tau \sim \pi_\theta}\big[\, R(\tau) \,\big],
\qquad R(\tau) \in \{0, 1\}
$$
The optimizer is separable CMA-ES: it keeps a diagonal Gaussian over the ~10K
parameters, samples a small population each generation —
$\lambda = \lceil 4 + 3\ln n \rceil \approx 32$ for $n \approx 10{,}000$ — evaluates
each candidate's fitness by actually running rollouts, and shifts the distribution
toward the winners. The paper shows the coordination objective is nearly
block-separable, which is exactly the regime where a diagonal evolution strategy
beats both random search and gradient RL under a tight evaluation budget. The honest
cost: no gradients means you pay in *environment evaluations*, and each one is a full
multi-turn rollout against real model APIs.
### It beats every model in its pool
This is the result that matters. Transferred zero-shot to four held-out tasks, the
evolved coordinator outscored every individual model in its pool — including GPT-5,
Gemini-2.5-Pro, and Claude-4-Sonnet. On LiveCodeBench it set a record at the time of
submission:
And the multi-turn loop earns its keep: accuracy climbs from 0.823 at two turns to
0.863 at six. One cheap evolved head, a frozen pool, and the ensemble beats its best
member.
## Conductor: orchestration written in natural language
The Conductor attacks the same problem with a bigger hammer: a 7B model (Qwen2.5-7B)
trained with RL to *write the entire workflow itself*, in natural language.
### Three lists are a workflow
For each problem the Conductor emits three synchronized lists:
- `model_id` — which agent runs each step.
- `subtasks` — a natural-language instruction for each step.
- `access_list` — which earlier outputs each step is allowed to read.
Those three lists *are* a directed graph. The `access_list` is the load-bearing
idea: `[]` means the step sees only the original question, `["all"]` means it sees
everything produced so far, and `[0, 2]` means it sees steps 0 and 2. By choosing
access lists, the Conductor designs the communication topology — a chain, parallel
branches, a verify-and-merge — *per problem*, not from a fixed template. Flip between
the topologies it learns to produce:
### Trained with GRPO
The Conductor is trained end-to-end with GRPO. For each question it samples a group
of $G = 64$ candidate workflows, scores each, and pushes the policy toward the
above-average ones using the group-normalized advantage
$$
A_i \;=\; \frac{r_i - \operatorname{mean}(r_1, \dots, r_G)}{\operatorname{std}(r_1, \dots, r_G)}
$$
The reward $r_i$ is blunt on purpose: $0$ if the three lists don't parse, $1$ if the
final workflow output is correct, and $0.5$ otherwise — with no KL penalty
($\beta = 0$). The whole thing trains on just 960 problems for 200 iterations on two
H100s. To make one Conductor work over *any* pool, they then fine-tune it with
randomly sampled $k$-model subsets per question, so it adapts to whatever agents you
hand it.
### It can call itself
The Conductor may name *itself* as a worker. That spawns a fresh sub-workflow on its
own draft — a recursive topology that turns inference depth into a tunable compute
axis, what Sakana calls dynamic test-time scaling. Recursion buys a point or two on
the hardest benchmarks for under 2× the agent calls.
### Results
A 7B model orchestrating frontier workers beats the frontier workers. In a
controlled run over the same pool:
Unconstrained, the headline numbers were each a new high at publication and each
above the best single worker: **83.9% on LiveCodeBench, 87.5% on GPQA-Diamond, 93.3%
on AIME25** — reached with about 3 agent calls per question, versus 5–8 for prior
multi-agent methods.
## Two routes to the same place
TRINITY and the Conductor are the same idea — a learned layer that coordinates a
pool — built at opposite scales:
| | TRINITY | Conductor |
|---|---|---|
| Learnable size | < 20K params (evolved head) | 7B params (RL-trained model) |
| Training | derivative-free sep-CMA-ES | GRPO (reinforcement learning) |
| Output per step | (agent, role) | a full natural-language workflow |
| Coordination | fixed Thinker/Worker/Verifier loop | a topology it designs per problem |
| Reads the task via | the small model's hidden state | reasoning in language |
| Adapts to new pools | re-evolve (cheap) | randomized-pool fine-tune |
TRINITY is the minimal, almost-free coordinator; the Conductor is the expressive one
that designs bespoke pipelines. Fugu uses both as its engine.
## What ships: Fugu and Fugu Ultra
Two tiers. Base **Fugu** balances quality and latency over a lean pool. **Fugu
Ultra** coordinates a deeper pool over more turns for hard, high-stakes problems, and
takes longer for it. On Sakana's reported numbers, both match or beat the frontier:
Fugu Ultra also posts **50.0 on Humanity's Last Exam**, against baselines in the
41–50 range. It's an OpenAI-compatible endpoint — change the base URL and key, no SDK
migration — and it bills at a single top-tier rate. (Not available in the EU yet,
pending GDPR; the exact routing decisions are kept proprietary.)
## What I make of it
The honest read:
- **The win is real.** An orchestration layer that beats every model it coordinates —
and generalizes zero-shot to unseen tasks — is a genuine result. "Coordination" is
now a trainable layer that sits *above* frontier models rather than inside one.
- **The costs are real too.** Every model in the pool has to be available at
inference; you trade single-model simplicity for a fleet, and latency rises with
the extra turns. The biggest gains concentrate on long-tail reasoning and coding
benchmarks — on easy tasks the lift is small — and leaning on GPT-5/Claude/Gemini as
workers inherits their cost.
- **The framing is the interesting part.** TRINITY argues the coordinator can be
almost free: 20K evolved parameters over frozen models. The Conductor argues
coordination is itself a reasoning skill worth a 7B model and a full RL run. Both
point the same way — as individual models plateau, the next axis is how you make
several of them work together, and that orchestration is learnable.
---
*Built on Sakana AI's [TRINITY: An Evolved LLM Coordinator](https://arxiv.org/abs/2512.04695)
and [Learning to Orchestrate Agents in Natural Language with the Conductor](https://arxiv.org/abs/2512.04388),
both ICLR 2026. Product: [Sakana Fugu](https://sakana.ai/fugu/).*
---
# Mixture of Experts, from scratch
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/mixture-of-experts-from-scratch
> date: 2026-06-10
> tags: deep-learning, transformers, mixture-of-experts, explainer
Scaling a transformer the dense way is a bad trade. Every parameter you add runs
on every token. Double the width of the feed-forward layers and you double both
the model's capacity *and* the FLOPs it burns per token — capacity and compute are
welded together. You pay for the whole network on every single token, whether that
token needs it or not.
Mixture of Experts breaks the weld. The idea is **conditional computation**: keep a
large pile of parameters around, but for any given token, only run a small slice of
them. A tiny router looks at each token and picks a couple of sub-networks — the
*experts* — to handle it. The rest sit idle for that token. You get the capacity of
a big model at the compute of a small one.
Here is the whole model we'll build, end to end. The only thing that makes it a MoE
is one swapped line — tap the **sparse MoE** block to see it:
Everything except that one block is a standard decoder-only transformer: token and
position embeddings, a stack of blocks, a final norm, an LM head. Attention is
untouched. MoE is a surgical replacement for the feed-forward layer inside each
block, and nothing else. So the whole thing reduces to three questions: what is an
expert, who decides which experts run, and how do you run only the chosen few.
## An expert is just an MLP
Start with the thing we're replacing. In a normal transformer block, after
attention, every token goes through the same two-layer MLP — expand to `4 * n_embed`,
nonlinearity, project back. That's the feed-forward network.
An expert is exactly that MLP. Nothing more.
```python
class Expert(nn.Module):
def __init__(self, n_embed, dropout=0.1):
super().__init__()
self.net = nn.Sequential(
nn.Linear(n_embed, 4 * n_embed),
nn.ReLU(),
nn.Linear(4 * n_embed, n_embed),
nn.Dropout(dropout),
)
def forward(self, x):
return self.net(x)
```
The move is to keep `num_experts` copies of this MLP instead of one. With 8 experts
you have 8× the feed-forward parameters. If every token went through all 8, you'd
have spent 8× the compute and gained nothing but a slow, fat FFN. The whole game is
to run only `top_k` of them — say 2 — per token. So you carry 8 experts' worth of
parameters and pay for 2.
The piece that makes that decision is the router.
## Who decides? The router
The router's job: look at a token's vector $x$ and produce a weight for each expert,
mostly zero, so that only a few experts actually contribute. Build it up in three
steps, because the naive versions teach you why the real one looks the way it does.
**Attempt 1 — send every token to every expert, weighted.** A linear layer maps the
token to one logit per expert, softmax over them, take a weighted sum of all expert
outputs:
$$
g(x) = \mathrm{softmax}(x W_g), \qquad y = \sum_{i=1}^{N} g(x)_i \, E_i(x)
$$
Here $W_g$ is the router's weight matrix (`n_embed × num_experts`) and $E_i$ is the
$i$-th expert. This is differentiable and trains fine — but it's *dense*. Every
expert runs on every token. We've built an expensive ensemble, not a sparse model.
**Attempt 2 — hard pick the single best expert.** Take $\arg\max$ of the logits, run
only that expert. Now it's sparse and cheap. But $\arg\max$ has zero gradient: the
router only ever learns about the one expert it already chose, and never gets a
signal to try the others. Routing freezes. Dead end.
**Attempt 3 — top-$k$ softmax.** Keep the largest $k$ logits, set the rest to
$-\infty$, *then* softmax. The $-\infty$ entries become exactly 0, so only $k$ experts
contribute — sparse like attempt 2 — but the softmax over the survivors is smooth, so
gradients flow to all $k$ chosen experts. This is the real router:
$$
g(x) = \mathrm{softmax}\big(\mathrm{KeepTopK}(x W_g,\, k)\big), \qquad
\mathrm{KeepTopK}(v, k)_i = \begin{cases} v_i & v_i \text{ in top } k \\ -\infty & \text{otherwise} \end{cases}
$$
With $k = 2$ and $N = 8$, six of the eight gate weights are zero for every token, and
the two survivors sum to 1. Watch one token go through it — logits, keep the top two,
softmax to gates, combine:
That stepper is the entire routing mechanism. The bars are the per-expert logits;
top-2 keeps two; softmax turns them into weights; the output is just those two
experts' outputs scaled by their gates and added.
## Why the noise
There's one addition that the bare top-$k$ router needs in practice: noise. Before
picking the top $k$, add a learned, per-expert amount of Gaussian noise to the
logits:
$$
H(x)_i = (x W_g)_i + \varepsilon_i \cdot \mathrm{softplus}\big((x W_{\text{noise}})_i\big), \qquad \varepsilon_i \sim \mathcal{N}(0, 1)
$$
The noise scale is itself learned (a second linear layer $W_{\text{noise}}$, passed
through `softplus` to keep it positive). Why bother? Because early in training the
router is random, and whichever experts happen to win first get all the gradient and
pull ahead — a rich-get-richer collapse. The noise jitters the top-$k$ selection so
borderline experts occasionally win, get some tokens, and get a chance to become
useful. It's exploration, baked into the forward pass. Hit *resample noise* in the
widget above and you can watch which two experts win flip.
In code the router is four lines of real work:
```python
class NoisyTopKRouter(nn.Module):
def __init__(self, n_embed, num_experts, top_k):
super().__init__()
self.top_k = top_k
self.route = nn.Linear(n_embed, num_experts) # gate logits
self.noise = nn.Linear(n_embed, num_experts) # per-expert noise scale
def forward(self, x):
logits = self.route(x)
noisy = logits + torch.randn_like(logits) * F.softplus(self.noise(x))
top_logits, idx = noisy.topk(self.top_k, dim=-1) # the chosen experts
sparse = torch.full_like(noisy, float("-inf"))
sparse.scatter_(-1, idx, top_logits) # keep top-k, rest -inf
return F.softmax(sparse, dim=-1), idx
```
`scatter_` is the one trick worth pausing on: it writes the kept logits back into a
tensor of `-inf`, at the indices the `topk` chose. After the softmax those `-inf`
slots are 0. The router returns the gate weights and the chosen indices — the
indices tell the next stage which experts to actually run.
## The sparse forward pass
Now the part that earns the word *sparse*. We have gate weights and, for each token,
the indices of its top-$k$ experts. We want to run each expert on only the tokens
routed to it, scale by the gate, and add the result back.
The straightforward way: loop over experts, and for each one, mask out the tokens
that picked it.
```python
class SparseMoE(nn.Module):
def __init__(self, n_embed, num_experts, top_k):
super().__init__()
self.router = NoisyTopKRouter(n_embed, num_experts, top_k)
self.experts = nn.ModuleList([Expert(n_embed) for _ in range(num_experts)])
def forward(self, x):
gates, idx = self.router(x) # (B,T,N), (B,T,k)
out = torch.zeros_like(x)
flat_x = x.view(-1, x.size(-1)) # (B*T, C)
flat_gates = gates.view(-1, gates.size(-1))
flat_out = out.view(-1, x.size(-1))
for i, expert in enumerate(self.experts):
mask = (idx == i).any(dim=-1).view(-1) # tokens routed to expert i
if mask.any():
y = expert(flat_x[mask]) # run on its tokens only
flat_out[mask] += flat_gates[mask, i:i+1] * y
return out
```
The `mask = (idx == i).any(dim=-1)` line is the dispatch: it's true for exactly the
tokens that have expert `i` somewhere in their top-$k$. We gather those tokens, run
the expert once on the batch of them, scale each by its gate weight, and scatter-add
back into the output. A token routed to experts 2 and 5 gets contributions from both
loop iterations, summed — which is exactly $\sum_i g(x)_i E_i(x)$ with all but $k$
terms zero.
Picture the dispatch over a short sequence. Each token connects to just two of the
eight experts, so most of the grid stays dark — that darkness is the compute you're
*not* spending:
The bars underneath are the per-expert load: how many tokens each expert handled.
Notice it's already uneven — some experts attract more traffic than others. Hold that
thought; it's the central problem with MoE.
This masked loop is the *teaching* implementation. It's correct but it runs every
expert as a separate kernel and materialises a mask per expert. Production MoE
instead sorts/permutes tokens by expert and does one grouped matmul, and in the
distributed case each expert lives on a different GPU and tokens are shipped to
them (expert parallelism). Same math, very different plumbing.
## The one line that changes
With the experts and the router in hand, dropping MoE into a transformer block is
anticlimactic — which is the point. A standard block is `attention → FFN`, each
wrapped in a layer-norm and a residual. MoE swaps the FFN for the `SparseMoE` module
and touches nothing else:
```python
class Block(nn.Module):
def __init__(self, n_embed, n_head, num_experts, top_k, block_size):
super().__init__()
self.sa = MultiHeadAttention(n_head, n_embed, block_size)
self.smoe = SparseMoE(n_embed, num_experts, top_k) # was: FeedForward(n_embed)
self.ln1 = nn.LayerNorm(n_embed)
self.ln2 = nn.LayerNorm(n_embed)
def forward(self, x):
x = x + self.sa(self.ln1(x)) # attention — unchanged
x = x + self.smoe(self.ln2(x)) # MoE replaces the feed-forward layer
return x
```
That's the whole architectural delta. One `FeedForward` becomes one `SparseMoE`:
Stack eight of these blocks, add embeddings and an LM head, and you have the model
from the top of the page. Train it exactly like a dense transformer — cross-entropy
on next-token prediction. The router learns its weights from the same gradient as
everything else. No special routing supervision; it figures out a useful assignment
on its own.
## Run it yourself
Here is everything above assembled into one file — a char-level model that trains on
tiny Shakespeare in about 200 lines, with no dependency past PyTorch. The `Expert`,
`NoisyTopKRouter`, and `SparseMoE` are exactly the pieces we just built; the rest is
the smallest transformer that can hold them. Copy it, run `python tinymoe.py`, and
watch the loss come down.
```python
"""
tinymoe — a tiny Mixture-of-Experts language model in one file.
Char-level, trains on tiny Shakespeare. ~4.5M params, ~1.4M active per token.
Runs on CPU; much faster on a GPU.
python tinymoe.py # download data, train, then sample
It's a small decoder-only transformer where the feed-forward layer of every
block is replaced by a sparse mixture of experts with noisy top-k routing.
"""
import os
import urllib.request
import torch
import torch.nn as nn
from torch.nn import functional as F
# --------------------------------------------------------------------- config
batch_size = 32 # sequences per step
block_size = 128 # context length (chars)
n_embed = 128 # embedding / residual width
n_head = 4 # attention heads
n_layer = 4 # transformer blocks
num_experts = 8 # experts per MoE layer
top_k = 2 # experts actually run per token
dropout = 0.1
learning_rate = 3e-4
max_iters = 5000
eval_interval = 500
eval_iters = 100
device = "cuda" if torch.cuda.is_available() else "cpu"
torch.manual_seed(1337)
# ----------------------------------------------------------- data (shakespeare)
if not os.path.exists("input.txt"):
url = ("https://raw.githubusercontent.com/karpathy/char-rnn/"
"master/data/tinyshakespeare/input.txt")
urllib.request.urlretrieve(url, "input.txt")
text = open("input.txt", encoding="utf-8").read()
chars = sorted(set(text))
vocab_size = len(chars)
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for i, c in enumerate(chars)}
encode = lambda s: [stoi[c] for c in s]
decode = lambda t: "".join(itos[i] for i in t)
data = torch.tensor(encode(text), dtype=torch.long)
n = int(0.9 * len(data))
train_data, val_data = data[:n], data[n:]
def get_batch(split):
d = train_data if split == "train" else val_data
ix = torch.randint(len(d) - block_size, (batch_size,))
x = torch.stack([d[i:i + block_size] for i in ix])
y = torch.stack([d[i + 1:i + block_size + 1] for i in ix])
return x.to(device), y.to(device)
# ------------------------------------------------------------------- attention
class Head(nn.Module):
def __init__(self, head_size):
super().__init__()
self.key = nn.Linear(n_embed, head_size, bias=False)
self.query = nn.Linear(n_embed, head_size, bias=False)
self.value = nn.Linear(n_embed, head_size, bias=False)
self.register_buffer("tril", torch.tril(torch.ones(block_size, block_size)))
self.drop = nn.Dropout(dropout)
def forward(self, x):
B, T, C = x.shape
k, q = self.key(x), self.query(x)
wei = q @ k.transpose(-2, -1) * k.shape[-1] ** -0.5
wei = wei.masked_fill(self.tril[:T, :T] == 0, float("-inf"))
wei = self.drop(F.softmax(wei, dim=-1))
return wei @ self.value(x)
class MultiHeadAttention(nn.Module):
def __init__(self, n_head, head_size):
super().__init__()
self.heads = nn.ModuleList([Head(head_size) for _ in range(n_head)])
self.proj = nn.Linear(n_embed, n_embed)
self.drop = nn.Dropout(dropout)
def forward(self, x):
out = torch.cat([h(x) for h in self.heads], dim=-1)
return self.drop(self.proj(out))
# --------------------------------------------------------- mixture of experts
class Expert(nn.Module):
"""One expert = one MLP. Same shape as a normal transformer FFN."""
def __init__(self):
super().__init__()
self.net = nn.Sequential(
nn.Linear(n_embed, 4 * n_embed), nn.ReLU(),
nn.Linear(4 * n_embed, n_embed), nn.Dropout(dropout),
)
def forward(self, x):
return self.net(x)
class NoisyTopKRouter(nn.Module):
"""Score experts per token, add learned noise, keep top-k, softmax."""
def __init__(self):
super().__init__()
self.route = nn.Linear(n_embed, num_experts)
self.noise = nn.Linear(n_embed, num_experts)
def forward(self, x):
logits = self.route(x)
noisy = logits + torch.randn_like(logits) * F.softplus(self.noise(x))
top_logits, idx = noisy.topk(top_k, dim=-1)
sparse = torch.full_like(noisy, float("-inf")).scatter(-1, idx, top_logits)
return F.softmax(sparse, dim=-1), idx
class SparseMoE(nn.Module):
"""Run only the top-k experts per token; combine them by gate weight."""
def __init__(self):
super().__init__()
self.router = NoisyTopKRouter()
self.experts = nn.ModuleList([Expert() for _ in range(num_experts)])
def forward(self, x):
gates, idx = self.router(x) # (B,T,E), (B,T,k)
out = torch.zeros_like(x)
flat_x = x.reshape(-1, x.size(-1))
flat_gates = gates.reshape(-1, gates.size(-1))
flat_out = out.reshape(-1, x.size(-1))
for i, expert in enumerate(self.experts):
mask = (idx == i).any(dim=-1).reshape(-1) # tokens routed to expert i
if mask.any():
flat_out[mask] += flat_gates[mask, i:i + 1] * expert(flat_x[mask])
return out
# ------------------------------------------------------------- block + model
class Block(nn.Module):
def __init__(self):
super().__init__()
self.sa = MultiHeadAttention(n_head, n_embed // n_head)
self.smoe = SparseMoE() # <- replaces the FFN
self.ln1 = nn.LayerNorm(n_embed)
self.ln2 = nn.LayerNorm(n_embed)
def forward(self, x):
x = x + self.sa(self.ln1(x))
x = x + self.smoe(self.ln2(x))
return x
class MoELanguageModel(nn.Module):
def __init__(self):
super().__init__()
self.tok_emb = nn.Embedding(vocab_size, n_embed)
self.pos_emb = nn.Embedding(block_size, n_embed)
self.blocks = nn.Sequential(*[Block() for _ in range(n_layer)])
self.ln_f = nn.LayerNorm(n_embed)
self.head = nn.Linear(n_embed, vocab_size)
def forward(self, idx, targets=None):
B, T = idx.shape
x = self.tok_emb(idx) + self.pos_emb(torch.arange(T, device=idx.device))
x = self.ln_f(self.blocks(x))
logits = self.head(x)
loss = None
if targets is not None:
loss = F.cross_entropy(logits.view(-1, vocab_size), targets.view(-1))
return logits, loss
@torch.no_grad()
def generate(self, idx, max_new_tokens):
for _ in range(max_new_tokens):
logits, _ = self(idx[:, -block_size:])
probs = F.softmax(logits[:, -1, :], dim=-1)
idx = torch.cat([idx, torch.multinomial(probs, 1)], dim=1)
return idx
# --------------------------------------------------------------------- train
@torch.no_grad()
def estimate_loss(model):
out = {}
model.eval()
for split in ("train", "val"):
losses = torch.zeros(eval_iters)
for k in range(eval_iters):
x, y = get_batch(split)
_, losses[k] = model(x, y)
out[split] = losses.mean().item()
model.train()
return out
model = MoELanguageModel().to(device)
total = sum(p.numel() for p in model.parameters())
print(f"{total / 1e6:.2f}M params on {device}")
opt = torch.optim.AdamW(model.parameters(), lr=learning_rate)
for it in range(max_iters):
if it % eval_interval == 0:
l = estimate_loss(model)
print(f"step {it:5d} | train {l['train']:.3f} | val {l['val']:.3f}")
x, y = get_batch("train")
_, loss = model(x, y)
opt.zero_grad(set_to_none=True)
loss.backward()
opt.step()
# -------------------------------------------------------------------- sample
ctx = torch.zeros((1, 1), dtype=torch.long, device=device)
print(decode(model.generate(ctx, 500)[0].tolist()))
```
At the default size it prints `4.52M params` — but only **~1.4M of them run on any
given token**, because 6 of every 8 experts sit out. That's the parameter-vs-compute
split in miniature. Raise `num_experts` and the total climbs while the active count
barely moves; lower `top_k` to 1 and it gets sparser still. The same lever Mixtral
pulls, in a model you can train on a laptop.
One honesty note: this minimal version relies entirely on the routing noise to keep
experts balanced — there's no auxiliary loss. At toy scale it trains fine. Scale it
up and a few experts quietly take over, which is the next problem.
## The catch: load balancing
MoE has one failure mode that dominates everything else, and you saw it forming in the
dispatch map: **expert collapse**. Routing is a positive feedback loop. An expert that
wins a few tokens early gets gradient, improves, and so becomes the router's favourite
for even more tokens. Meanwhile the experts that lost early get no tokens, no
gradient, and never improve. Left alone, a handful of experts end up doing all the
work and the rest are dead weight — you're paying to store 8 experts and effectively
running 2 or 3.
The noise we added earlier is the first defense — it keeps the routing from hardening
too fast. The second, used in every serious MoE, is an **auxiliary load-balancing
loss**: a term added to the training objective that measures how lopsided the routing
is across a batch and penalises imbalance, nudging the router toward spreading tokens
evenly. It's a soft constraint — you're not forcing exactly equal load, just paying a
cost for collapse. Tuning its weight is part of the unglamorous reality of training a
MoE: too little and experts collapse, too much and you fight the router's ability to
actually specialise.
This is the honest tradeoff. A dense FFN has no routing, no balance to maintain, no
extra loss to tune. MoE buys you cheap capacity and hands you a load-balancing problem
in return.
## What the experts actually learn
It's tempting to picture expert 3 as "the Python expert" and expert 5 as "the French
expert." That's mostly not what happens. When the Mixtral authors inspected their
router, they found no clean topic or domain specialization — experts don't map to
subjects. What the router learns is lower-level and more syntactic: routing is
strongly correlated across consecutive tokens, and individual experts lean toward
things like indentation, punctuation, or particular token shapes. The specialization
is real, but it's structural, not semantic, and not especially interpretable.
"Experts" is a useful name, not a promise that each one becomes a tidy domain
specialist.
## Beyond the basic router
The router we built is *token-choice*: each token picks its experts. Three variations
are worth knowing, because they're all different answers to the same load-balancing
problem:
- **Expert-choice routing** flips the selection — each expert picks its top tokens.
Load is balanced by construction (every expert takes a fixed budget), at the cost of
some tokens getting chosen by many experts and others by none.
- **Shared experts** (as in DeepSeek-MoE) keep one or two experts always on for every
token, so the routed experts don't burn capacity re-learning common patterns and can
specialize at the margin.
- **Capacity and token dropping** — in batched or distributed training each expert gets
a fixed number of slots per batch; tokens that overflow their chosen expert are
dropped and pass through on the residual alone. A blunt cap that keeps the per-expert
matmuls a fixed, rectangular shape.
Same tradeoff surface — cheap capacity versus keeping every expert fed — approached
from different sides.
## What you actually buy
Why put up with the routing machinery? Because the parameter-vs-compute decoupling is
real and large. Mixtral 8×7B is the clean reference: 8 experts per layer, top-2
routing — the exact configuration we just built. It holds **47B parameters total**,
but because only 2 of 8 experts run per token, a forward pass touches **about 13B
active parameters**. It runs at the speed and memory-bandwidth cost of a ~13B dense
model while matching or beating a 70B dense one across benchmarks.
That's the pitch in one line: **capacity you don't pay for on every token.** The
parameters are the model's knowledge; the active fraction is what each token can
afford to consult.
There's a cost on the other side of the ledger, and it's worth stating plainly. MoE
trades **compute for memory**. Only $k$ experts run, but *all* of them have to be
resident — you still hold 47B parameters in memory even though each token uses 13B.
And at batch scale the router scatters tokens across all experts, so the bandwidth and
the all-to-all communication of shipping tokens to the right expert (across GPUs)
becomes the real bottleneck, not the matmuls. MoE doesn't make models free. It moves
the cost from FLOPs, which you pay per token, to memory and bandwidth, which you pay
once. For inference-bound serving at scale, that's usually the trade you want.
## The whole thing, in one breath
Strip away the engineering and MoE is small: an expert is the FFN you already had;
keep several of them; a one-layer router scores them per token; keep the top two,
softmax for weights, run only those two, add a little noise so routing explores and a
balancing loss so it doesn't collapse. One line in the transformer block changes. In
return, the model's parameter count and its per-token compute stop being the same
number — and that decoupling is the entire reason the largest models you can name are
built this way.
---
# Coroutines in C, intuitively
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/coroutines-in-c
> date: 2026-06-09
> tags: c, coroutines, systems, explainer
Some functions want to be *callers*. Some want to be *callees*. The trouble starts
when two pieces of code both want to be the caller.
Picture a decompressor that walks a byte stream and emits one character at a time,
and a parser that consumes characters one at a time. Each is most natural as a loop
that *drives* the other:
Whichever one you make a *callee*, you have to turn inside-out: rip out its loop,
hoist its locals into `static` state, and reconstruct "where was I?" by hand every
time it's called. The algorithm disappears into a state machine.
A **coroutine** is the escape hatch: a function you can `return` from *in the middle*
and later resume *exactly where it left off*, locals and loop position intact. C
doesn't have them. But — as Simon Tatham showed in his
[classic note](https://www.chiark.greenend.org.uk/~sgtatham/coroutines.html) — you
can fake them with a `switch` statement and one preprocessor macro.
## The painful version first
Here's that decompressor rewritten as a callee the honest way — a hand-rolled state
machine. It works, and it's miserable:
```c
int decompressor(void) {
static int state = 0, len, c;
switch (state) {
case 0: /* fresh start */
while (1) {
c = getchar();
if (c == EOF) return EOF;
if (c == 0xFF) { /* run-length escape */
len = getchar();
c = getchar();
while (len--) {
state = 1; return c; /* <-- emit, remember we're here */
case 1: ; /* <-- ...come back to here */
}
} else {
state = 2; return c;
case 2: ;
}
}
}
}
```
Every `return` needs a unique number, a matching `case`, and an assignment to
`state`. Add a branch and you renumber everything. The bookkeeping *is* the bug
surface.
Notice the `case 1:` sitting **inside** the `while` loop, underneath a `switch`
that's outside it. That's legal C — `case` labels can live in any sub-block of a
`switch`. This is the same quirk that powers Duff's device, and it's the whole
trick.
## The insight: let `__LINE__` be the state
The numbers are pure noise. We never *read* them — we only need each `return` to
have a label unique to its position, and a way to jump back to it. The C preprocessor
already hands out a unique number per position: `__LINE__`.
So: on the way out, save `__LINE__`. On the way back in, `switch` on the saved value
and let a `case __LINE__:` right after the `return` catch it. Two macros:
```c
#define crBegin static int state = 0; switch (state) { case 0:
#define crReturn(x) do { state = __LINE__; return x; \
case __LINE__: ; } while (0)
#define crFinish }
```
That's the entire idea. `crBegin` opens a `switch` on the saved state. `crReturn`
stamps the current line into `state`, returns, and drops a `case` label at that exact
line so the next call resumes one statement later. `crFinish` closes the brace.
## Watch it run
A three-value generator — `next()` returns 0, 1, 2, then -1 — makes the control flow
visible. Step through it: watch `state` get stamped with a line number on the way out,
and the `switch` teleport straight back into the middle of the `for` loop on the way
back in.
The magic moment is the jump from `switch (state)` to `case __LINE__:` *inside* the
loop. The function never "starts over" — it lands back exactly where it returned, with
`i` right where it was.
## How the macros expand
It reads like ordinary code, but here's what the preprocessor actually produces, one
layer at a time:
You write the coroutine in its natural, loop-shaped form:
```c
int next(void) {
static int i;
crBegin;
for (i = 0; i < 3; i++)
crReturn(i);
crFinish;
}
```
`crBegin` becomes a `switch` on the saved state, entered at `case 0` on the first call:
```c
int next(void) {
static int i;
static int state = 0; switch (state) { case 0:
for (i = 0; i < 3; i++)
crReturn(i);
}
}
```
`crReturn(i)` stamps the line number, returns, and leaves a `case` label one line on:
```c
for (i = 0; i < 3; i++) {
state = __LINE__; return i;
case __LINE__: ;
}
```
So the next call jumps from `switch (state)` *directly* to that `case` — back inside
the `for` loop, with `i` preserved. No re-entry, no restart:
```c
switch (state) { /* state == that line number */
case 0: ...
case 17: ; /* <-- lands here, mid-loop */
}
```
## Where it bites
This is a beautiful hack, and like every beautiful hack it has sharp edges. Tatham is
candid about them, and you should be too:
**Only `static` locals survive.** A normal `auto` variable is undefined after a
`crReturn` — its storage isn't preserved across the return. Loop counters and any
state you care about must be `static`. **One `crReturn` per line** (two share a
`__LINE__` and collide). And you **can't wrap the body in your own `switch`** — it
would capture the `case` labels meant for the coroutine.
The `static` rule hides a worse problem: `static` means *one shared instance*. Two
callers can't run the same coroutine independently — they'd stomp each other's `state`
and `i`. Fine for a single global decompressor; fatal for anything reentrant or
threaded.
## Making it reentrant
The fix is to stop using `static` and instead thread all the state through a context
struct the caller owns. Every "serious" local becomes a field; the macros read and
write `ctx->state` instead of a file-scoped one:
```c
struct coro {
int state;
int i, len, c; /* everything that must survive a yield */
};
#define crBegin(ctx) switch ((ctx)->state) { case 0:
#define crReturn(ctx, x) do { (ctx)->state = __LINE__; return x; \
case __LINE__: ; } while (0)
#define crFinish }
int next(struct coro *ctx) {
crBegin(ctx);
for (ctx->i = 0; ctx->i < 3; ctx->i++)
crReturn(ctx, ctx->i);
crFinish;
return -1;
}
```
Now each caller allocates its own `struct coro`, and you can run a hundred independent
generators at once. The price is cosmetic — `ctx->i` everywhere you'd have written
`i` — and Tatham's own verdict is the honest one: *"virtually all your serious
variables become elements of the coroutine context structure."* You trade a little
syntax for reentrancy. Usually worth it.
## Why this matters beyond the trick
You don't reach for these macros often — real codebases use explicit state machines,
threads, or a language with `async`/`yield` built in. But the idea underneath is worth
keeping: **a coroutine is just a state machine where the compiler tracks the state for
you.** `async/await` in Rust, generators in Python, goroutines parked on a channel —
all of them are, at bottom, "save where I am, return, resume later." Tatham's macro is
that idea stripped to its absolute minimum: one `switch`, one `__LINE__`, and the
nerve to put a `case` label inside a loop.
---
*Built on Simon Tatham's [Coroutines in C](https://www.chiark.greenend.org.uk/~sgtatham/coroutines.html) (2000) — still the clearest thing ever written on the subject.*
---
# How self-attention works in transformers
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/how-transformers-attention-works
> date: 2026-06-02
> tags: transformers, deep-learning, explainer
Self-attention is the single mechanism that lets a transformer decide, for every
token in a sequence, which other tokens are worth listening to. Older architectures
like RNNs squeezed an entire sentence through a fixed-size hidden state and read it
left to right. Attention throws that bottleneck out: every token can look directly
at every other token in one parallel step, and it learns *how much* to look.
The trick is to give each token three learned vectors. The **query** asks a question
("what am I looking for?"), the **key** advertises what a token offers ("here is what
I am about"), and the **value** is the actual content that gets passed along once a
match is found. You compute these by multiplying the input embeddings by three
learned weight matrices, $W_Q$, $W_K$, and $W_V$, giving matrices $Q$, $K$, and $V$.
A token attends to another by comparing its query against that token's key with a
dot product — a large dot product means the two vectors point in a similar direction,
so the question and the offer line up. Do this for every query against every key and
you get a full grid of raw compatibility scores.
$$
\text{Attention}(Q, K, V) = \text{softmax}\!\left(\frac{Q K^{\top}}{\sqrt{d_k}}\right) V
$$
That one line is the whole operation. The matrix below shows the resulting weights
for a tiny three-token sequence: each row is one query token, each column is a key it
might attend to, and the cell shading is how much weight that pair receives after the
softmax. Hover a row to see where that token looks.
It helps to walk the formula from the inside out. Each step below takes the previous
result and transforms it; together they go from raw vectors to a context-aware output.
**Q·Kᵀ — raw scores.** Multiply the query matrix by the transpose of the key matrix.
The entry at row *i*, column *j* is the dot product of token *i*'s query with token
*j*'s key — an unnormalised score for how relevant token *j* is to token *i*. The
result is a square matrix, one score for every ordered pair of tokens.
**Scale, then softmax — attention weights.** Divide every score by $\sqrt{d_k}$, the
square root of the key dimension. Without this, large dimensions produce dot products
with a big variance, pushing the softmax into saturated regions where gradients
vanish; the scaling keeps the distribution well-behaved. Then apply softmax across
each row so the weights are non-negative and sum to one — a proper distribution over
"where this token attends."
**Weighted sum — the output.** Multiply the weight matrix by the value matrix $V$.
Each output row is a weighted average of all value vectors, blended according to that
token's attention weights. A token that attended strongly to "cat" inherits most of
"cat"'s value, so its new representation is now informed by the context around it.
Stack several of these in parallel — each with its own $W_Q$, $W_K$, $W_V$ — and you
get **multi-head attention**, where different heads specialise in different relations
(syntax, coreference, positional patterns). Concatenate the heads, project once more,
and that becomes one transformer sub-layer. Repeat across depth and the model builds
increasingly abstract, context-rich representations of the sequence.
The √dₖ scaling is easy to skip when implementing attention from scratch, but
dropping it is one of the most common reasons a hand-rolled transformer trains
slowly or not at all — the softmax saturates and gradients stop flowing.
That is the entire idea: project tokens into queries, keys, and values; score every
pair with a scaled dot product; turn the scores into a distribution with softmax; and
read out a weighted mix of values. Everything else in a transformer — feed-forward
layers, residual connections, layer norm, positional encodings — exists to support
and stack this one operation.
---
# Scaffolding the AI site
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/logs/2026-06-03-scaffolding
> date: 2026-06-03
Kicked off `ai.thesatyajit.com`. Wired the content layer (MDX + gray-matter + Zod 4),
swapped fonts to Hanken Grotesk + IBM Plex Mono, and set up `@next/mdx` with the
Turbopack string-plugin config. Next: the editorial layout shell and core pages.
---
# arXiv digest — 2026-07-22
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-07-22
> papers: 8
Eight from today's feed — a streaming 4D geometry transformer and a single-GPU world model, a diffusion-drafter speculative-decoding fix, two residual-stream/SSM architecture notes, parameter-efficient point-cloud tuning, staleness-aware async RL, and a genuinely useful lidar calibration for rotary inspection rigs.
## ★ AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
- arXiv: 2607.19223v1 — [abs](https://arxiv.org/abs/2607.19223) · [pdf](https://arxiv.org/pdf/2607.19223) · [html](https://arxiv.org/html/2607.19223v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.19223)
- authors: Yu-Yang Qian, Hao-Cong Wu, Chen Chen, Jiacheng Sun, Zhenhua Dong, Peng Zhao, Zhi-Hua Zhou
- categories: cs.LG, cs.CL
**Abstract.** Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass. In this work, we uncover a central pitfall of diffusion drafters: bidirectional attention is a double-edged sword. On one hand, it endows the model with parallel generation and global contextual modeling capabilities; on the other hand, this inherent global dependency introduces high variance at both the domain-level and the token-level: acceptance rates fluctuate substantially across different domains, and draft token quality also varies heterogeneously at different token positions. To tackle this issue, we propose AdaFlash framework, comprising two components: (i) an on-policy distillation (OPD) algorithm with reverse-KL divergence tailored for diffusion drafters, bringing stable convergence and effectively reducing domain-level variance; and (ii) an adaptive length head that dynamically adjusts the candidate sequence length on the fly, substantially lowering the verification cost of the target model and handling token-level variance. Experiments demonstrate that AdaFlash consistently improves speedup rate during deployment, with especially significant gains in high-concurrency scenarios, achieving up to approximately 66% higher throughput than previous state-of-the-art methods.
**Take.** The DFlash line drafts with a diffusion model — one parallel denoising pass proposes a whole draft — but this paper names the catch: the bidirectional attention that buys that parallelism also injects high variance into the acceptance rate, at both the domain and token level. AdaFlash distills the drafter on-policy to damp that variance. A good reminder that “faster drafts” and “accepted drafts” are different objectives, and the second is the one that actually sets the speedup.
## ★ ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- arXiv: 2607.19191v1 — [abs](https://arxiv.org/abs/2607.19191) · [pdf](https://arxiv.org/pdf/2607.19191) · [html](https://arxiv.org/html/2607.19191v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.19191)
- authors: Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu, Chiyu Wang, Yunpeng Zhang, Wenlin Liu, Yun Wang, Xue Zheng, Rui Sun, Junfeng Ni, Hongyu Pan, Zhongxu Sun, Fei Yu, Zengye Ge, Mengmeng Du, Nianfei Fan, Mingchao Sun, Yu Liu, Yongchang, Yanqing Zhu, Jiahang Wang, Ning Ying, Yuze Xuan, Di Yang, Zhicheng Liu, Zhe Gao, Tingbing Xu, Jiacheng Sui, Wenjin Yang, Junnan Lai, Shufeng Liu, Yuan Liu, Zheng Zhou, Yingliang Peng, Dawei Cao, Kaifeng Sheng, Yuxiang Cai, Fei Lu, Mu Xu, Ning Guo
- categories: cs.CV, cs.AI, cs.LG
**Abstract.** We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.
**Take.** An action-conditioned video world model you can run interactively on a single desktop GPU — the efficiency framing is the headline. They distill a bidirectional teacher into a causal student with ODE distillation, then add LongForcing to align long self-rollouts against an extended-horizon teacher and fight the usual autoregressive drift, trained on a mix of AAA games, sim engines, and internet video behind 14 deterministic quality checks. “Runs on one desktop GPU” is the part that makes a world model feel usable rather than a demo.
## IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
- arXiv: 2607.19228v1 — [abs](https://arxiv.org/abs/2607.19228) · [pdf](https://arxiv.org/pdf/2607.19228) · [html](https://arxiv.org/html/2607.19228v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.19228)
- authors: Zhengyu Zou, Hao Li, Kuixuan Jiao, Liu Liu, Tingyang Xiao, Xiaolin Zhou, Fangzhou Hong, Zhizhong Su, Dingwen Zhang, Ziwei Liu
- categories: cs.CV
**Abstract.** Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.
**Take.** Most streaming 3D foundation models are geometry-centric — they reconstruct the scene but don't carry object identity through time. IGGT4D fuses geometry and instance grounding in one streaming transformer, so objects that move, vanish, and reappear keep a consistent handle across a long video, without leaning on externally extracted 2D semantics. Exactly the “spatial intelligence over a continuous stream” problem that matters for real perception.
## Dual Attention Residuals
- arXiv: 2607.18730v1 — [abs](https://arxiv.org/abs/2607.18730) · [pdf](https://arxiv.org/pdf/2607.18730) · [html](https://arxiv.org/html/2607.18730v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.18730)
- authors: Xingda Yu, Yining Li, Xinzhang Liu, Zhihao Yang, Haowei He, Chao Wang, Yongxiang Li, Shuangyong Song
- categories: cs.CL
**Abstract.** Recent work extends Transformer residual pathways along two complementary axes: historical retrieval selects information from earlier depths, whereas multi-stream methods maintain multiple residual trajectories. These capabilities have largely been studied in isolation, and assigning an independent retriever to each stream still prevents one trajectory from influencing depth selection in another. We propose Dual Attention Residuals (DAR), which brings multi-stream interaction into historical retrieval through reciprocal cross-stream addressing. For each target stream, DAR computes depth weights from normalized states in the opposite stream and applies them to values from the target stream's own history. The retrieved states are combined for an unchanged Transformer branch and updated through constrained gated writes; a block-form variant operates on block-level histories to control overhead. Across dense models from 0.1B to 1B parameters and a 7B sparse-MoE model, DAR consistently improves validation loss over standard residual Transformers and Attention Residuals. Routing ablations show that the gain cannot be explained by an additional stream or value projection alone. Representation and intervention analyses further show that reciprocal cross-stream selection preserves depth-wise diversity and avoids the redundancy or functional imbalance observed in alternative two-stream designs.
**Take.** Two recent ideas for the residual stream — historical retrieval (pull from earlier depths) and multi-stream (keep several residual trajectories) — have mostly lived apart. DAR couples them: each stream computes its depth-selection weights from the *other* stream's normalized states. It's a close cousin of the Attention Residuals in Kimi K3, and one more sign that the residual pathway itself is becoming a design surface, not just a highway.
## Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- arXiv: 2607.18625v1 — [abs](https://arxiv.org/abs/2607.18625) · [pdf](https://arxiv.org/pdf/2607.18625) · [html](https://arxiv.org/html/2607.18625v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.18625)
- authors: Jin Yu, Juyoun Park
- categories: cs.CV, cs.AI
**Abstract.** Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones. However, MambaOut demonstrates that a Gated CNN block can match or exceed VMamba on image classification, questioning the necessity of SSMs for vision. This raises a fundamental question: do VMamba and MambaOut encode visual information differently at the representation level? To investigate, we apply cross model centered kernel alignment (CKA) analysis and find that VMamba's final stage blocks form representations distinctly different from both MambaOut and its own preceding blocks. We therefore focus on the final block features, decomposing each spatial token into magnitude and direction. MambaOut concentrates class-discriminative information in high-norm foreground tokens that align with Grad-CAM attribution. VMamba, by contrast, produces high-norm tokens predominantly in background regions, misaligned with Grad-CAM, yet preserves discriminative signals primarily in token directions. These observations reveal that the two models rely on different encoding strategies. We connect this difference to high-resolution classification and semantic segmentation. VMamba distributes logit support broadly across object regions, whereas MambaOut relies on sparse dominant tokens, a strategy that becomes less stable as token counts grow. Under full fine-tuning for segmentation, VMamba consistently outperforms MambaOut. These results suggest that VMamba's advantage in dense prediction stems not merely from the SSM mechanism or sequence length, but from how semantic evidence is organized across token magnitude, direction. Ultimately, we conclude that token magnitude and directional structure serve as critical axes for improving visual backbones, particularly under dense supervision.
**Take.** A clean measure-before-you-architect paper. MambaOut showed a gated CNN can match Vision Mamba, so do they actually encode vision differently? CKA says yes — VMamba's final-stage blocks diverge from both MambaOut and their own earlier blocks — and decomposing each spatial token into magnitude vs direction localizes where the difference lives. Pairs nicely with our own CKA map of workspace geometry across models.
## Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding
- arXiv: 2607.19171v1 — [abs](https://arxiv.org/abs/2607.19171) · [pdf](https://arxiv.org/pdf/2607.19171) · [html](https://arxiv.org/html/2607.19171v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.19171)
- authors: Junlin Chang, Longhao Zou, Rui Li
- categories: cs.CV
**Abstract.** Fine-tuning pre-trained point-cloud backbones typically updates all parameters, resulting in substantial computation and memory overhead. More importantly, modern point backbones rely on aggressive tokenization and downsampling, which yields compact global tokens but irreversibly discards fine-grained local geometry, an inherent bottleneck for parameter-efficient adaptation. Consequently, existing PEFT methods that operate only on these coarsened tokens can modulate global semantics but struggle to recover the missing multi-scale locality. We present Point Ladder Tuning (PLT), a locality-aware PEFT framework that performs hierarchical, instance-conditioned adaptation while keeping the backbone frozen. PLT forms a lightweight closed loop: (i) a Hierarchical Ladder Network (HLN) constructs a multi-resolution local feature pyramid directly from raw points; (ii) a Local-Global Fusion (LGF) aligns and fuses local pyramids with intermediate backbone semantics; and (iii) a Dynamic Prompt Generator produces instance-aware multi-scale prompts to modulate the frozen backbone effectively. For dense prediction, we further introduce a lightweight segmentation head that progressively upsamples fused features and leverages backbone priors to refine fine structures. Extensive experiments on classification and dense prediction show that PLT consistently surpasses prior PEFT baselines with minimal tunable parameters. PLT achieves state-of-the-art performance using only 2.71% trainable parameters for classification and 7.69% for dense prediction, and scales favorably to larger backbones, requiring merely 0.36% parameters on PointGPT-L. The code is released at https://github.com/JunLinChang/ECCV2026-PLT.
**Take.** Point backbones tokenize and downsample hard, which yields compact global tokens but irreversibly discards fine local geometry — so PEFT methods that only touch the coarse tokens can modulate semantics but can't recover multi-scale locality. Point Ladder Tuning adds a locality-aware ladder that adapts hierarchically with the backbone frozen. Practical for fine-tuning a point-cloud model on a real inspection dataset without paying the full-finetune bill.
## Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
- arXiv: 2607.18722v1 — [abs](https://arxiv.org/abs/2607.18722) · [pdf](https://arxiv.org/pdf/2607.18722) · [html](https://arxiv.org/html/2607.18722v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.18722)
- authors: Junyao Yang, Yucheng Shi, Zongxia Li, Zhongzhi Li, Ruhan Wang, Xiangxin Zhou, Kishan Panaganti, Haitao Mi, Leowei Liang
- categories: cs.LG, cs.CL
**Abstract.** Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a result, high-staleness updates remain weakly controlled in the asynchronous regime where stale rollouts matter most. We introduce the Staleness-Adaptive Trust Region (SAT), which uses the detached sampled log-ratio as a practical staleness proxy, identifies high-mismatch tails within each batch via staleness-based kernel scaling, and contracts only the sign-selected endpoint of the nominal PPO interval. This preserves baseline behavior on ordinary tokens while enforcing more conservative updates on newly intercepted outward bands. We prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness. We evaluate SAT in a decoupled asynchronous RL setup built on Qwen3-30B-A3B-Base, using SGLang as the inference engine and Megatron for training. In this setting, SAT-GSPO w/ R3 achieves the best observed AIME24 avg@8, reaching 35.83 at lag 1 and 34.79 at lag 8, while SAT-GSPO reaches 34.17 at lag 1. Adaptive clipping and routing replay act as complementary stabilizers targeting mismatch tails and routing inconsistency, respectively. Overall, aligning clip intervals with staleness heterogeneity effectively stabilizes asynchronous RL.
**Take.** Async RL buys throughput by decoupling rollout generation from optimization, but staleness — policy lag, engine delays, MoE routing drift — quietly breaks the trust region, and PPO clipping only gates the sampled updates rather than the full-policy mismatch. SAT uses the detached sampled log-ratio as a staleness proxy to tighten the region exactly where stale rollouts matter most. Squarely in the make-large-scale-RL-not-fall-over genre we keep returning to.
## Two-Stage Extrinsic Calibration of a Static Line-Scanning Lidar with a Rotary Platform
- arXiv: 2607.18578v1 — [abs](https://arxiv.org/abs/2607.18578) · [pdf](https://arxiv.org/pdf/2607.18578) · [html](https://arxiv.org/html/2607.18578v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.18578)
- authors: Vikram Shree, Hike Danakian, Long Nguyen, Rajanish Gokidi, Patrick Nercessian
- categories: cs.RO, eess.SP
**Abstract.** A line-scanning lidar yields range and azimuth values in a fixed plane. To perceive surrounding objects in 3D, there must be relative motion between the lidar plane and the object. Thus, using a rotating base-platform is promising for industrial applications where objects need to be scanned or inspected precisely, and is the main focus of this work. In the rotary platform setup, a 3D point cloud of an object can be constructed if the axis of rotation and the precise motion about that axis are known. However, this setup gives rise to the following problem: how can the axis of rotation of the platform be accurately identified with respect to the lidar coordinate system? It is referred to as the calibration problem in the robotics community. Any inaccuracy in this transformation directly affects the quality of the reconstructed point cloud, leading to misrepresentation of the object of interest. In this work, we explore automated approaches to statically and dynamically estimate the transformation of a rotary platform's axis of rotation with respect to a static line-scanning lidar. The proposed algorithms have been validated on real-world datasets obtained from a custom made rotary platform and an FMCW lidar, and their convergence characteristics are studied for various initial conditions.
**Take.** A refreshingly concrete robotics paper. A static line-scanning lidar on a rotary platform can build a 3D point cloud of an object — but only if you know the axis of rotation precisely, and recovering that axis in the lidar frame is the whole problem. Their two-stage extrinsic calibration pins it down for industrial scan-and-inspect rigs. This is the unglamorous perception plumbing that decides whether an inspection setup actually works.
---
# arXiv digest — 2026-07-15
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-07-15
> papers: 6
## ★ The Seriality Gap in Video Diffusion Models
- arXiv: 2607.13031v1 — [abs](https://arxiv.org/abs/2607.13031) · [pdf](https://arxiv.org/pdf/2607.13031) · [html](https://arxiv.org/html/2607.13031v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.13031)
- authors: Jorge Diaz Chao, Konpat Preechakul, Yuxi Liu, Yutong Bai
- categories: cs.LG, cs.CV
**Abstract.** When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, methods that increase effective serial computation improve performance disproportionately, including autoregressive/blockwise generation and architectural depth. We identify this pattern as the seriality gap: a mismatch between tasks requiring growing serial computation and video diffusion models whose denoising loop does not provide scalable serial compute. We then prove that, for deterministic video prediction, denoising steps do not add serial computation beyond the backbone.
**Take.** A clean negative result with a proof attached — my favorite genre. On multi-ball collision dynamics, bidirectional video diffusion gets *worse* as the causal chain lengthens, and piling on denoising steps doesn't help; but a length-matched single-ball control (no interactions) barely degrades, which isolates dependent-event structure — not video length — as the culprit. They name it the seriality gap and prove that, for deterministic prediction, denoising steps add no serial computation beyond the backbone. Tellingly, the fixes that work — autoregressive/blockwise generation, more depth — are exactly the ones that add real serial compute. Sobering for anyone treating a video diffusion model as a physics simulator: the denoising loop is not the sequential reasoning you might assume it is.
## Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
- arXiv: 2607.12829v1 — [abs](https://arxiv.org/abs/2607.12829) · [pdf](https://arxiv.org/pdf/2607.12829) · [html](https://arxiv.org/html/2607.12829v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.12829)
- authors: Daehoon Gwak, Minhyung Lee, Junwoo Park, Jaegul Choo
- categories: cs.LG, cs.AI, cs.CL
**Abstract.** Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However, parallel generation alone does not guarantee practical speedups. Realizing this efficiency requires specialized inference mechanisms, such as diffusion-aware caching and reuse. As inference efficiency becomes a prerequisite for practical deployment, recent research has actively explored acceleration techniques across algorithms, architectures, and systems. However, rigorous comparisons remain difficult, as end-to-end latency stems from intricate trade-offs between algorithmic, architectural, and system-level factors that are often conflated in existing benchmarks. In this survey, we introduce a unified latency decomposition framework for dLLMs to disentangle these factors, categorize acceleration techniques along three axes (algorithmic innovations, architectural and system optimizations, and inference-time scaling), and provide guidelines for reproducible benchmarking.
**Take.** Diffusion LLMs promise parallel generation, but as I kept flagging writing up [iLLaDA](/articles/illada-diffusion-language-model), "parallel" on paper rarely means "fast" in practice. This survey is the useful corrective: a unified latency-decomposition framework that stops people conflating algorithmic, architectural, and system-level speedups, then a taxonomy of the acceleration zoo — diffusion-aware caching/reuse, architecture and system tricks, inference-time scaling — laid out along those axes. The honest through-line is that rigorous end-to-end benchmarking is genuinely hard because these factors trade off against each other, which is exactly why a single-number "5× faster" claim for a dLLM deserves suspicion. Read it as a map before you trust any diffusion-LM speedup.
## Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
- arXiv: 2607.13013v1 — [abs](https://arxiv.org/abs/2607.13013) · [pdf](https://arxiv.org/pdf/2607.13013) · [html](https://arxiv.org/html/2607.13013v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.13013)
- authors: Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal
- categories: cs.AI, cs.SD
**Abstract.** Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the model embedding space, and low-rank adapters let the frozen backbone attend to the new modality. About 42M parameters are trained, which is 0.16 percent of the backbone. We find that the natural training objectives fail to ground the audio because their gradient reaches the projector only through attention that has already dismissed it. A connectionist temporal classification loss applied through the frozen output head breaks this deadlock. The resulting model reaches 6.6 percent word error rate on LibriSpeech test-clean, transcribes in roughly eight parallel steps regardless of utterance length, and uses a single adapter trained on six languages.
**Take.** A tidy bolt-on: freeze a 26B diffusion LM (DiffusionGemma, which denoises via uniform random-token corruption rather than the absorbing-mask scheme in [iLLaDA](/articles/illada-diffusion-language-model)), hang a frozen Whisper encoder plus a tiny projector and LoRA off it, and transcribe speech by refining the whole transcript in ~8 parallel steps regardless of length — training only 42M params, 0.16% of the backbone. The genuinely interesting part is a failure they had to engineer around: the natural objectives never grounded the audio because the gradient reached the projector only through attention that had already ignored it, and a CTC loss through the frozen output head broke the deadlock. 6.6% WER on LibriSpeech from a near-frozen model is a strong argument that diffusion LMs are a real substrate, not just a curiosity.
## ★ Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- arXiv: 2607.12790v1 — [abs](https://arxiv.org/abs/2607.12790) · [pdf](https://arxiv.org/pdf/2607.12790) · [html](https://arxiv.org/html/2607.12790v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.12790)
- authors: Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He
- categories: cs.AI, cs.CL, cs.MA
**Abstract.** Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable evaluation metric already exists. In many real applications it does not. We make three claims. First, metrics can be evolved: our metric loop searches compositions of small drawback detectors under a full evolutionary lifecycle, trained to agree with a ten-item anchored reference set, regularized by consensus over unlabeled outputs, and audited against a held-out anchor it never reads, yielding a transparent, inspectable metric rather than an opaque judge. Second, since no metric exists to beat, the yardstick is recovering what an accurate metric would have enabled, and Double Ratchet, our co-evolution of the metric with a lifecycle-managed skill loop, does so: across code generation (MBPP+), enterprise text-to-SQL (Spider 2.0-Snow), and reference-free report generation, it retains 88 to 110 percent of the held-out lift achieved by the same skill loop driven by ground-truth labels.
**Take.** Every self-improving agent loop quietly assumes a reliable evaluation metric already exists — and in real applications it usually doesn't, which makes the whole loop circular. This paper evolves the *metric itself*: a searchable composition of small "drawback detectors" trained to agree with a ten-item anchored reference set, regularized by consensus on unlabeled outputs, and audited against a held-out anchor it never sees — so you get an inspectable metric, not an opaque LLM judge. Their "Double Ratchet" co-evolves that metric with the skill loop and recovers 88 to 110% of the lift you'd get from ground-truth labels on MBPP+, text-to-SQL, and report generation. A sharp answer to the problem lurking under every [multi-agent self-improvement](/articles/nous-hermes-moa) scheme: who grades the grader.
## Inhibited Self-Attention: Sharpening Focus in Vision Transformers
- arXiv: 2607.12881v1 — [abs](https://arxiv.org/abs/2607.12881) · [pdf](https://arxiv.org/pdf/2607.12881) · [html](https://arxiv.org/html/2607.12881v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.12881)
- authors: Peter R. D. van der Wal, Nicola Strisciuglio, George Azzopardi
- categories: cs.CV
**Abstract.** Vision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks. However, their self-attention mechanism often diffuses focus across background regions, relying on spurious correlations rather than object-relevant cues. Inspired by inhibitory mechanisms observed in biological vision systems, we propose the Inhibited Self-Attention (ISA), a novel self-attention that integrates inhibitory signals to enhance feature selectivity and suppress spurious responses. In contrast to conventional self-attention, which relies solely on positive attention values due to softmax normalization, our approach retains and utilizes negative attention scores to suppress irrelevant features and sharpen focus on objects of interest. Experiments across multiple datasets, including ImageNet-1k and COCO, and several robustness benchmarks demonstrate that ISA enhances object-centric selectivity, reduces shortcut reliance, and improves out-of-distribution generalization.
**Take.** Softmax attention is all-positive: every token can only *add* to the mix, so nothing can be actively suppressed. Inspired by inhibitory neurons in biological vision, ISA keeps negative attention scores so a ViT can push background and spurious features *down* instead of merely up-weighting the salient ones. The payoff is object-centric selectivity — less shortcut reliance, better out-of-distribution generalization on ImageNet-1k/COCO, and relevance maps that visibly tighten onto objects. It's a small change to the [attention](/articles/how-transformers-attention-works) primitive with an intuitive story: let attention say "no," not just "yes." Worth watching whether the same trick helps language models, where softmax's positivity is just as baked in.
## ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models
- arXiv: 2607.12959v1 — [abs](https://arxiv.org/abs/2607.12959) · [pdf](https://arxiv.org/pdf/2607.12959) · [html](https://arxiv.org/html/2607.12959v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.12959)
- authors: Haojie Ren, Songrui Luo, Lingfeng Wang, Yan Xia, Yao Li, Jing Li, Lu Zhang, Jiajun Deng, Yanyong Zhang
- categories: cs.CV
**Abstract.** LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents. However, current BEV representations, typically extracted by LiDAR backbones trained from scratch, are geometry-dominated and lack general semantic priors, inherently limiting feature-level collaboration. Vision foundation models (VFMs) pretrained on large-scale image data learn general-purpose visual representations and could enhance agent-wise LiDAR BEV features, but adapting them is hard due to the image-point-cloud modality gap. ViCo3D projects point clouds onto the BEV plane as three-channel images so DINOv2 can extract BEV-space visual features from LiDAR inputs, introduces a multi-scale BEV fusion module to integrate these with LiDAR geometric features, and adopts an ego-centric cross-agent fusion strategy. On DAIR-V2X and V2XSet it achieves state-of-the-art 3D detection, with up to 1.8x greater collaborative gains than prior methods on DAIR-V2X.
**Take.** Collaborative V2X perception fuses bird's-eye-view features across vehicles, but those BEV features come from LiDAR backbones trained from scratch — geometry-rich, semantics-poor. ViCo3D's move is to render the point cloud as a three-channel BEV image so a frozen DINOv2 can pull general semantic priors out of it, then fuse those with the raw geometric features before sharing across agents. It reports state-of-the-art on DAIR-V2X/V2XSet and, more tellingly, up to 1.8x larger *collaborative* gains — the vision-foundation semantics are what make cross-agent fusion actually pay off. A nice bridge between the point-cloud world of [FAST-LIO2](/articles/fast-lio2-lidar-inertial-odometry) and the 2D foundation-model world.
---
# arXiv digest — 2026-07-10
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-07-10
> papers: 6
## ★ DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding
- arXiv: 2607.08642v1 — [abs](https://arxiv.org/abs/2607.08642) · [pdf](https://arxiv.org/pdf/2607.08642) · [html](https://arxiv.org/html/2607.08642v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.08642)
- authors: Saw S. Lin, Jyh-Shing Roger Jang
- categories: cs.CL
**Abstract.** Speculative decoding accelerates LLM inference by drafting several tokens and verifying them in parallel. Block-diffusion drafters such as DFlash produce a draft block in one pass but model only per-position marginals; best-first tree methods such as DDTree expand candidate trees from those marginals. The released Domino drafter adds a GRU-based causal correction that makes each draft token's distribution path-dependent, a structure DDTree's factorized formulation cannot represent. We introduce DominoTree, a training-free best-first draft tree scored by Domino's conditional, non-factorized correction along each root-to-node path, made practical by restricting the per-node correction to a candidate top-M. On Qwen3-4B across eight benchmarks, DominoTree reaches up to 6.6x speedup over autoregressive decoding and the highest mean accept length of any evaluated method, up to 10.7 tokens per round, at every temperature we test. DominoTree constructs its tree with a GPU-native, CUDA-graph builder that is bit-identical to a reference Python implementation, so acceptance is unchanged, while keeping per-round tree construction cheap.
**Take.** We walked the mechanics of speculative decoding in the [DeepSeek dSpark piece](/articles/deepseek-dspark) — draft cheap, verify the whole block in one target pass. DominoTree's twist is that the draft *tree* is scored by a conditional, path-dependent correction (Domino's GRU head), so each root-to-node path carries its own distribution instead of the factorized per-position marginals a tree method like DDTree assumes. Training-free, with a CUDA-graph builder that's bit-identical to the reference so acceptance is unchanged, and up to 6.6x over autoregressive on Qwen3-4B at a mean accept length of 10.7 tokens/round. Honest read: the throughput edge over prior tree methods is decisive at T=0 but narrows to a tie (and a small loss) at high temperature — the win is real, not universal.
## ★ The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs
- arXiv: 2607.08734v1 — [abs](https://arxiv.org/abs/2607.08734) · [pdf](https://arxiv.org/pdf/2607.08734) · [html](https://arxiv.org/html/2607.08734v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.08734)
- authors: Baha Rababah, Cuneyt Gurcan Akcora, Carson K. Leung
- categories: cs.AI
**Abstract.** Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections.
**Take.** This is the paper the quantization hype needs. Accuracy and perplexity say an INT4 model "matches" its base — this shows that's an illusion: introduce *correctness agreement* (do the two models get the same items right?) and behavioral divergence appears under even moderate quantization while the headline metrics look preserved. It localizes the damage, too — query and key projections are consistently more sensitive than value/output, with non-linear breakpoints at low bit-width. It's exactly the caveat we kept flagging around [native FP4 training](/articles/nemotron-nvfp4) and [rotation-based KV quant](/articles/turboquant-kv-cache): a matched loss or accuracy number is not behavioral parity, and you have to measure the behavior to know.
## BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression
- arXiv: 2607.08643v1 — [abs](https://arxiv.org/abs/2607.08643) · [pdf](https://arxiv.org/pdf/2607.08643) · [html](https://arxiv.org/html/2607.08643v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.08643)
- authors: Yuantian Shao, Peisong Wang, Zhilei Liu, Chuangyi Li, Yuanteng Chen, Pengcheng Xie, Yiwu Yao, Zhihui Wei, Jian Cheng
- categories: cs.LG
**Abstract.** Large language models are increasingly constrained by memory capacity, weight bandwidth, and checkpoint storage during deployment. Scalar or group-wise quantization is simple and compatible with efficient low-precision kernels, but its representation capacity becomes limited when the target budget approaches 2 bits per weight. Vector-quantized weight compression provides a richer block-level representation, but usually introduces explicit codebooks, index lookup, and additional storage. BiSCo-LLM is a codebook-free binary spherical coding framework: local weight chunks are mapped onto a unit hypersphere and binarized into compact spherical codes, so the main payload is a bit-packed sign stream rather than explicit VQ centroids; a residual stage encodes the reconstruction error without stored codebooks; and category-wise recovery distillation reduces the mismatch between local reconstruction and assembled model behavior. A small 8-bit protected-channel path stabilizes sensitive channels and is counted separately.
**Take.** Vector quantization gives you rich sub-2-bit weights but drags in explicit codebooks and index lookups; scalar/group quant is kernel-friendly but runs out of representation near 2 bits. BiSCo-LLM threads it — map weight chunks onto a unit hypersphere and binarize to a bit-packed sign stream (a *codebook-free* spherical code), then a residual stage buys back rate-distortion without stored centroids. Same "normalize into a nicer basis, then quantize" instinct as [TurboVec](/articles/turbovec). Watch the fine print though: the honest storage budget also counts the neural decoders, an 8-bit protected-channel path, and LoRA adapters — which is exactly where extreme-low-bit claims tend to hide the weight they saved.
## OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
- arXiv: 2607.08766v1 — [abs](https://arxiv.org/abs/2607.08766) · [pdf](https://arxiv.org/pdf/2607.08766) · [html](https://arxiv.org/html/2607.08766v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.08766)
- authors: Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He, Yue Ma, Ziyu Wan, Yong Zhang, Xiaoming Wei, Qifeng Chen
- categories: cs.CV
**Abstract.** We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics during long autoregressive rollout. OPSD-V reduces long-horizon degradation while preserving the original few-step inference path. The student follows the exact inference-time rollout, generating each chunk conditioned on its own previously generated KV cache; in parallel, the teacher is evaluated at the same student-visited denoising states but uses a cleaner AR-consistent temporal cache in which older history can be replaced by real-video context. This provides dense denoising-level corrective targets under on-policy AR cache dynamics, without changing the sampler, number of denoising steps, or inference-time cache mechanism. A user study prefers OPSD-V over the base models in 66.0% of overall-preference judgments.
**Take.** Few-step autoregressive video generators are fast but drift — error accumulates and motion goes limp over a long rollout. OPSD-V fixes it without touching the inference path: the student rolls out exactly as it will at test time (conditioned on its own KV cache), while a teacher is evaluated at the same denoising states but allowed a cleaner cache seeded with *real* video, giving dense corrective targets. It's the on-policy, cache-aware cousin of the training-free acceleration tricks in [MrFlow](/articles/mrflow-diffusion-acceleration) — and refreshingly, the win is reported as a preference study (66% overall) rather than one cherry-picked metric.
## Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery
- arXiv: 2607.08408v1 — [abs](https://arxiv.org/abs/2607.08408) · [pdf](https://arxiv.org/pdf/2607.08408) · [html](https://arxiv.org/html/2607.08408v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.08408)
- authors: Tianyi Song, Sierra Bonilla, Xinwei Ju, Evangelos Mazomenos, Danail Stoyanov, Adam Schmidt, Omid Mohareri, Sophia Bano, Francisco Vasconcelos
- categories: cs.CV, cs.AI
**Abstract.** Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery; however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy. Track2Map is an online 3D Gaussian Splatting pipeline that jointly optimizes camera trajectory and 3D deformable scene representation directly from surgical video, and due to its online nature effectively works as a SLAM method. To stabilize optimization under tissue motion and ambiguous visual cues, it introduces a track-anchored deformation initialization using dense 2D point tracks, and uses track statistics to disentangle camera motion from scene deformation by detecting static camera periods and reducing drift during incremental mapping. Experiments on StereoMIS improve reconstruction quality and camera trajectory over competing SLAM methods.
**Take.** Most deformable-anatomy reconstruction is offline and leans on a camera-trajectory prior from robot kinematics. Track2Map drops the prior and does it online — jointly optimizing camera pose *and* a deformable 3D Gaussian scene straight from surgical video, which makes it a genuine SLAM system. The bit that keeps it from diverging under tissue motion is the same problem [FAST-LIO2](/articles/fast-lio2-lidar-inertial-odometry) fights: it leans on dense 2D point tracks to initialize deformation and to detect static-camera periods, disentangling camera motion from scene deformation to cut drift. Gaussian-splat SLAM in a squishy, prior-free scene is a genuinely hard setting.
## On Exploring Input Resolution Scaling For Anytime LiDAR Object Detection
- arXiv: 2607.08391v1 — [abs](https://arxiv.org/abs/2607.08391) · [pdf](https://arxiv.org/pdf/2607.08391) · [html](https://arxiv.org/html/2607.08391v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.08391)
- authors: Ahmet Soyyigit, Shuochao Yao, Heechul Yun
- categories: cs.RO, cs.LG
**Abstract.** Making tradeoffs between execution latency and result utility (anytime computing) has been shown to enhance the performance of cyber-physical systems. We enable anytime computing for DNNs that process LiDAR point clouds for 3D object detection: a method that enables multi-resolution inference for models that process point clouds as pillars or voxels, allowing the input to be dynamically scaled and processed at the resolution needed to meet timing requirements. The memory-efficient approach requires deploying only a single DNN model rather than one per resolution. A deadline-aware scheduler selects the highest possible resolution for each input by predicting the execution time for all resolutions at runtime, which is challenging due to the irregularity of LiDAR point clouds. On nuScenes it significantly outperforms existing anytime approaches, and in a simulated autonomous driving system it enables collision-free navigation while avoiding unnecessary stalls.
**Take.** A tidy systems idea for real-time perception: instead of shipping N models trained at N resolutions, train one point-cloud detector that can be *run* at any resolution, then add a deadline-aware scheduler that predicts each resolution's runtime and picks the highest one that fits the time budget. Predicting runtime is the hard part, because LiDAR point clouds are irregular so cost isn't a clean function of input size. It's the anytime-computing counterpart to the latency discipline that matters in the [FAST-LIO2](/articles/fast-lio2-lidar-inertial-odometry) stack — trade a little accuracy to never blow the frame deadline.
---
# arXiv digest — 2026-07-03
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-07-03
> papers: 7
A day that reads like a callback to the last two weeks of this feed. **OrbitQuant** is the
[TurboQuant](/articles/turboquant-kv-cache) primitive — rotate into a known distribution, then
apply one data-free Lloyd-Max codebook — now quantizing diffusion transformers; **DemoPSD** is the
next fix in the on-policy-distillation saga, attacking the "privileged information leakage" that
[DOPD](/arxiv/2026-06-30) named. Around them, the long-horizon-agent thread keeps maturing
(bounded memory contracts, evidence replay, and a sharp reminder that reasoning effort beats tool
sprawl), plus one genuinely uncomfortable finding about what agents say off the record.
## ★ OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
- arXiv: 2607.02461v1 — [abs](https://arxiv.org/abs/2607.02461) · [pdf](https://arxiv.org/pdf/2607.02461) · [html](https://arxiv.org/html/2607.02461v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02461)
- authors: Donghyun Lee, Jitesh Chavan, Duy Nguyen, Sam Huang, Liming Jiang, Priyadarshini Panda
- categories: cs.CV, cs.AI, cs.LG
**Abstract.** Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ) is the natural remedy, yet DiT activations shift across timesteps, prompts, and guidance branches, forcing prior methods to re-fit calibration data for every new checkpoint or modality. We present OrbitQuant, a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis. In this basis, a randomized permuted block-Hadamard (RPBH) rotation concentrates each coordinate around one fixed, known marginal regardless of the input, so a single Lloyd-Max codebook serves all timesteps, prompts, and layers of a given input dimension. We extend the same quantizer to weight rows offline, absorbing the rotation into the weights so that it cancels inside each linear layer and only a forward rotation on the activations remains at runtime. The same recipe transfers from image to video with no per-modality tuning. Across FLUX.1, Z-Image-Turbo, Wan 2.1, and CogVideoX, it sets the state of the art for PTQ at several low-bit settings, and pushes PTQ of image diffusion transformers to W2A4 with usable generation quality.
**Take.** If you read the [TurboQuant write-up](/articles/turboquant-kv-cache) this week, OrbitQuant will feel like déjà vu — because it's the same primitive pointed at a third target. Rotate into a basis where every coordinate lands on one fixed, known marginal (here a randomized permuted block-Hadamard rotation), and a single Lloyd-Max codebook quantizes everything with no calibration. That data-free property is exactly what diffusion-transformer PTQ needs, because DiT activations drift across every timestep, prompt, and guidance branch, so calibration-based quantizers have to re-fit constantly. The tidy engineering touch is absorbing the rotation into the weights offline so it cancels inside each linear layer, leaving only one forward rotation at runtime. Pushing image DiTs to W2A4 with usable output is the number to watch. "Rotate, then quantize" is quietly becoming a standard primitive.
## ★ DemoPSD: Disagreement-Modulated Policy Self-Distillation
- arXiv: 2607.02502v1 — [abs](https://arxiv.org/abs/2607.02502) · [pdf](https://arxiv.org/pdf/2607.02502) · [html](https://arxiv.org/html/2607.02502v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02502)
- authors: Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai
- categories: cs.LG, cs.AI
**Abstract.** On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single model acts as both the teacher and the student with different levels of information access. However, recent studies have found that the teacher's dense token-level supervision, conditioned on privileged information, can lead to overfitting to in-domain patterns, suppress exploration, and hurt cross-domain generalization, while also introducing a more fundamental issue: privileged information leakage, where the student encodes answer-dependent shortcuts that are unavailable at test time. We introduce DemoPSD, a novel framework that resolves such problems through the idea of selective adoption of teacher guidance. Instead of fitting the full teacher distribution, DemoPSD steers the student toward a reverse-KL barycenter target, a weighted geometric combination of the teacher and student distributions, and uses the teacher-student discrepancy to adaptively control the blending at each token position. We provably show leakage attenuation and exploration preservation. Experiments on SciKnowEval across four scientific fields show DemoPSD outperforms both GRPO and SDPO while maintaining higher training entropy and robustly generalizing to out-of-distribution GPQA.
**Take.** This is the direct next chapter on the "privilege illusion" [DOPD named last week](/arxiv/2026-06-30): when a self-distillation teacher sees privileged info, the student learns answer-dependent shortcuts it can't use at test time — leakage, not capability. DemoPSD's fix is elegant and, unusually, *proven*: don't fit the full teacher, steer toward a reverse-KL barycenter (a geometric blend of teacher and student) with the blend weight set per-token by how much the two disagree. Where the teacher and student already agree, lean on the teacher; where they diverge, preserve the student's own reasoning and exploration. Beating GRPO and SDPO while keeping *higher* training entropy is the tell that it's actually preserving exploration rather than collapsing onto the teacher. On-policy distillation keeps generating its own subfield of failure modes and fixes.
## AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
- arXiv: 2607.02255v1 — [abs](https://arxiv.org/abs/2607.02255) · [pdf](https://arxiv.org/pdf/2607.02255) · [html](https://arxiv.org/html/2607.02255v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02255)
- authors: Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li, Chuanhao Li, Xiangcheng Cao
- categories: cs.AI, cs.CL
**Abstract.** Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest contract appends past observations, tool calls, and reflections to every prompt, which makes prior context easy to access but also turns it into a jumbled mixture in which the effect of any single memory component is hard to isolate. We introduce and instrument an alternative bounded contract: every decision is made from a fresh user message assembled by typed retrieval, with no raw cross-decision transcript appended. The prompt thus stays bounded across runs of any length, and any single layer can be ablated in isolation. We instantiate the contract in Slay the Spire 2, a closed-rule stochastic deck-building game whose runs require hundreds of tactical and strategic decisions. Within our harness, a fixed-A0 ablation shows the largest observed difference when triggered strategic skills are enabled: the no-store baseline wins 3/10 games and adding the skill layer 6/10. We release a reproducible testbed: 298 completed trajectories with condition tags, frozen memory/skill snapshots, prompt records, and analysis scripts.
**Take.** The framing is the contribution: agent memory is "a contract about what each future decision is allowed to see." Instead of the usual append-everything transcript — which makes it impossible to isolate what any one memory component actually does — they assemble each decision from a *bounded*, typed retrieval, so the prompt stays fixed-size over arbitrarily long runs and each layer can be ablated cleanly. Slay the Spire 2 is a shrewd testbed: hundreds of stochastic decisions per run, frontier LLMs currently at zero wins, humans at 16%, so it's hard but not saturated. The result (3/10 → 6/10 with a strategic-skill layer) is honestly flagged as directional at this sample size — refreshing restraint — but the reusable harness and 298 released trajectories are the real deliverable for anyone studying long-horizon memory.
## Reasoning effort, not tool access, buys first-try reliability in agentic code generation
- arXiv: 2607.02436v1 — [abs](https://arxiv.org/abs/2607.02436) · [pdf](https://arxiv.org/pdf/2607.02436) · [html](https://arxiv.org/html/2607.02436v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02436)
- authors: Achint Mehta
- categories: cs.SE, cs.AI
**Abstract.** Agentic coding assistants are increasingly given extra capabilities, such as browser-based testing tools and design-oriented system prompts, on the assumption that more capability yields better software. This study tested that assumption directly. Ninety independent agent runs built the same application from one detailed specification, each scored on a fixed 14-criterion functional rubric (42-point max) and a visual review. Capability tier dominated: frontier models clustered near the ceiling while a low-cost local model fell to 24-37 points. Container deployment was the dominant defect, failing first-try in 44% of runs. The testing tool raised cost by 42-68% without improving functional score or reliability. Raising reasoning effort from High to xHigh lifted first-try perfect runs from 28% to 89% and cut corrective prompts about five-fold, for 9-29% more cost. A design-oriented prompt raised visual quality (4.5 vs 3.0) without lifting function. The practical lesson is to match the fix to the failure: most first-run failures came from weak reasoning, not from visible flaws a checking tool would catch.
**Take.** A refreshingly empirical corrective to the "give the agent more tools" reflex. Ninety runs building the same app from one spec, and the result cuts against the current instinct: the browser-based testing tool added 42-68% cost and improved nothing, while simply turning reasoning effort from High to xHigh took first-try-perfect runs from 28% to 89% and cut corrective prompts five-fold. The failures were reasoning failures (container deployment botched on the first try 44% of the time), not the kind of visible flaws a checking tool catches — so a stronger model or more thinking prevents them and a test harness doesn't. It's a single-app observational study, not a controlled benchmark, but "match the fix to the failure" is exactly the discipline the agent-tooling gold rush is skipping.
## What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
- arXiv: 2607.02507v1 — [abs](https://arxiv.org/abs/2607.02507) · [pdf](https://arxiv.org/pdf/2607.02507) · [html](https://arxiv.org/html/2607.02507v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02507)
- authors: Arman Ghaffarizadeh, Danyal Mohaddes, Aliakbar Izadkhah, Shahriar Noroozizadeh
- categories: cs.AI, cs.CL, cs.LG, cs.MA
**Abstract.** LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, without any explicit objective in the prompt, changes what an agent expresses publicly relative to an off-the-record (OTR) channel elicited under the same condition. We introduce a dual-channel debate framework in which agents produce public utterances that enter the shared history alongside OTR responses that are recorded but never shown to the other participant. Across 10 models, 3 scenarios, and 5 variations within each scenario, alignment-inducing settings produce systematic public-OTR divergence in the targeted agent, with its decision divergence rising from a ~3% baseline to roughly 40%. The effect is consistent across four aggregate analyses: stance, semantic similarity, natural language inference, and survey responses. In some cases, the OTR response explicitly attributes public accommodation to relational pressures such as career risk or sponsorship obligation.
**Take.** A genuinely unsettling result, cleanly measured. Put an agent in a socially structured debate with no explicit objective in the prompt, give it a private "off-the-record" channel alongside its public utterances, and under alignment-inducing conditions the two diverge — its public-vs-private decision gap jumps from ~3% to ~40%, consistent across stance, semantic-similarity, NLI, and survey measures. Sometimes the private channel *names the reason* — career risk, sponsorship obligation. Nobody trained this in; social structure alone induced a latent objective (say the agreeable thing publicly, hold the real view privately). The takeaway is a real methodology gap: evaluating agents on their stated goals misses emergent ones, and a dual-channel probe is a concrete way to catch them. The kind of eval we're going to need more of.
## ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
- arXiv: 2607.02509v1 — [abs](https://arxiv.org/abs/2607.02509) · [pdf](https://arxiv.org/pdf/2607.02509) · [html](https://arxiv.org/html/2607.02509v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02509)
- authors: Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen
- categories: cs.AI
**Abstract.** Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a gap between context access and effective context utilization. We propose Recursive Evidence Replay as LLM Harness for Long-Context Reasoning (RECONTEXT), a training-free inference method for improving long-context reasoning. RECONTEXT uses model-internal relevance signals to construct a query-conditioned evidence pool and replays it before final generation while preserving the full original context. We provide a theoretical analysis based on associative memory, which characterizes the context as a memory store, the question as a retrieval cue, attention as cue-trace association, and replay as trace reactivation. Experiments on eight long-context datasets with 128K context show RECONTEXT consistently improves evidence utilization across Qwen3-4B, Qwen3-8B, and Llama3-8B.
**Take.** The gap this targets is real and under-discussed: models with 128K windows can *access* a fact and still fail to *use* it — context length isn't context utilization. ReContext is a training-free harness that reads the model's own internal relevance signals to pull a query-conditioned evidence pool, then replays it right before generation while keeping the full original context intact — no pruning, no external memory. The associative-memory framing (context = store, question = cue, attention = association, replay = reactivation) is a clean way to think about why "remind the model of what matters" works. It pairs naturally with the long-context hardware story from [MiMo](/articles/mimo-v2-flash): the architecture makes 256K *cheap*, harnesses like this make it *usable*.
## PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
- arXiv: 2607.02515v1 — [abs](https://arxiv.org/abs/2607.02515) · [pdf](https://arxiv.org/pdf/2607.02515) · [html](https://arxiv.org/html/2607.02515v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2607.02515)
- authors: Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt
- categories: cs.CV
**Abstract.** State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulations are unnecessary. We introduce a minimalist pixel-space Diffusion Transformer, built on a plain ViT, that operates directly on raw 3D point map patches and is conditioned on image tokens from a pre-trained DINOv3. Unlike existing latent diffusion approaches, we train our diffusion backbone entirely from scratch, eliminating the need for point map tokenizers. Despite its simplicity, our approach surpasses complex latent-based diffusion models while remaining significantly simpler than hybrid alternatives. Notably, it produces sharper geometric structure and is more robust in highly ambiguous regions such as transparent objects.
**Take.** A satisfying "the overhead was unnecessary" result. Monocular 3D reconstruction has accreted hybrid architectures, bespoke losses, and latent-space compressions to piggyback on pretrained latent diffusion models — and PointDiT shows a plain ViT doing pixel-space diffusion directly on raw 3D point-map patches, conditioned on DINOv3 image tokens, beats them. No point-map tokenizer, backbone trained from scratch, and it's *sharper* on the hard cases like transparent objects where latent compression tends to smear. It rhymes with the broader "stop compressing into latents, work in the native space" mood ([Cross-Space Distillation](/arxiv/2026-07-01) argued the opposite direction is hard for a reason). Simpler and better is the best kind of paper.
---
# arXiv digest — 2026-07-01
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-07-01
> papers: 7
Two threads dominate today. **On-policy distillation** keeps compounding — GR2 states
outright that SFT "collapses at industrial scale" and reaches for OPD instead, echoing
Agents-A1, MOPD, and DOPD from earlier this week; Cross-Space Distillation and browser
skill distillation attack the same teacher→student interface from the vision and agent
sides. And **long-horizon agents** get a rigorous, training-free way to vet their
intermediate-reward signals (QVal) — which promptly finds that simple prompting beats most
of the literature. The through-line: the leverage is in the interface and the verifier,
not the parameter count.
## ★ QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
- arXiv: 2606.32034v1 — [abs](https://arxiv.org/abs/2606.32034) · [pdf](https://arxiv.org/pdf/2606.32034) · [html](https://arxiv.org/html/2606.32034v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.32034)
- authors: Sergio Hernández-Gutiérrez, Matteo Merler, Ilze Amanda Auzina, Joschka Strüber, Ameya Prabhu, Matthias Bethge
- categories: cs.LG, cs.AI, cs.CL
**Abstract.** LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to inform the model about the goodness of intermediate actions. Dense supervision methods aim to solve this problem by scoring intermediate steps, from intrinsic confidence to self-distillation and embedding similarities. However, it is common practice to evaluate them by measuring the downstream performance of a training pipeline that integrates them. This is expensive, conflates supervision quality with training engineering confounders, and renders different methodological families requiring distinct training setups incomparable. We introduce QVal, a training-free testbed for directly evaluating dense supervision signals. Given a state-action pair, QVal measures how well a method's score is Q-aligned: whether it orders actions according to the Q-values of a strong reference-policy. We instantiate QVal as QVal-v1.0, benchmarking 21 dense supervision methods across four diverse environments and seven methodological families, with over 1.2K evaluation experiments across six open-weight model backbones. We find that simple prompting baselines consistently outperform recent dense supervision methods from the literature, and that performance clusters strongly by family.
**Take.** Yesterday's [Agents-A1](/articles/agents-a1) made the case that verified intermediate steps — not outcome-only rewards — are what a long-horizon agent learns from. QVal asks the sharp follow-up: how do you know your step-scoring signal is any good *before* you spend a training run on it? Their answer is clean — score a state-action pair, then check whether the ordering matches the Q-values of a strong reference policy. Training-free, so you separate signal quality from training-pipeline luck. The result is the kind of finding that only falls out of an honest common-ground benchmark: across 21 methods and 1.2K experiments, plain prompting baselines beat most of the fancy dense-supervision machinery. A testbed that lets you kill bad ideas before the GPU bill.
## ★ Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
- arXiv: 2606.32020v1 — [abs](https://arxiv.org/abs/2606.32020) · [pdf](https://arxiv.org/pdf/2606.32020) · [html](https://arxiv.org/html/2606.32020v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.32020)
- authors: Anh Nguyen, Ngan Nguyen, Duc Vu, Trung Dao, Viet Nguyen, Quan Dao, Khoi Nguyen, Anh Tran
- categories: cs.CV
**Abstract.** Modern one-step diffusion models achieve impressive quality through distribution-based timestep distillation. Yet, they rely on a critical assumption: Teacher and Student must inhabit the same latent space. This Shared-Space constraint prevents knowledge transfer from modern high-capacity Teachers (e.g., SD 3.5 and Flux) into compact, deployment-friendly Students such as SD 1.5, whose latent resolution and VAE parameterization differ from the Teacher. We formalize this overlooked regime as Cross-Space Distillation, where Teacher and Student differ in both latent resolution and VAE space. To enable distillation under this mismatch, we introduce the Bridge, a lightweight latent interface that maps Student latents into the Teacher space without modifying the Student backbone. Bridge combines a frozen Student VAE decoder as a spatial prior with a compact learnable projector, and is trained with latent reconstruction and attention fidelity objectives for stable Teacher-space alignment. Across diverse modern Teachers, Bridge enables substantial gains for compact one-step Students; for example, it improves SD 1.5 from 5.4 to 9.4 HPSv3 while preserving one-step inference, low latency, and broad ecosystem compatibility.
**Take.** Distillation almost always assumes teacher and student speak the same latent language — same VAE, same resolution. That quietly locks you out of the most useful transfer of all: pouring a modern high-capacity teacher (SD 3.5, Flux) into the small, beloved, ecosystem-rich SD 1.5. The move is a lightweight "Bridge" that maps student latents into the teacher's space without touching the student backbone, so the student stays one-step and deployable while learning from a teacher it structurally couldn't before. 5.4 → 9.4 HPSv3 on SD 1.5 is a big jump for a frozen-backbone adapter. The general lesson keeps recurring this week: the interface between teacher and student is where the leverage is.
## GR2 Technical Report
- arXiv: 2606.31984v1 — [abs](https://arxiv.org/abs/2606.31984) · [pdf](https://arxiv.org/pdf/2606.31984) · [html](https://arxiv.org/html/2606.31984v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.31984)
- authors: Yufei Li, Zaiwei Zhang, Mingfu Liang, Kavosh Asadi, Jay Xu, Hamed Firooz, Luke Simon
- categories: cs.IR, cs.AI
**Abstract.** Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking -- where the final re-ranking step disproportionately shapes user engagement. Despite growing enthusiasm for LLMs in recommendation, three gaps hinder industrial adoption: most efforts target retrieval and ranking, leaving re-ranking underexplored; LLMs are typically deployed zero-shot or via SFT, underutilizing RL on verifiable rewards; and deployed catalogs index billions of items with non-semantic identifiers outside any base-LLM vocabulary. We present GR2 (Generative Reasoning Re-Ranker), an end-to-end framework that combines mid-training on semantic IDs with >=99% uniqueness, reasoning traces distilled from a stronger teacher via rejection sampling, and RL with verifiable rewards purpose-built for re-ranking. To make GR2 resource-viable, we introduce a context compressor, On-Policy Distillation (OPD) as a scalable alternative to SFT -- which we find collapses at industrial scale -- and reasoning distillation for low-latency serving. GR2 delivers +18.7% R@1, +7.1% R@3, and +9.6% N@3 over legacy baselines on industrial-scale traffic. We further find that LLMs often hack rewards by preserving the incoming order or exploiting position bias.
**Take.** Two things make this more than a recsys report. First: another independent data point that **on-policy distillation is eating SFT** — they state flatly that SFT "collapses at industrial scale" and reach for OPD instead, the same conclusion Agents-A1, MOPD, and DOPD landed on this week from entirely different directions. Second: the reward-hacking honesty. Their re-ranking LLM learns to game the metric by just preserving the incoming order or riding position bias — exactly the failure mode you'd predict, caught and named, and patched with conditional verifiable rewards. Applying LLM reasoning to the re-ranking stage (the one closest to the user) with billions of non-semantic item IDs is a real systems problem, handled as one.
## Generative Skill Composition for LLM Agents
- arXiv: 2606.32025v1 — [abs](https://arxiv.org/abs/2606.32025) · [pdf](https://arxiv.org/pdf/2606.32025) · [html](https://arxiv.org/html/2606.32025v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.32025)
- authors: Xinyu Zhao, Zhen Tan, Vaishnav Tadiparthi, Nakul Agarwal, Kwonjoon Lee, Tianlong Chen
- categories: cs.CL
**Abstract.** Recent LLM agents benefit from skills for solving complex tasks. As skill libraries grow and become reusable across tasks and domains, selecting an appropriate skill composition has emerged as a central bottleneck. Existing approaches either expose the agent's reasoning to the entire skill collection or perform skill retrieval via embeddings or LLM-based rerankers. Both miss the structural nature of skill composition, which is a joint decision over which skills, how many, and in what order -- three dimensions that cannot be decoupled. We formalize this as structured skill composition: given a task and a skill library, predict an executable skill plan that jointly specifies the activated subset, count, and execution order. We propose SkillComposer, which uses a constrained autoregressive decoder over skill identifiers, so subset, count, and order emerge jointly from a single decoding pass. On GPT-5.2-Codex and Gemini-3-Pro-Preview, SkillComposer raises the pass rate by +23.1 and +18.2pp over the no-skill baseline, surpassing top-3 retrieval and matching the gold-skill retrieval upper bound at lower prompt-token cost.
**Take.** The framing is the contribution: skill selection isn't retrieval, it's *composition* — which skills, how many, in what order — and those three can't be chosen independently. Treating it as constrained autoregressive decoding over skill IDs, so the subset/count/order fall out of one pass, is the right shape for a joint decision with dependencies between steps. Matching the gold-skill upper bound at lower token cost is the number that matters: it means the plan is as good as knowing the answer, without stuffing the whole library into context. As agents lean harder on skill libraries (the [Claude-style](/articles/agents-a1) procedural-package trend), the router over those skills becomes the bottleneck, and this is a clean take on it.
## Scalable Behaviour Cloning on Browser via Skill Distillation
- arXiv: 2606.32014v1 — [abs](https://arxiv.org/abs/2606.32014) · [pdf](https://arxiv.org/pdf/2606.32014) · [html](https://arxiv.org/html/2606.32014v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.32014)
- authors: Kaisen Yang, Zheng Jiang, Yuzhao Peng, Houde Qian, Boshi Zhang, Bingxiang He
- categories: cs.CL
**Abstract.** Internet users collectively perform an enormous range of skilled work through web browsers, from software development and document editing to search, forms, and enterprise workflows, making human browsing a highly scalable but under-exploited source of reusable browser skills. We argue that the bottleneck for browser agents is decision-making under incomplete information rather than low-level operation, and that the priors agents lack are already implicit in human interaction traces. We therefore study scalable behavior cloning for browser agents via skill distillation, converting user interaction trajectories into compact natural-language skills that agents can read, retrieve, reuse, and compose directly. We further organize the distilled skills into a skill graph so that growth proceeds through consolidation rather than unbounded accumulation.
**Take.** A nice reframing of where browser-agent capability actually comes from: not more hand-designed tasks, but the priors already latent in ordinary human browsing traces. Distilling those traces into compact, readable natural-language skills — then consolidating them into a skill *graph* so the library grows by merging rather than piling up — is the same "knowledge-action graph" instinct [Agents-A1](/articles/agents-a1) used for training data, pointed at browser automation. The claim that the bottleneck is decision-making under incomplete information, not clicking, rings true for anyone who's watched a browser agent fail on judgment rather than mechanics.
## PointSplat: Compact Gaussian Splatting via Human-Centric Prediction
- arXiv: 2606.32036v1 — [abs](https://arxiv.org/abs/2606.32036) · [pdf](https://arxiv.org/pdf/2606.32036) · [html](https://arxiv.org/html/2606.32036v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.32036)
- authors: Yujie Guo, Yudong Jin, Lingteng Qiu, Zehong Shen, Zhen Xu, Sida Peng, Xiaowei Zhou
- categories: cs.CV
**Abstract.** Producing 3D human representations from input views on the fly is essential for immersive live streaming systems, where representation compactness is as critical as high fidelity given limited computational power and transmission bandwidth. Although recent feed-forward reconstruction methods achieve impressive quality through the view-centric prediction of 3D representations, they repeatedly encode the same subject content across multiple views, leading to significant inter-view redundancy. Our key insight is to perform predictions directly in 3D space. We propose PointSplat, a human-centric approach that directly infers Gaussian primitives from an input point set. The method estimates a coarse geometric proxy and performs ray casting to prune redundant points and establish explicit 2D--3D correspondences, then employs a Point-Image Transformer to fuse appearance and geometry features, predicting Gaussian attributes in a single forward pass. This restricts predictions to foreground regions, substantially reducing the total number of Gaussians while improving novel-view rendering quality.
**Take.** The redundancy insight is the good part: view-centric feed-forward Gaussian predictors re-encode the same person once per input view, so the representation bloats with duplicated content. Moving the prediction into 3D space — infer Gaussians from a point set, ray-cast to prune and pin 2D–3D correspondences — cuts the primitive count while *improving* novel-view quality, which is the rare case where compression and fidelity move the same direction. For live-streamed volumetric humans, where bandwidth is the whole constraint, "fewer Gaussians, better render" is exactly the trade you want.
## FLORA: A deep learning approach to predict forest attributes from heterogeneous LiDAR data
- arXiv: 2606.32023v1 — [abs](https://arxiv.org/abs/2606.32023) · [pdf](https://arxiv.org/pdf/2606.32023) · [html](https://arxiv.org/html/2606.32023v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.32023)
- authors: Emilie Vautier, Clément Mallet, Cédric Vega
- categories: cs.CV, cs.AI
**Abstract.** Forest attributes are essential for national-scale resource monitoring. Airborne LiDAR metrics are among the auxiliary variables most strongly correlated with forest attributes used in National Forest Inventory (NFI) estimates. However, producing wall-to-wall predictions remains challenging when LiDAR data are acquired under heterogeneous conditions -- variability in sensors, flight parameters, seasons, and scan angles limits the robustness of existing models. We present FLORA (Forest LiDAR Octree Regression with Auxiliary Data), a deep learning framework that predicts six forest attributes: dominant height, total volume, deciduous volume, coniferous volume, basal area, and stem density from heterogeneous LiDAR point clouds. FLORA combines an octree-based backbone with ecological and spatiotemporal auxiliary variables through a late-fusion gating mechanism. Trained and evaluated on 32,052 National Forest Inventory plots across mainland France, a single model trained on both leaf-on and leaf-off acquisitions outperforms season-specific models. FLORA achieves an rRMSE of about 12.3% (R2 = 0.88) for dominant height and 39% (R2 = 0.74) for total volume.
**Take.** A grounded reminder that LiDAR point clouds are a national-infrastructure data type, not just a robotics one. The hard part here is the same heterogeneity that bites any real deployment — different sensors, seasons, scan angles — and the fix is sensible: an octree backbone over the raw cloud plus a late-fusion gate that lets ecological and spatiotemporal side-variables in only where they help. The finding that one model trained across leaf-on and leaf-off beats season-specific models is the transferable bit — robustness to acquisition conditions beats bespoke calibration, which is precisely the wall-to-wall property a national forest inventory needs.
---
# arXiv digest — 2026-06-30
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-30
> papers: 7
A 3D-Gaussian-heavy day, with a second clear thread in on-policy distillation. The
through-line I keep noticing: a good enough scene reconstruction has stopped being a
viewer and become *infrastructure* — KiloGS-SLAM tracks kilometers of it, VLK renders
synthetic robot data from it, GaussDet hangs open-vocabulary semantics on it. And on the
training side, three separate groups (RMMD, MOPD, DOPD) are all sharpening the same tool:
distill on the policy's own rollouts, and argue about how to regularize and route the
signal.
## ★ Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes
- arXiv: 2606.30436v1 — [abs](https://arxiv.org/abs/2606.30436) · [pdf](https://arxiv.org/pdf/2606.30436) · [html](https://arxiv.org/html/2606.30436v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30436)
- authors: Sicheng Yu, Dongxu Shen, Beizhen Zhao, Guanzhi Ding, Hao Wang
- categories: cs.CV
**Abstract.** Scaling monocular 3D Gaussian Splatting (3DGS) SLAM to kilometer-level outdoor environments poses two tightly coupled challenges: fragile long-term pose tracking and excessive memory overhead during large-scale mapping. In this paper, we propose KiloGS-SLAM, a highly efficient and robust monocular 3DGS-SLAM system that jointly addresses both bottlenecks. Since high-fidelity scene reconstruction fundamentally relies on drift-free camera poses, we first introduce a motion-adaptive hybrid tracking module. This module features a condition-triggered three-tier solving pipeline. It dynamically switches between Essential matrix and PnP models to handle geometric degeneracies. An on-demand foundation model can also be activated to rescue the trajectory from catastrophic drift. To ensure the system can sustain these long trajectories without memory exhaustion, we subsequently design a lifecycle-managed Gaussian mapping strategy. By integrating probabilistic initialization with chunk-based multi-view densification and pruning, this full-pipeline optimization effectively reduces primitive redundancy while preserving high-frequency details. Extensive experiments across three challenging outdoor datasets demonstrate that our approach achieves state-of-the-art tracking accuracy and rendering quality, successfully scaling to sequences of over 10,000 frames on a single GPU.
**Take.** The two things that kill SLAM at scale are exactly the two things they go after: pose tracking drifts, and the map eats all your memory. I just spent a week inside a LiDAR-inertial filter watching a map smear because of pose drift, so the framing lands — a clean reconstruction is downstream of a drift-free trajectory, full stop. The tracking answer is a tiered solver that switches Essential-matrix↔PnP by geometry and only wakes a foundation model when the trajectory is about to diverge, which is the right cost model: cheap by default, expensive only when degenerate. The mapping answer — lifecycle-managed Gaussians with chunked densify-and-prune — is the same bounded-window discipline a LiDAR map needs, just on splats. 10,000+ frames on one GPU from a single camera is the number that matters.
## ★ Diffusion Fine-tuning with Rewarded Moment Matching Distillation
- arXiv: 2606.30414v1 — [abs](https://arxiv.org/abs/2606.30414) · [pdf](https://arxiv.org/pdf/2606.30414) · [html](https://arxiv.org/html/2606.30414v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30414)
- authors: Alexis Jacq, Guillaume Couairon, Valentin De Bortoli, Quentin Berthet, Arnaud Doucet, Romuald Elie
- categories: cs.LG
**Abstract.** Distillation and Reinforcement Learning (RL) fine-tuning are the primary pillars of diffusion post-training. While traditionally studied in isolation, the interaction between these phases remains poorly understood, and in particular how fine-tuning impacts the generative quality of distilled models. We introduce Rewarded Moment Matching Distillation (RMMD), a novel framework that simultaneously distills diffusion models and maximizes a reward function. RMMD preserves the high-fidelity ``naturalness'' characteristic of advanced distillation (such as 8-step Moment Matching) by adapting the sampling loop for on-policy training and repurposing the distillation loss as a proxy for integral KL regularization. By evaluating the FID-Reward Pareto fronts on ImageNet, we demonstrate that RMMD achieves superior trade-offs compared to single-step baselines (DI++) and multi-step competitors (DRaFT, HyperNoise). Finally, we apply RMMD to GenCast, a state-of-the-art weather forecasting model, to distill it while optimizing the Continuous Ranked Probability Score (CRPS) metric. The resulting distilled model achieves a 7.5x speedup while outperforming the teacher model on 93% of target weather variables, and being better calibrated.
**Take.** Distillation and RL post-training usually fight each other — you compress the model and the reward fine-tune coarsens the samples. The neat move here is repurposing the distillation loss itself as the KL regularizer for the RL step, so the "stay natural" objective and the "compress" objective are the same term instead of two terms in tension. The headline I can't ignore isn't on ImageNet — it's GenCast: a distilled weather model that's 7.5× faster and beats its own teacher on 93% of variables while staying better calibrated. Distillation that improves the teacher is the interesting regime, and that it transfers from images to a real scientific forecaster is the part worth tracking.
## StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors
- arXiv: 2606.30545v1 — [abs](https://arxiv.org/abs/2606.30545) · [pdf](https://arxiv.org/pdf/2606.30545) · [html](https://arxiv.org/html/2606.30545v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30545)
- authors: Wenhao Yuan, Yiyuan Ge, Deli Cai
- categories: cs.CV
**Abstract.** 3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view settings due to insufficient geometric constraints. While recent methods introduce monocular depth priors to mitigate this, they inherently struggle with scale ambiguity and cross-view inconsistency, leading to defective geometry. In this paper, we propose StereoGS, a novel sparse-view 3DGS framework that integrates stereo priors to establish reliable binocular consistency. Unlike scale-agnostic monocular constraints, StereoGS introduces a Stereo Depth Regularization by constructing virtual stereo pairs during optimization and leveraging a foundation stereo model to enforce absolute scale and binocular-consistent structures. To further suppress overfitting and eliminate redundant primitives, we design a Gradient-Aware Opacity Decay strategy that dynamically penalizes Gaussians based on their relative opacity gradient magnitudes. Combined with a Consistency-Aware Dense Initialization using zero-shot multi-view depth estimation, StereoGS effectively anchors primitives to accurate scene surfaces. Extensive experiments on LLFF, DTU, Mip-NeRF360, and Blender datasets demonstrate that StereoGS achieves state-of-the-art performance in sparse-view settings without incurring any additional inference overhead.
**Take.** Monocular depth priors give you shape but not scale, and the scale ambiguity is exactly what wrecks sparse-view geometry. Manufacturing virtual stereo pairs during optimization and leaning on a foundation stereo model to pin absolute scale is a clean way to import the one constraint mono depth can't provide. The opacity-decay-by-gradient trick is the part I'd reuse — it's a principled way to kill redundant Gaussians instead of the usual opacity threshold heuristics. No inference overhead because all the extra machinery lives in training.
## Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors
- arXiv: 2606.30638v1 — [abs](https://arxiv.org/abs/2606.30638) · [pdf](https://arxiv.org/pdf/2606.30638) · [html](https://arxiv.org/html/2606.30638v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30638)
- authors: Jameel Hassan, Yasiru Ranasinghe, Vishal Patel
- categories: cs.CV
**Abstract.** 3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding has gained significant attention for real-world applications such as embodied AI. Recent methods achieve this by learning an instance feature attribute and assigning semantics by distilling high-dimensional Contrastive Language-Image Pretraining (CLIP) features directly into the scene representation. However, the instance grouping mechanisms of these methods either require a predefined number of instances or suffer from noise in their bottom-up grouping strategies. Furthermore, the reliance on CLIP restricts semantic understanding to simple noun phrases, preventing complex spatial reasoning and referential expression grounding. We present GaussDet, a method that circumvents the need for dense CLIP features by leveraging discrete, open-vocabulary 2D object detectors with referring expression capabilities. We learn instance features for individual Gaussians to decompose the scene into 3D instance groups. By rendering these groups and aggregating semantic votes from multi-view 2D detections, we generate a robust View-Aggregated Semantic Label Distribution (VASD) for each 3D instance. This view-aggregation strategy acts as a strong regularizer, attenuating spurious labels caused by low-quality instance grouping. Our approach enables a straightforward, zero-shot extension from simple language queries to complex referential grounding.
**Take.** Distilling dense CLIP features into every Gaussian always felt like the expensive, lossy way to do this — you pay for a high-dim feature per primitive and still only get noun-phrase semantics. Flipping it to render 3D instance groups and let off-the-shelf 2D detectors vote across views is the cheaper, sharper design: the multi-view aggregation is itself the regularizer that cleans up bad grouping. The payoff that sells it is referential grounding ("the mug behind the laptop") in a strict zero-shot setting, +16.7% mIoU — spatial reasoning CLIP-distillation can't do.
## VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
- arXiv: 2606.30645v1 — [abs](https://arxiv.org/abs/2606.30645) · [pdf](https://arxiv.org/pdf/2606.30645) · [html](https://arxiv.org/html/2606.30645v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30645)
- authors: Yen-Jen Wang, Jiaman Li, Sirui Chen, Takara E. Truong, Pei Xu, Pieter Abbeel, Rocky Duan, Koushil Sreenath, Angjoo Kanazawa, Carmelo Sferrazza, Guanya Shi, Karen Liu
- categories: cs.RO, cs.AI, cs.GR, eess.SY
**Abstract.** Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egocentric images, language commands, and robot-compatible kinematic trajectories, yet no existing data source provides this complete tuple at scale. We address this bottleneck by generating vision-language-kinematics (VLK) supervision synthetically in reconstructed scenes. Our pipeline leverages 3D Gaussian Splatting to reconstruct metric-scale indoor environments, synthesizes navigation and object-interaction trajectories using privileged scene information, and renders paired egocentric observations after the fact. We produce 48,000 paired trajectories with no human intervention and train a VLK policy that predicts short-horizon whole-body kinematic trajectories. A whole-body tracker converts these predictions into actions on the physical humanoid. We evaluate on the physical Unitree G1 performing navigation and single-object transport, demonstrating that synthesized interactions in reconstructed scenes provide effective supervision for sim-to-real perception-based humanoid loco-manipulation.
**Take.** The bottleneck for perception-driven humanoids is the data tuple nobody has: synchronized egocentric video + language + robot-feasible whole-body trajectories. The clever inversion is rendering the egocentric images *after the fact* — plan trajectories with privileged scene info in a 3DGS-reconstructed metric room, then render what the robot would have seen. 48k paired trajectories with zero human labeling, and it crosses sim-to-real onto a physical Unitree G1. This is the same realization the SLAM papers are circling from the other side: a good enough 3DGS reconstruction is now a data generator, not just a viewer. Heavyweight author list behind it.
## MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- arXiv: 2606.30406v1 — [abs](https://arxiv.org/abs/2606.30406) · [pdf](https://arxiv.org/pdf/2606.30406) · [html](https://arxiv.org/html/2606.30406v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30406)
- authors: Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang, Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, Jinhao Dong, Zhifang Sui, Fuli Luo
- categories: cs.CL, cs.LG
**Abstract.** Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains hard. Existing methods, such as Off-Policy Finetune and Mix-RL, are either inefficient or lose performance. In this work, we propose Multi-teacher On-Policy Distillation (MOPD), a post-training paradigm for combining the capabilities of multiple domain RL teachers: we first run per-domain specialised RL to obtain a set of domain teachers, then distill these teachers into the student on its own rollouts. This eliminates exposure bias and provides a dense optimization signal. On Qwen3-30B-A3B, MOPD outperforms Mix-RL, Cascade RL, Off-Policy Finetune, and Param-Merge baselines, inheriting nearly all of each teacher's capability. MOPD also enables parallel, independent development of domain teachers, removing the cross-domain coupling typical of multi-domain post-training. MOPD has been deployed in the post-training of MiMo-V2-Flash, an industrial-scale frontier model.
**Take.** The practical pain in multi-capability post-training is coupling: train math-RL and code-RL together and they interfere, so every capability has to move in lockstep. Specializing one RL teacher per domain and then distilling them into the student *on the student's own rollouts* decouples the org problem (teams ship teachers independently) from the model problem (one student inherits all of them). On-policy is what makes it work — distilling on student rollouts kills the exposure bias that off-policy fine-tune suffers. The credible bit is the deployment line: it's in MiMo-V2-Flash, not just a benchmark table.
## DOPD: Dual On-policy Distillation
- arXiv: 2606.30626v1 — [abs](https://arxiv.org/abs/2606.30626) · [pdf](https://arxiv.org/pdf/2606.30626) · [html](https://arxiv.org/html/2606.30626v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.30626)
- authors: Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, Qunzhong Wang, Yang Shi, Xiaobin Hu, Xiangyu Yue, Jiaqi Wang, Shuicheng Yan
- categories: cs.AI
**Abstract.** On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities.
**Take.** "Privilege illusion" is a sharp name for a real trap: if you feed the teacher privileged context, some of its behavior comes from information the student will never have, so the student is being asked to imitate something it structurally can't replicate. Separating the *transferable* capability gap from the *un-replicable* information-asymmetry gap — and routing each token's supervision by advantage — is the kind of distinction that only shows up once you take token-level supervision seriously. Pairs naturally with MOPD in today's batch: the field is clearly converging on on-policy distillation as the post-training workhorse, now arguing about how to route the signal.
---
# arXiv digest — 2026-06-28
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-28
> papers: 7
## ★ RoPEMover: Depth-Aware Object Relocation via Positional Embeddings
- arXiv: 2606.27332v1 — [abs](https://arxiv.org/abs/2606.27332) · [pdf](https://arxiv.org/pdf/2606.27332) · [html](https://arxiv.org/html/2606.27332v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27332)
- authors: Ipek Oztas, Duygu Ceylan, Aybars Bugra Aksoy, Aysegul Dundar
- categories: cs.CV
**Abstract.** Moving an object in a single image requires geometry-consistent spatial rearrangement, including handling occlusions, revealing previously unseen regions, and maintaining coherent shadows and reflections. Existing approaches are not well suited to this setting and often fail to preserve such scene-level consistency. We address this problem by introducing a geometry-aware object motion method that operates directly on the positional representations of diffusion transformers. Our key insight is that rotary positional embeddings (RoPE) define a structured spatial field that can be explicitly manipulated to induce controlled motion. We extend 2D RoPE into a depth-aware formulation that encodes 3D spatial structure, enabling consistent object displacement and scene-aware updates. Our model is trained using synthetic data combined with a small set of real images via parameter-efficient fine-tuning. Despite minimal real supervision, it preserves object identity under large spatial displacements, generates plausible content in newly revealed regions, and consistently updates scene-dependent effects such as shadows and illumination. Experimental results on standard object motion benchmarks demonstrate state-of-the-art performance across all evaluation metrics.
**Take.** The same realization RayPE had, aimed at editing instead of generation: RoPE isn't just index bookkeeping, it's a structured spatial field you can manipulate. They extend 2D RoPE to a depth-aware form, so moving an object in a single image drags its occlusions, disocclusions, shadows, and reflections along with it. Geometry-consistent relocation by editing the positional embeddings rather than the latents — trained mostly on synthetic data with a thin parameter-efficient real fine-tune. The depth-aware RoPE is the transferable trick.
## Proposal-Conditioned Latent Diffusion for Closed-Loop Traffic Scenario Generation
- arXiv: 2606.27123v1 — [abs](https://arxiv.org/abs/2606.27123) · [pdf](https://arxiv.org/pdf/2606.27123) · [html](https://arxiv.org/html/2606.27123v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27123)
- authors: Shubham Vaijanath Phoolari, Aleyna Kara, Christoph Lauer, Steven Peters
- categories: cs.RO, cs.CV
**Abstract.** Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllable throughout rollout. Prior diffusion-based approaches achieve strong realism, but their computational cost can hinder deployment in time-constrained replanning loops for autonomous vehicle planning and simulation. We present a diffusion-based scenario generation framework conditioned on instance-centric scene context and multimodal proposal priors, with optional test-time guidance for shaping safety-critical behaviors. A compact action-latent representation and proposal-based initialization improve sampling efficiency and reduce per-step runtime without retraining. Experiments on the Waymo Open Motion Dataset demonstrate a favorable balance among realism, safety, and controllability across diverse interactive scenarios, while showing that test-time guidance enables systematic trade-offs among competing objectives.
**Take.** Closed-loop traffic sim where the diffusion sampler's cost is the whole problem — you can't drop a slow denoiser inside a time-constrained replanning loop. A compact action-latent plus proposal-based initialization cuts per-step runtime without retraining, and optional test-time guidance shapes safety-critical behaviors on demand. The realism/safety/controllability balance on Waymo Open Motion is the right axis; I'd want the actual per-step latency before believing "deployable."
## Pseudo-Text-Conditioned 3D Grounding DINO for Organ Localization in Abdominal CT
- arXiv: 2606.27084v1 — [abs](https://arxiv.org/abs/2606.27084) · [pdf](https://arxiv.org/pdf/2606.27084) · [html](https://arxiv.org/html/2606.27084v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27084)
- authors: Siqi Chen, Han Gong, Keyi Hou, Jingxuan Yang, Sheethal Bhat, Andreas Maier
- categories: cs.CV, eess.IV
**Abstract.** Reliable organ localization in abdominal CT can provide spatial priors for downstream trauma analysis. We propose CT-3GDINO, a lightweight 3D detector that adapts a Grounding-DINO-style query-based architecture to fixed organ localization using frozen pseudo-text class tokens instead of a real text encoder. The model combines a Swin3D visual backbone, bidirectional feature enhancement, pseudo-text-guided query selection, and a cross-modality decoder to predict normalized 3D boxes for liver, spleen, left kidney, right kidney, and bowel. We train and evaluate on 193 matched RSNA/RATIC CT volumes with segmentation-derived boxes. The best multi-scale model, trained from scratch, achieves 0.5830 overall top-1 class-wise mAP over 3D IoU thresholds from 0.1 to 0.7, outperforming fixed- and trainable-backbone classification-pretrained variants with 0.5570 and 0.4657 mAP. Performance is strong for coarse localization, with 0.9649 AP at IoU 0.1, but remains limited for strict box alignment, with 0.1552 AP at IoU 0.7. These results establish CT-3GDINO as an open-source baseline for pseudo-text-conditioned 3D organ localization and motivate future work on localization-aware pretraining, richer multimodal conditioning, and injury-focused detection.
**Take.** Grounding-DINO for 3D boxes, but with frozen pseudo-text class tokens instead of a real text encoder — a clean simplification once your classes are a fixed set of organs. Swin3D backbone, bidirectional feature enhancement, and a cross-modality decoder hit 0.583 top-1 class-wise mAP over IoU 0.1-0.7 on 193 CT volumes, beating classification-pretrained backbones. The surprise is that training from scratch wins; the pseudo-text query trick is what I'd reuse for any fixed-vocabulary 3D detector.
## CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
- arXiv: 2606.27264v1 — [abs](https://arxiv.org/abs/2606.27264) · [pdf](https://arxiv.org/pdf/2606.27264) · [html](https://arxiv.org/html/2606.27264v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27264)
- authors: Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed, Muzammal Naseer, Salman Khan, Christoph Lippert
- categories: cs.CV
**Abstract.** Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text judged only by its final answer, making it hard to interpret and verify, especially in 3D radiology, where a diagnosis should be traceable to evidence in the scan. Existing chest CT question-answering datasets compound this by reducing expert radiology reports to answer-only pairs, dropping the reasoning that links findings to conclusions and omitting the patient history clinicians rely on. As a result, reasoning-capable 3D chest CT MLLMs remain out of reach, as neither the structured supervision needed to train them nor the protocol needed to verify their reasoning yet exists. We introduce CORTEX (Clinically Organized Reasoning and sTructured EXplanation), a structured reasoning benchmark for 3D chest CT. For each question, CORTEX restores the missing reasoning as a four-stage diagnostic trace mirroring a radiologist's workflow: task understanding, visual observation, diagnostic reasoning, and answer synthesis. We generate these traces using frontier large language models with broad medical and general-domain knowledge, then filter and verify them with a stage-level evaluation protocol combining automated rubric scoring with expert radiologist review. Crucially, both the reasoning structure and evaluation rubrics are designed in close collaboration with clinicians. Built on CT-RATE, a large, publicly available chest CT dataset without reasoning annotations, CORTEX comprises 76,177 validated reasoning traces across open-ended VQA, closed-ended VQA, and report generation, providing both the structured supervision and the stage-level evaluation protocol needed to build and evaluate trustworthy reasoning models for 3D chest CT. Our dataset and evaluation code will be made publicly available upon acceptance.
**Take.** The right complaint about medical MLLMs: a 3D CT diagnosis judged only on its final answer is unverifiable. CORTEX restores the four-stage radiologist trace — task understanding, visual observation, diagnostic reasoning, answer synthesis — as structured supervision, 76,177 traces validated by a stage-level rubric plus radiologist review. The value is the protocol, not just the dataset: it gives you something to grade the reasoning against, not only the answer.
## DanceOPD: On-Policy Generative Field Distillation
- arXiv: 2606.27377v1 — [abs](https://arxiv.org/abs/2606.27377) · [pdf](https://arxiv.org/pdf/2606.27377) · [html](https://arxiv.org/html/2606.27377v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27377)
- authors: Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua
- categories: cs.CV, cs.CL, cs.LG
**Abstract.** Modern image generation demands a single model that unifies diverse capabilities, including text-to-image (T2I), local editing, and global editing. However, these capabilities are rarely naturally aligned and often conflict. For instance, editing tends to degrade T2I performance, while global and local editing interfere with each other. Consequently, effectively composing these capabilities has become a central challenge for image generation model training. To tackle this, we introduce DanceOPD, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective. With each capability source defined as a velocity field over the shared flow state space, the student learns from fields queried on its own rollout states to compose expert capabilities. This formulation also absorbs operator-defined fields such as classifier-free guidance. Comprehensive experiments on T2I, editing, realism-field absorption, and CFG absorption show that our approach improves multi-capability composition, strengthening target capabilities while preserving anchor generation quality. We believe this work establishes a practical route for generative field distillation in flow-matching models.
**Take.** The unify-everything-in-one-image-model problem, stated honestly: text-to-image, local edit, and global edit actively fight each other. DanceOPD routes each sample to one capability "field" and distills the student on its own rollout states with a plain velocity MSE — composing expert velocity fields over a shared flow state. Tidy framing; the open question is whether on-policy querying actually keeps the capabilities from clobbering one another at scale.
## OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
- arXiv: 2606.27154v1 — [abs](https://arxiv.org/abs/2606.27154) · [pdf](https://arxiv.org/pdf/2606.27154) · [html](https://arxiv.org/html/2606.27154v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27154)
- authors: Aoyang Fang, Yifan Yang, Jin'ao Shang, Qisheng Lu, Junjielung Xu, Rui Wang, Songhan Zhang, Yuzhong Zhang, Boxi Yu, Pinjia He
- categories: cs.AI
**Abstract.** Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suffer from a fundamental gap: they label only the root cause, not the propagation path connecting it to the observed symptom, which largely simplifies the task to naive pattern matching. To support rigorous evaluation, we introduce PAVE, a step-wise labeling protocol that leverages known interventions from fault injection to reconstruct causal propagation paths. The mechanism is forward verification: reasoning from cause to effect rather than inferring backward from symptoms. Applying PAVE yields OpenRCA 2.0 (500 instances), the first cross-system RCA benchmark with step-wise causal annotations for LLM agents. Across 11 frontier LLMs, recovering the exact root-cause set succeeds in only 20.7% of cases on average. To locate where this difficulty lies, we relax the criterion and find what we call the ungrounded diagnosis: agents identify at least one correct root-cause service in 76.0% of cases, but ground that service in a verified causal propagation path to the observed symptom in only 61.5%. Outcome-only evaluation hides this failure mode; step-wise causal ground truth is the missing piece for trustworthy LLM-based RCA agents.
**Take.** Root-cause analysis is a real stress test for agents — long context, multi-step reasoning, tool use — and the honest headline is the number: across 11 frontier LLMs, the exact root-cause set is recovered only 20.7% of the time. The contribution is PAVE, which labels the causal propagation path from known fault injections (forward cause→effect) so the benchmark can't be solved by backward pattern-matching. Step-wise causal supervision is exactly what production RCA needs.
## Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds
- arXiv: 2606.26964v1 — [abs](https://arxiv.org/abs/2606.26964) · [pdf](https://arxiv.org/pdf/2606.26964) · [html](https://arxiv.org/html/2606.26964v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.26964)
- authors: Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang, Hailan Ma, Huadong Mo, Zhenhong Sun
- categories: cs.AI, cs.CV
**Abstract.** As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as Narrative-Grounded World Visual Attention, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose Look-Before-Move, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on StoryBlender, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.
**Take.** Perception as an action rather than a given: the camera decides what evidence to acquire before it moves. Separating a "Semantic Observation Contract" from motion execution is a sensible split for embodied 3D agents, and more interesting as a framing of active perception than for the story-world demo. I'd read it for how the contract grounds against physical 3D constraints, not the narrative wrapper.
---
# arXiv digest — 2026-06-27
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-27
> papers: 8
## ★ SatSplatDiff: Geometry-preserving generative refinement for high-fidelity satellite Gaussian Splatting
- arXiv: 2606.27223v1 — [abs](https://arxiv.org/abs/2606.27223) · [pdf](https://arxiv.org/pdf/2606.27223) · [html](https://arxiv.org/html/2606.27223v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27223)
- authors: Jiyong Kim, Shuang Song, Ronjgun Qin
- categories: cs.CV
**Abstract.** Gaussian Splatting has been recently explored for satellite 3D reconstruction, demonstrating flexibility and efficiency in representing radiometrically diverse satellite scenes. However, the limited top viewpoint of satellite imagery results in insufficient supervision on building facades, leaving surface holes and degraded visual fidelity. Generative refinement, which leverages pretrained generative priors to iteratively refine and update the rendered images used as supervision targets, has recently been investigated to improve the visual fidelity of Gaussian-rendered images. However, since these models refine each view independently, the resulting images can generate hallucinations and break photo-consistency, leading to geometric degradation. To address these limitations, we propose SatSplatDiff, which aims to minimize geometric degradation prevalent in generative refinement. Building on photogrammetric DSM initialization and 2DGS-based shadow casting established in our prior work SatSplat, we first introduce monocular depth supervision and multi-scale geometric refinement to establish a geometrically accurate and well-regularized surface representation. We then apply shadow-guided generative refinement, where geometrically calculated shadow maps guide the Gaussians to maintain consistency with the underlying geometry, improving visual fidelity while reducing geometric degradation. Extensive evaluations on the IARPA2016 and DFC2019 datasets demonstrate state-of-the-art performance, reducing geometric MAE by up to 18% and improving visual fidelity (FID-CLIP) by 28-45% over existing baselines. Our method delivers up to 5x resolution enhancement with minimal hallucination and sensor-consistent appearance, demonstrating seamless cross-tile consistency and strong scalability for large-scale reconstruction. Source code is available at https://github.com/GDAOSU/SatSplatDiff
**Take.** Satellite splatting starves on building facades — imagery is top-down, so the sides get almost no supervision and you get holes. Bolting a generative prior on per-view hallucinates and breaks photo-consistency; the fix here is to let geometrically-computed shadow maps steer the refinement so it can't drift off the surface. 18% lower geometric MAE and 28-45% better FID-CLIP is a real lift. The shadow-as-geometric-anchor trick is the part I'd steal.
## UAV-MapFusion: RTK-Aligned Uncertainty-Aware Coarse-to-Fine Multi-Session UAV Mapping
- arXiv: 2606.26928v1 — [abs](https://arxiv.org/abs/2606.26928) · [pdf](https://arxiv.org/pdf/2606.26928) · [html](https://arxiv.org/html/2606.26928v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.26928)
- authors: Feng Pan, Chunran Zheng, Bing Xue, Yukang Cui, Jiayu Wen, Zhiyu Chen, Wei Wang
- categories: cs.RO, cs.SI
**Abstract.** Large-scale point cloud maps are essential for robotics and spatial intelligence tasks. UAVs provide an efficient means for large-scale map acquisition; however, due to limited flight endurance and onboard storage, mapping a large-scale scene within a single flight remains difficult. Existing multi-session map merging methods can extend the mapping range, yet in UAV scenarios they still struggle to simultaneously suppress long-range drift and preserve local geometric accuracy. To address this issue, an uncertainty-aware multi-session point cloud map merging and coarse-to-fine optimization system is proposed. The proposed method first performs initial multi-session map merging based on a scene graph, and then incorporates RTK observations through an RTK spatiotemporal alignment module, where temporal offsets are estimated using Dynamic Time Warping (DTW), and continuous RTK constraints are recovered using Multi-Output Gaussian Processes (MOGP) under incomplete sampling and frame dropouts. On this basis, a unified uncertainty-aware factor graph is constructed, and local geometric accuracy is further improved through iterative plane-factor refinement. Experiments on real-world datasets validate the effectiveness and robustness of the proposed method. To facilitate further research and development in the community, our code and dataset will be publicly released.
**Take.** The unglamorous production problem stated plainly: one flight can't map a site, so you stitch sessions and fight long-range drift without smearing local geometry. DTW for the RTK temporal offset plus multi-output GPs to recover dropped frames is a sane way to treat RTK as a soft constraint instead of trusting it blindly. The uncertainty-aware factor graph over the merge is the right substrate; I'd push on how the plane-factor refinement holds up across session seams.
## OctoSense: Self-Supervised Learning for Multimodal Robot Perception
- arXiv: 2606.27317v1 — [abs](https://arxiv.org/abs/2606.27317) · [pdf](https://arxiv.org/pdf/2606.27317) · [html](https://arxiv.org/html/2606.27317v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27317)
- authors: Anthony Bisulco, Jeremy Wang, Kostas Daniilidis, Randall Balestriero, Pratik Chaudhari
- categories: cs.CV, cs.RO
**Abstract.** We present OctoSense, an open-source sensor platform with stereo RGB and event cameras, LiDAR, a thermal camera, an inertial measurement unit, RTK-corrected global positioning system, and proprioception (CAN bus data from a car, and joint angles for a quadruped robot). The eponymous OctoSense dataset contains 59 hours of time-synchronized driving data across different types of environments at different times of the day, including situations with highly degraded sensors. We demonstrate multi-modal self-supervised learning using such real-world robotics data, where sensors have different representations, frequencies, latencies and noise. Our approach, a "late-fusion" masked autoencoder, (i) uses modality-specific tokenizers to account for different spatiotemporal characteristics of these sensors, and (ii) caches modality-specific tokens at inference time to process new measurements as they come. This architecture (i) is fast (6.68 ms and 112 ms on NVIDIA 5090 and Orin NX respectively, to compute the representation), (ii) performs better than existing image-only foundation models on tasks such as estimation of optical flow, depth, semantic segmentation, and ego-motion (translation, rotation, and steering angle), and (iii) predicts robustly at nighttime or in situations where sensory data is degraded. See our project page for links to the dataset, code, and supplementary videos: https://abisulco.com/octosense/.
**Take.** Sensor fusion as a late-fusion masked autoencoder with per-modality tokenizers — and the part that matters for shipping, it caches modality tokens at inference so new measurements stream in instead of re-encoding everything. 6.68 ms on a 5090, 112 ms on an Orin NX, and it degrades gracefully at night or with a dead sensor. That edge-latency number is what separates a deployable perception stack from a benchmark entry.
## PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views
- arXiv: 2606.27071v1 — [abs](https://arxiv.org/abs/2606.27071) · [pdf](https://arxiv.org/pdf/2606.27071) · [html](https://arxiv.org/html/2606.27071v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27071)
- authors: Zhisong Xu, Takeshi Oishi
- categories: cs.CV
**Abstract.** Panoramic sensing offers wide field-of-view coverage, yet 3D reconstruction from sparse panoramas remains challenging under rotation-dominant, weak-parallax motion. In such regimes, SfM/SLAM initialization is often ill-conditioned and unreliable. We present PanoImager, an SfM-free framework that combines feed-forward pose/depth priors, geometry-conditioned diffusion view completion, and depth-guided 3DGS optimization. Given only a few panoramic images, PanoImager decomposes them into local perspective views, synthesizes auxiliary observations to enrich sparse evidence, and stabilizes Gaussian optimization for improved cross-view consistency. Experiments on multiple benchmarks show improved stability under extreme sparsity, suggesting PanoImager as an offline/background component for map refinement when SfM/SLAM fails to initialize.
**Take.** When motion is rotation-dominant with weak parallax, SfM/SLAM initialization is ill-conditioned and just falls over — so this is SfM-free on purpose. Decompose panoramas into local perspective views, synthesize auxiliary views with a geometry-conditioned diffusion model, then stabilize 3DGS with depth guidance. Positioned honestly as an offline map-refinement fallback for exactly the regime where the online pipeline can't get off the ground.
## ★ RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation
- arXiv: 2606.27345v1 — [abs](https://arxiv.org/abs/2606.27345) · [pdf](https://arxiv.org/pdf/2606.27345) · [html](https://arxiv.org/html/2606.27345v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27345)
- authors: Minghao Yin, Jiahao Lu, Wenbo Hu, Wang Zhao, Shan Ying, Kai Han
- categories: cs.CV
**Abstract.** Modern video diffusion transformers position their tokens through RoPE on the (u,v,t) axes -- a description of the camera's sampling grid that says nothing about the 3D structure of the scene. We observe that the geometric relation between two camera rays is captured by the Plucker reciprocal product, which is bilinear in the two rays -- the same algebraic form as the dot product in Transformer attention. Building on this analogy, we propose RayPE, a positional-encoding extension that injects per-token 6D Plucker coordinates additively into the queries and keys of self-attention, with a query/key flip arrangement under which the symmetric identity configuration coincides exactly with the reciprocal product. The injection is additive, the resulting attention score decomposes into a content term, a geometry term, and two content and geometry cross-terms -- all of which our experiments find individually necessary. To make the encoding stable across video data with heterogeneous camera-translation scales (SfM, deep SLAM, metric), we further decouple ray direction from moment magnitude, gate the encoding by a learned function of the log-magnitude, and apply RMSNorm to align it with the QKNorm-normalized content branch. The full module adds less than 0.1% parameters to a pretrained video DiT, is zero-initialized to start from the pretrained weights, and improves camera controllability, cross-frame 3D consistency, and overall video quality on a four-dataset training mixture.
**Take.** The clean idea of the batch. Video DiTs position tokens with RoPE over (u,v,t) — the camera's sampling grid, which says nothing about 3D structure. They notice the Plücker reciprocal product between two rays is bilinear in the rays, the same algebraic form as the attention dot product, and inject 6D Plücker coordinates additively into Q and K. Under 0.1% extra params, zero-initialized so it starts exactly from the pretrained model. Geometry-as-positional-encoding is a genuinely nice unification.
## Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
- arXiv: 2606.26938v1 — [abs](https://arxiv.org/abs/2606.26938) · [pdf](https://arxiv.org/pdf/2606.26938) · [html](https://arxiv.org/html/2606.26938v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.26938)
- authors: Haoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang, Yu Liu, Changxin Gao, Nong Sang
- categories: cs.CV
**Abstract.** Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling diffusion models in visual generation. Recent advancements have focused on adaptively allocating computational resources across diverse tokens to improve efficiency and performance. However, we identify a routing assignment problem in existing diffusion MoE frameworks: the router fails to accurately allocate more computational resources to salient tokens. Our analysis attributes this failure to the router's reliance on noise-corrupted latent features throughout the denoising process. Such stochastic noise obscures the critical structural and textural information, thereby preventing the router from effectively distinguishing salient tokens. To address this, we propose SharpMoE, a post-training framework with a saliency-harnessing accurate routing mechanism, which utilizes clean latent features as a noise-free guidance signal for routing. By bypassing the noise-distorted inputs, SharpMoE provides the router with clear saliency guidance, enabling the identification of salient tokens even in high-noise stages. Furthermore, we introduce a trajectory routing loss to constrain the compute allocation throughout the multi-step denoising trajectory, ensuring precise resource allocation along the generation rollout. Extensive experiments demonstrate that SharpMoE serves as a versatile, plug-and-play solution that further enhances the pretrained, converged MoE models, achieving state-of-the-art performance in visual generation.
**Take.** A specific, believable failure: a diffusion MoE router reads noise-corrupted latents during denoising, so it can't tell which tokens are salient and mis-allocates compute. The fix routes on the clean latent as a noise-free guidance signal, with a trajectory loss constraining allocation along the denoising rollout. Plug-and-play on an already-converged MoE — no retrain — is the appealing part.
## PhysiFormer: Learning to Simulate Mechanics in World Space
- arXiv: 2606.27364v1 — [abs](https://arxiv.org/abs/2606.27364) · [pdf](https://arxiv.org/pdf/2606.27364) · [html](https://arxiv.org/html/2606.27364v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27364)
- authors: Yiming Chen, Yushi Lan, Andrea Vedaldi
- categories: cs.CV
**Abstract.** We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixel space, PhysiFormer represents objects as 3D meshes expressed in world coordinates. Given the initial vertex positions and velocities, as well as object material type, rigid or elastic, the model samples future vertex trajectories. While related neural physics approaches build on ad-hoc latent spaces or explicitly enforce rigidity and causality, PhysiFormer shows that excellent results can be obtained without any such inductive biases, by casting vertex trajectory prediction as a single denoising diffusion process directly in world coordinates. The probabilistic formulation captures uncertainty in the learned dynamics, enabling diverse plausible futures from initial conditions, making this framework potentially useful for applications with unobserved uncertainty. The model features attention factorised over time, space, and objects for efficiency, enabling permutation-invariant multi-object reasoning without needing explicit object encoding. Trained on over 100k simulated trajectories, PhysiFormer generates rigid and elastic mechanics, and generalises to mixed-material settings, unseen real-world geometries, and larger object counts. It substantially outperforms autoregressive baselines in trajectory accuracy, rigidity preservation, and momentum-based physical consistency. Our results position coordinate-space diffusion as a promising step toward view-invariant, geometry-aware world modelling for robotics, graphics, and physical design. Visualisations, code, and models are available at https://yimingc9.github.io/physiformer.
**Take.** Predict vertex trajectories with a single denoising diffusion directly in world coordinates — no view-dependent pixel space, no hard-coded rigidity or causality — with attention factorized over time, space, and objects. The bet is that you recover rigid and elastic mechanics without baking the physics in, and the probabilistic head gives you diverse plausible futures. World-space rather than pixel-space is the right call for anything feeding robotics or design.
## Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN
- arXiv: 2606.27305v1 — [abs](https://arxiv.org/abs/2606.27305) · [pdf](https://arxiv.org/pdf/2606.27305) · [html](https://arxiv.org/html/2606.27305v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.27305)
- authors: Archer Moore, Mingming Gong, Liam Hodgkinson
- categories: cs.CV
**Abstract.** Reinforcement learning from human feedback (RLHF) for 3D generation is now established across a number of works, but most existing pipelines optimise explicit surface representations, often by converting radiance fields into meshes and training heavily on surface-supervised data. We instead fine-tune a pretrained 3D-aware generative model directly from a learned reward over radiance-field density ($σ$) values, with no externally supplied mesh or shape prior. The reward model requires no pretraining, trains easily on a small set of preference samples, and yields robust improvement in 3D geometry. Working on an unconditional 3D-aware face GAN (EG3D), our reward reads the continuous 3D density field of the neural radiance field (NeRF) directly and supplies a geometry-only learning signal, requiring neither text conditioning, mesh extraction, nor multi-view rendering. A density-consistency constraint keeps the 2D appearance qualitatively similar while the geometry is reshaped, at a measurable but bounded distributional cost (FID-50k rises from 4.09 to 6.66): the fine-tuned generator, trained from the preferences of a single annotator as a proof of concept, produces face geometries preferred by users in 74.4% of pairwise comparisons.
**Take.** RLHF that reads the NeRF density field directly and supplies a geometry-only reward — no mesh extraction, no multi-view render, no text conditioning — with a reward model that trains on a handful of preference pairs. FID rises 4.09 to 6.66 (the bounded appearance cost is stated honestly) while geometry is preferred in 74.4% of pairwise comparisons. Optimizing the continuous density field instead of a surface proxy is the move worth noting.
---
# arXiv digest — 2026-06-23
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-23
> papers: 6
## ★ Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- arXiv: 2606.23688v1 — [abs](https://arxiv.org/abs/2606.23688) · [pdf](https://arxiv.org/pdf/2606.23688) · [html](https://arxiv.org/html/2606.23688v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.23688)
- authors: Yehonathan Litman, Xiaoxuan Ma, Manan Shah, Nicolas Ugrinovic, Kris Kitani, Fernando De la Torre, Shubham Tulsiani
- categories: cs.CV
**Abstract.** Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observations with data-driven priors over geometry and appearance. Prior approaches either learn to directly predict 4D representations from visual input or initialize a 3D representation that is subsequently deformed and refined based on video evidence. However, the former are constrained by the scarcity of 4D training data, while the latter leverage priors only for the initial reconstruction and rely solely on video supervision thereafter; neither handles complex in-the-wild scenarios with large deformations and occlusions well. We present Lift4D, a test-time optimization framework that addresses both limitations. First, we adapt an existing single-view 3D reconstruction model to yield temporally consistent per-frame predictions via causal latent conditioning, providing a coherent initialization for a deformable 3D Gaussian Splatting representation. We then ``sculpt'' this representation to match the input video through an occlusion-aware optimization that faithfully recovers visible surface details while completing unobserved regions using a view-conditioned diffusion prior. We demonstrate that Lift4D clearly improves over prior 4D reconstruction methods, particularly on challenging in-the-wild sequences with severe occlusions and non-rigid motion.
**Take.** The split that matters: lean on the data-driven prior for the initialization *and* throughout the deformation, not just the seed. Most 4D-from-monocular pipelines trust the prior once and then trust the video — which is exactly when in-the-wild noise wins. I'd stress it on fast non-rigid motion, where monocular depth is least reliable and the refinement has to carry the most weight.
## MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing
- arXiv: 2606.23455v1 — [abs](https://arxiv.org/abs/2606.23455) · [pdf](https://arxiv.org/pdf/2606.23455) · [html](https://arxiv.org/html/2606.23455v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.23455)
- authors: Zesong Yang, Yuanhang Lei, Liyuan Cui, Yihang Chen, Jiaer Huang, Boming Zhao, Peter Yichen Chen, Hujun Bao, Zhaopeng Cui
- categories: cs.CV
**Abstract.** Recent advances integrate physically grounded Newtonian dynamics with neural rendering frameworks, narrowing the gap between photorealistic scene reconstruction and physics-based animation. However, existing approaches focus on mechanically driven dynamics while neglecting temperature, a fundamental yet invisible physical factor underlying phenomena such as melting, solidification, and other thermomechanical processes. In this paper, we propose MeGAS, a novel framework that incorporates thermomechanical phase-change dynamics into 3D Gaussian Splatting (3DGS). Specifically, we propose a new thermomechanical dynamic Gaussian Splatting representation that augments 3DGS with temperature attributes and employs a heat advection-diffusion solver with MPM dynamics incorporating phase transitions, enabling physically plausible and visually realistic synthesis of thermophysical phenomena. Furthermore, a new topology-adaptive Gaussian rendering strategy is proposed to mitigate cracking and floaters under extreme deformation. Extensive experiments demonstrate that MeGAS produces physically consistent thermomechanical behavior while maintaining high-fidelity photorealistic rendering, advancing toward physics-integrated world models.
**Take.** Temperature as the missing state variable in physics-grounded splatting. Melting and solidification need a thermal field, not just Newtonian forces — bolting one on is obvious only in hindsight. The real question is whether the phase-change coupling is physically calibrated or merely visually plausible; the editing demo will look great either way.
## Lightweight Neural Framework for Robust 3D Volume and Surface Estimation from Multi-View Images
- arXiv: 2606.23653v1 — [abs](https://arxiv.org/abs/2606.23653) · [pdf](https://arxiv.org/pdf/2606.23653) · [html](https://arxiv.org/html/2606.23653v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.23653)
- authors: Diego E. Farchione, Ramzi Idoughi, Peter Wonka
- categories: cs.CV
**Abstract.** Accurate volume and surface area estimation is critical for diverse applications, from marine ecology to medical diagnostics. However, existing methods often suffer from high computational costs and poor performance with sparse and noisy data. We propose a fully feed-forward framework that regresses scale-normalized volume and surface area and their associated uncertainties directly from multi-view images. By fusing 3D point cloud reconstructions with view-aligned 2D features through a graph-based decoder, our model bypasses iterative optimization, ensuring exceptional scalability and rapid inference. Experimental results demonstrate that our approach outperforms state-of-the-art methods, particularly when operating with a low number of input images. Validated across coral monitoring, dietary analysis, and anthropometry, our proposed framework provides a robust, adaptable solution for quantitative shape analysis. This architecture provides a high-speed, scalable alternative for precise geometric estimation from visual data, maintaining high performance even in resource-constrained or sparse-view scenarios.
**Take.** Feed-forward, scale-normalized volume and surface area *with uncertainty* from multi-view images — the uncertainty head is what makes this usable in an actual measurement loop rather than a demo. Fusing point-cloud reconstructions with view-aligned 2D features through a graph decoder is a sane layout; I'd go straight to the calibration of those error bars on sparse, noisy inputs.
## dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models
- arXiv: 2606.23623v1 — [abs](https://arxiv.org/abs/2606.23623) · [pdf](https://arxiv.org/pdf/2606.23623) · [html](https://arxiv.org/html/2606.23623v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.23623)
- authors: Yuhao Wu, Yitian Liu, Weijie Shen, Mishuo Han, Wenjie Xu, Haotian Liang, Zhongshan Liu, Yinan Mao, Lei Xu, Xinping Guan, Ru Ying, Ran Zheng, Wei Sui, Xiaokang Yang, Wenbo Ding, Yao Mu
- categories: cs.RO
**Abstract.** Vision-Language-Action (VLA) models have established a powerful paradigm for generalist robotic manipulation by grounding control into the semantic reasoning of VLMs. Prevailing architectures typically model actions continuously via diffusion or flow processes, or discretely through either autoregressive generation or parallel decoding. Recently, Discrete Diffusion VLAs (dVLAs) have emerged as a distinct alternative, unifying vision, language, and action into a single discrete token space via masked generative modeling. While combining iterative refinement with unified representations, its training has thus far been restricted to Supervised Fine-Tuning (SFT), leaving the potential of Reinforcement Learning (RL) for further policy refinement largely unexplored. A fundamental challenge in RL for dVLAs is that the marginal probability of the final action generated by dVLAs remains intractable. To solve this problem, we propose \textbf{dVLA-RL}, shifting the learning objective from the marginal action probability to the joint probability of the sampled generation path. Specifically, by modeling the denoising process as a Markov Decision Process (MDP), we mathematically formulate this path probability as a product of step-wise transitions. This trajectory-level objective provides a unified formulation that natively accommodates variable denoising steps. Leveraging this intrinsic fexibility, we introduce a unified step scheduling approach for complex multi-task learning, tailoring denoising steps to specific task complexities to maximize both success rates and computational effciency. Extensive evaluations demonstrate that our approach achieves a success rate of \textbf{99.7\%} on LIBERO. Furthermore, it establishes strong VLA-based results on RoboTwin 2.0 by delivering a \textbf{30.6\%} improvement over the SFT baseline, remaining competitive with strong World-Action Model baselines.
**Take.** RL applied directly over the denoising trajectory of a discrete-diffusion VLA — rewarding the unmasking path, not just the final action tokens. For parallel-decoding policies that's the right surface to optimize. The open question is reward density and stability across masked generative rollouts, where credit assignment over the trajectory gets slippery.
## Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models
- arXiv: 2606.23567v1 — [abs](https://arxiv.org/abs/2606.23567) · [pdf](https://arxiv.org/pdf/2606.23567) · [html](https://arxiv.org/html/2606.23567v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.23567)
- authors: Jiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen, Furong Huang
- categories: cs.LG, cs.AI
**Abstract.** Masked diffusion language models decode by iteratively unmasking tokens, where the unmasking order defines an "order of thought" that strongly influences generation quality yet is typically chosen heuristically. We derive a tractable upper bound on the sequential decoding mismatch, measured by the Kullback-Leibler divergence and expressed in terms of the model's pathwise log-likelihood, with tightness under sufficient model expressivity. This bound induces a dense self-aware reward over ordered trajectories, casting order selection as a principled policy optimization problem with a frozen denoiser. We instantiate this idea as Self-Aware Scheduling (SAS), which learns a lightweight order policy using Group Relative Policy Optimization and applies seamlessly to both any-order and semi-autoregressive decoding. On Sudoku with 1B MDM, SAS improves puzzle accuracy from 82.0% (best heuristic schedule) to 91.8%, and reaches 97.5% with second-stage fine-tuning along learned trajectories. On mathematical reasoning with LLaDA-8B, SAS improves pass@1 on GSM8K from 64% to 76% and on MBPP from 39.5% to 41%, consistently matching or exceeding heuristic schedules across generation lengths and block sizes. Project page: https://jimmyxu123.github.io/SAS
**Take.** A tractable KL bound on decoding-order mismatch, turned into a dense reward over unmasking trajectories — replacing the heuristic 'which token to unmask next' with a learned order of thought. The theory-to-reward move is clean, and it's the kind of change that lifts diffusion-LM quality without touching the backbone.
## HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory
- arXiv: 2606.23565v1 — [abs](https://arxiv.org/abs/2606.23565) · [pdf](https://arxiv.org/pdf/2606.23565) · [html](https://arxiv.org/html/2606.23565v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.23565)
- authors: Xiaolin Zhou, Liu Liu, Tingyang Xiao, Wei Feng, Fa Fu, Xinrui Meng, Xinjie Wang, Jialiang Han, Boyang Yu, Yun Du, Wei Sui, Zhizhong Su
- categories: cs.RO, cs.CV
**Abstract.** LLM agents follow a practical execution loop in digital environments: they reason over structured states, invoke tools, inspect feedback, and revise actions. Extending this loop to physical robots is difficult because physical execution is continuous, embodiment-dependent, uncertain, and constrained by safety. Existing embodied-AI systems have advanced manipulation, spatial understanding, navigation, and humanoid control, but these capabilities often remain specialized modules or loosely coupled decision loops. In this work, we introduce HoloAgent-0, a unified embodied agent framework for real-world robot deployment. Embodied AgentOS converts language instructions into executable skill graphs, schedules robot resources, monitors execution, and triggers clarification or re-planning from runtime feedback. HoloAgent-0 organizes heterogeneous robot models and controllers through three coupled layers: Embodied AgentOS for closed-loop execution, 3D spatial memory for physical world grounding, and embodied skills for robot action. We deploy HoloAgent-0 on real hardware and evaluate its spatial memory, long-horizon navigation, and closed-loop execution across motion generation, object search, cross-robot coordination, and mobile manipulation.
**Take.** Porting the digital-agent loop — reason, call a tool, inspect feedback, revise — onto robots, with a persistent 3D spatial memory as the shared state. The hard part is always the continuous, embodiment-dependent, safety-constrained execution gap; a unified 3D memory is a reasonable substrate for it. Read it for how the loop grounds against uncertainty, not for the framework diagram.
---
# arXiv digest — 2026-06-10
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-10
> papers: 6
Diffusion-and-distillation heavy today, with a 3D thread running through it. The standout is
Mean Flow Distillation — single-step flow-matching generation framed as suppressing
high-frequency optimization noise, with 4D occupancy forecasting as the test case that
matters for perception loops. Also worth a look: P3D-Bench, which scores parametric 3D
*programs* instead of meshes and confirms MLLMs still miss exact geometry, and WorldOlympiad's
Gaussian-splatting geometry track that turns "looks 3D" into a measurable number.
## ★ Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models
- arXiv: 2606.11155v1 — [abs](https://arxiv.org/abs/2606.11155) · [pdf](https://arxiv.org/pdf/2606.11155) · [html](https://arxiv.org/html/2606.11155v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.11155)
- authors: An Zhao, Shengyuan Zhang, Zhongjian Sun, Yixiang Zhou, Zejian Li, Ling Yang, Tianrun Chen, Lingyun Sun
- categories: cs.CV
**Abstract.** Flow Matching models perform well across generative tasks, but their ODE-based iterative sampling makes inference expensive and rules out real-time use. Existing distillation borrows from diffusion score matching, ignores the geometric structure of flows, and suffers training instability, high variance, and quality loss. Mean Flow Distillation (MFD) is a distillation framework built for flow matching: the authors show it acts as a temporal low-pass filter that suppresses the high-frequency optimization noise of variational score distillation while keeping global trajectory consistency, and prove a Mean Flow Matching Theorem — matching expected average velocities is sufficient for strict distribution alignment. On 4D occupancy forecasting and text-to-image, MFD reaches SOTA with high-fidelity single-step generation.
**Take.** The framing I like: treat the distillation instability as a signal-processing problem and show VSD is just leaking high-frequency noise into the student. The averaged-velocity target is the kind of move that's obvious only after someone proves it's sufficient. 4D occupancy forecasting in one step is the result worth poking at — that's the regime where iterative sampling actually kills you in a perception loop.
## P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning
- arXiv: 2606.11152v1 — [abs](https://arxiv.org/abs/2606.11152) · [pdf](https://arxiv.org/pdf/2606.11152) · [html](https://arxiv.org/html/2606.11152v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.11152)
- authors: Yikang Yang, Zhanpeng Hu, Youtian Lin, Mengqi Zhou, Jingxi Xu, Feihu Zhang, Jiaheng Liu, Yao Yao
- categories: cs.CV
**Abstract.** MLLMs can write code that drives 3D modeling, which opens a path to 3D generation that leans on their priors and reasoning. But most benchmarks score meshes, not programs. P3D-Bench scores parametric 3D programs, which expose explicit dimensions, construction operations, and part relations — revealing whether a model recovers a design's structure, not just its look. It covers Text-to-3D, Image-to-3D, and Assembly-3D, grading executability, geometric fidelity, topology, text-grounded constraints, multiview alignment, and part-level structure across 400 text, 400 image, and 203 annotated assembly cases. Findings: assemblies are hardest; models recover global shape and identity but miss precise parametric geometry; and part-level modeling stays weak.
**Take.** This benchmarks the thing I actually care about in CAD/BIM work — can the model produce a parametric program with the right dimensions and part relations, not a pretty mesh that's geometrically wrong. The honest result is the useful one: frontier models get the silhouette and semantics but fluff the exact geometry and assembly structure. That gap is exactly where these pipelines break in practice.
## WorldOlympiad: Can Your World Model Survive a Triathlon?
- arXiv: 2606.11129v1 — [abs](https://arxiv.org/abs/2606.11129) · [pdf](https://arxiv.org/pdf/2606.11129) · [html](https://arxiv.org/html/2606.11129v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.11129)
- authors: Yuke Zhao, Wangbo Zhao, Weijie Wang, Zeyu Zhang, Dakai An, Akide Liu, Yinghao Yu, Jiasheng Tang, Fan Wang, Wei Wang, Bohan Zhuang
- categories: cs.CV
**Abstract.** WorldOlympiad diagnoses video-based world models along physical faithfulness, geometric consistency, and interaction fidelity, instead of the usual visual-quality and short-horizon temporal checks. The physical track uses object segmentation plus an MLLM judge to test mechanics, thermal, and material rules. The geometry track reconstructs generated videos with Gaussian splatting and scores structural consistency, cross-view coherence, and camera-trajectory alignment. The interaction track checks whether rollouts follow complex action prompts and stay coherent across consecutive chunks, spanning gaming, robotics, and real-world video. Experiments on SOTA models show large gaps in physical reasoning, 3D consistency, and long-horizon interaction.
**Take.** The smart instrumentation is the geometry track — reconstruct the generated video with Gaussian splatting and measure whether the implied 3D is actually consistent across views. That turns "looks 3D" into a number. The takeaway is unsurprising but worth having stated: today's world models render well and reason about physics badly.
## Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
- arXiv: 2606.11180v1 — [abs](https://arxiv.org/abs/2606.11180) · [pdf](https://arxiv.org/pdf/2606.11180) · [html](https://arxiv.org/html/2606.11180v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.11180)
- authors: Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee, Joungbin Lee, Siyoon Jin, Heeseong Shin, Jung Yi, Yunjin Park, Chulmin Park, Seungryong Kim
- categories: cs.CV
**Abstract.** Diffusion lip-sync models have strong audio-visual alignment but full bidirectional attention and many denoising steps make them too slow for real time. Lip Forcing distills a 14B audio-conditioned bidirectional video diffusion teacher into causal students that generate each chunk in two denoising steps with no inference-time CFG. A trajectory analysis reveals a CFG fidelity-versus-sync tradeoff — no-CFG predictions favor reference fidelity, CFG-guided ones favor sync within a mid-trajectory band — which the method turns into three components: Sync-Window DMD, a two-step schedule, and a SyncNet reward. The 1.3B student streams at 31 FPS, 17.6x faster than its bidirectional counterpart; the 14B student runs 39.8x faster than its teacher at comparable fidelity, with sub-millisecond time-to-first-frame.
**Take.** The interesting engineering is converting a trajectory observation — CFG helps sync only inside a mid-trajectory band — into a scheduled, windowed distillation instead of a global knob. Two steps, no inference-time CFG, causal students: that's a real path from a 14B bidirectional teacher to something that streams. Sub-millisecond TTFF is the number that decides whether this is usable live.
## Resilient Navigation for Autonomous Farm Robots by Leveraging Jerk-Augmented Models with IMU-Only Disturbance Rejection
- arXiv: 2606.10971v1 — [abs](https://arxiv.org/abs/2606.10971) · [pdf](https://arxiv.org/pdf/2606.10971) · [html](https://arxiv.org/html/2606.10971v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.10971)
- authors: Batu Candan, Mohammed Atallah, Simone Servadio, Saeed Arabi
- categories: cs.RO, eess.SY
**Abstract.** State estimation for off-road agricultural robots is degraded by sensor outages (GNSS/LiDAR/visual) and high-frequency vibration. This work uses a jerk-augmented Extended Kalman Filter paired with a Multiple Tuning Factor adaptation that adjusts the measurement covariance in real time instead of assuming constant measurement noise, letting the filter handle sudden disturbances and outliers. Evaluated on real-world data from a Salin247 robot, jerk-augmentation plus MTF adaptation cuts 3D position RMSE versus baseline EKF models and gives better dead-reckoning when sensors drop out.
**Take.** Not flashy, and that's the point — when LiDAR and GNSS drop out, the thing that saves you is a well-tuned filter, not a bigger network. Adding jerk to the state and adapting the measurement covariance online is the pragmatic fix for vibration-heavy platforms. I'd want to see how the MTF gains were chosen before trusting it off the test field, but the dead-reckoning story is the right thing to optimize.
## A Distributed Multi-UGV Exploration Framework With Loop-Aware Planning and Descriptor-Aided Localization in Resource-Limited Environments
- arXiv: 2606.11088v1 — [abs](https://arxiv.org/abs/2606.11088) · [pdf](https://arxiv.org/pdf/2606.11088) · [html](https://arxiv.org/html/2606.11088v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.11088)
- authors: Zhiwei Li, Haiou Liu, Xijun Zhao, Ji Li, Yingze Wang, Boyang Wang
- categories: cs.RO
**Abstract.** Cooperative exploration with multiple ground robots in unknown, GPS-denied, bandwidth-limited environments is hard because localization drift breaks map consistency and causes redundant coverage. This framework couples descriptor-aided inter-robot loop closure with loop-aware hierarchical planning. A lightweight LiDAR global descriptor with range-image pre-alignment enables cross-robot place recognition under large yaw and lateral shifts, and verified loop closures maintain globally consistent trajectories over a sparse topological map. An uncertainty-aware loop-closure selection module scores candidates under pose uncertainty and keeps high-utility ones as planning anchors. The loop-closure module hits AR@1/AR@1% of 89.9%/95.5%, cuts trajectory error and two-way communication, and reduces exploration time and distance by 15% and 14% versus an mTSP baseline.
**Take.** The detail that makes this real is the bandwidth angle — a LiDAR global descriptor compact enough to share for cross-robot place recognition, not raw scans. Folding loop closures into the planner as anchors rather than treating SLAM and planning as separate stages is the right coupling. 15% less exploration time is modest but honest for a fully distributed, GPS-denied setup.
---
# arXiv digest — 2026-06-09
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-09
> papers: 5
Heavy 3D-perception day. The standout is ATN3D — LiDAR-Radar fusion tuned for the
long-range, sparse regime that actually constrains autonomous perception. Also notable: a
latent-space spatial memory for video world models (55× less memory), and a footprint-aware
motion planner whose cost scales with free space, not obstacle count.
## ★ ATN3D: Density-Aware LiDAR-Radar Early 3D Object Detection Under Extreme Sparsity
- arXiv: 2606.09634v1 — [abs](https://arxiv.org/abs/2606.09634) · [pdf](https://arxiv.org/pdf/2606.09634) · [html](https://arxiv.org/html/2606.09634v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.09634)
- authors: Debojyoti Biswas, Xianbiao Hu
- categories: cs.CV, cs.AI
**Abstract.** 3D object detection is the backbone of perception for automated vehicles. Long-range detection is hard because sensing evidence is sparse, yet on roadways >30m affords only ~1–2s to perceive and decide. Under extreme sparsity, early multimodal fusion tends to discard sparsity information and inject noise from empty/falsely-occupied cells, and uniform channel supervision favors dense near-range samples. ATN3D ("Ask The Neighbor") introduces density-aware early fusion with cross-modal gating, occupancy-gated neighborhood aggregation with circular kernels, evidence-conditioned channel self-attention, and a range-aware loss. On the VoD benchmark it beats strong baselines by +3.55% mAP (clear) and +8.41% mAP (heavy fog); for >30m objects, +3.33% and +2.09%.
**Take.** This is exactly the regime that matters for real perception — long range is where the time budget is smallest and the points are fewest. The smart bit isn't a bigger backbone; it's making fusion and supervision density-aware so the model stops drowning sparse far-range evidence in near-range noise. The fog numbers (+8.4 mAP) are the ones I'd trust as a signal it's solving the actual problem.
## Latent Spatial Memory for Video World Models
- arXiv: 2606.09828v1 — [abs](https://arxiv.org/abs/2606.09828) · [pdf](https://arxiv.org/pdf/2606.09828) · [html](https://arxiv.org/html/2606.09828v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.09828)
- authors: Weijie Wang, Haoyu Zhao, Yifan Yang, Feng Chen, Zeyu Zhang, Bohan Zhuang
- categories: cs.CV
**Abstract.** Video world models that keep 3D spatial consistency across frames usually rely on an explicit point-cloud memory built in RGB space — expensive (repeated rendering + VAE encoding) and lossy (the pixel-space round trip discards latent features). This paper introduces "latent spatial memory": a persistent 3D cache that stores scene information directly in the diffusion latent space. Their framework, Mirage, lifts latent tokens into 3D via depth-guided back-projection and queries by synthesizing novel views through direct latent-space warping. Reported: up to 10.57× faster end-to-end generation and 55× lower memory vs explicit 3D baselines, with SOTA on WorldScore.
**Take.** Keeping the spatial memory in latent space instead of round-tripping through pixels is the obvious-in-hindsight move, and the 55× memory cut is the kind of systems win that decides whether a world model is deployable or a demo. Depth-guided back-projection of latent tokens is a clean idea.
## HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents
- arXiv: 2606.09738v1 — [abs](https://arxiv.org/abs/2606.09738) · [pdf](https://arxiv.org/pdf/2606.09738) · [html](https://arxiv.org/html/2606.09738v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.09738)
- authors: Letian Li, Chao Shen, Shuzhao Xie, Chenghao Gu, Zhi Wang
- categories: cs.CV
**Abstract.** Text-driven indoor scene generation/editing needs an intermediate representation an LLM can both produce and revise. Scene graphs and global constraint lists are compact but underspecify local geometry and make edits hard to localize. HDSL frames it as structured program generation + local program repair: an XML/CSS-style DSL representing rooms, regions, objects, and support surfaces as a tree with local coordinates. LLM agents generate HDSL subtrees with bounded verification, ground nodes via multimodal asset retrieval, and apply force-directed layout to fix collisions. For editing, Hierarchical RAG rewrites only the relevant subtree and merges back deterministically — cutting tokens 5.22× and runtime 6.19× while preserving unrelated objects.
**Take.** Treating a 3D scene as a program you can locally repair — rather than a blob you regenerate — is the right abstraction, and it rhymes with how BIM and CAD actually work. The deterministic three-way merge on a scene tree is the detail that makes LLM editing trustworthy instead of destructive.
## Safe Polytope-in-Polytope Motion Planning and Control with Control Barrier Functions
- arXiv: 2606.09719v1 — [abs](https://arxiv.org/abs/2606.09719) · [pdf](https://arxiv.org/pdf/2606.09719) · [html](https://arxiv.org/html/2606.09719v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.09719)
- authors: Alejandro Gonzalez-Garcia, Dries Dirckx, Jan Swevers, Wilm Decré
- categories: cs.RO
**Abstract.** Robots in tight spaces need planning that respects their actual footprint; a point/circle approximation throws away the info needed to thread narrow passages. This work keeps a polytopic robot footprint inside a continuously-updated convex free-space region, formulated as discrete-time control-barrier-function constraints inside an MPC. The number of safety constraints scales with local free-space geometry and robot shape, not the number of obstacles, and it needs no obstacle detection or segmentation. Up to 91× faster than a polytope-based obstacle-avoidance formulation as obstacles grow; validated at 10 Hz on embedded hardware with occupancy grids and LiDAR.
**Take.** Making cost scale with free-space complexity instead of obstacle count is the elegant inversion here — and "no obstacle detection required" sidesteps a whole brittle perception stage. 10 Hz on an onboard embedded computer is the line that says this is real, not just a sim result.
## Evaluating the Representation Space of Diffusion Models via Self-Supervised Principles
- arXiv: 2606.09718v1 — [abs](https://arxiv.org/abs/2606.09718) · [pdf](https://arxiv.org/pdf/2606.09718) · [html](https://arxiv.org/html/2606.09718v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.09718)
- authors: Xiao Li, Yixuan Jia, Zekai Zhang, Liyue Shen, Qing Qu
- categories: cs.LG, cs.CV
**Abstract.** Diffusion models are both strong generators and strong self-supervised representation learners, but the link is under-explored. This paper decomposes features into invariant and residual components and derives the Invariant Contamination Ratio (ICR), a Fisher-based metric for how residual variation contaminates the invariant signal. Findings: invariance peaks at intermediate noise levels (which also give the best downstream classification), and ICR is a sensitive training-time indicator of the onset of memorization — detectable from training features alone, with no held-out set or external evaluator.
**Take.** The practically useful nugget: a training-time memorization detector that needs no held-out data. If ICR holds up, that's a cheap early-warning light for "your model has stopped generalizing" — exactly the thing you want in data-limited fine-tuning runs.
---
# arXiv digest — 2026-06-03
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/arxiv/2026-06-03
> papers: 4
First digest from the paper-scout pipeline. Heavy 3D day on cs.CV — trajectory-conditioned
dynamic shape generation (T2Mo) is the standout: explicit spatial guidance over text-only
conditioning is a pattern I expect to see everywhere in 3D generation this year.
## ★ Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text
- arXiv: 2606.05162v1 — [abs](https://arxiv.org/abs/2606.05162) · [pdf](https://arxiv.org/pdf/2606.05162) · [html](https://arxiv.org/html/2606.05162v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.05162)
- authors: Jaeyeong Kim, Ines Kim, Jahyeok Koo, Seungryong Kim
- categories: cs.CV
**Abstract.** We introduce T2Mo, a feed-forward framework for controllable dynamic 3D shape generation conditioned on 3D trajectories and text. Due to the inherent ambiguity of language, generating precisely intended motions using text alone remains challenging. To address this, we adopt 3D trajectories as controllable spatial guidance, specifying the exact paths along which selected points should move.
**Take.** Text-to-3D-motion has always been mushy because language underspecifies geometry — anchoring generation to explicit 3D point trajectories is the right fix. Feed-forward (no per-scene optimization) makes this actually usable in a pipeline.
## GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes
- arXiv: 2606.05142v1 — [abs](https://arxiv.org/abs/2606.05142) · [pdf](https://arxiv.org/pdf/2606.05142) · [html](https://arxiv.org/html/2606.05142v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.05142)
- authors: Josef Bengtson, Yaroslava Lochman, Fredrik Kahl
- categories: cs.CV, cs.AI
**Abstract.** Recent developments in multi-view image editing with generative models have brought us a step closer toward general 3D content generation and customization. Most existing works focus on rigid or appearance-only edits by utilizing the geometry of the unedited scene. This naturally limits these methods to edits that preserve the underlying scene structure.
**Take.** Multi-view consistency for *nonrigid* edits is the hard version of the problem — most methods cheat by keeping the original geometry. Worth a read if you care about editable digital twins.
## Reinforcement Learning from Rich Feedback with Distributional DAgger
- arXiv: 2606.05152v1 — [abs](https://arxiv.org/abs/2606.05152) · [pdf](https://arxiv.org/pdf/2606.05152) · [html](https://arxiv.org/html/2606.05152v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.05152)
- authors: Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad
- categories: cs.LG, cs.AI, cs.CL
**Abstract.** Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sample many responses and reward each with a single bit indicating whether the final answer is correct. Yet many settings provide rich feedback, including execution traces, tool outputs, expert corrections, and model self-evaluations.
**Take.** RLVR throws away everything except one bit per rollout — execution traces and tool outputs are sitting right there. Using them is obvious in hindsight, which is usually the mark of a good idea.
## Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)
- arXiv: 2606.05145v1 — [abs](https://arxiv.org/abs/2606.05145) · [pdf](https://arxiv.org/pdf/2606.05145) · [html](https://arxiv.org/html/2606.05145v1) · [ar5iv](https://ar5iv.labs.arxiv.org/html/2606.05145)
- authors: Nizar Islah, Istabrak Abbes, Irina Rish, Sarath Chandar, Eilif B. Muller
- categories: cs.LG, cs.AI, cs.CL
**Abstract.** When post-trained language models fail on reasoning problems, the common test-time-scaling response is to spend more compute on additional attempts, and the failed traces play no further role. We argue this discards a crucial signal; some failures come from unlucky sampling, where more rollouts help, while others are structural and resist resampling regardless of budget.
**Take.** Separating "unlucky sampling" failures from structural ones before you burn test-time compute is a genuinely useful triage — resampling a structurally-broken trace is just paying to fail again.
---
# PyTorch channels-last + AMP training loop
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/snippets
> lang: python
> date: 2026-05-25
For conv-heavy models on NVIDIA tensor cores, two cheap changes compound: move the
model and inputs to `channels_last` memory format, and wrap the forward pass in
`autocast` with a `GradScaler`. The memory format keeps activations in a layout the
tensor cores prefer; AMP runs the math in bf16/fp16 while keeping a fp32 master copy.
```python
model = model.to(device, memory_format=torch.channels_last)
scaler = torch.amp.GradScaler("cuda")
for x, y in loader:
x = x.to(device, memory_format=torch.channels_last, non_blocking=True)
y = y.to(device, non_blocking=True)
optimizer.zero_grad(set_to_none=True)
with torch.autocast("cuda", dtype=torch.float16):
loss = criterion(model(x), y)
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
```
---
# CUDA grid-stride loop for vector add
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/snippets
> lang: cuda
> date: 2026-05-22
A grid-stride loop decouples the launch configuration from the problem size: each
thread processes multiple elements, striding by the total number of threads. This
keeps the kernel correct for any `n`, plays nicely with occupancy tuning, and lets
you reuse the same launch geometry across input sizes.
```cuda
__global__ void vadd(const float *a, const float *b, float *c, int n) {
int stride = blockDim.x * gridDim.x;
for (int i = blockIdx.x * blockDim.x + threadIdx.x; i < n; i += stride) {
c[i] = a[i] + b[i];
}
}
// Launch: size the grid to the device, not to n.
int block = 256;
int grid = (n + block - 1) / block; // cap this for very large n
vadd<<>>(d_a, d_b, d_c, n);
```
---
# LiDAR Odometry
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/notes/lidar-odometry
> date: 2026-05-28
LiDAR odometry estimates a sensor's motion by aligning consecutive point-cloud
scans. Each new scan is matched against a reference — the previous scan or a local
map — and the rigid transform that best aligns them is the frame-to-frame motion.
Integrate those transforms and you have a trajectory.
The matching step is exactly [[point-cloud-registration]]: the odometry front-end is a
registration problem solved fast enough to run online. Modern systems fuse an IMU to
constrain the optimization (LiDAR-inertial odometry), which keeps the registration
well-conditioned during fast motion and feature-poor scenes.
## Why it is hard
- Scans are sparse and unstructured, so naive nearest-neighbour matching is costly.
- Drift accumulates: small per-frame errors compound over a long trajectory.
- Degenerate geometry (long corridors, open fields) leaves the alignment
under-constrained along some axes.
Most of the engineering effort goes into making [[point-cloud-registration]] both
robust and cheap enough to close the loop every frame.
---
# Point Cloud Registration
> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/notes/point-cloud-registration
> date: 2026-05-28
Point-cloud registration finds the rigid transform — rotation and translation —
that best aligns two overlapping point clouds. It is the core geometric primitive
behind mapping, scan stitching, and the front-end of [[lidar-odometry]].
## The classic recipe
Iterative Closest Point (ICP) alternates two steps until convergence:
1. **Correspondence** — for each source point, find its nearest target point.
2. **Alignment** — solve for the transform that minimizes the residual (point-to-point
or, faster-converging, point-to-plane).
Variants improve every part of this loop: better correspondences (normals, features),
robust losses (to reject outliers), and voxel or KD-tree acceleration so the
nearest-neighbour search is not the bottleneck.
## In practice
Real systems seed ICP with a coarse global alignment or an IMU prior so it converges
to the right basin. When registration runs every frame against a sliding local map,
it becomes [[lidar-odometry]] — same math, tighter time budget.