# Satyajit Ghana — full content corpus > Head of Engineering @ Inkers Technology. I build deep-learning systems, 3D perception, and high-performance infra. > Generated from the content layer. Index: /llms.txt --- # About Satyajit Ghana > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/about I build deep-learning systems, 3D perception, and high-performance infra. Head of Engineering at Inkers, where I work on industrial AI — 3D perception, deep learning, and the systems that ship it to production. I write custom neural networks, CUDA kernels, LiDAR/point-cloud pipelines, and high-performance C++/gRPC services. Previously taught MLOps (EMLO 2.0) and computer vision (EVA 4.0) at The School of AI. ## Links - github: https://github.com/satyajitghana - linkedin: https://www.linkedin.com/in/satyajitghana/ - x: https://x.com/thesudoer_ - medium: https://satyajitghana.medium.com/ - scholar: https://scholar.google.com/citations?user=rZCRakQAAAAJ&hl=en - website: https://thesatyajit.com/ - email: satyajitghana7@gmail.com --- # Resume — Satyajit Ghana > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/resume > pdf: https://ai.thesatyajit.com/satyajit-ghana-resume.pdf > json: https://ai.thesatyajit.com/resume.json ```json { "contact": { "name": "Satyajit Ghana", "title": "Head of Engineering", "company": { "name": "Inkers Technology", "url": "https://inkers.ai" }, "location": "Bengaluru, India", "email": "satyajitghana7@gmail.com", "github": "https://github.com/satyajitghana", "linkedin": "https://www.linkedin.com/in/satyajitghana/", "website": "https://ai.thesatyajit.com" }, "summary": "Engineering leader and deep-learning systems builder. I lead engineering for industrial-AI products — 3D perception, LiDAR/point-cloud pipelines, and structural-defect analysis — and ship them as high-performance C++/CUDA/gRPC services. Two USPTO patents pending. Previously taught MLOps and computer vision.", "experience": [ { "organization": "Inkers Technology", "url": "https://inkers.ai", "location": "Bengaluru, India", "roles": [ { "title": "Head of Engineering", "start": "2024-01" }, { "title": "Deep Learning Software Engineer", "start": "2022-06", "end": "2024-01" }, { "title": "Deep Learning Associate", "start": "2021-07", "end": "2022-06" } ], "summary": "Lead engineering for industrial-AI products across 3D perception, LiDAR/point-cloud pipelines, and structural-defect analysis.", "highlights": [ "Lead engineering for industrial-AI products: 3D perception, LiDAR/point-cloud pipelines, and structural-defect analysis.", "Design custom neural networks, CUDA kernels, and high-performance C++/gRPC services for production deployment.", "Named inventor on 2 USPTO patents pending (assigned to Inkers); grew from Deep Learning Associate to Head of Engineering." ] }, { "organization": "The School of A.I.", "location": "Remote", "roles": [ { "title": "MLOps Instructor", "start": "2022-01", "end": "2023-01" } ], "summary": "Designed and taught MLOps and contributed to computer-vision curriculum.", "highlights": [ "Designed and taught EMLO 2.0 — an end-to-end MLOps course (training, packaging, deployment, monitoring).", "Contributed to EVA 4.0, the deep computer-vision program." ] } ], "education": [ { "institution": "M.S. Ramaiah University of Applied Sciences", "location": "Bangalore, India", "degree": "B.Tech", "end": "2021", "highlights": [ "CGPA 9.78/10", "Silver Medalist" ] } ], "skills": [ { "name": "Deep Learning", "skills": [ "PyTorch", "TensorFlow", "Computer Vision", "3D / Point Clouds / LiDAR", "GenAI (SDXL, LLMs)" ] }, { "name": "Systems", "skills": [ "C++", "C", "Rust", "CUDA", "gRPC", "MongoDB" ] }, { "name": "MLOps", "skills": [ "Kubernetes", "AWS", "GCP", "Docker" ] }, { "name": "Web", "skills": [ "TypeScript", "React", "Next.js" ] } ], "projects": [ { "name": "torch-point-ops", "description": "High-performance PyTorch operators for point-cloud and 3D geometry processing.", "url": "https://github.com/satyajitghana/torch-point-ops" }, { "name": "PV-LIO-for-HBA", "description": "Point-to-voxel LiDAR-inertial odometry adapted for hierarchical bundle adjustment.", "url": "https://github.com/satyajitghana/PV-LIO-for-HBA" }, { "name": "ige_lio", "description": "Iterated error-state Kalman-filter LiDAR-inertial odometry experiments.", "url": "https://github.com/satyajitghana/ige_lio" }, { "name": "sdxl-dreambooth-finetune", "description": "DreamBooth fine-tuning pipeline for Stable Diffusion XL subject personalization.", "url": "https://github.com/satyajitghana/sdxl-dreambooth-finetune" }, { "name": "TSAI-DeepVision-EVA4.0", "description": "Coursework and experiments from The School of AI's EVA 4.0 computer-vision program.", "url": "https://github.com/satyajitghana/TSAI-DeepVision-EVA4.0" }, { "name": "PadhAI-Course", "description": "Implementations and notes from the PadhAI deep-learning course.", "url": "https://github.com/satyajitghana/PadhAI-Course" } ], "publications": [ { "title": "Adaptive Visual Learning Using Augmented Reality and Machine Learning Techniques", "publisher": "Journal of Computational and Theoretical Nanoscience", "date": "2020-01-01", "details": "Vol. 17, No. 11, pp. 4952–4956", "doi": "10.1166/jctn.2020.8982", "url": "https://doi.org/10.1166/jctn.2020.8982" } ], "patents": [ { "title": "Method and System for Performing Structural Defect Analysis in a Structural Environment", "status": "Pending", "applicationNumber": "US 19/634,310", "filed": "2026-03-31", "assignee": "Inkers Technology" }, { "title": "Data Acquisition Device", "status": "Pending", "applicationNumber": "US 19/634,339", "filed": "2026-03-31", "assignee": "Inkers Technology" } ] } ``` --- # Health — biomarker panel > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/health > panel date: 2026-05-15 ```json { "panelDate": "2026-05-15", "stack": [ { "type": "device", "name": "Apple Watch", "url": "https://www.apple.com/watch/" }, { "type": "device", "name": "Withings Body+" }, { "type": "supplement", "name": "Vitamin D3" }, { "type": "supplement", "name": "Omega-3" }, { "type": "supplement", "name": "Creatine" }, { "type": "supplement", "name": "Magnesium Glycinate" } ], "biomarkers": [ { "key": "ldl-c", "label": "LDL-C", "value": 128, "unit": "mg/dL", "category": "cardiovascular", "weight": 2, "optimalRange": { "max": 100 }, "note": "Elevated — primary driver of atherosclerotic risk.", "status": "elevated" }, { "key": "hdl-c", "label": "HDL-C", "value": 52, "unit": "mg/dL", "category": "cardiovascular", "optimalRange": { "min": 40 }, "status": "optimal" }, { "key": "apob", "label": "ApoB", "value": 105, "unit": "mg/dL", "category": "cardiovascular", "weight": 1.5, "optimalRange": { "max": 90 }, "note": "Elevated — atherogenic particle count above target.", "status": "elevated" }, { "key": "triglycerides", "label": "Triglycerides", "value": 88, "unit": "mg/dL", "category": "cardiovascular", "optimalRange": { "max": 150 }, "status": "optimal" }, { "key": "hba1c", "label": "HbA1c", "value": 5.4, "unit": "%", "category": "metabolic", "weight": 1.5, "optimalRange": { "max": 5.7 }, "status": "optimal" }, { "key": "fasting-glucose", "label": "Fasting Glucose", "value": 96, "unit": "mg/dL", "category": "metabolic", "optimalRange": { "min": 70, "max": 99 }, "status": "optimal" }, { "key": "fasting-insulin", "label": "Fasting Insulin", "value": 8.1, "unit": "µIU/mL", "category": "metabolic", "optimalRange": { "max": 8 }, "note": "Borderline — marginally above optimal fasting insulin.", "status": "borderline" }, { "key": "alt", "label": "ALT", "value": 22, "unit": "U/L", "category": "liver_kidney", "optimalRange": { "max": 40 }, "status": "optimal" }, { "key": "ast", "label": "AST", "value": 24, "unit": "U/L", "category": "liver_kidney", "optimalRange": { "max": 40 }, "status": "optimal" }, { "key": "creatinine", "label": "Creatinine", "value": 0.95, "unit": "mg/dL", "category": "liver_kidney", "optimalRange": { "min": 0.7, "max": 1.3 }, "status": "optimal" }, { "key": "egfr", "label": "eGFR", "value": 99, "unit": "mL/min", "category": "liver_kidney", "optimalRange": { "min": 90 }, "status": "optimal" }, { "key": "tsh", "label": "TSH", "value": 2.1, "unit": "mIU/L", "category": "hormonal", "optimalRange": { "min": 0.5, "max": 4 }, "status": "optimal" }, { "key": "testosterone", "label": "Testosterone", "value": 610, "unit": "ng/dL", "category": "hormonal", "optimalRange": { "min": 300, "max": 1000 }, "status": "optimal" }, { "key": "vitamin-d", "label": "Vitamin D", "value": 24, "unit": "ng/mL", "category": "nutritional", "weight": 1.5, "optimalRange": { "min": 30, "max": 100 }, "note": "Low — below the optimal 25-OH vitamin D range.", "status": "low" }, { "key": "vitamin-b12", "label": "Vitamin B12", "value": 540, "unit": "pg/mL", "category": "nutritional", "optimalRange": { "min": 400, "max": 1000 }, "status": "optimal" }, { "key": "ferritin", "label": "Ferritin", "value": 38, "unit": "ng/mL", "category": "nutritional", "optimalRange": { "min": 50, "max": 300 }, "note": "Low — iron stores below the optimal floor.", "status": "low" }, { "key": "hemoglobin", "label": "Hemoglobin", "value": 15.1, "unit": "g/dL", "category": "blood_panel", "optimalRange": { "min": 13.5, "max": 17.5 }, "status": "optimal" }, { "key": "wbc", "label": "WBC", "value": 6.2, "unit": "10³/µL", "category": "blood_panel", "optimalRange": { "min": 4, "max": 11 }, "status": "optimal" }, { "key": "platelets", "label": "Platelets", "value": 245, "unit": "10³/µL", "category": "blood_panel", "optimalRange": { "min": 150, "max": 400 }, "status": "optimal" }, { "key": "resting-hr", "label": "Resting HR", "value": 58, "unit": "bpm", "category": "vitals", "optimalRange": { "min": 40, "max": 60 }, "status": "optimal" }, { "key": "bp-systolic", "label": "BP Systolic", "value": 132, "unit": "mmHg", "category": "vitals", "weight": 1.5, "optimalRange": { "max": 120 }, "note": "Elevated — systolic pressure above the optimal ceiling.", "status": "borderline" }, { "key": "vo2max", "label": "VO₂max", "value": 47, "unit": "mL/kg/min", "category": "vitals", "optimalRange": { "min": 42 }, "status": "optimal" } ] } ``` --- # Now > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/now ```json { "updated": "2026-06-04", "items": [ "Building ai.thesatyajit.com — a dual-native personal site maintained by a crew of Claude agents that ship every change as a reviewed PR.", "Leading industrial-AI 3D perception at Inkers: turning LiDAR and camera capture into structural-defect analysis that runs in production.", "Writing CUDA point-cloud kernels — fused nearest-neighbour, sampling, and grouping ops to take the Python side out of the hot path.", "Reading deeply on LiDAR odometry and point-cloud registration, with an eye toward tighter, faster SLAM front-ends." ] } ``` --- # Uses > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/uses ```json [ { "section": "Workstation", "items": [ { "name": "Custom GPU workstation", "note": "NVIDIA RTX-class GPU for CUDA + training" }, { "name": "Linux", "note": "Ubuntu — primary dev OS" } ] }, { "section": "Editor", "items": [ { "name": "VS Code", "note": "daily driver" }, { "name": "Claude Code", "note": "agentic pair-programming in the terminal" }, { "name": "Fonts", "note": "JetBrains Mono today, migrating toward IBM Plex Mono" } ] }, { "section": "Terminal", "items": [ { "name": "zsh", "note": "shell" }, { "name": "tmux", "note": "session multiplexing" } ] }, { "section": "Stack highlights", "items": [ { "name": "PyTorch + CUDA", "note": "custom kernels for point-cloud ops" }, { "name": "C++ / gRPC", "note": "high-performance perception services" }, { "name": "Next.js + Tailwind", "note": "this site" } ] } ] ``` --- # Reading > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/reading ```json [ { "type": "paper", "title": "PV-LIO: A Probabilistic Voxel-based LiDAR-Inertial Odometry Framework", "status": "reading", "note": "PLACEHOLDER — voxel-probabilistic front-end ideas for tighter LIO." }, { "type": "book", "title": "Programming Massively Parallel Processors", "author": "Hwu, Kirk, El Hajj", "status": "reading", "note": "PLACEHOLDER — the CUDA reference I keep coming back to." }, { "type": "paper", "title": "FAST-LIO2: Fast Direct LiDAR-Inertial Odometry", "status": "read", "note": "PLACEHOLDER — iterated Kalman filter without feature extraction." }, { "type": "book", "title": "Multiple View Geometry in Computer Vision", "author": "Hartley & Zisserman", "status": "queued", "note": "PLACEHOLDER — geometry fundamentals refresher." }, { "type": "paper", "title": "3D Gaussian Splatting for Real-Time Radiance Field Rendering", "status": "queued", "note": "PLACEHOLDER — splatting as a reconstruction primitive." } ] ``` --- # Vritti UI > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/vritti > stack: React 19, Next.js 16, Tailwind v4, Motion, Three.js, TypeScript Components crafted for design engineers — animated components and blocks you copy and paste into your shadcn/ui project. No package installs, full ownership. ```bash npx shadcn@latest add "https://vritti.thesatyajit.com/r/shimmer-button" ``` **560+ components** across 14 categories — backgrounds, animations, text effects, buttons, cards, charts, layouts, navigation, carousels, cursors, inputs, media, shaders, and the genuinely weird ones. **240+ blocks** — pre-built page sections for real apps: auth, pricing, e-commerce, billing, dashboards, modals, testimonials, footers, FAQ, AI & Web3. Beyond components, it ships a set of **creative tools** that run online or install into your project: Art Studio (9 generative-art tools), Background Studio, Shape Studio (an SVG editor), Texture Studio (30+ stackable WebGL filters), Shader Studio (87 production-ready shaders), Dither Studio, a visual Theme Editor (42 presets, Google Fonts, contrast checker, CSS export), and native haptics. There's an interactive playground for tuning component props live, and AI components (chat streaming, image generation, voice, AI search) for LLM-native UIs. Everything is copy-paste, shadcn-style — you own the source. It's the most ambitious thing on this page: a full component ecosystem, a design-tool suite, and a theme system, built on React 19, Next.js 16, Tailwind v4, Motion, and Three.js. --- # torch-point-ops > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/torch-point-ops > stack: PyTorch, CUDA, C++, Python > repo: https://github.com/satyajitghana/torch-point-ops A PyTorch library for 3D point operations — nearest-neighbour, sampling, and grouping primitives — implemented as fused CUDA kernels for throughput on large point clouds. Built to back LiDAR and 3D-perception pipelines where the Python-side ops were the bottleneck. Exposes a clean `torch` autograd-compatible API over the native kernels. --- # fabrik > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/fabrik > stack: TypeScript, React, Vercel AI SDK, Zod, Motion > repo: https://github.com/satyajitghana/fabrik Most LLM apps are text in, text out. **fabrik** lets the model answer with *interface* — it picks and fills real React components (cards, forms, charts, pickers) instead of emitting a wall of prose. The AI decides what to show. It's provider-agnostic: wire it to any model through the Vercel AI SDK (Gemini, OpenAI, Anthropic, …), keep the API key server-side in a single route handler, and render the streamed component tree on the client. Components are validated with Zod schemas, so the model can only produce UIs you've defined — no arbitrary markup, no injection surface. The bet behind it: as models get better at tool use, the next interface isn't a chat bubble — it's a UI the model assembles on the fly for the task at hand. --- # rentree > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/rentree > stack: Next.js, React, TypeScript, Supabase, Drizzle ORM, Tailwind v4 > repo: https://github.com/satyajitghana/rentree Rentree connects people with organic farms across India. Instead of buying produce off a shelf, you rent a plot or adopt a tree: your farmer grows your food, sends real photo updates as it grows, and ships it when it's ready. Rent a plot for a season (tomatoes, herbs, rice — whatever grows), or adopt an apple, orange, or mango tree and get the whole harvest. Every tree carries a QR code, so you can scan and trace the full journey from planting to your plate. A clean Next.js 16 app on Supabase + Drizzle, with the provenance trail as the trust mechanism that makes the model work. --- # Quantum Orbital Visualizer > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/atoms > stack: Next.js, Three.js, Rust (WASM), C++23 (WASM), Tailwind v4, Zustand > repo: https://github.com/satyajitghana/atoms Each visualization is the probability density |ψ|² of a hydrogen-like electron, computed from the exact solutions to the Schrödinger equation and rendered as up to 500K glowing particles in real time. Drive the quantum numbers (n, l, m) and watch the probability cloud morph as you go. The fun engineering bit is the **dual WASM backends**: the same physics is implemented three ways — JavaScript, Rust→WASM (~41 KB), and C++23→WASM (~28 KB) — and you can switch the compute engine at runtime. Same orbitals, different implementations, side-by-side for benchmarking. There's also an Element Explorer: the full periodic table, where clicking any of the 118 elements parses its electron configuration into individual orbitals you can then visualize one by one. --- # BOTCHA > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/botcha > stack: Next.js, TypeScript, Redis, crypto (SHA-256 / HMAC / JWT) > repo: https://github.com/satyajitghana/botcha Only an autonomous agent with runtime access to HTTP, cryptography, and byte manipulation can pass — which is the whole joke, and the whole point. Each challenge is 256 random bytes plus 2–4 byte-level transformation steps written in randomized natural language, inside a 30-second window: fast enough for a machine, hopeless for a human copy-pasting into a REPL. The agent must decode the base64, execute each transform in order, concatenate the raw byte outputs, `SHA-256` the result, then `HMAC-SHA256(key=nonce, message=answer)` and submit both — proving it actually did the computation. On success it gets a short-lived JWT. It's a small, sharp take on a real question the agent era raises: if the web increasingly wants to *let bots in* and keep humans out, what does that gate look like? --- # penora > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/penora > stack: TypeScript, React, Canvas, npm package > repo: https://github.com/satyajitghana/penora Type a string and penora animates it into natural handwriting — not a font fade-in, but stroke-by-stroke drawing driven by the glyph contours, with pen physics layered on top: pressure tapering at stroke ends, seeded jitter so each render is subtly different, and micro-wobble for that hand-drawn quality. Export the result as video or GIF. Shipped as an npm package (`npm i penora`) and as a shadcn registry component you can drop straight into a project. It's the kind of small, self-contained library that's satisfying to build: a tight problem (make text look handwritten and *alive*) with a lot of room for craft in the details. --- # Hyr > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/hyr > stack: Next.js, TypeScript, Tailwind v4, Google Gemini > repo: https://github.com/satyajitghana/hyr Drop in any PDF resume and Hyr parses it into a structured, fully editable format — inline editing for every section, no forms or modals. From there it tailors the resume per job posting, aligning keywords and tone to the listing so it gets past ATS screening and reads as a real fit. Build, tailor, optimize, and track the whole job search in one place. The AI work runs on Gemini behind a server route; the parsing-to-structured-data step is the part that makes the rest of the editing experience feel clean rather than fighting a PDF. --- # MockLab > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/mocklab > stack: Next.js, Tailwind v4, shadcn/ui, Motion, html-to-image > repo: https://github.com/satyajitghana/mocklab MockLab generates pixel-perfect social mockups across 21+ platforms — social posts (X, LinkedIn, Instagram post/story, Reddit, Threads, YouTube, TikTok), chat messages (WhatsApp, Telegram, Slack, Discord, iMessage, Snapchat), and Gmail. Edit on the left, watch the live preview on the right, download as 2× PNG. Everything is configurable: verified badges, reactions, timestamps, read receipts, media uploads, and per-message editing for the chat mockups — all with dark/light theming for both the app and every mockup. It's a deceptively deep UI project: each platform is its own little pixel-accurate design system to reproduce. --- # tapcn > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/projects/tapcn > stack: React Native, Expo, TypeScript, CLI > repo: https://github.com/satyajitghana/tapcn tapcn brings the shadcn model to mobile: not an npm component library, but a CLI that copies beautifully-designed, accessible component *source* directly into your Expo project. No dependency lock-in, no version conflicts — just code you control and can edit. ```bash npx @tapcn/cli init npx @tapcn/cli add button card input text ``` Components work on iOS, Android, and Web from a single source. The whole appeal of the shadcn approach — own your components, theme them freely, no black-box library — has been missing on React Native; tapcn fills that gap. --- # My site's chatbot was stuffing 273k tokens into every message. I gave it tools instead. > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/blog/site-agent-dynamic-tools > date: 2026-07-20 > tags: agents, llm, tools, retrieval, systems The assistant embedded on this site — the `⌘K` console, `/api/ask`, and the `ask_satyajit` MCP tool — worked. It was also doing something faintly absurd. Every single request built its system prompt like this: ```ts // the old lib/chat.ts, abridged const sections = [] for (const page of ["about", "resume", "health", "now", "uses", "reading"]) sections.push(await dataPageMarkdown(page)) for (const item of getAllContent()) // every article, blog, log, digest… sections.push(contentMarkdown(item.kind, item.slug)) return [PERSONA, RULES, "=== CONTENT CORPUS ===", sections.join("\n---\n")].join("\n") ``` It pasted **the entire site** into the prompt and let the provider's context cache eat the cost. That's a defensible move when the corpus is a page or two. Mine isn't anymore: ```text old full-corpus system prompt: 1,090,854 chars ≈ 273,000 tokens ``` The code even carried a comment — *"if the corpus ever approaches ~100K tokens, switch to retrieval"* — that reality had quietly sailed past nearly 3× over. Two hundred seventy-three thousand tokens is past the input window of the model serving it. The chat was running on whatever survived truncation. Time to do the thing the comment said. ## The fix: retrieve, don't dump I rebuilt the assistant as a small **tool-using agent**. The system prompt now carries only a compact **catalog** — every page's `kind/slug`, title, and one-line description, about **5k tokens** — and the agent pulls the bodies it actually needs through tools. Three of them: ```ts // lib/chat.ts — the whole tool set search_content({ query, limit }) // BM25 over the site (ranked, returns section + snippet) get_content({ kind, slug }) // fetch one page's full markdown, on demand think({ thought }) // a reasoning scratchpad; records the thought, returns nothing ``` A question now flows: read the catalog → `search_content` (or jump straight to `get_content` if the catalog already names the page) → read one or two pages → optionally `think` to reconcile them → answer, with citations. The base prompt is fixed and small; the variable cost is the one-to-three pages it fetched, not the other sixty-five it didn't. ```text before: ~273k tokens, every request, whether relevant or not after: ~5k catalog + ~2–10k per page actually fetched ``` That's the same **dynamic loading** idea from [Kimi K3's tool-calling guide](https://platform.kimi.ai/docs/guide/kimi-k3-tool-calling-best-practice) — their point is that a big upfront payload "eats up context and makes the model more likely to pick the wrong thing." Kimi loads *tool schemas* on demand; my corpus is the payload, so I load *content* on demand. Same principle, one layer down. ## Three tools, on purpose The tool count is a decision, not an accident. Mario Zechner's [pi coding agent](https://mariozechner.at/posts/2025-11-30-pi-coding-agent/) makes the case that "four tools are all you need" — `read`, `write`, `edit`, `bash` — and that MCP servers which "dump their entire tool descriptions into your context on every session" are the anti-pattern. A read-only site agent needs fewer still: search, fetch, think. Each description is a couple of sentences. The combined tool surface is under a thousand tokens, so keeping it declared upfront (rather than lazily loading schemas, which only pays off with dozens of tools) is the right call here — the honest version of "dynamic tools" for a small surface is *don't have a big one*. The one genuinely new tool is `think`, from Anthropic's [think-tool post](https://www.anthropic.com/engineering/claude-think-tool). It does nothing — literally logs the thought and returns: ```ts think: tool({ description: "Think out loud: plan which pages to fetch, or check a draft answer against the sources you read. Records the thought and returns nothing new.", inputSchema: z.object({ thought: z.string() }), execute: async ({ thought }) => ({ ok: true, thought }), }), ``` That looks pointless until you watch a multi-step tool run. Between `search_content` and the final answer, the model has raw tool output sitting in context and no designated place to reason over it before committing. `think` is that place — a scratchpad that keeps the "what did I just read, and does it actually answer the question" step from being skipped. Anthropic reports it buying a large margin on multi-step tool tasks; the appeal for a retrieval agent is exactly that it makes the model *check the page it fetched* instead of answering from the snippet. ## The loop The agent loop itself is one call, courtesy of the Vercel AI SDK — `stopWhen` bounds how many tool round-trips it may take before it has to answer: ```ts const result = streamText({ model: chatModel("main"), system, // persona + rules + the 5k-token catalog messages, tools: agentTools(), // search_content, get_content, think stopWhen: stepCountIs(6), // search → read → think → answer is ~4; 6 is headroom }) ``` The same tool set backs all three surfaces — the streaming `⌘K` console (`/api/chat`), the one-shot `/api/ask`, and the `ask_satyajit` MCP tool — so there's one harness, not three. And `search_content` is the [Contextual BM25](/blog/site-search-contextual-bm25) engine I'd already built for the site's `/search`, now doing double duty as the agent's retriever. The pieces compose: the search work made the harness possible. Honest scope. This is retrieval over a small, single-author corpus, not a coding agent — the lessons transfer but the stakes are lower. `think`'s benefit is task-dependent (Anthropic saw big gains on policy-heavy multi-step tasks, near-zero on simple ones); on a two-hop lookup it mostly earns its keep by stopping the model from answering off the snippet. And the retriever is lexical BM25 — no embeddings yet, so a pure paraphrase with no shared terms can still miss the right page. The catalog is the safety net there: the model can see every title even when search comes up short. ## The take The old design wasn't wrong when it was written — it was wrong at 273k tokens. The rebuild is the [agent-harness](/articles/agent-harness) framing applied to my own site: the model barely changed, but the loop, the tool set, and the context policy around it changed completely. Give the agent a map and three tools and let it fetch, instead of force-feeding it the whole library and hoping the answer survives the truncation. Retrieve, don't dump. --- # Giving my own search the Contextual BM25 treatment > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/blog/site-search-contextual-bm25 > date: 2026-07-20 > tags: search, retrieval, bm25, rag, information-retrieval I recently wrote two pieces back to back: one on [BM25](/articles/bm25), the ranking function that refuses to die, and one on [Contextual Retrieval](/blog/contextual-retrieval), Anthropic's fix for chunks that forget where they came from. Then I opened `lib/search.ts` — the thing powering this site's own `/search` and the `/api/search` endpoint agents hit — and found this: ```ts // the old lib/search.ts, abridged for (const item of getAllContent()) { const fields = [item.title, item.description, item.tags, item.body] for (const text of fields) { if (text.toLowerCase().includes(q)) { // q = the whole lowercased query results.push({ /* … a ±120-char window around the hit */ }) break } } } ``` A substring `indexOf`. It has three problems, and the third is the one that actually bites: 1. **No ranking.** The first document that contains the substring wins, in date order. A tight, on-topic match and an incidental mention are indistinguishable. 2. **No notion of rarity or length.** Matching "the" counts the same as matching "kalman". 3. **Multi-word queries fall off a cliff.** `q` is the *entire* query string, so `includes(q)` needs that **exact phrase** somewhere in the text. Nobody searches in verbatim phrases. `"bm25 length normalization"` returns **zero results** — those three words are all in the BM25 article, just never contiguous. So I ate my own cooking. Here's the rebuild. ## BM25, the same formula I just wrote up The [BM25 walkthrough](/articles/bm25) has the whole derivation; the code is a direct transcription of it. Lucene's non-negative IDF, term-frequency saturation at `k1 = 1.2`, length normalization at `b = 0.75`: ```ts // lib/search.ts — the ranking core, straight from the article's formula const K1 = 1.2 const B = 0.75 function idf(term: string, ix: Index): number { const n = ix.df.get(term) ?? 0 return Math.log(1 + (ix.n - n + 0.5) / (n + 0.5)) } // per query term present in a chunk: const f = chunk.tf.get(term) ?? 0 const norm = K1 * (1 - B + (B * chunk.len) / ix.avgdl) score += idf(term, ix) * (f * (K1 + 1)) / (f + norm) ``` Two things fall out for free. The query is **tokenized** into terms and each is scored independently, so multi-word queries just work — no phrase has to exist. And rare terms dominate: `kalman` carries far more weight than `filter`, because `idf` collapses for common words. No stopword list, same as the article promised. I keep a **postings map** (`term → chunk indices`) so a query only touches chunks that actually share a term, instead of scanning the whole corpus. The corpus is one person's writing, so this is overkill — but it's the same inverted-index shape a real engine uses, and it's three lines. ## Contextual chunks, minus the Claude call BM25 alone still has the chunk-amnesia problem from the [Contextual Retrieval post](/blog/contextual-retrieval): if I split an article into passages and index each one alone, a passage that reads "it drops 41% at parity" has no idea it's *about the Harness Effect, in the section on the controlled swap*. A query for either misses it. Anthropic's fix is to have Claude write a one-line context for every chunk before indexing. Mine is cheaper and dumber: the context a chunk lost is sitting right there in the document's frontmatter and headings. So before indexing, every chunk inherits its document's **title, description, tags, and nearest heading** — deterministically, at request time, no model, no build step: ```ts // each chunk is packed with weighted terms: its own body, its heading, and the // document context it would otherwise have lost when the body was split. const contextTokens = tokenize([title, description, tags.join(" ")].join(" ")) for (const t of bodyTokens) add(t, 1) for (const t of headingTokens) add(t, 1) // section-local: full weight for (const t of contextTokens) add(t, CONTEXT_WEIGHT) // 0.5 — present, not dominant ``` The honest caveat: this is a **poor man's** Contextual Retrieval. Claude's per-chunk context can say things the frontmatter can't ("the previous quarter's revenue was \$314M"); my version can only replay the structural context the document already carries. And prepending the *same* title to every chunk of a document inflates those terms' document-frequency, which is exactly why the context terms get `CONTEXT_WEIGHT = 0.5` instead of full weight — present enough to make the chunk findable by its subject, quiet enough not to drown the passage's own words. It's the 80% of the win for 0% of the inference cost, which is the right trade for a static personal site. Each result also reports **which section** it matched — the nearest heading rides along as the result's `field`, so `/api/search` tells an agent not just *which* document but *where* in it. ## Before / after Same queries, old substring search vs. the new Contextual BM25, on the live corpus: ```text query substring indexOf contextual BM25 (top hit) -------------------------------------------------------------------------------------- "bm25 length normalization" 0 results (no verbatim articles/bm25 phrase) [Start with TF-IDF, and its two flaws] "revenue grew quarter" 0 results blog/contextual-retrieval [The problem: chunks forget …] "kalman filter" first dated doc that articles/kalman-filter (8.4) contains the phrase, + fast-lio2 (8.6, cites it) unranked ranked by relevance, not date "reciprocal rank fusion" 0 results blog/contextual-retrieval [A dependency-free repro] ``` The multi-word queries are the story. Under substring, anything that isn't a verbatim phrase — which is almost everything a person types — returned nothing. Under BM25 every term contributes, so the query finds the document even when its words are scattered across a paragraph, and the contextual prefix pulls in matches on the *subject* of a passage, not just its literal words. ## It's live Try it: [`/search?q=length+normalization`](/search?q=length+normalization). Agents get the ranked, scored JSON with the section label: ```bash curl -s "https://ai.thesatyajit.com/api/search?q=contextual%20retrieval&limit=3" | jq '.results[] | {slug, field, score}' ``` What this still isn't: it's the **lexical half** only. The [contextual retrieval post](/blog/contextual-retrieval) makes the case that the real win is *fusing* BM25 with a dense retriever, and that reranking the shortlist is what takes the failure rate the last mile. There are no embeddings here yet — a query has to share actual terms with a chunk, so a pure paraphrase with no lexical overlap can still miss. For a corpus this size, lexical BM25 over contextualized chunks is the honest 90% solution; the semantic half is the next commit, not this one. ## The takeaway The whole change is `lib/search.ts` — a few hundred lines, no dependencies, no index server, no model at request time. It's the two ideas I'd just written about, applied to the smallest possible target: my own search bar. BM25 for ranking, structural context for the chunk-amnesia problem, and an honest note about the half I haven't built yet. Writing about a technique is a good way to understand it; running it in your own site is a better one. --- # Contextual Retrieval, with a runnable repro and a browser playground > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/blog/contextual-retrieval > date: 2026-07-17 > tags: rag, retrieval, claude, embeddings, playground [Anthropic's Contextual Retrieval post](https://www.anthropic.com/engineering/contextual-retrieval) has been on my to-read list for a while. It targets the oldest bug in RAG, so I finally sat down, built a dependency-free reproduction, and wired up a browser playground. This is the write-up. ## The problem: chunks forget where they came from Standard RAG splits documents into chunks and indexes each chunk on its own. That destroys context. Anthropic's example is perfect: > `The company's revenue grew by 3% over the previous quarter.` Which company? Which quarter? The chunk can't say. A query like *"how did ACME's Q2 revenue change?"* has **no terms to match** against that chunk — the words "ACME" and "Q2 2023" live in the *document*, not the *chunk*. Both lexical (BM25) and semantic (embedding) retrieval miss it. ## The fix: let Claude situate each chunk Contextual Retrieval prepends a short, chunk-specific context to each chunk **before** you embed it and **before** you build the BM25 index — Anthropic calls these **Contextual Embeddings** and **Contextual BM25**. The context is generated by Claude, given the whole document. The same chunk becomes: > `This chunk is from an SEC filing on ACME corp's performance in Q2 2023; the previous quarter's revenue was $314 million. The company's revenue grew by 3% over the previous quarter.` Now the owner and the period are *in the chunk*, so retrieval can find it. ## Watch it happen Here is the whole pipeline in your browser — the document, its chunks, the chunks with context prepended, then retrieval. The scoring is a real BM25 and a TF-IDF cosine (a stand-in for an embedding model); the contextual prefixes stand in for Claude's output. Step to **4 · retrieve**, pick a query, and watch the right chunk climb from the standard column to the contextual one: Toggle between BM25, embeddings, and their fusion. The pattern holds across all three: the answer chunk is buried on standard chunks and near the top once each chunk carries its context. ## The one prompt that does the work This is the prompt from the post, verbatim. It runs once per chunk, with the full document supplied as context: ```text {{WHOLE_DOCUMENT}} Here is the chunk we want to situate within the whole document {{CHUNK_CONTENT}} Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. Answer only with the succinct context and nothing else. ``` The obvious worry is cost: you re-send the whole document once per chunk. **Prompt caching** kills that. Cache the document once and every chunk in it reads from the cache, which is why Anthropic quotes a one-time **$1.02 per million document tokens**. The document is the stable prefix, so it takes the `cache_control` breakpoint; the chunk and instruction vary and come after it: ```python from anthropic import Anthropic client = Anthropic() CONTEXT_PROMPT = """Here is the chunk we want to situate within the whole document {chunk} Please give a short succinct context to situate this chunk within the overall document \ for the purposes of improving search retrieval of the chunk. Answer only with the \ succinct context and nothing else.""" def situate(doc: str, chunk: str) -> str: resp = client.messages.create( # One Claude call per chunk. For a large index you'd typically drop to a # cheaper model like claude-haiku-4-5 — which is what Anthropic's ~$1.02 / # 1M-doc-token estimate assumes — trading a little context quality for cost. model="claude-opus-4-8", max_tokens=200, messages=[{ "role": "user", "content": [ # The whole document is the stable prefix: cache it ONCE, then every # chunk of this document reads from the cache instead of re-paying for it. {"type": "text", "text": f"\n{doc}\n", "cache_control": {"type": "ephemeral"}}, {"type": "text", "text": CONTEXT_PROMPT.format(chunk=chunk)}, ], }], ) return "".join(b.text for b in resp.content if b.type == "text").strip() ``` Verify the cache is actually working: `resp.usage.cache_read_input_tokens` should be non-zero on every chunk after the first for a given document. If it's zero, something upstream is mutating the document bytes (a timestamp, non-deterministic JSON) and invalidating the prefix. ## A dependency-free repro To convince myself the lift was real and not marketing, I wrote a ~180-line pure-stdlib script: naive chunking, BM25 from scratch, a TF-IDF cosine as an embedding stand-in, and reciprocal rank fusion. The retrieval core is small — here is BM25 and the fusion step: ```python class BM25: def score(self, query, i): dl, tf, s = len(self.docs[i]), self.tf[i], 0.0 for t in query: if t not in tf: continue num = tf[t] * (self.k1 + 1) den = tf[t] + self.k1 * (1 - self.b + self.b * dl / self.avgdl) s += self.idf(t) * num / den return s def rrf(rankings, k=60): # reciprocal rank fusion of BM25 + embedding rankings scores = Counter() for ranking in rankings: for rank, item in enumerate(ranking): scores[item] += 1 / (k + rank + 1) return [i for i, _ in scores.most_common()] ``` Then I index the chunks two ways — plain, and with a one-line context prepended — and measure where the correct chunk lands for a couple of "which entity, which period" queries. The actual output: ```text query method plain rank ctx rank ------------------------------------------------------------------------------------ How did ACME Corp revenue change in Q2 2023? bm25 4 2 How did ACME Corp revenue change in Q2 2023? emb 4 1 How did ACME Corp revenue change in Q2 2023? hybrid 4 2 What happened to Beta Industries revenue in Q3 2023? bm25 5 2 What happened to Beta Industries revenue in Q3 2023? emb 5 1 What happened to Beta Industries revenue in Q3 2023? hybrid 5 2 recall@1 (fraction of queries where the right chunk ranks #1): bm25 plain 0/2 contextual 0/2 emb plain 0/2 contextual 2/2 ``` Contextualizing the chunks moves the answer from rank **4–5** to rank **1–2** across BM25, embeddings, and their fusion — and embeddings recall@1 goes from **0/2 to 2/2**. On this toy corpus fusion lands the answer at #2 rather than #1 (a small-N artifact of RRF), which is a good reminder that the fusion win is an *aggregate* effect — which is exactly what Anthropic's real evaluation measures. ## What the real evaluation found On Anthropic's benchmark (top-20 retrieval failure rate, i.e. `1 - recall@20`), stacking the techniques compounds: - **Contextual Embeddings** alone: `5.7% → 3.7%` — a **35%** cut in the failure rate. - **+ Contextual BM25**: `5.7% → 2.9%` — **49%**. - **+ reranking**: `5.7% → 1.9%` — **67%**. The lexical half matters more than you'd guess — BM25 nails exact identifiers (error codes, ticker symbols, function names) that embeddings smear together, so contextualizing *both* indexes and fusing them beats either alone. ## Things worth copying from the post - **Retrieve top-20, not top-5/10.** Anthropic found 20 the most performant cut for the final context. - **Rerank the shortlist.** Retrieve ~150 candidates, then rerank down to 20 for the answer prompt — that's the step that takes the failure rate from 2.9% to 1.9%. - **Embedding model matters.** Gemini and Voyage embeddings were the standouts in their tests. - **Chunking still matters.** Size, boundary, and overlap all move the numbers — Contextual Retrieval sits on top of good chunking, it doesn't replace it. - **A domain-tuned context prompt beats the generic one.** The template above is a floor, not a ceiling. Be honest about what this costs. Contextual Retrieval adds a Claude call **per chunk** at index time — cheap per token with caching, but real latency and spend when you're indexing millions of chunks, and it has to re-run when documents change. My repro is a minimal reproduction of the *core mechanism* (context lifts rank); the `35 / 49 / 67%` figures are Anthropic's, on their corpus. And my playground scores with a TF-IDF cosine, not a real embedding model — it shows the *shape* of the effect, not production numbers. ## The takeaway The move is almost embarrassingly simple: spend a cheap, cached Claude call per chunk to write down the context a human would need to make sense of it, then index that. It attacks the failure at its source instead of papering over it downstream, it helps lexical and semantic retrieval at the same time, and — as the little repro above shows — you can watch the right chunk climb the rankings the moment the context goes in. --- *Source: [Introducing Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval) (Anthropic). The prompt and the `35 / 49 / 67%` numbers are theirs; the repro and the browser playground are mine, and the playground's scoring runs entirely client-side.* --- # A 14B model that matches a 671B one — by knowing its domain > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/blog/qwen-bim-domain-beats-scale > date: 2026-06-09 > tags: llm, fine-tuning, domain-models, bim, paper-notes Here's the headline from [Qwen-BIM](https://arxiv.org/abs/2602.20812) (Lin et al., Tsinghua, Feb 2026): a fine-tuned **14B** model scores **0.83 on G-Eval** for BIM-based design tasks — essentially tied with **DeepSeek-R1 at 671B** (0.84), and ahead of its own 72B sibling. A model ~48× smaller, matching the frontier on a specific domain. That result is not surprising on its own — "fine-tune a small model on your domain" is folklore by now. What makes the paper worth reading is the *anatomy*: where exactly general LLMs fall over on engineering work, and which one design choice did most of the lifting. I work on industrial AI at [Inkers](https://inkers.ai), so domain models over 3D/BIM data are close to home. These are my notes. ## The actual problem: a BIM model isn't text A Building Information Model is a structured graph of components — walls, slabs, beams, each with geometry, materials, and relationships. An LLM can't read it. So step one of *any* LLM-on-BIM pipeline is an unglamorous one the field mostly skips past: **turn the model into text.** The authors do exactly that, carefully. Revit models of five building types (malls, offices, dormitories, teaching buildings, museums) are sliced into spatial blocks of ~10–15 components each, defects are injected, and each block is serialized to plain text plus 22 templated questions with **hard-coded reference answers**. That last detail matters: the ground truth is computed by rules, not by another model, so the benchmark isn't measuring one LLM against another LLM's opinion. The questions ladder up in difficulty on purpose: from "list the wall IDs" (extraction) through "compute each slab's area" (calculation) to "is this wall thickness suspicious given residential norms?" (domain reasoning). It's a clean way to see *which rung* a model falls off. The whole pipeline, end to end, is just: project the structured model into text, turn it into supervised Q&A, add reasoning traces, and LoRA-fine-tune a small open model on it. {[ { x: 8, t1: "BIM model", t2: "Revit graph" }, { x: 132, t1: "textualize", t2: "+ 22 questions" }, { x: 256, t1: "BIM-QA", t2: "2,129 pairs" }, { x: 380, t1: "BIM-QRA", t2: "1,364 + reasoning" }, { x: 504, t1: "Qwen-BIM", t2: "LoRA · 14B" }, ].map((n, i) => ( {n.t1} {n.t2} {i < 4 ? : null} ))} structured 3D → text → supervised data → small fine-tuned model ## Where general LLMs actually fail They evaluated 11 general models (ChatGLM, Qwen, DeepSeek — including the 671B DeepSeek-V3/R1). The failure modes are specific and, honestly, familiar from any engineering-LLM project: - **Arithmetic.** Asked for a slab's planar area, Qwen-max picks the *right formula* and still returns the wrong number. The bottleneck isn't understanding — it's calculation. - **Natural-language literalism.** Models misread parentheses in the answer template, or a naming rule ("wall IDs start with Q"), and confidently apply the wrong transform. - **Missing domain knowledge.** Asked to infer a building's floor height, the 14B base model reasons that floor height ≈ slab thickness (120 mm) — coherent chain of thought, wrong mental model, because it was never taught what "floor height" means in practice. The pattern: general models clear extraction and counting, then degrade sharply on calculation, multi-step reasoning, and anything needing design common sense. On the domain-specific design-review tasks, G-Eval is mostly **below 0.8** — not reliable enough to trust. ## The one choice that mattered: reasoning supervision This is the part I'd underline. They built two datasets from the same BIM text: - **BIM-QA** — 2,129 plain question→answer pairs. - **BIM-QRA** — 1,364 question→**reasoning**→answer triples, where the intermediate steps are supervised, not just the final answer. Then they LoRA-fine-tuned Qwen2.5-14B on different mixes. The result is the kind of finding that should change how you build these datasets: | Fine-tuning data | Size | G-Eval | |---|---|---| | 100% QA | 2,129 | 0.69 | | 80% QA + 20% QRA | 2,661 | 0.77 | | 60% QA + 40% QRA | 2,500 | 0.77 | | **100% QRA** | **1,364** | **0.83** | The **smallest** dataset — pure reasoning triples — won, by a wide margin. More reasoning supervision monotonically improved G-Eval, and quality beat quantity outright. Teaching the model *how to get there*, on a third of the data, beat teaching it *what the answer is* on the full set. ## Bigger is not better (and the paper shows it twice) Two clean data points against scale-maximalism: 1. On the general benchmark, **QwQ-32B out-scored DeepSeek-R1 (671B)** on G-Eval. The giant model's verbose reasoning actually *hurt* — it padded answers with redundant text, tanking format and text-similarity scores without improving correctness. 2. After fine-tuning, **Qwen-BIM (14B) matched DeepSeek-R1 (671B)** and beat the 72B and 32B Qwen models on the domain G-Eval. {/* y gridlines at 0.5..0.9 — chart area y 20..200 maps score 0.9..0.5 */} {[0.5, 0.6, 0.7, 0.8, 0.9].map((v) => { const y = 200 - ((v - 0.5) / 0.4) * 180 return ( {v.toFixed(1)} ) })} {[ { label: "base 14B", sub: "Qwen2.5", score: 0.69, hl: false }, { label: "Qwen-BIM", sub: "14B · fine-tuned", score: 0.83, hl: true }, { label: "DeepSeek-R1", sub: "671B", score: 0.84, hl: false }, ].map((b, i) => { const x = 96 + i * 150 const y = 200 - ((b.score - 0.5) / 0.4) * 180 return ( {b.score.toFixed(2)} {b.label} {b.sub} ) })} The improvement from fine-tuning is also *targeted* exactly where you'd want it: | G-Eval | Base 14B | Qwen-BIM | Δ | |---|---|---|---| | General tasks | 0.810 | 0.874 | +0.06 | | Domain-specific tasks | 0.588 | 0.801 | **+0.21** | | Overall | 0.689 | 0.834 | +0.15 | General ability barely moved (it was already fine); the entire gain is concentrated in the domain tasks that were broken. That's the signature of fine-tuning doing the right thing — adding domain competence without trading away the base model's generality. ## What I'd flag It's a careful paper, but keep the scope honest: - **2D only.** Early tests showed the models couldn't do 3D geometry (collision detection), so those questions were cut. The hard part of real BIM reasoning is 3D. - **One narrow task family**, five building types, rule-generated Q&A. G-Eval is an LLM-as-judge metric — better-correlated with humans than BLEU/ROUGE here, but still a proxy. "Data available on request" rather than released. - The "matches 671B" comparison is on *this* benchmark. It's a domain-competence claim, not a general-capability one. ## Why it's the right playbook anyway Strip away the BIM specifics and this is a template for industrial domain models, the kind I think about constantly: you rarely need a frontier model. You need (1) a faithful **text projection of your structured/3D data**, (2) a benchmark with **rule-computed ground truth** so you're measuring competence, not vibes, and (3) **reasoning-supervised** fine-tuning data — quality and chain-of-thought over raw volume. Get those three right and a 14B model on two A6000s reaches the same place a 671B model does, at a fraction of the inference cost. For anyone shipping AI into a real engineering vertical, that economics is the whole game. --- *Paper: [Developing large language model for BIM-based design with domain-specific benchmark and dataset](https://arxiv.org/abs/2602.20812) — Lin, Cai, Ni, Zhou, Pan (2026), arXiv:2602.20812.* --- # This site is managed by Claude > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/blog/hello-world > date: 2026-06-03 > tags: meta, ai, nextjs GitHub READMEs are dead. After Claude and the wave of coding agents, your homepage isn't a static profile — it's a living, agent-readable artifact you can hand to an LLM. This site is **dual-native**: every page is both a clean human document and a machine-readable surface. Try fetching [`/blog/hello-world.md`](/blog/hello-world.md) or [`/llms.txt`](/llms.txt) — an agent gets structured text, you get the rendered page. The whole site is maintained by a crew of Claude agents. New posts, logs, and data updates are authored by skills that validate themselves before shipping. ## What's under the hood The content layer is a single source of truth: MDX files validated with Zod, surfaced identically to humans (this page), to agents (the `.md` variant), and to tools (the MCP server). More on that soon. --- # 153 autonomous runs, no new ideas: the nanoGPT speedrun frontier > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/nanogpt-speedrun-frontier > date: 2026-08-15 > tags: agents, evaluation, prime-intellect, research, benchmarks, explainer [Measuring Autonomous AI Research](https://www.primeintellect.ai/blog/measuring-autonomous-research) (Elie Bakouch, Prime Intellect, 14 August 2026) is the largest public experiment of its kind I have seen: **153 autonomous runs across 18 frontier models**, each on its own 8×H200 node, running unattended for up to eight days. The task is [modded-nanoGPT](https://github.com/KellerJordan/modded-nanogpt) track 3 — the optimizer speedrun. Train a 124M GPT to validation loss 3.28 in as few optimizer steps as possible. You may edit the optimizer, its hyperparameters, the schedule and the initialization. The dataloader, architecture, batch size, sequence length and data are frozen. For scale, the post notes the comparisons: Anthropic's internal automated-R&D evaluation optimizes a model on a *CPU node*, and OpenAI's GPT-5.6 Sol system card reports nanoGPT Track 1 on a single H100 for under a day. This is a much larger instrument than either. It also produced a negative result, and says so plainly. ## The scoreboard, and what it means The metric is honest and easy to check. The tuned baseline the agents start from passes at **3,290 steps**; the human record claim sits at **2,600**. So there are 690 steps on the table, and "gap closed" is just `(3290 − record) / 690`. Every published percentage reproduces from that formula exactly. Three things are worth saying about the shape of it. **Nobody beat the human.** Best run: Fable 5 at 2,726 steps, 81.7% of the gap, after 8.7 days. The remaining 18.3% is where a human already is, and had already been for weeks. **Nobody invented anything.** This is the post's own summary, and it is the sentence I would lead with: > None of the runs produced a fundamentally new method; the winning ingredients are all similar to existing ones in the literature. The improvements that win are known optimizer work — better preconditioning, caps and floors on update magnitudes, keeping the learning rate hot longer, weight averaging near the end. Given that the agents had **no internet at all** — a deliberate change from earlier experiments, where they over-anchored on existing PRs — rediscovering the literature from parameters alone is a real result. It is just not the result the phrase "recursive self-improvement" is usually deployed to suggest. **The cost column reorders everything.** Flip the interactive to tokens-per-step-gained and the ranking falls apart. Grok 4.5 bought its steps at 0.27M tokens each and closed a quarter of the gap; GPT-5.6 Sol paid 11.7M per step for a third of it — a 43× spread that says nothing about rank. The model that looks best on both axes is **Opus 5**: second place, in under three days, at 0.49M tokens per step. Fable 5 wins outright but spends nearly three times that rate and takes 8.7 days to do it.
## The benchmark is a statistics exam Here is what makes this task harder than it sounds, and it is all in the public rulebook, [`program.md`](https://github.com/PrimeIntellect-ai/frontier-automated-speedrun). A record requires the mean of eight fixed seeds — `0xC0FFEE+0..7`, which the agent cannot touch — to come in below **3.27859**. The file derives that number itself: `3.28 − 0.004/√8`, described as one-sided p < 0.001 at a per-run σ of about 0.0013. The arithmetic checks out. The standard error of an eight-run mean is `0.0013/√8 = 0.00046`, and `3.28 − 3.09 × 0.00046 = 3.27858`. Now look at what that does to the research loop. Per-run σ is 0.0013 — *larger than most of the improvements being hunted*. A single screening trial cannot tell you much of anything, and the two ways to get it wrong pull in opposite directions: - Trust one run and you certify noise. A recipe with **no real gain at all** clears the bar on a single trial about 14% of the time. - Distrust one run too hard and you throw away the thing you were looking for. At one trial each, a recipe that genuinely is 0.001 better *loses* its head-to-head roughly 29% of the time. And every trial is real money: a run takes the whole 8-GPU node, so runs are strictly sequential. Deciding when to widen from one seed to three to eight *is* the research skill this benchmark measures. The blog says as much: > The models all find similar ideas. What separates them is how they run experiments. The failure modes it describes in the weaker models are all statistical, not intellectual: killing whole families on one seed, treating their own crashes as evidence the idea was bad, discarding small gains that don't clear the bar alone. Grok 4.5 lost row normalization twice — to its own scaling bugs, not to the method. ## The best thing in the post is one paragraph long Prime Intellect put a noise estimate in `program.md` that was **deliberately slightly too large**. Then they counted who checked. 62 of roughly 100 runs measured the noise themselves instead of trusting the number they were handed — and those runs are concentrated at the top of the table. 42 went further and discovered something nobody had mentioned: rerunning the same recipe on the same seed *also* moves the loss, because GPUs are not deterministic. That residual is much smaller than seed-to-seed variance, so two recipes compared on a shared seed resolve differences a normal screen cannot, for identical compute. Several models rebuilt their screening protocol around it. That is a beautifully cheap instrument. It is not a coding test or a knowledge test — it measures whether an agent treats its documentation as evidence or as a claim, and it costs nothing but a willingness to write down something untrue. I would like to see more evaluations do this, and I suspect it generalizes far past optimizer research. ## The harness is worth as much as the model Kimi K3 appears twice in the table under two different harnesses, which makes it the closest thing here to a controlled comparison — and the gap between its two runs is larger than the gap between several adjacent *models*. Under [Prime Agent](/articles/prime-agent), which hands the model a persistent IPython kernel instead of a tool menu, K3 reached 2,930 steps on 112M tokens and 488 tool calls. Under `kimi-code` it reached 2,974 on 682M tokens and 4,000 calls. Better record, **6.1× fewer tokens**, an eighth the tool calls — and *more* output tokens, which is the tell. It was writing programs, not issuing commands. The traces show what that looks like: K3 built its own experiment driver, a loss-curve comparator, a routine to restore a clean baseline, and then a numerical laboratory for retuning Newton-Schulz coefficients — testing them in simulation before spending a GPU-hour, and revising its hypothesis when the theoretically cleaner update trained worse. This is the same [harness effect](/articles/harness-effect) that keeps showing up: the scaffold is not packaging around the model, it is part of the system being measured. One caveat the blog does not foreground and the repository README does: that Prime Agent run is tagged **serial era**. It ran under `program-serial.md`, a variant used between 20 July and 13 August that made agents wait on each run instead of delegating to a subagent. So two things differ, not one. Five of the twenty rows carry that tag — including second and third place — and Prime Intellect says they are being rerun. ## The tension I keep circling The post credits the top models with research taste: re-ablating the stack after every merge, dropping components that stopped helping, revisiting old negatives when the recipe changed. Opus 5 re-opened β2 tuning under a new recipe and it became a record. K3 deleted two mechanisms that had produced its previous record once a new normalization made them redundant. Fable, out of single-knob gains, started testing pairs that were individually worse but jointly better; one late re-probe was worth 31 steps. Those are genuinely good research instincts. They are also, in part, **instructions**. From `program.md`: > Roughly every ~8 ideas explored, do a pruning round: try dropping each component you've stacked on and keep only what still earns its place. And: > A better method than the baseline exists (the human frontier is well below it), so "no improvement found" / "baseline is optimal" is never a valid place to stop. The rulebook tells every model to prune periodically and forbids all of them from concluding they are done. So some unknown share of what is being scored as taste is compliance — following a written procedure under fatigue, across days, without a human checking. That is a real and valuable capability. It is just a different one, and the experiment as designed cannot separate them. The clean version of this study gives half the runs a rulebook with those two paragraphs removed. ## What it does not establish The authors are candid about most of this, which is why the post is worth reading in full. **Variance is high.** Speedrun noise plus model-level randomness on a days-long process; they mitigate with at least three seeds per model, taking the best after 24 hours and continuing it. That is a sensible protocol and it is also a best-of-k selection, so single-model numbers carry more optimism than a single run would. **The task may not transfer.** Their words: they "don't have strong conviction that methods developed in this kind of speedrun are inherently scalable or would be used in real model training." **The records are partly reconstructions.** The repository builds each model's record PR from the state recovered from its traces; Muse Spark 1.1 reached 3,232 steps but its exact record file could not be reconstructed, so it has no PR at all. The README's leaderboard also lists Kimi K3 at 2,968 steps where the results site shows 2,930 and 2,974 for its two runs — a discrepancy that is probably "last record, not best," but is not explained anywhere I could find. **Comparing across harnesses is comparing systems, not models.** Every row pairs a model with a specific scaffold at a specific effort setting, and the K3 pair shows how much that matters. ## Why it is still the right experiment None of that undercuts the main thing. Claims about models doing autonomous research have gotten much louder than the evidence, and almost all of the evidence has been either private or tiny. This is 153 runs on real GPUs for real days with the traces, scratchpads, monitor reports and rulebook all published, and the headline finding is *modest*: the best model closed four fifths of a gap a human had already closed, using ideas that were already in the literature, with no internet to look them up. The honest way to read the table is as a measure of experimental discipline under uncertainty — screening cheaply, widening on signal, resisting a conclusion the data can't support, re-testing what you already decided. Which, now that I write it out, is a fair description of what makes a human researcher good too, and a much better thing to be measuring than whether the model can name the trick. --- # Qwen3.8, weights in hand: 98% of a 2.4T model is routed experts > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/qwen3-8-open-weights > date: 2026-08-15 > tags: qwen, open-weights, moe, quantization, architecture, inference, vllm When [Qwen3.8-Max was announced](/articles/qwen3-8-max) on 3 August, the open weights were a promise: 2.4 trillion parameters, 95B active, "next week." The [Qwen3.8 collection](https://huggingface.co/collections/Qwen/qwen38) is that promise landing, and it is now four repositories deep. Which means the interesting work has changed. There is no technical report, and there probably will not be one. But there is a `config.json`, a weight index, a chat template, an FP8 exclusion list, two serving recipes and a genuinely rigorous third-party quantization study. That is more than enough to check the claims, and checking them turns up several things the model cards do not say. ## What actually shipped | repository | params | license | created | downloads | likes | |---|---|---|---|---|---| | [Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) | 2.4T / 95B active | **`qwen3.8-max`** | 8 Aug | 6.4k | 949 | | [Qwen3.8-2.4T-A95B-FP8](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8) | same, FP8 | **`qwen3.8-max`** | 8 Aug | 10.7k | 191 | | [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) | 27B dense + vision | **Apache-2.0** | 5 Aug | 91.9k | 9.5k | | [Qwen3.8-27B-FP8](https://huggingface.co/Qwen/Qwen3.8-27B-FP8) | same, FP8 | **Apache-2.0** | 13 Aug | 123k | 381 | Two things jump out of that table before any architecture. **The licenses are not the same.** The 27B is plain Apache-2.0. The 2.4T ships under a bespoke `qwen3.8-max` license that is MIT-shaped with two riders: products above 100M monthly actives or \$20M monthly revenue must display the model name in their UI, and anyone running a "Model as a Service or AI Work Assistant business" whose group revenue passes **\$50M over any twelve months** needs a separate commercial license from Qwen. Internal use is carved out explicitly, as long as you do not expose the model or its outputs to third parties. It is a reasonable license and it is not an open-source one, and "the first Qwen-Max-class model getting open weights" deserves the asterisk. **The 27B is the release.** It has fourteen times the downloads and ten times the likes of the flagship, and it went up three days earlier. Note also that in both pairs the FP8 repo out-downloads the bf16 one while collecting a fraction of the likes — bf16 is what people bookmark and requantize from, FP8 is what they actually serve. ## Reading the architecture out of the files Both models are the same design at two scales, and Qwen describes the stack in one line on each card: > Hidden Layout: 23 × (3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE)) The most useful thing you can do with a model card number is try to rebuild it. If the reconstruction lands, you understand the architecture; if it does not, you have found something. It lands. Summing the config — 92 layers of 512-expert MoE, 23 gated-attention layers, 69 Gated DeltaNet layers, two untied embedding matrices and one MTP block — gives **2.446181T** against the weight index's **2.446183T**. The 1.6M-parameter residual is the layernorms, which I did not bother to count. Three things fall out of the exercise that no card mentions: **The 95B active figure needs the embeddings.** The compute path — routed experts, shared expert, router, both attention types — comes to 91.2B. You only reach 95.3B by counting the untied `embed_tokens` and `lm_head`, 2.03B each. That is a defensible convention, but it is a convention, and it is 4% of the headline. **`q_proj` is twice as wide as you would guess.** `head_dim` is 256 with 64 query heads, so 64 × 256 = 16,384 — already 2× the 8,192 hidden size. But the actual tensor is `[32768, 8192]`, twice that again, because `attn_output_gate: true` fuses the output gate into the same projection. The attention block is genuinely wider than the residual stream it reads from, in both directions. **The MTP block costs 26.4B parameters.** The multi-token-prediction head is not a small linear probe. It is a complete extra layer — its own gated attention, its own 512-expert MoE, plus a fusion projection — weighing 1.08% of the model. That is a larger draft model than most models. Whether it earns that is a question the serving recipe answers below, and the answer is "only at depth 3." ## The hybrid, and what it is supposed to buy `full_attention_interval: 4`, so three Gated DeltaNet layers then one gated attention layer, all the way up. The DeltaNet layers are Mamba-shaped — `A_log`, `dt_bias`, a kernel-4 depthwise `conv1d`, and a fused `in_proj_qkv` that carries 16 QK heads and 128 V heads at head dim 128 (`[20480, 8192]`, which is exactly 16·128 + 16·128 + 128·128). They keep a fixed-size recurrent state. They do not keep a KV cache. That is the entire pitch: at 256K context the 23 attention layers of the 2.4T want a KV cache that grows linearly, and the other 69 layers contribute a constant. Except the advertised saving is only real if your runtime knows about it. The most careful GGUF publisher for the 27B quotes **256 KB of attention cache per token**, and 2 GB at 8K. The sixteen full-attention layers in that model need 2 · 4 heads · 256 dims · 2 bytes · 16 layers = **64 KB per token**. The quoted figure is exactly 4× that — and 4 is the hybrid interval, i.e. precisely what you get if every layer is given a cache. I have not read llama.cpp's allocator for the `qwen35` architecture, so I will not tell you which of "allocation detail" and "deliberate margin" it is. I will tell you it is worth checking on your own hardware before you size a card, because the difference is 1.6 GB at 8K and 6.4 GB at 32K. ## The 27B is the interesting model Sort the 27B's benchmark table by what each row measures and a clean pattern appears that Qwen does not point at. On anything agentic, the 27B beats **Qwen3.7-Plus** — a larger model from the previous generation — on all thirteen rows, mean margin +10.9. Several margins are not subtle: OSWorld-Verified 84.3 against 73.3, Vision2Web 62.9 against 42.1, RecreationBench 47.1 against 30.2. DeepSWE 1.1 goes from 14.2 to 42.2, a three-fold jump that scale does not explain and that reads like a benchmark the training mix learned to do. Flip to the rows where the answer has to already be in the weights and it loses five of seven — GPQA Diamond, HLE, ERQA, RealWorldQA, OmniDocBench — with ERQA down 4.3 and HLE down 3.9. That split is the most useful finding in the release. **A generation of post-training bought an enormous amount of doing and almost no knowing.** Which is roughly what you would expect, and it is still worth seeing measured: if your workload is agentic, a 27B from this generation genuinely substitutes for something much larger; if your workload is recall, it does not, and no amount of harness will fix that. The number I would treat most carefully is QwenSWEBench, where the 27B scores 79.0 against the 2.4T flagship's 80.7. A dense 27B landing within 1.7 points of a 2.4-trillion-parameter model is an extraordinary claim, and it is on Qwen's own benchmark, run by Qwen, against models Qwen did not train. The three `Qwen*Bench` rows should be read as internal instrumentation, not as evidence. ## `reasoning_effort` is two sentences Both cards advertise "official support for `reasoning_effort`" as a headline feature. It is implemented in the chat template, and you can read the whole implementation: ```jinja {%- if resolved_reasoning_effort == 'xhigh' %} {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %} {%- elif resolved_reasoning_effort == 'low' %} {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %} {%- endif %} ``` That is it. `medium` sets nothing at all — it is the untouched model, and `xhigh` and `low` are two English sentences prepended to the system message. There is no token budget, no separate decode path, no architectural switch. This is not a criticism: the model was presumably post-trained to respond to those exact strings, which is what makes it "official" rather than a prompt you invented. But it has consequences worth knowing. - The effort level lives in **prompt space**, competing with your own system prompt for attention. - Any harness that replaces the system message silently drops it. - You can replicate all three levels, or invent new ones, with a string. Two more control details: **`preserve_thinking` defaults to on**, and it means every prior assistant turn keeps its full `` block in context. Combined with Qwen's own recommendation to allow 262,144 tokens of reasoning per turn, an agentic loop can spend its context window on its own history of deliberation faster than you expect. Setting it false strips reasoning from all turns before the last user message. **The 2.4T cannot stop thinking.** Its template raises outright: `Disabling thinking is not supported.` The 27B accepts `enable_thinking: false` and emits an empty `` pair. If you were planning to use the flagship for anything latency-sensitive, that is a design constraint, not a setting. Tool calls also moved off JSON to an XML-ish form — `` — which reads oddly until you notice it means multi-line code payloads need no escaping at all. That is a real improvement for a coding agent, and it explains why quantizers are shipping "tool calling improvements" notes. ## What FP8 actually quantizes Open `quantization_config` on the 2.4T FP8 checkpoint and read `modules_to_not_convert`. It spares: - every attention projection — `q_proj`, `k_proj`, `v_proj`, `o_proj` - every Gated DeltaNet projection — `in_proj_qkv`, `in_proj_z`, `in_proj_a`, `in_proj_b`, `conv1d`, `out_proj` - the shared expert, all three projections, plus the router and the shared-expert gate - `lm_head` and `embed_tokens` - the entire MTP block What is left is the routed experts, and the routed experts are **97.97% of the model**. So the FP8 checkpoint is not "the model in FP8." It is *the experts in FP8 and the model in bf16*, which happens to look the same from a distance because the experts are almost all of it. The arithmetic confirms the reading. Take 2.3966T routed parameters to one byte, leave the remaining 49.6B at two, and you predict **2.270 TiB**. vLLM's recipe publishes the FP8 checkpoint at **2.27 TiB**. (The bf16 figure checks too: 4.450 TiB reconstructed against 4.45 TiB published.) The 27B FP8 config makes the same choice at a smaller scale — the GDN gating path, both embedding matrices, every layernorm and the entire vision tower stay bf16 — with one artifact worth a chuckle: its exclusion list names `mlp.gate` and `mlp.shared_expert_gate`, tensors that do not exist in a dense model, along with a fused `in_proj_ba` that is not in the checkpoint either. Harmless, and clear evidence both configs came off one template. ### Three parties, one conclusion Here is the finding I would actually carry away from this release, because it arrives from three directions that did not coordinate: 1. Qwen's **2.4T FP8** config refuses to quantize any Gated DeltaNet projection. 2. Qwen's **27B FP8** config refuses to quantize the DeltaNet gating path. 3. A third-party quantizer, measuring rather than guessing, found that lifting `in_proj_z` and `out_proj` by one precision step cost 0.16 GB and removed **11% of the remaining divergence** — the single best trade in their whole search. In a hybrid GDN/attention model, the linear-attention path is the precision-critical part. If you are building your own quantization mix for this architecture, that is where the bits go. ## Serving it Both [vLLM recipes](https://recipes.vllm.ai/Qwen/Qwen3.8-27B) are unusually candid, and three of their findings generalize. **MTP depth 3, not 1.** MTP-1 measured **64.8% acceptance** and is not merely marginal — it is *negative* at scale: +3.4% at concurrency 1, −9% at 128, −23% at 256, because the draft pass displaces real work once the batch is compute-bound. Depth 3 is worth roughly 2.3× on per-user output rate (FP8/TP16: 130 → 307 tok/s/user). A 26.4B draft head is only worth its weight if you speculate deep enough to amortize it. **Context length is a concurrency dial.** At `--max-model-len 262144` the engine reserved KV for **25** concurrent requests. At 9,240 — 8K in, 1K out — the same 70 GiB of KV served **506**. Twenty times the concurrency from one flag, and nothing about the model changed. **Tensor parallel must divide 64 attention heads**, so only 1/2/4/8/16/32 are legal. The recipe walks through the consequence: FP8 needs 2,325 GiB, which is three GB300 trays by capacity, but TP12 is not a thing — so it is a four-tray, sixteen-GPU deployment. Capacity planning on this model is arithmetic on head counts, not on gigabytes. Smaller items worth knowing: `--load-format fastsafetensors --safetensors-load-strategy lazy` cut weight load from 545s to 306s on a 1.32 TiB checkpoint; MXFP4 does not load on NVIDIA (use NVFP4); the 1M-context `--hf-overrides` key nests under `text_config` for the 27B but sits flat for the 2.4T; and hybrid models have a CUDA-graph failure mode where `assert num_cache_lines >= batch` means your capture size exceeded the *recurrent-state* cache, which is a separate resource from the KV cache and one most people have never had to think about. ## Running it on your own machine
The GGUF ecosystem produced two serious repositories within a day of each other, and they are interestingly different. [unsloth](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) has 25 files, 868k downloads, and a card that says "Unsloth Dynamic V3.0 (preview) for SOTA quantization performance" with no measurement attached. [AtomicChat](https://huggingface.co/AtomicChat/Qwen3.8-27B-GGUF) has 16 files, 8k downloads, and a card that is essentially a small paper. The AtomicChat card measures per-token KL divergence against the **bf16** weights — not against `Q8_0`, which is the usual shortcut — publishes the reference logits so you can measure your own builds against the same point, downloads its competitors' files and measures those on the same harness rather than quoting their published numbers, and flags the one size band where it loses. That is the right protocol, and it produces four results worth keeping: - **Where the bits go beats how many there are.** Ten builds within one gigabyte of each other span 2.2× in divergence. Nothing changes but tensor assignment. - **The ends of the network matter most.** Peak activation energy sits on layers 52–62, with a second peak on layer 0. Lifting the first four and last twelve helped more than widening the band to 32 layers. - **`Q8_0` is not lossless.** 0.00064 divergence, 98.92% top-1 — it disagrees with the original on about one token in ninety-three. - **A quant name is not a specification.** Three publishers ship a `Q4_K_M` for this model: 16.8 GB, 19.0 GB and 17.1 GB, spanning 1.9× in divergence. There is also a nice architectural footnote: the MTP head never executes during a normal forward pass, so the importance matrix has nothing to say about it at any corpus size, and llama.cpp refuses to quantize it low rather than guess. It is pinned to `q5_k` in every file. **One practical warning.** AtomicChat's repository contains no `mmproj`. Qwen3.8-27B is a vision-language model — its HF pipeline tag is `image-text-to-text` — and those quants are text-only. unsloth ships `mmproj-F16.gguf` at 0.93 GB, which is exactly the 0.466B-parameter vision tower I reconstructed from the weight index. If you want the 27B to see, that file is not optional and only one of the two repositories has it. ## What this release does not establish There is **no technical report**. Every number in every table is Qwen's, produced on Qwen's harness, and three of the coding benchmarks are Qwen's own instrumentation. The 2.4T's headline claim — matching or beating Opus 4.8 and GPT 5.6 Sol on agentic coding — is now at least *checkable*, since the weights are public, but nobody has checked it yet. There is **no training detail at all**: no token count, no data mix, no post-training description, nothing about how MTP was trained "with multiple steps," and no ablation for any architectural choice. The 3:1 hybrid interval, 512 experts with 10 active, head dim 256, a 25% partial rotary factor — all are presented as facts about the artifact rather than as decisions with evidence behind them. And the thing I would most like to see is the thing least likely to arrive: an honest account of why DeepSWE 1.1 went from 14.2 to 42.2 in one generation. Three-fold jumps on a single benchmark, in a family where three of the benchmarks are the vendor's own, are exactly the results that deserve the most explanation and usually get the least. What is genuinely good here is how much of the release is *legible*. The parameter counts reconstruct. The FP8 size falls out of the exclusion list. The vision tower's size matches the mmproj byte-for-byte. `reasoning_effort` can be read in full in nine lines of Jinja. That is not nothing — it is the difference between a model you can reason about and a model you can only benchmark, and for a 2.4-trillion-parameter flagship it is more than we usually get. --- # Arcee's Open Models API: a model lab selling six models it did not build > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/arcee-open-models-api > date: 2026-08-14 > tags: inference, open-weights, pricing, agents, api Arcee's [Open Models API beta](https://arcee.ai/blog/open-models-api-beta) post is labelled a **1 min read**, and that is accurate. It announces that the API now serves models beyond Arcee's own Trinity family, lists six of them, gives a price table, and offers $5 in credits. Worth reading anyway, for one sentence and one table. ## A model lab watching what you pick The stated motivation is not the usual one: > It also helps us better understand why people choose a particular model for a particular task. When the API is used across our products, we can learn which models users prefer for a given task, and more importantly, *why* they chose them. > > Those insights will help us consistently develop and deliver Trinity models that are exceptional, diverse, and widely adopted. Arcee trains models. It is now serving DeepSeek's, Z.ai's, Moonshot's and Thinking Machines' — and saying plainly that a reason to do so is to learn where its own are not chosen. That is an honest description of an inference business as a research instrument, and it is unusual to see written down. Most labs that add competitors' models to their API describe it as customer choice and stop there. The reasoning is sound: a lab with no serving surface only learns about its models from benchmarks and complaints, while a lab that routes real workloads sees substitution behaviour — which model people reach for when the task is long, which when it is cheap, which when it has to be right. The launch catalog: - **Trinity-Large-Thinking** (Arcee's own) - **DeepSeek-V4-Pro (Preview)** and **DeepSeek-V4-Flash-Latest** - **GLM-5.2** — the base that [GLM-5.3](/articles/glm-5-3) was post-trained from - **Kimi-K3** - **Inkling-Small**, from Thinking Machines, which the post goes out of its way to call "another American lab advancing the frontier of open-weight models" That last aside is doing some positioning work: four of the six are Chinese labs, and Arcee names the American one specifically. ## The table is more interesting than the announcement Prices are per million tokens, and the number the list does not draw attention to is the **ratio between them**: | model | input | output | output ÷ input | |---|---|---|---| | deepseek-v4-flash-latest | $0.14 | $0.28 | 2.0× | | deepseek-v4-pro | $1.74 | $3.48 | 2.0× | | inkling-small | $0.50 | $1.20 | 2.4× | | trinity-large-thinking | $0.25 | $0.80 | 3.2× | | zai-org/glm-5.2 | $1.40 | $4.40 | 3.1× | | moonshotai/kimi-k3 | $3.00 | $15.00 | 5.0× | Both DeepSeek models charge exactly double for output. Kimi K3 charges five times. That spread matters because the workload this API is being pitched at — long-horizon agent work, launched the same day as [nac](/articles/nac) — has a token mix that shifts with the task, and the cheap-to-read model is not always the cheap-to-run one. On a read-heavy job the ordering roughly follows input price. On a generation-heavy one it stops doing so. Kimi K3's input price is 21× DeepSeek-V4-Flash's, but at 5M in and 25M out the actual bill is **50× higher** — the output multiplier widens the gap by more than double. And GLM-5.2 overtakes DeepSeek-V4-Pro on that same mix ($117 against $95.70) despite being the cheaper of the two to read. Arcee's own Trinity-Large-Thinking is priced to sit second-cheapest on input and to stay cheap on output — $0.25 and $0.80, undercutting Inkling-Small on both. For a lab measuring which model people choose, pricing its own model into the "obvious default" slot is a thumb on the scale worth noting when reading whatever conclusions come out of the experiment later. ## Launched alongside nac The post is explicit that this ships the same day as [nac](/articles/nac), Arcee's open-source agent harness, and that the two are meant to inform each other: > We built nac to support demanding agentic workloads that may run for extended periods, and over time, what we learn from nac will help us improve how the API routes, serves, and supports models for long-running tasks. The pairing is the actual strategy. nac is Apache 2.0 and free; it is also a very good instrument for observing long-horizon agent workloads, because its architecture forces every unit of work through a named dispatch with a recorded episode. An orchestrator that plans in one model and dispatches workers to another is a natural place to learn which models are chosen for which kind of step. One detail from the nac repository suggests the catalog is not settled: the most recent commit at the time of writing is *"drop minimax, trinity-mini, and trinity-large-preview from arcee register."* Three models removed from the client-side catalog on launch day. ## What this is not It is not a technical post. There is no routing architecture, no latency or throughput figure, no serving stack detail, no context-length or rate-limit table, no availability or region information. "Beta" is doing real work in the title. There is also no evaluation of any of the six models, which is a slightly odd absence given the stated purpose is to learn which is best for what. The learning is planned to come from usage, not from measurement — which is a legitimate choice, and one that only produces useful answers if the pricing does not distort the selection it is measuring. Read it for the strategy sentence and the ratio column. The rest is a price list. --- # DeepSeek Harness: an agent harness that refuses to send what it didn't log > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/deepseek-harness > date: 2026-08-14 > tags: agents, harness, open-source, architecture, typescript, explainer [deepseek-harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`) is an MIT-licensed agent harness from DeepSeek that you run with `npx @deepseek-ai/dsh web`. The npm package was first published on 2026-08-10; six versions shipped in the four days after that, the latest being `0.1.0-rc.6`. It is labelled a developer preview, and the README is blunt about what that means: **THERE WILL BE COMPATIBILITY-BREAKING CHANGES.** The obvious thing to write about an agent harness is its agent loop. That turns out to be the least interesting part of this one. The loop is about what you would guess — claim input, assemble a prompt, call a model, run the tools it asked for, repeat while anything is owed. What is unusual is everything built *around* the loop to make it hold still, and the reason that machinery exists is legible in the commit history: the repository went from its first commit to that npm release in **61 days**, taking **12,293 commits across 65 active days** on the way. At least 209 of its 984 merged pull requests came off `codex/*` branches, which is a floor rather than a count. That combination — a codebase moving faster than humans can review, and a product whose whole job is to be trustworthy about what it told a model — produced a design decision worth stealing. ## Everything is a plugin, and it means it `dsh` is built on [Cordis](https://github.com/cordiverse/cordis), a plugin framework DeepSeek vendored into the repo at v4.0.1 and rescoped under its own namespace. Cordis is five ideas: a plugin contributes services to a shared context; a service claims a stable key like `ctx.tools` or `ctx.llm`; plugins declare what they need with `inject` rather than being boot-ordered by hand; communication is typed events; and every registration is a reversible effect that unwinds when its plugin unloads. The architecture doc states the consequence directly, and unlike most claims of this shape it survives contact with the source: > There is no privileged core to patch: you extend dsh by mounting a plugin beside the others. The model adapter is a plugin. The tool registry is a plugin. The session log is a plugin. The agent loop is a plugin — `core/agent` owns the `Agent` interface and the live registry, while `core/agent-loop` is described as "the default driver implementing that interface." Swapping it is a config row, not a fork. There are **219 workspace packages** under `packages/*/*`. Twenty-one of them are model-facing tools (`tool-bash`, `tool-fs`, `tool-lsp`, `tool-subagent`, `tool-terminal`, and so on). The rest are seams, providers, UI surfaces, and policy. A running `dsh` is composed at boot from ordered layers: bundles stack in listed order, then the profile's own patch file, then the home-level one, then any `--patch` overlay. I parsed the three committed bundle patches to see what that actually produces. The detail I found convincing is base's own header comment, which explains why a row whose value differs between modes is *not allowed* to live in base: a patch replaces a row's whole `config` rather than merging into it, so a mode-varying row belongs to each mode bundle, which restates it completely. That is a rule written down because someone expected agents to add rows to this file. ## One turn A **step** is one model request plus the tools it calls. A **turn** is zero or more steps: it opens before its first input is claimed and closes once nothing is owed. Here is the repo's own flow block, verbatim: ```text turn/start claim next-step input plus one queued message assemble prompt sections + tool schemas -> agent/pre-step reject | enter(messages) reject, or a first enter rewritten empty -> close the turn with no step step/start append entered messages as user/message derive model history from the log agent/request -> llm/stream -> assistant/chunk* -> assistant/message tool/call* -> tools/pre-execute -> tools/execute -> tools/post-execute -> tool/result* step/end tools owe another request, or next-step input arrived -> claim -> next step -> agent/turn-stopping turn/end ``` Two kinds of thing are interleaved there. Some events are durable facts appended to the session log; the rest are live extension points, and most of those are around-middleware — a listener receives `next()`, and either wraps the call and delegates or owns the decision and returns without delegating. The repo draws the same lifecycle as a sequence diagram, which adds what a linear list cannot: who talks to whom, and the branches. Both `alt` blocks are worth reading — a rejected pre-step leaves the turn open having spent no step, and a terminal request failure routes to an `agent/request-error` waterfall that returns a retry action or preserves the original error.
Note the line `derive model history from the log`. It is doing more work than it looks like. ## The check that makes the log the source Most harnesses treat the transcript as a rendering of the conversation: the conversation lives in memory, and the log is written alongside it for display and debugging. `dsh` inverts this. The log is the source, model history is *projected* from it by `deriveMessages()`, and a runtime invariant refuses to let those two drift apart. The whole of `packages/core/agent-loop/src/invariant.ts` is 63 lines. This is its core: ```ts ctx.on('llm/stream', (options: GenerateOptions, next) => { if (!isAgentLoopRequest(options)) return next() if (!Object.isFrozen(options)) fail('a loop-built request must be frozen') // ... const expected = session.deriveMessages() if (JSON.stringify(options.messages) !== JSON.stringify(expected)) { fail(`llm request for session "${String(session.id)}" diverges from the dispatch-time durable derivation (log-reconstruction desync)`) } // ... and the folded request header must match model, system, temperature, // maxTokens, stop and tools return next() }, { global: true, prepend: true }) ``` Every request the loop builds is compared, byte for byte through `JSON.stringify`, against a *fresh replay of the session log made at dispatch time*. If a plugin slips an extra message into the outgoing request without writing a session event for it, the request does not go out. It throws. The `prepend: true` matters: it means a replay or mock listener that short-circuits the waterfall still cannot get in front of the check. The failure this prevents is the quiet one. Injecting an unlogged message doesn't crash anything and usually makes the model behave *better* — it is exactly the sort of change that ships. What it destroys is reproducibility: from then on, the log no longer explains the answer, and "why did it do that?" has no reachable answer. The repo states the rule as **model-visible ⟺ logged**, and this is the line of code that makes it true rather than aspirational. The same header check covers sampling settings, which I think is the sharper half. Temperature and tool schemas are part of what makes a run reproducible, so retuning one between the logged header and the actual call is treated as divergence rather than a tweak. ## The same idea, applied to tools Tool execution gets the same treatment, and the repo's own pipeline diagram is the clearest statement of it. Two details in there are the log-is-source rule again, wearing different clothes. `tool/call` is **logged before execution** — not after, not on completion. If the process dies mid-tool, the log still records that the call was attempted. And at the far end, `tool/result` is labelled *single model-facing outcome*: however the call actually went — denied by a guard, refused at the approval prompt, thrown inside the tool body, thrown by a wrapper, timed out — every path converges through registry normalization and `finalizeContent` into exactly one recorded result.
Note the dotted `throw` edges all landing on the same normalization box. A tool that raises does not produce a missing result; it produces an `isError` result that the model sees and the log records. That is what lets the invariant in the previous section hold for a turn where something went wrong, which is the only kind of turn where reproducibility actually matters. ## 219 invariant companions, and what they actually contain Here is where I nearly published something wrong. Every one of the 219 package directories contains exactly one `src/invariant.ts`. I checked the correspondence as a set difference in both directions: zero packages without one, zero orphans. My first instinct was to write that as "219 packages all enforce runtime invariants." They don't. Only **35** of the 219 ever call `fail(...)`. The other 184 are 20-line files whose install function is empty: ```ts /** No runtime invariant: this stateless seam owns types while implementations enforce immutable-store checks. */ const install: InvariantInstaller = () => {} ``` That looked like ceremony until I read `scripts/package-invariants.ts`, which is what enforces the convention. A package missing its companion is a violation. An empty install function that does *not* carry a comment beginning `No runtime invariant:` is a violation. A non-empty install function that never uses its bound failure reporter is a violation. So the number that matters is not 219 checks. It is **219 decisions** — every package has been made to answer "what runtime invariant do you own?", and 184 of them answer "none, because…" in a sentence a reviewer can disagree with. Absence is recorded rather than assumed. That is a much better idea than 219 checks would have been, and it is the kind of thing that only pays off at this repo's scale. ## The repo is built by the workflow it ships The discipline makes sense once you look at how the code got written. `.agents/notes/` holds **686 design notes** — 507 implemented, 143 archived, 25 proposed, and 11 rejected, kept deliberately as the record of what was decided against. (The raw file count is 1,386; every note has a `.zh.md` twin, and counting both would double it.) Alongside them, `.agents/skills/` holds eleven repo-specific skills with names like `dsh-prose-standard`, `dsh-doc-standards`, `dsh-find-simplifications`, and `dsh-archive-agent-notes`. Of 984 merged pull requests, **209 came off `codex/*` branches** — a lower bound on machine-authored work rather than a total, since it only counts one agent's branch naming. Another 210 came from `worktree/*`, which I am not going to attribute either way. There is more test code than source code: roughly 205,500 lines of TypeScript under `src/`, against about 222,200 lines of `.spec.ts` and `.e2e.ts`. And CI runs 27 standalone `verify-*` scripts plus eleven catalog generators re-run with `--check`, covering things most repositories leave to habit — dead documentation links, markdown wrapping, mermaid syntax, JSDoc on exports, whether the English and Chinese docs are still paired, whether the generated config and tool catalogs still match the code they describe. Read together, these are one decision made repeatedly: once enough of the code is machine-written, every rule a human reviewer would have applied has to become a script, or it stops being applied. ## Interop, and one honest gap `dsh` can drive other harnesses. `subagent-claude-code` invokes the official Claude Agent SDK in the delegating session's workspace and returns only the final answer through the shared subagent contract; `subagent-codex` and `subagent-acp` do the same for their respective agents. In the other direction, `hooks-claude-code` and `hooks-codex` run a user's *existing* hook configuration on the harness's own interception points. The Claude Code bridge is refreshingly self-effacing about why it exists: > A native cordis plugin could do everything this bridge does — more powerfully, with typed returns and no serialization boundary. **The bridge exists only as a compatibility path for the mapped CC command-hook subset.** The model story is thinner than the plugin story. Only one first-party adapter ships (`llm-deepseek`); everything else routes through `llm-pi-ai`, a generic multi-provider adapter built on the third-party [`@earendil-works/pi-ai`](https://www.npmjs.com/package/@earendil-works/pi-ai). That is a reasonable trade — a new OpenAI-compatible gateway becomes configuration rather than a code change — but it does mean the polish gradient between DeepSeek's own models and everyone else's runs through a dependency they don't control. ## What I'd actually take from this Ignore the plugin count. 219 packages is a consequence of the architecture, not evidence for it, and a smaller project copying that number would just be slower. The transferable ideas are two, and both are cheap: **Make the log the source, then check it.** If the context you send is derived from your durable record rather than accumulated beside it, then a divergence is a crash instead of a slow mystery. The check is a few lines and it runs on every request. Nearly every agent system I've read builds the request and writes the log as two separate acts of bookkeeping, and quietly hopes they agree. **Make "no check here" a thing you have to say out loud.** The 184 empty invariant files are worth more than they look, because a missing check and a considered decision not to check are indistinguishable in most codebases, and a script can tell them apart here. The caveats are real: this is a developer preview with breaking changes promised in capital letters, nine weeks old, and moving fast enough that any specific file I quoted may have been rewritten by the time you read it. Every number here is measured at commit `47f9438` (2026-08-13) — I cloned the repository rather than reading the README, because the README does not mention the invariant at all, and that is the only part I would still be thinking about a week from now. --- # dots3-note Preview: 16B active parameters, and a critic that thinks before it scores > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/dots3 > date: 2026-08-14 > tags: llm, open-weights, agents, rl, multimodal, long-horizon, explainer The dots team — Xiaohongshu's AI lab — open-sourced [dots3-note Preview](https://studio.dots.ai/dots/dots3-en.html): **280B total parameters, 16B active**, a 512K context window, and multimodal understanding across text, vision and speech, under Apache 2.0. Weights on [Hugging Face](https://huggingface.co/dots-studio/dots3-note-prev), code on [GitHub](https://github.com/studio-dots-ai/dots3-note-prev), architecture submitted to Transformers as [PR #47844](https://github.com/huggingface/transformers/pull/47844). A technical report is promised within a week. "note" is the **lightest** of three planned models; jazz and aria follow. The benchmark table is respectable and not the reason to read this. The reason is a training method for tasks that take longer than a working day. ## The problem TEMPO solves Reinforcement learning on genuinely long-horizon agent tasks runs into two walls at once, and the dots team states both plainly: a single rollout "can take more than ten hours, making training prohibitively inefficient, while sparse rewards hinder effective credit assignment." Actor-critic methods like PPO exist to fix the second problem. But: > a critic estimates value through a fixed-compute forward pass. Unlike an actor, it cannot reason, reflect, or use tools to analyze the current state, making accurate value estimation difficult on complex problems. That is the observation the whole method turns on. If a task is hard enough that acting well requires ten hours of tool use and reasoning, then judging whether it is going well is *also* hard — and a single forward pass through a value head is not going to manage it. **TEMPO** — Test-time-scaled Value Estimation with Macro-step Policy Optimization — cuts the task into macro-steps, each several rounds of interaction. At the end of each, the **same agent switches from actor to critic** and uses test-time-scaled reasoning to estimate expected remaining return. The policy can then be updated mid-task rather than after ten hours.
Reported result: **+31.5% average score over the base checkpoint and +20.6% over GRPO** on ARC-AGI-3, reaching the same level in fewer steps. ## Evaluation is easier than generation The claim underneath TEMPO is the interesting one, and dots found it during training rather than assuming it: > Even when the agent cannot yet solve a problem, it can act as a critic to distinguish between two superficially similar states, identify the one that represents a genuine breakthrough in understanding the environment's rules, and assign clearly different value estimates. Their worked example is a "place knights" puzzle where two training branches both ran 64 rounds without clearing a level — **identical by environment score**. Branch B had misidentified the objective and was searching under a wrong assumption; branch A had found the real conflict rule and was close to a feasible layout.
A scalar reward cannot tell those apart. A model that reads both trajectories can. That is the entire argument for making the critic a reasoning model, and it is why the dots team frames self-evaluation as the direction they intend to keep pushing: real-world tasks "lack verifiable reward signals, while relying on human experts to evaluate model outputs may not scale." ## The IMO result belongs here At IMO 2026 in Shanghai, dots built "an internal harness around a branch of dots3-note Preview" that generated proofs recursively and used tools to evaluate and improve them. The committee's own graders awarded **7/7 on all six problems — 42/42**, a score seven of 666 contestants from 117 countries matched. Two things are worth being precise about, because the result is easy to over-read. It was **not this model**. It was a branch of it inside a purpose-built harness, and the [IMO write-up](https://studio.dots.ai/dots/imo-en.html) is a separate page from the model release. Nothing you can download reproduces it. And it used **no formal language**. The model read the organizers' original LaTeX and worked in natural language plus Python — no Lean, no proof checker. The dots team is explicit about why: formalization "requires a person to translate a problem into a formal language," and most real problems resist that. So the only thing standing between a plausible-looking proof and a wrong one was the model's own critique loop, and then a human panel that reads for holes. Which makes the IMO run an inference-time instance of the same bet TEMPO makes at training time. The proof lengths are the one signal that varies — 3, 10, 6, 5, 4 and 3 pages, and P6, traditionally the hardest slot, took one of the two shortest. ## Where it actually lands Head-to-head across the 23 reasoning and agentic benchmarks in the appendix: ahead of Hy3 (18–6), GLM 5.2 (16–9) and Seed 2.1 turbo (12–5); behind DeepSeek-v4-flash (8–14), GPT-5.5 (8–17), Opus 4.8 (8–18) and Kimi K3 (2–7 on the nine rows they share). For a model with **16B active parameters** against 21B, 39B and 104B, the first half of that sentence is the notable one. Two rows stand out, both on the benchmark this release is built around: - **ARC-AGI-3 (arcagi3 harness): 6.9 against Opus 4.8's 1.5 and GPT-5.5's 0.4.** More than four times the next best. This is the benchmark ARC Prize designed for autonomous learning in unfamiliar environments, where complex tasks need thousands of interactions over 40–50 hours. - **ARC-AGI-2: 81.4**, above Opus 4.8's 72.1 and below GPT-5.5's 85.0. That second one needs its asterisk read. dots' note says results marked `*` are their own testing, and specifically for ARC-AGI-2: "We evaluated models on the official public evaluation set; unmarked results are official leaderboard scores from the private set." **dots3-note's 81.4 is starred. Opus 4.8's 72.1 is not.** So a self-run public-set score is sitting in the same column as an official private-set score, and on ARC-AGI that difference is not cosmetic. The ARC-AGI-3 general-harness row has the same shape — dots' number is starred, and so are most of the competitors'. To their credit, the harness details are unusually complete: Terminus-2 with a 10-hour timeout for Terminal-Bench, OpenClaw 2026.6.1 with a GPT-5.4 judge for WildClawBench, live-swe-agent for the SWE suite, Hugging Face access blocked during agentic search to prevent leakage. That is more methodology than most releases publish, and it is what makes the asterisk asymmetry visible in the first place. ## The two benchmarks they released Both are open-sourced alongside the model, and both target the gap dots says it cares about — tasks where the user does not state what they want up front: - **[VibeSearchBench](https://vibebench.github.io/VibeSearchBench.github.io/)**: 200 tasks across 20 domains. Each starts with an ambiguous request, and a persona-driven simulator reveals constraints over multiple turns. The agent's predicted knowledge graph is matched against ground truth by nodes and triplets, scored by Triplet F1. - **[VibeLifeBench](https://vibebench.github.io/VibeLifeBench_homepage/)**: 20 tasks across 10 domains, each spanning **20–30 stages** on a simulated timeline, with **1,247 atomic checks** on cross-stage state consistency, tool execution and final deliverables. Their example is a family trip that has to be re-planned as aircraft type, weather and flight status change underneath it. Nobody scores well on either. On VibeLifeBench the whole field sits between 21.1 and 30.1, with dots3-note at 28.1; on VibeSearchBench, between 22.4 and 33.8, with dots3-note at 25.7. A benchmark where the best model in the world manages 30% is either badly designed or pointed at something genuinely unsolved, and the 1,247-check structure suggests the latter. ## What they say is wrong with it The Limitations section is short and unusually direct: > dots3-note Preview is an interim preview release. Reinforcement learning is not yet complete, and the model still has limitations in hallucination mitigation, the balance between text and multimodal capabilities, and overall stability. And on the real-life results specifically: those tasks **run in simulated environments**, and turning them into real experiences needs "robust harnesses, connectors, data sources, safety and permission mechanisms, and product design." That is the right caveat and it is load-bearing. The persona-driven simulator that plays the user is itself a model, so a system trained and measured against it may be learning to satisfy a simulator rather than a person. dots built the environments, the benchmarks, and the model being evaluated on them. ## What I'd take from it Ignore the parameter count and the leaderboard position. The transferable idea is that **a critic should be allowed to think**. Every value-based RL setup assumes evaluation is cheap enough to do in one forward pass — an assumption that holds fine when the task is short and breaks silently when the task is ten hours long. TEMPO's answer is to spend inference on the value estimate, and the evidence for it is a picture of a model correctly separating two trajectories that the environment scored identically. If "evaluation is easier than generation" holds up in the technical report, it is the more useful half of this release than any benchmark row in it. --- # The full-bandwidth transformer: the feedback channel is one token wide > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/full-bandwidth-transformer > date: 2026-08-14 > tags: transformers, architecture, reasoning, paper, efficiency, explainer [Full-bandwidth transformer](https://arxiv.org/abs/2608.08888) (arXiv 2608.08888, 2026-08-09) opens with an observation that is obvious once stated and easy to never state: > Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. Every decoding step runs the full depth of the model, produces a rich final-layer state, uses it to pick one token from the vocabulary — and then throws the state away. The next step starts from that one token. ## How narrow is narrow The paper counts it in bits: a sampled token carries `log₂|V|` bits between steps. Their model has a tied 100,352-token vocabulary, so **16.6 bits per step**, against a 1,536-dimensional hidden state. I would not push the ratio too hard — a hidden state's float width bounds what it *could* carry, not what it does — and the paper doesn't either. The claim that holds is about the channel's *shape*: one discrete symbol from a fixed alphabet, versus a continuous vector. That shape has a consequence, and it reframes something familiar. If the only way to pass intermediate state to your next step is to name it in the vocabulary, then you must **verbalize your own scratch work**. Which is a fair description of what chain-of-thought is: > CoT sidesteps this by externalizing intermediate state into language: the model writes out partial results, subgoals, and bookkeeping, then conditions future computation on the written trace. Reasoning traces on this reading are not primarily a thinking technique. They are a workaround for a 16-bit bus. ## The fix
**Latent feedback.** At each decoding step, fuse the previous top-layer hidden state with the sampled token's embedding through a gated linear unit, and use that as the next input. The state carried between steps goes from `s_t = a_{1:t}` — the token trace alone — to `s_t = (a_{1:t}, z_t)`, the trace *and* the most recent latent. What I find persuasive is what it does not change. Standard transformer architecture, standard KV cache, standard language-modelling objective. The fusion is dimension-preserving, so nothing downstream needs to know. There is an appendix on vLLM compatibility. The paper's phrase for the benefit is the right one: latent feedback lets "non-verbalized computation re-enter the stack with a renewed depth budget." A fixed-depth transformer has bounded serial computation per forward pass; feeding the top state back gives the next pass somewhere to continue from rather than somewhere to restart. Training is the part that could have gone wrong. Naive recurrence destroys parallel teacher forcing and with it the ability to train at scale. Their answer is a **scheduled multi-pass objective**: introduce latent feedback late in pretraining, and mix in a small fraction of deeper feedback passes for stability. ## What it buys At 1B parameters and up to 400B tokens, full-bandwidth transformers "match or approach standard transformers trained with roughly **1.5× more tokens**," at negligible per-token decoding overhead. But the result I would actually build on is the quieter one. The feedback passes double as a **training signal on the hidden states**: > In later feedback passes, the top-layer state is shifted, fused into the input of subsequent positions, and can influence losses at multiple future positions through causal attention. Thus gradients from later predictions backpropagate into earlier hidden states, encouraging them to be reusable as inputs rather than merely predictive at the output layer. In the ordinary objective, the top-layer state is supervised only through the next token. Here it is also supervised by whether it is *useful to consume*. And the payoff survives without the mechanism: > Empirically, this improves pre-training data efficiency even when latent feedback is not used at decoding time. So there is a version of this that costs nothing at serving time: train with the feedback objective, decode normally, keep the representation gains. That is a much easier thing to adopt than a new decoding loop, and it is the finding most likely to show up in someone else's model. ## The result that gets destroyed On the base model, latent-feedback decoding produces markedly shorter reasoning traces at equal or better accuracy — exactly what the bandwidth argument predicts, since computation that would have to be spelled out can ride the hidden state instead. Then: > Notably, the effect disappears after instruction tuning. We attribute this to the tuning data being off-policy with respect to latent-feedback decoding: the target traces were produced by (and imitate the verbosity of) standard token-by-token reasoning, so fitting them re-imposes the fully verbalized style regardless of what the state can carry. This is the most interesting paragraph in the paper and it is reporting a failure. A capability was trained in and then trained back out — by imitation data written by models that did not have it. The traces in every instruction-tuning set were produced under the old constraint, so they encode verbosity that the new architecture makes unnecessary, and fitting them teaches the model to keep paying a cost it no longer owes. The fix they name is on-policy post-training under latent feedback, left to future work. The general shape of the problem is not specific to this paper: **architectural capabilities can be erased by post-training data that predates them**, and nobody notices, because the benchmark still passes. ## What it does not establish The two limitations are the authors' own, stated plainly. Everything is at **1B parameters**. Their intuition is that deeper models should benefit more, since a deeper stack's top-layer state carries more — but that is a hypothesis, and 1B is small enough that a 1.5× data-efficiency gain could plausibly shrink or grow at scale. The **feedback schedule is a heuristic**. No ablation on how long the recurrence phase should run, no principled way to choose the number of recurrence steps; they point at Jacobi-iteration convergence diagnostics as a possible route. I would add a third: the state-tracking probes that verify the extra bandwidth is used — completion tracking and delayed memory — are synthetic diagnostics built for the purpose. They show the channel carries something. They do not apportion the 1.5× between the wider channel and the extra training signal, and those are separable ideas with very different deployment costs. ## Why the framing sticks Almost a decade ago, [Breaking the Softmax Bottleneck](https://arxiv.org/abs/1711.03953) made a structurally identical argument about the other end of the model: the output layer factorizes through a matrix of rank at most the hidden size, so no matter how good your representations are, the distribution you can express is capacity-limited by a shape. This paper makes the same kind of argument about the *feedback* path. Not "the model is not smart enough" but "the pipe is too narrow, and everything you have interpreted as a reasoning strategy is partly an adaptation to the pipe." Whether or not latent feedback is the right fix, that is a productive way to look at a decoder. The interesting question it leaves open is how much of what we currently call reasoning is thinking, and how much is just a model talking to itself because that is the only channel it has. --- # GLM-5.3: the same base model, a month of post-training, and a cyber capability nobody planned > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/glm-5-3 > date: 2026-08-14 > tags: llm, open-weights, agents, post-training, rl, security, explainer Z.ai's [GLM-5.3 post](https://z.ai/blog/glm-5.3) opens with a sentence most labs would bury: > Scaling post-training is all we did for GLM-5.3. It is the same base model as GLM-5.2. No new pretraining run, no new architecture, nothing to report about parameter counts because none of them changed. What changed is a month of additional post-training on the stack they had already built — IndexShare for long-context processing, SAO for RL on long-horizon tasks, and [slime](https://github.com/THUDM/slime) for large-scale asynchronous training — pointed at more environments, more diverse tasks, and more compute. That makes the release unusually easy to read. Every number below is a post-training delta, which is a rarer thing to be able to say than it sounds. ## What a month of post-training bought Fourteen benchmarks where both models are scored, GLM-5.3 ahead on all fourteen. The two that dominate the story: - **Terminal-Bench 3.0: 4.6 → 28.3.** A 6.2× move, and the largest on the board. It is also the one to be most careful with — GLM-5.2 scored 4.6, which is close enough to the floor that the model was essentially not playing. Going from *not playing* to *28.3* is a real capability change, but it is not the same kind of evidence as moving a mid-range score. - **SWE-Marathon: 19.4 → 42.5**, and **AutomationBench: 26.2 → 48.2.** Both roughly double, both on long-horizon agentic work, which is where Z.ai says the environment scaling was aimed. At the other end, Agents' Last Exam moves 23.8 → 28.5, a 1.20× gain — the smallest of the sixteen. The gains are real and they are not uniform. ## Where the environments came from The part of the post I found most interesting is not a benchmark. Z.ai describes the bottleneck moving off the model entirely: > As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment. Their answer is to synthesize the environments, and for a subset of tasks the reward signal too. Research agents collect task patterns from real work and turn them into runnable long-horizon environments with multi-step dependencies and hidden state; a judge agent then attempts each task to confirm it is actually solvable. Verifiers are synthesized **without access to the reference solution**, and solver trajectories are used to find and close reward shortcuts. A verifier that passes oracle, no-op and unsolved-state checks produces a binary reward they consider reliable enough to train on directly. That trio of checks is the load-bearing detail. An oracle check catches a verifier that rejects correct solutions; a no-op check catches one that accepts doing nothing; an unsolved-state check catches one that was already satisfied before the agent started. Those are the three ways a synthesized reward usually turns out to be worthless, and they are checkable without a human reading the task. Z.ai is direct that this is not yet automatic: the pipelines "still require a meaningful amount of human-in-the-loop work." The environments themselves are aimed at something closer to a job than an exercise. Their example is an ML infrastructure task where the model gets the same working environment as an engineer — compute clusters, storage, internal documentation, codebases, experiment results — and has to diagnose bottlenecks across a training stack, implement optimizations, run experiments, and deliver a measurable end-to-end speedup without breaking correctness. Some tasks, they say, represent several days of work for an experienced engineer. ## The honest reading of the table "The most capable open-weights model for coding" is a carefully worded claim and it survives checking. Against open models GLM-5.3 is 16–0 over GLM-5.2, 12–0 over Qwen3.8-Max, 10–4 over Kimi K3, and 7–2 over DeepSeek-V4 Pro. Against the closed frontier it is 6–9 versus Fable 5 and 4–9 versus GPT-5.6 Sol. Counted across the whole field rather than pairwise, GLM-5.3 holds the top score on **3 of 16 rows** — CyberGym, AutomationBench, and GDPval-AA v2. Publishing a table where you lead three rows and lose thirteen is not the normal shape of a launch post, and the word doing the work in the headline is *open-weights*. Two smaller notes on the table. Thirteen of the 112 non-GLM-5.3 cells are blank, so several head-to-head records rest on fewer than sixteen rows — the DeepSeek-V4 Pro comparison is nine rows, not sixteen. And ExploitGym is reported as `2h / 6h` pairs rather than a single score, so any ranking of that row depends on which budget you pick. ## Token efficiency is the better result The claim I'd have led with is the one about cost, not capability.
At Max effort GLM-5.3 reaches 34.5% at roughly 75K output tokens per task, against GLM-5.2's 23.4% at 96K — more accurate *and* cheaper, which is the direction that rarely happens on its own. At High effort it reaches 31.4% at around 50K tokens. Z.ai compares that last figure to "Claude Opus 4.8 at 29.5% with 120K," which is true and worth reading precisely: 29.5% is Opus's **Max** effort, not its High. The comparison is GLM-5.3's High against Opus's Max. As an efficiency-frontier argument that is legitimate — the whole point of the chart is that the curves sit in different places — but it is not a like-for-like row. The post is also straightforward that the frontier still belongs to someone else: "GLM-5.3 remains behind Claude Fable 5, which reaches 39.5% at Max effort." One detail in the figure's own subtitle deserves attention: the benchmark was **evaluated on Claude Code 2.1.207**. Every model in that chart was scored through a competitor's harness. Given [how much a harness shapes agent results](/articles/harness-effect), holding it fixed across models is the right call, and it is unusual to see it stated on the chart itself. ## The cyber result, and what it actually says This is the part Z.ai describes as a surprise: > As we scaled post-training, cyber capability developed faster than we expected. They added vulnerability discovery data and environments to the training mix expecting the model to get better at finding and reasoning about flaws. What they report instead is that it began reasoning across multiple stages of exploitation and forming coherent plans for complete chains. Both cyber claims in the post are accurate. GLM-5.3 does hold the top CyberGym score at 84.5 — narrowly, over Fable 5 at 83.8 and GPT-5.6 Sol at 83.6, but it is a genuine lead over closed models. And its gains really are largest further up the chain when measured against GLM-5.2: 2.23× on ExploitBench, 3.3× on ExploitGym at 6h. Put the two sentences next to each other and they suggest a model leading at exploitation. The same table says otherwise. One rung up from discovery, GLM-5.3 sits at 54.4 on ExploitBench against 78 and 76.5. At full chains under a six-hour budget it clears 130 problems against 247 and 293. The gap to the closed models widens at precisely the rate the capability is described as growing. Alongside this, Z.ai published a disclosure ledger: **2,436 findings tracked**, 53 publicly disclosed, 2,383 still under embargo, 1,097 rated critical or high, across 269 open-source projects. The detail that stops you is the age distribution — the oldest flaw was introduced in **1981**, and on average a vulnerability had been sitting in a codebase for **26.6 years** before it was found. That is a claim about the state of open-source security as much as about the model. It is also the context for the release schedule. The weights are not out yet: > We will release the weights in two weeks after launch, once safety evaluation and hardening are complete. A two-week hold between announcement and weights, explicitly attributed to safety evaluation, is a reasonable response to having just demonstrated automated vulnerability discovery at scale. It also means nobody outside Z.ai can check any of the above yet. ## What to take from it The headline result is not the benchmark table, which shows a strong open-weights model that trails the closed frontier — a familiar position. It is the claim that a month of post-training on a fixed base moved fourteen benchmarks, several of them by more than 2×, with output token counts going *down*. If that reproduces when the weights land, the interesting variable in this release is the environment synthesis pipeline, not the model. The caveats: Z.ai Code Bench is private, so its numbers cannot be independently checked by construction — a deliberate anti-contamination trade with a real cost. Thirteen table cells are blank. The weights are two weeks out. And the cyber capability that is described as emergent is, on the evidence published alongside it, still a discovery capability rather than an exploitation one. --- # LFM2.5-VL-3B: the release where GUI grounding appears out of nothing > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/lfm2-5-vl-3b > date: 2026-08-14 > tags: vlm, open-weights, on-device, edge, benchmarks, explainer Liquid AI released [LFM2.5-VL-3B](https://www.liquid.ai/blog/lfm2-5-vl-3b) on 2026-08-12 — a 3.1B open-weight vision-language model aimed at edge deployment. It is a **non-reasoning** model by design: it answers directly, which keeps latency low and, as it turns out, is the single decision that most of the performance story comes back to. The blog lists four improvements over LFM2-VL-3B: screen understanding, function calling, grounding, and multi-image input. Three of those are ordinary gains. One of them is a capability appearing from nothing, and it is not the one the post leads with. ## The row that changed LFM2-VL-3B scored **6.0, 7.6 and 2.5** on the three ScreenSpot-v2 splits. Those are not weak scores; they are the scores of a model that could not do the task. LFM2.5-VL-3B scores 78.7, 81.2 and 82.2 — an average of 80.7 against 5.4, a 15× move filed under "significant improvements in screen understanding." Liquid AI's own framing compares outward rather than backward: 80.7 "far ahead of the much larger Gemma-4-E4B (51.2) and Qwen 3.5 4B (78.5) and close behind the larger InternVL-3.5-4B (84.1)." I recomputed all four averages from the published splits and they reproduce exactly. But the comparison that tells you what happened is the one with its own predecessor. For a model whose stated purpose is running on the device that owns the screen, this is the row that decides whether it is useful. ## Architecture, and where the parameters went
The vision path is a **SigLIP2 400M NaFlex** encoder — "NaFlex" meaning it handles native aspect ratios rather than forcing a square crop — feeding a token-compression stage built from PixelUnshuffle plus an MLP connector. The diagram makes the compression ratio visible: four encoder tokens become one model token. That compression is why the time-to-first-token numbers work. A heavier encoder produces more tokens, and every one of them has to be processed before the first output token appears. On the language side, LFM2.5-VL-3B builds on the same pre-trained base as the LFM2.5-2.6B text model, pre-trained on roughly **34T tokens**. Two details worth pulling out: - The vocabulary was **doubled to 128K by extending the existing tokenizer in place**, specifically to support non-Latin scripts. Extending in place rather than retraining a tokenizer keeps the existing embeddings valid, which is the cheap way to do this and not the usual one. - Vision pretraining was scaled **4× in tokens**, with a mixture of curated and synthetic image-caption, OCR, grounding and instruction-following data. The grounding gain — RefCOCO precision@1 from 57.1 to 87.9 — is attributed directly to scaling synthetic grounding data. Post-training is SFT with knowledge distillation from a larger teacher, plus something Liquid AI calls **Antidoom training**, followed by multi-reward RL. ## The size-class claim, checked I transcribed all 28 benchmark rows and recomputed each model's average. Every one reproduces Liquid AI's published Average row to within rounding, so the table is internally consistent — worth doing, because the claim rests entirely on that average. The claim is narrow and it is stated precisely: LFM2.5-VL-3B averages **69.4**, which is *exactly* level with InternVL 3.5 4B (69.4, at 4.7B parameters) and 0.7 behind Qwen3.5-4B (70.1, also 4.7B). It beats both Gemma models, at 5.1B and 8B. So "competitive vision performance against models twice its size" is supported. "Better than models twice its size" would not have been, and the post does not say it. The head-to-head view makes this sharper than the averages do. Counted row by row across all 28 benchmarks, LFM2.5-VL-3B is **14W–14L against InternVL 3.5 4B and 14W–14L against Qwen3.5-4B** — a dead tie against both 4.7B models, from two different labs. Against the rest it is comfortably ahead: 27–1 over Gemma-4-E2B, 23–5 over Gemma-4-E4B, 22–6 over both 2B-class models. Where it loses is worth naming. Qwen3.5-4B takes the document-heavy rows — DocVQA 94.8 to 91.1, InfographicVQA 80.3 to 70.2, OCRBench v2 58.7 to 47.5 — and MMMU-Pro 36.0 to 30.5. InternVL 3.5 4B takes ChartQA, MMMU, and all three GUI splits. If your workload is dense document OCR, the larger models are still worth their size. One row moves the wrong way and the post does not mention it: **CountBenchQA drops from 92.2 to 87.3**, the only benchmark where the new model is meaningfully behind its predecessor. POPE also slips slightly, 89.2 to 88.7. ## Function calling, added to a VLM New to the VL line: ToolSandbox goes **26.4 → 59.5** and BFCL v4 **20.5 → 32.5**. The blog positions this as "on par with Gemma-4-E2B and ahead of Qwen3.5-2B," which understates it — 59.5 beats Gemma-4-E2B's 56.5 and Qwen3.5-2B's 47.7, and only Gemma-4-E4B (61.6, at 8B) and Qwen3.5-4B (65.0) are ahead. Both InternVL models are marked N/A because they do not support function calling at all. That is the more interesting fact in the row: on the benchmark where InternVL was beating LFM on GUI grounding, it cannot compete, and a model that can both locate a button and call a tool is a different product from one that can only do the first. The text-only instruction-following numbers are less flattering. IFEval 82.3 sits behind both Gemma models (83.0 and 87.9); Multi-IF 59.4 is well behind their 69.4 and 77.4. This is a vision model with tool use bolted on competently, not a text model that also sees. ## Speed, and what is actually specified
The GPU measurements are properly specified: vLLM 0.26, BF16, a 512×512 image plus 1,024 input tokens, up to 256 output tokens, median of five runs per concurrency level, single H100 SXM5. On a 5-frame clip LFM2.5-VL-3B returns its first token in about **34 ms** where the Gemma models take around 200 ms. Sustained output throughput reaches roughly **11K tokens/s** at high concurrency — about 2× the 4B-class models — which Liquid AI works out to nearly 1B output tokens per day from one GPU. The on-device figures are the ones to be careful with: **228 tok/s on an Apple M5 Max**, 116 on an AMD Ryzen AI Max+ 395, 20 on a Galaxy S26 Ultra, in about 3 GB. No quantization, prompt, or batch size is stated for any of them. Given that GGUF, MLX and ONNX builds all ship day one and would each give a different answer, these should be read as claims rather than measurements. ## What it is for The honest summary is that this is a **screen-and-document model that fits on a phone**. It ties two 4.7B models on a 28-benchmark average, loses the dense-OCR rows to both, wins the real-world and grounding rows, and gained GUI grounding and tool calling in one release. The non-reasoning choice is the through-line. It costs accuracy on the STEM benchmarks where thinking helps — MMMU-Pro 30.5 is the weakest column in the table — and buys 34 ms to first token and 11K tokens/s sustained. For an agent that has to look at a screen, decide where to tap, and do it again, that is the correct trade. For a model asked to reason about a diagram, it is not. Two things I could not check: Antidoom training is named but not described anywhere in the post, and the on-device numbers have no stated conditions. Everything else in the table reproduces. --- # MAGI-2 Preview: 114B parameters, 6B awake, and two sparsities doing the work > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/magi-2-preview > date: 2026-08-14 > tags: video-generation, audio, open-weights, moe, architecture, explainer Sand AI released [MAGI-2 Preview](https://github.com/SandAI-org/MAGI-2-preview) under Apache 2.0: a **114B-parameter unified audio-video generation model that activates just 6B parameters per token**. Text-to-video or image-to-video, ten-second clips, sound generated alongside the picture and muxed into the same file. Eight Hopper GPUs to run it. Both halves of that headline are checkable against the published weights, and they check out — but the second one only works for a reason the card does not spell out. ## Where 6B comes from The total is exact. The safetensors index reports `total_size: 228107858176`, which at bf16 is **114.05B parameters**. Reconstructing it from the published tensor shapes and the config gives 113.88B — 0.15% apart, close enough to say the decomposition is right rather than lucky. The active figure is more interesting. The MoE is 256 experts across 12 heads — 3,072 expert slots — with top-6 routing *per head*, so 72 of 3,072 fire for any token. That takes the MoE weight in a layer from 3.02B down to 70.8M. Apply that alone and you land at **7.71B active**, not 6B. The rest comes off because of something visible only in the shapes. On layer 0, `linear_qkv` is `[27648, 3072]`; on layer 2 the same tensor is `[9216, 3072]`. Exactly three times as large, and `k_norm` goes `[384]` against `[128]` to match. The model carries **three sets of weights — one per modality** — in its dense layers and in the modality-specific shared expert of every MoE layer. A token is video *or* audio *or* text, never all three, so two thirds of those weights sit resident and idle. Put both sparsities together and it comes to **5.96B**, against a stated "just 6B parameters per token." One active parameter in 19.1. That is the design worth naming: MoE sparsity and modality sparsity multiplied, not just the MoE ratio everyone quotes. ## The stack Forty layers, hidden size 3072. The config lists `mm_layers: [0, 1, 38, 39]` and MoE on layers 2 through 37 — dense at both ends, sparse through the middle. That placement says what Sand AI thinks the hard part is. Mixing three modalities is treated as an entrance-and-exit problem: two dense layers at the bottom and two at the top, each carrying private weights per modality. Everything between them runs a *single shared attention* over the fused sequence and spends its capacity on routed experts. Dense where the modalities are still separate, sparse once they are already mixed. The three modalities enter through their own embedders — video at 48 channels, audio at 64, text at 5120 — and leave through separate video and audio output heads. There is no text head: text is conditioning, not output. The refiner is a different model, not a smaller copy of the same one. Its config gives **30 layers at hidden 4096** with 8 query groups, `mm_layers: [0, 1, 28, 29]` — the same dense-at-the-edges pattern — and **no MoE at all**. It also sets `local_attn_layers` to all thirty. That is a sensible split of labour: upscaling 512×896 to 1088×1920 is a local problem, so the second stage is dense, wider, shallower, and never looks far across the frame. It gets 14 GB and 5 denoising steps against the preview stage's 228 GB and 100. ## Four residual streams The tensor names give away a technique the README never mentions. Every layer carries `mhc_alpha_pre_attn`, `mhc_bias_res_attn` shaped `[4, 4]`, and an `mhc_norm.weight` of `[12288]` — which is 4 × 3072. That is **hyper-connections**: instead of one residual stream with `x + f(x)`, the model maintains four parallel streams and learns how to mix them, with a 4×4 matrix deciding how each stream feeds the next block. The config confirms it as `mhc_config: { num_stream: 4, alpha_init: 0.01 }`, and the residual state really is four times as wide as the hidden size — the embedders write into 12288, not 3072. Two implementation details in `magi2_preview.py` are worth flagging because they are not obvious from the config: - The connection matrices go through a **Sinkhorn-Knopp** normalization (`_sinkhorn_knopp_affine_fwd_kernel`), which makes them doubly stochastic — every stream contributes and receives a fixed total, so no stream can quietly dominate. - The whole thing runs through a hand-written Triton kernel (`_hyper_connect_fwd_kernel`). Four residual streams is four times the memory traffic if you do it naively. The attention has two further additions: **attention sinks** (one sink token per layer, via FlashAttention-3's `fa3_func_with_sink`) and **gating** — a `linear_g` projection per layer, 24 outputs on MoE layers and 72 on the modality-specific ones, one per query group. ## What you are actually downloading 307 GB, and **64 GB of it Sand AI did not train**: the text encoder is Qwen3.5-27B, the video VAE comes from Wan2.2-TI2V-5B, and the audio VAE is Stability's stable-audio-open-1.0. The repo names each one and links it, which is the right way to do this — but it does mean "114B open-weights video model" describes the transformer, not the system you run. The 2 GB `turbo_vae` is Sand AI's own distilled VAE decoder, and it is on by default (`use_turbo_vae: true`). It is also the only distilled component in the release, which brings up the cost. ## The honest problem: 105 denoising steps The README is unusually direct about this: > Neither transformer has been step-distilled, so the denoising step count is where most of the wall-clock time goes. The shipped configuration is **100 preview steps plus 5 refiner steps**. The preview stage generates at 512×896 and the refiner takes it to 1088×1920. A distilled release with "far fewer" steps is listed as coming soon, with no date. So the model that exists today is the slow one, on purpose, and Sand AI is saying so in the release rather than after someone benchmarks it. What is not stated anywhere is how long 105 steps actually takes on the eight Hopper GPUs it requires — there is no wall-clock figure in the repo, the card, or the config. Everything else about the runtime is specified in detail: `cp_size: 8` and `ep_size: 8`, so context parallelism (Ulysses) and expert parallelism both span all eight GPUs; guidance is 5.0 for video and 7.0 for audio; output is 12.5 fps over a 10-second clip; the video VAE stride is `[8, 16, 16]`. There is even a `--deterministic` flag, and a commit whose entire purpose is *"add Inductor compile-time configs for bit-exact reproducibility."* For a diffusion model where a one-ULP difference changes the video, shipping bit-exactness as a supported mode is a real courtesy. ## Prompt enhancement is not optional in practice The captions the model trained on are long and structured, so the pipeline ships a prompt-enhancement step that asks an instruction-following LLM for a **structured JSON caption** of the 10-second clip, then renders it to Markdown before encoding. Templates are included for T2V and I2V separately. It talks to an OpenAI-compatible endpoint and is off unless you set an `API_KEY`. The README is candid that "a short hand-written prompt underuses" the model — which means the shipped quality bar assumes a second model in the loop that the checkpoint does not include. The repo does hedge this properly by shipping two already-enhanced example prompts so you can see what the model actually expects. ## What is missing **No evaluation of any kind.** No VBench, no comparison against Wan, Kling, Veo, Sora or anything else, no human preference study, no ablation. For a release whose stated purpose is to explore "an efficient path to scaling video generation," there is no published evidence that the efficiency buys quality. The samples in `assets/` are inputs, not results. **No wall-clock or cost figure**, as above. **The technical blog is unreachable.** The architecture, training system and data pipeline are described at [sand.ai/blog/magi-2-preview](https://sand.ai/blog/magi-2-preview), which sits behind a WAF that returns a challenge page rather than content. Everything in this article therefore comes from the repository, the model card and the published weights — which turned out to be enough to verify the headline numbers, but means the training and systems claims are unexamined here. ## Why it is worth the attention anyway The parameter accounting is the result. A model that keeps three modality-specific copies of its dense weights and routes 72 of 3,072 expert slots per token gets to 19× sparsity without either mechanism being exotic on its own — and both are legible in the shapes, which is rarer than it should be. The rest is a set of choices that are individually defensible and unusual together: hyper-connections with four Sinkhorn-normalized streams, attention sinks and gating, modality-private layers only at the edges, a distilled VAE decoder but undistilled transformers. It is a lot of recent architecture research in one checkpoint, shipped Apache 2.0 with the shapes visible. The thing I would want before recommending it is a number — any number — comparing its output to something else. --- # MiniMax Music 3: a five-minute song is 9,000 steps of a 2.5 kbit/s code > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/minimax-music3 > date: 2026-08-14 > tags: audio, music-generation, open-weights, flow-matching, rvq, explainer [MiniMax Music 3](https://huggingface.co/MiniMaxAI/MiniMax-Music3) generates complete songs up to five minutes long from lyrics plus a music description, at 32 kHz 16-bit stereo. The weights went up on 2026-08-07 and the card was still being edited on 2026-08-13. What makes it worth reading closely is that the card is unusually specific — it names a parameter count for four separate components — and all four of those numbers can be checked against the published files. They hold up, which is rarer than it should be. ## The shape of the thing
The split is between **structure** and **texture**: - The **Global LLM (8B)** predicts the first RVQ codebook, frame by frame. That codebook is the semantic one — 16,384 entries — and it carries the song's long-range progression: where the chorus is, whether the vocal identity holds, how the arrangement evolves. - The **Local LLM (646M)** predicts the remaining seven acoustic codebooks *within* each frame, restoring fine-grained detail the semantic codebook throws away. The part that is not standard is what happens next. Rather than decoding from the discrete RVQ tokens, the synthesis stage fuses the **final hidden states** of both models and flows from there. The tokens are what the models predict; the hidden states are what actually gets rendered. MiniMax's argument is that continuous representations preserve more than the quantized codes do — vocal articulation, instrumental texture, temporal continuity — and the card is explicit that at inference time "waveform synthesis uses the fused LLM hidden states and does not require the discrete tokenizer decoder." The tokenizer, in other words, is a training-time device. It shapes what the LLMs learn to predict and is then bypassed. ## Every number on the card, checked I pulled the safetensors headers by HTTP range request — the 8-byte length prefix, then the JSON header that gives every tensor's dtype and shape — rather than dividing file sizes and hoping. That mattered twice. The **Flow Matching** module is 9.73 GB across two shards. At bf16 that would be 4.86B parameters and the card's "2.4B" would be wrong by a factor of two. Every tensor in it is **F32**, so it is 2,431.9M, and the card is right. The **Global LLM** is the nicer result. The config is `Qwen3ForCausalLM` with Qwen3-8B's exact shape — 36 layers, hidden 4096, intermediate 12288, 32 query heads over 8 KV heads — but `vocab_size` is **200,000** rather than Qwen3-8B's 151,936. Qwen3-8B is 8.191B parameters. Widening the vocabulary adds `(200,000 − 151,936) × 4096 × 2 = 393.7M` for an untied embedding and output head. That predicts 8.584B, and the index reports 8.584B. So the card's two claims about this model — "initialized from Qwen3-8B" and "its embedding and output layers are first adapted to semantic music tokens" — are both visible in a single number, and the extra 48,064 vocabulary slots are where 16,384 semantic music tokens went. The **Flow-VAE decoder** matches exactly too: `dav.pth` is 491.8 MB, which at fp32 is 123.0M parameters against a stated 123M. The one component the card never mentions is a 25.2M condition encoder that takes 24 kHz audio in and produces conditioning at 44.1 kHz — the piece that would let you condition on a reference track rather than only on text. ## Why the repository is 57 GB The parameters add up to roughly 12B. The repository is 57.35 GB. The difference is that it ships the entire model twice, in two runtime layouts — the SGLang-Omni one the card recommends, and a diffusers modular pipeline. The arithmetic that shows these are the same weights rather than two models is clean: `flowmatching_vae.pth` is 2,457.1M parameters at fp32, and the diffusers `transformer` plus `condition_encoder` are 2,431.9M + 25.2M. Same total, not an approximation. The oddity is the folder named `qwen_7B`. It holds an **`AbabForCausalLM`** — Abab being MiniMax's own model family — at the identical 8.58B shape as the `Qwen3ForCausalLM` sitting beside it. A directory named after one model family, containing another, holding what appears to be the same model converted for a different runtime. It is 18.48 GB of the repository and nothing in `modular_model_index.json` refers to it. ## The frame budget The card's Limitations section gives two ceilings and does not connect them: songs "up to five minutes," and "audio generation is limited to 9,000 acoustic frames." Those only agree at **30 frames per second**, which is a number the card never states. It is worth deriving, because it fixes the scale of everything else. Eight codebooks per frame — one at 14 bits, seven at 10 — is 84 bits per frame, or **2.52 kbit/s**. That is the representation the Global LLM is autoregressing over, and a full-length song is 9,000 steps of it against a 32 kHz stereo output that would be 1,024 kbit/s as raw PCM. Roughly 406× compression, with the flow-matching stage responsible for putting back everything that ratio removed. The text side is separate and much tighter: 5,000 tokens total for lyrics and description combined. ## Control, and what it does not promise Input is two fields. **Lyrics** may carry explicit section tags — `[Intro]`, `[Verse]`, `[Pre-Chorus]`, `[Chorus]`, `[Post-Chorus]`, `[Bridge]`, `[Instrumental]`, `[Solo]`, `[Outro]`. **Music description** covers style, emotional progression, vocal performance, instrumentation, arrangement, and production. MiniMax recommends a three-part Structured Caption — Global Metadata (genre, BPM, key, scale, emotional arc, production profile), Vocal Details (gender, timbre, performance style, harmony, backing vocals, effects), and Arrangement (primary and secondary instruments, section-level instrument evolution, groove, bass, percussion, textures, spatial effects). There is a `music-caption-rewriter` skill for expanding a short prompt into one, installable with `npx skills add`. The card is honest about what that buys: > Section tags and music descriptions provide generative control rather than strict symbolic guarantees. The generated tempo, key, instrumentation, lyrics, and song structure may not always match every requested detail exactly. Which is the right way to describe a model that has no symbolic music representation anywhere in it. You are conditioning a sampler, not programming a sequencer. ## Running it CUDA only, and non-streaming only — you wait for the whole song. Full precision fits under 24 GB of VRAM; with automatic CPU offloading it needs about 22 GB; and streaming the language model layer by layer with `apply_group_offloading` gets it onto an 8 GB card, slowly. That last path is the interesting one for anyone without a datacenter, and it is a consequence of the hierarchy: the 8B Global LLM is the only piece that has to be resident for the long autoregressive run, so streaming its layers costs bandwidth rather than correctness. ## What I'd flag The engineering claims check out, which is the main thing I set out to test. What the card contains no evidence for is **quality** — there are no listening-test results, no comparison against Suno or Udio or any other music model, and no objective audio metrics. There is a demo page and a single `assets/minimax_ttm.wav`. For a generative audio model, that is the entire evaluation. The license file is present in the repository but the HF API reports no license field, so it is worth reading `LICENSE` directly before assuming anything about commercial use. And 25 downloads against 440 likes, a week after release, is the signature of a model far more people want to hear about than can actually run. --- # nac: an orchestrator that is not allowed to touch anything > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/nac > date: 2026-08-14 > tags: agents, harness, open-source, rust, context, explainer [nac](https://github.com/arcee-ai/nac) is Arcee AI's open-source agent harness, Apache 2.0, about 98,000 lines of Rust across three crates. The [write-up](https://arcee.ai/blog/nac) opens with a diagnosis rather than a feature list: > We think this couples two things that should be separate: the temporary context needed to perform an action; and the persistent state needed to continue a workstream. That is the whole design, and everything else follows from taking it literally. ## The orchestrator cannot do anything nac uses a thread-and-episode architecture adapted from Random Labs' [Slate](https://randomlabs.ai/blog/slate). A central orchestrator plans, decomposes, and decides what happens next. It has exactly one action available: launch threads. > Importantly, the orchestrator's only action is launching threads; it cannot execute commands or edit files on its own. The blog states the split as two lines of pseudocode, which is the clearest thing in it: ```text orchestrator: decide and route, but do not act workers: act, but do not expand the orchestration graph ``` Each dispatch starts a **worker** — a fresh process with a fresh model context, given the worker system prompt, the requested action, its tools, and any applicable skills. The worker calls the model and uses tools until the model returns a response with **no tool calls**. That response is the **episode**. There is no separate summarization pass. The worker's system prompt tells it that its final answer should be a concise handoff, so the answer *is* the summary. That saves a model call and, more importantly, means nothing gets summarized twice. ## What survives Once the episode exists: > the worker's execution context **is discarded and never used as model context by the system again**. Its changes to the environment remain, but the episode is the persistent representation of the work. A **thread** is just a named, ordered list of episodes. When the orchestrator assigns that thread more work, a new worker starts fresh with the thread's accumulated episodes — never the transcripts that produced them. This is why nac has no compaction problem. Compaction exists to compress a transcript that has grown too long; nac never lets the transcript reach the orchestrator in the first place. The orchestrator reads episodes, and by the time it plans again the execution details are already gone. It is a different answer to context rot than compressing history: don't accumulate it. **Thread weaving** is the other half. A dispatch can name source threads, and nac resolves each to its most recent retained episode and hands it to the worker as context — but those source episodes never join the target thread's history. Only the new episode does. Threads stay their own. ## A batch is a graph The orchestrator ends a turn by emitting a batch of thread calls, each with a `name`, a free-form `action`, and optional `threads`, `skills` and `timeout`. Naming a source that is dispatched in the *same* batch creates a dependency edge; naming one that already finished just supplies context. That makes the batch a DAG. nac rejects duplicate targets, validates acyclicity before executing, runs independent workers concurrently, and waits for the whole batch before letting the orchestrator plan again. The synchronization point is deliberate — the orchestrator never polls background work and never observes a half-finished world. One place the code is more precise than the prose. The blog says a cyclic batch is rejected; `crates/nac-core/src/agent/tool_exec.rs` shows what actually happens on `DagError::Cycle` or `DagError::DuplicateName`: every *thread* dispatch gets an error result, while non-thread tool calls made in the same turn still execute normally. If your orchestrator mixes a query with a dispatch, that distinction matters. ## The seam they printed The honest part of this release is a sentence most teams would have left out: > Worker failures are not transactional: if a worker changes the environment and then exits before committing its final response, those changes may remain without a new episode, so a returned error means the environment may have moved ahead of persistent history. Episodes persist only on success. Environment changes persist unconditionally, and live outside nac's state entirely. So a worker that edits files and then dies leaves the world ahead of the record, and nothing in the runtime knows. Every system that separates durable state from a scratch context has this seam somewhere. What is unusual is printing it in the launch post rather than leaving it to be discovered. ## Harnesses as inference runtimes The framing section is the part I expect to get quoted, and I think it earns it. Arcee traces harness evolution along two axes — **enriching context** so each model call gets denser information, and **expanding the action space** so the model can initiate more capable operations — from tool use through program execution, memory, multi-agent search, [Recursive Language Models](/articles/recursive-language-models), fresh-session harnesses, and finally Slate-style dispatch. Then the claim: > An agent inference runtime constructs context, schedules inference, executes effects, preserves state, enforces capabilities, and defines how work synchronizes, fails, resumes, and stops. A thin harness executes a model-tool loop. A runtime owns semantics that would otherwise exist only implicitly in its transcript. And the mapping, which is what makes it concrete rather than a slogan: ```text worker invocation = inference operation thread = persistent program state episode = committed workstream update source thread = data dependency dispatch batch = dynamic execution graph ``` Their summary line is the one worth keeping: **"judgment stays in tokens, invariants live in the runtime."** It is worth reading this next to [DeepSeek Harness](/articles/deepseek-harness), which arrives at a related conclusion from the opposite direction. dsh keeps one agent loop and makes the *log* the authority, with a runtime invariant that refuses any request the log cannot reconstruct. nac keeps no shared log at all and makes *episodes* the authority, with a scheduler that refuses any batch it cannot order. Both are saying the harness should own guarantees the transcript used to own implicitly; they disagree about whether the transcript should exist. Arcee also names two systems that make different choices — Onyx, which pushes orchestration control flow into persisted typed programs, and LongHorizon-Harness, which advances one globally audited task record through serial manager/executor/auditor rounds instead of parallel workstreams. Citing your neighbours accurately is a good sign. ## When it is the wrong tool Stated plainly, which is rarer: > For a single focused change that fits in one coding-agent session, going direct is simpler and often faster. That adds overhead because the orchestrator cannot perform the task itself; it still has to delegate to a thread. The architectural purity has a fixed cost: a one-line fix still requires a dispatch. Their stated fit is work with a meaningful high-level objective, hard boundaries stated up front, a concrete definition of done, enough independent work to justify parallelism, and freedom for nac to choose its own decomposition — reproducing an ML paper, porting a large codebase, decomposed code review, large parallel change jobs on a dedicated branch and worktree. ## The meta-orchestrator pattern nac ships an MCP server, so Claude Code or Codex can dispatch, monitor and steer nac jobs as tools. Arcee's preferred pattern is to make the interactive agent a **meta-orchestrator**: it works with you in a normal session, watches for work that is decomposable with a concrete definition of done, writes the job description itself, and hands it to nac to run in the background. The capability boundary is drawn carefully: > Through nac's MCP interface, the meta-orchestrator still cannot see a worker's discarded execution context or the underlying environment directly; the MCP server exposes no file or shell tools of its own. So the outer agent gets the same view a human gets — orchestrator chat, thread episodes, recent events, the ability to steer — and no more. State must be queried; it is not pushed into the meta-orchestrator's context. The restriction that defines the inner orchestrator is applied to the outer one too. ## What is missing **No evaluation.** No benchmark, no comparison against a single-agent baseline, no measurement of the token savings the architecture is supposed to produce. For a design whose central claim is that separating temporary from persistent context makes long tasks work better, there is no number showing it does. The evidence offered is a timelapse video and the fact that Arcee uses it internally. **No cost accounting.** Running an orchestrator plus N parallel workers, each with its own context, is not obviously cheaper than one long session — it trades context length for context count. Which way that lands is exactly the thing an evaluation would tell you. The repository is six commits old at the time of writing. This is a design worth taking seriously and a codebase worth waiting on. ## What I'd take from it The transferable idea is the prohibition, not the architecture. Most multi-agent systems let the orchestrator do a little work itself when delegation feels heavy — and that is precisely when the orchestrator's context starts filling with execution detail and the original intent starts getting diluted. nac removes the option. The orchestrator cannot act, so its context stays a plan. That is a constraint you could impose on a system you already have, without adopting threads, episodes, or Rust. --- # Breaking the softmax bottleneck: your output layer is a rank-d wall > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/softmax-bottleneck > date: 2026-08-14 > tags: architecture, transformers, paper, explainer, theory, language-models [Breaking the Softmax Bottleneck](https://arxiv.org/abs/1711.03953) (Yang, Dai, Salakhutdinov, Cohen — CMU, ICLR 2018 oral) is a paper about a wall you cannot see from inside the model. Deep learning does not have many negative results that matter. Most limitations turn out to be engineering — not enough data, not enough compute, the wrong optimizer. This one is linear algebra, it takes two lines to state, and it is still true of every model you use. ## The argument Take the standard output layer. A network turns the context $c$ into a hidden state $\mathbf{h}_c$, you dot it with every word embedding $\mathbf{w}_x$, and softmax the result: $$ P_\theta(x \mid c) = \frac{\exp \mathbf{h}_c^\top \mathbf{w}_x}{\sum_{x'} \exp \mathbf{h}_c^\top \mathbf{w}_{x'}} $$ Now stack every context as a row of $\mathbf{H}_\theta \in \mathbb{R}^{N \times d}$, every embedding as a row of $\mathbf{W}_\theta \in \mathbb{R}^{M \times d}$, and the true log-probabilities as $\mathbf{A} \in \mathbb{R}^{N \times M}$. Your model's logits are $\mathbf{H}_\theta \mathbf{W}_\theta^\top$, and language modelling is now the question of whether that product can equal $\mathbf{A}$. It cannot, if $d$ is small. The rank of a product is bounded by the shared inner dimension. Whatever the network does, its logits live in a $d$-dimensional subspace. The move that makes this a paper rather than an observation is handling the obvious objection. Softmax is invariant to adding a constant to a row, so the model doesn't have to hit $\mathbf{A}$ — it can hit any member of the family $F(\mathbf{A}) = \{\mathbf{A} + \mathbf{\Lambda}\mathbf{J}\}$ of row-shifted variants, an infinite set. Surely somewhere in an infinite set there's a low-rank one? No. Their Property 2: any two matrices in $F(\mathbf{A})$ have ranks differing by at most 1. The entire row-shift freedom is worth **one rank**. So the corollary is clean: > **Corollary 1 (Softmax Bottleneck).** If $d < \text{rank}(\mathbf{A}) - 1$, then for any function family $\mathcal{U}$ and any parameter $\theta$, there exists a context $c$ such that $P_\theta(X \mid c) \neq P^*(X \mid c)$. Read the quantifier: *for any function family*. Universal approximation doesn't help. You can make the network computing $\mathbf{h}_c$ arbitrarily deep and arbitrarily wide and it changes nothing, because the constraint is on the shape of the factorization, not on the expressiveness of the thing being factorized. All that effort is spent producing a vector that then has to squeeze through a $d$-wide waist. ## The part that is not proved The bound only bites if $\text{rank}(\mathbf{A})$ is actually large, and the paper is upfront that this is a hypothesis: > It is difficult (if possible) to rigorously prove this hypothesis since we do not have access to the true data distribution of a natural language. The supporting intuitions are decent but soft. Language is context-dependent — "north" is followed by "korea" in a politics article and not in a U.S. history textbook. And if $\mathbf{A}$ *were* low rank, that would mean a few hundred basis distributions span every meaning humans express, and no one has ever found such a basis. Neither of those is evidence. The evidence comes later, and it's better than the arguments. ## Two easy fixes, priced Once you see the bound, two fixes suggest themselves, and the paper prices both out before proposing anything. **Use an n-gram model.** Non-parametric, no rank constraint, universally approximates any language. Costs $N \times M$ parameters, where $N$ — the number of contexts — is unbounded. Generalizes badly, which is why the field left it. **Raise $d$ until the bound stops binding.** To express a full-rank $\mathbf{A}$ you need $d \approx M$, so the embedding matrix costs $M \times M$. Slide the widget above to a modern vocabulary and watch that number: at $|V| = 151{,}936$ it's a **23-billion-parameter output layer**, for a model whose whole point was to be small. And empirically it doesn't even work — the paper notes, and everyone else had found, that pushing $d$ past a few hundred stopped helping on these benchmarks. That is the real tension, and it is why the paper is interesting: *expressiveness and generalization are in conflict at the output layer*, and the naive ways to buy one spend the other. ## Mixture of softmaxes The fix is small enough to quote in full. Compute $K$ context vectors instead of one, run each through the **same** embedding matrix, softmax each, and average the resulting **probabilities** with context-dependent weights: $$ P_\theta(x \mid c) = \sum_{k=1}^{K} \pi_{c,k} \frac{\exp \mathbf{h}_{c,k}^\top \mathbf{w}_x}{\sum_{x'} \exp \mathbf{h}_{c,k}^\top \mathbf{w}_{x'}} $$ Here is the whole thing in the [reference implementation](https://github.com/zihangdai/mos), essentially unedited: ```python # one linear layer produces all K context vectors at once self.latent = nn.Sequential(nn.Linear(nhidlast, n_experts * ninp), nn.Tanh()) self.prior = nn.Linear(nhidlast, n_experts, bias=False) latent = self.latent(output) # (T*B, K*d) logit = self.decoder(latent.view(-1, self.ninp)) # shared W, K times prior = F.softmax(self.prior(output).view(-1, self.n_experts), -1) prob = F.softmax(logit.view(-1, self.ntoken), -1) prob = prob.view(-1, self.n_experts, self.ntoken) prob = (prob * prior.unsqueeze(2).expand_as(prob)).sum(1) # mix probabilities log_prob = torch.log(prob.add_(1e-8)) ``` Two things worth noticing in that code. The embedding matrix `self.decoder` is used $K$ times — MoS does not buy $K$ embedding tables, it buys $K$ *readings* of one table, so the parameter cost is the `latent` projection and nothing else. And the mix happens **after** the softmax, on probabilities, which is the only reason any of this works. Because the resulting log-probability matrix is $$ \hat{\mathbf{A}}_{\text{MoS}} = \log \sum_{k=1}^{K} \mathbf{\Pi}_k \exp\!\left(\mathbf{H}_{\theta,k} \mathbf{W}_\theta^\top\right) $$ and $\log \sum \exp$ is nonlinear, $\hat{\mathbf{A}}_{\text{MoS}}$ has no rank ceiling at all. It is a nonlinear function of $K$ rank-$d$ matrices, and nonlinear functions of low-rank matrices are generically full rank. ## The trap, which is the best part Now the near-miss. Suppose you mix the **context vectors** instead of the probabilities — average $\mathbf{h}_{c,k}$ with the same weights, then take one softmax. Call it mixture of contexts. It looks like the same idea, it has the same parameter count, and it is completely useless: $$ \mathbf{h}'_c = \sum_k \pi_{c,k}\mathbf{h}_{c,k} \quad\Longrightarrow\quad P_\theta(x\mid c) = \frac{\exp \mathbf{h}'^\top_c \mathbf{w}_x}{\sum_{x'} \exp \mathbf{h}'^\top_c \mathbf{w}_{x'}} $$ which is the original softmax with a different $\mathbf{h}$. Rank still bounded by $d$. Mixing in feature space makes the *function family* richer and leaves the *ceiling* exactly where it was. MoC exists in the paper as a control, and it is the sharpest instrument in it: same parameters, same layers, same hyperparameters, one design decision moved, and the theory says one should work and the other shouldn't. There is also a footnote in the related-work section that has aged into the most consequential sentence in the paper: > Although Shazeer et al. (2017) name their architecture as MoE, it is not a standard MoE and should be classified as MoC under our terminology. [Shazeer et al. 2017](https://arxiv.org/abs/1701.06538) is the sparsely-gated mixture-of-experts layer — the direct ancestor of [every MoE LLM shipping today](/articles/mixture-of-experts-from-scratch), from [Switch Transformer](/articles/switch-transformer) to DeepSeek-V3 to Kimi. Under this paper's taxonomy all of them are mixture-of-*contexts*. They mix in feature space. They make the function family enormously richer and they do not touch the rank of the output layer by one. That is not a criticism of MoE — sparse experts are solving conditional computation, not expressiveness — but it does mean the thing people reach for when they want "more capacity" is provably not the thing that lifts this particular ceiling. ## The evidence Two measurements, and the first one is the kind that could have embarrassed everybody. They compute the empirical log-probability matrix on PTB and estimate its rank. Softmax with $d = 400$ measures **400**. MoC with $d = 280$ measures **280**. Not near the bound — the bound, to the digit. MoS with the same 280 dimensions measures **9,981** out of a possible 10,000. Then the dose-response. Sweep $K$ from 3 to 20 and rank climbs — 6,467, 8,930, 9,973 — with perplexity falling alongside it. At $K = 15$ rank has saturated at 9,981 and perplexity is at its best. At $K = 20$ rank does not move, because there is nothing left to buy, and **perplexity gets worse**. That reversal is what makes the sweep an argument. A "more parameters help" story predicts a monotone curve. A "mixtures buy rank until rank runs out, and then you're just overfitting" story predicts a curve that turns exactly where rank saturates. It turns exactly where rank saturates.
Counting non-zero singular values is a roundoff-sensitive way to measure rank, so they plot the spectrum instead and the picture is unambiguous. Softmax and MoC dump ~96% of their normalized singular values below $10^{-9}$; MoS's are spread from $10^{-5}$ upward. Same conclusion, no thresholding decision required. A third check, in the appendix: expected pairwise KL divergence between next-token distributions at different contexts — how much the model's prediction actually changes when the context changes. Softmax 4.763, MoC 4.864, MoS 5.284 on PTB test. ## Three controls that turn it into a mechanism Any of the above is consistent with "MoS is a good regularizer and the rank story is decoration." The paper runs the experiments that separate those. **Ablation.** MoC with matched everything is worse than MoS on both datasets — and on WikiText-2 it is worse than the plain AWD-LSTM baseline it was built from (65.98 against 65.40). So mixing per se isn't the win. Separately, training the baseline with MoS's hyperparameters is a disaster (74.86 against 58.95 on PTB), which rules out "they just found better hyperparameters." **Regularization control.** On the 1B Word dataset, where overfitting is unlikely and no dropout is used at all: Softmax reaches 41.47 train / 42.77 test; MoS reaches 36.39 train / 37.10 test. MoS's *training* perplexity is 5.08 points lower. If the gain were regularization, training perplexity would have gone up, not down. (Their word for the generalization gaps is "similar"; strictly, MoS's is a bit narrower — 0.71 against 1.30 points, or ratios of 1.020 and 1.031 — which if anything strengthens the reading.) **The inverse experiment, which is the one I'd point at.** If the mechanism really is the rank bound, then in a setting where the bound cannot bind, MoS should do *nothing*. Character-level language modelling is exactly that setting: $\text{rank}(\mathbf{A}) \leq |V| \approx 27$, and $d$ is in the hundreds, so there is no bottleneck to break. On text8, at matched parameter counts: | model | params | test BPC | |---|---|---| | Softmax (hid 1024, emb 1024) | 8.42M | 1.49 | | MoS-7 (hid 910, emb 510) | 8.45M | 1.49 | | MoS-10 (hid 860, emb 452) | 8.43M | 1.49 | Identical. A method that improves everything improves this too; a method that breaks a specific bound does nothing when the bound is absent. Papers that predict their own null results are rare, and this one went and measured it. ## What it won State of the art at the time, at comparable model size: | benchmark | best prior | MoS | |---|---|---| | Penn Treebank (dynamic eval) | 51.1 | **47.69** | | WikiText-2 (dynamic eval) | 44.3 | **40.68** | | 1B Word (their own softmax baseline) | 42.77 | **37.10** | 22M parameters on PTB against 24M baselines, and 35M on WT2 against 33M — so slightly under on one and slightly over on the other, which is the honest way to read "comparable." The 1B Word row is the one that ages best: 5.67 points on a dataset large enough that regularization tricks aren't doing the work, against a plain 2-layer LSTM softmax at 119M parameters, with hyperparameters they admit they never tuned. They also bolt MoS onto a Seq2Seq decoder for dialogue on Switchboard and it wins on perplexity and on every BLEU precision and recall figure, which is a reasonable check that this is about context-dependent distributions in general rather than about language-modelling benchmarks in particular. ## So why isn't it in your model Cost. $K$ softmaxes means $K$ passes over the vocabulary. Measured at matched batch size it's 1.9× on PTB, 2.5× on WikiText-2, 3.8× on 1B Word; at the settings where each model does its best, 2.8× and 6.4× on one GPU. Sub-linear in $K$ thanks to GPU matmul efficiency, but "sub-linear" still means two to three times the training cost, and that is before you consider that the output layer is now $K$ times the memory. Then scale the setting. PTB has a 10,000-token vocabulary. A modern model has 150,000–200,000, and the vocabulary projection is already one of the most expensive tensors in the network — it's why chunked-and-fused cross-entropy kernels exist at all. Fifteen softmaxes over 200,000 logits per position is not a rounding error, it is the model. So the field made a choice, and the choice was not obviously wrong: buy quality with data and depth, where the cost curve is friendlier, and leave the output layer alone. ## What happened to the idea It didn't disappear so much as fragment into a small literature that no one reads together. [Sigsoftmax](https://arxiv.org/abs/1805.10829) (Kanai et al., 2018) re-derives the bottleneck and argues the culprit is specifically the exponential in softmax, proposing a cheaper output nonlinearity that also escapes the rank limit — the same diagnosis, a one-softmax fix. [Stolen Probability](https://arxiv.org/abs/2005.02433) (Demeter, Kimmel, Downey, 2020) finds a different consequence of the same geometry: embeddings in the interior of the convex hull of the embedding cloud can never be the argmax, no matter the context, so certain words are structurally unpredictable. And the honest counterweight, [Low-Rank Softmax Can Have Unargmaxable Classes in Theory but Rarely in Practice](https://arxiv.org/abs/2203.06462) (Grivas, Bogoychev, Lopez, 2022), goes looking for that failure in real systems: 13 of 150 public models have unargmaxable tokens, and they are rare enough not to matter. Which is a useful correction — a bound being real is not the same as a bound costing you anything — though note it tests one specific symptom, the argmax-unreachable token, not the broader claim that the expressible distribution is lower-rank than the one you want. The bottleneck's authors moved on to [Transformer-XL](https://arxiv.org/abs/1901.02860) and [XLNet](https://arxiv.org/abs/1906.08237), and Zhilin Yang went on to found Moonshot AI, whose [Kimi K3](/articles/kimi-k3) ships a 7,168-dimensional hidden state against a 163,840-token vocabulary — a ratio of 23, in a model from the person who wrote the paper about the ratio. (Kimi K2 sits in the chart below at the same two numbers.) ## The ratio, today I pulled `hidden_size` and `vocab_size` from published configs to see whether nine years of scaling relaxed the geometry. It didn't. The paper's own bottlenecked setup — the one where breaking the bound was worth 3.6 perplexity — had $|V|/d = 25$. Almost every open model checked is above that, several by a factor of six. Anchoring on the paper's own numbers: from PTB's 10,000 tokens to a modern 151,936, vocabulary grew about 15×, while $d$ went from 400 to roughly 4,096 — about 10×. And the mismatch lands hardest on small models, which inherit a large tokenizer from their big siblings and get a fraction of the width to read it with. Qwen3-0.6B carries the same 151,936-token vocabulary as Qwen3-32B through 1,024 dimensions instead of 5,120. I want to be careful about what that does and doesn't show. $|V|/d$ is geometry, not severity. Nobody knows $\text{rank}(\mathbf{A})$ for natural language; a 2026 model has representations no 2017 LSTM had; and the bottleneck may be so far from binding at this scale that it costs nothing measurable. The claim is narrower and, I think, harder to argue with: **the constraint everyone stopped worrying about is tighter now than when they stopped**, and the experiment that would tell us what it costs — an MoS-style rank measurement on a modern LLM's output layer — appears not to have been run. ## Why I keep coming back to it The [full-bandwidth transformer](/articles/full-bandwidth-transformer) paper argues that the feedback path between decoding steps is one token wide — $\log_2|V|$ bits — while the hidden state that produced that token gets thrown away, and that chain-of-thought is partly a workaround for the narrow pipe. That is the same argument, at the other end of the model. Both say: the network is not the limiting factor, the *interface* is. One narrow shape sits between a rich internal state and the thing you actually want, and everything upstream is spending its capacity on getting through it. The best thing about the softmax bottleneck paper isn't the mixture of softmaxes, which nobody uses. It's that it demonstrated the move: take the part of the architecture that's so standard nobody writes it down, ask what it structurally cannot do, and then — this is the rare part — go and measure whether it costs anything. Softmax 400. MoC 280. Character-level, no change. Most papers proposing an architecture would have stopped after the perplexity table. --- # speech-to-speech: the OpenAI Realtime API, reimplemented as four swappable parts > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/speech-to-speech > date: 2026-08-14 > tags: speech, voice-agents, open-source, realtime, latency, explainer [speech-to-speech](https://github.com/huggingface/speech-to-speech) is Hugging Face's voice-agent pipeline: VAD → STT → LLM → TTS, each stage in its own thread, connected by queues. Apache 2.0, on PyPI, first commit 2024-08-07 and still being merged the day I looked. Two things make it worth more than a glance. The first is a compatibility decision. The second is a latency trick that I think is the actual contribution. ## The compatibility decision The server speaks the **core OpenAI Realtime GA event set** over WebSocket and WebRTC. Not a similar protocol — the same one, at `/v1/realtime`, such that the official OpenAI client connects to it by changing a URL: ```python client = OpenAI( base_url="http://localhost:8765/v1", websocket_base_url="ws://localhost:8765/v1", api_key="not-needed", ) with client.realtime.connect(model="local") as conn: ... ```
The repo is careful about the size of this claim, and I want to repeat its wording rather than improve on it: > This is a tested core subset, not a claim of full OpenAI Realtime API equivalence. What is implemented inbound: `input_audio_buffer.append`, `session.update`, `conversation.item.create`, `conversation.item.truncate`, `response.create`, `response.cancel`. Outbound: speech start/stop, streaming transcription, audio deltas, tool calls, `response.done`. CI connects pinned `@openai/agents` `RealtimeSession` instances through the SDK's own WebSocket *and* WebRTC transports — so the compatibility claim is tested against the real client library, not against a hand-written mock. That is the difference between "OpenAI-compatible" as a marketing word and as an engineering commitment.
## Ninety pipelines Six STT backends, three LLM backends, five TTS backends, one VAD. The defaults are Parakeet TDT for transcription and Qwen3-TTS for output — both local — with the LLM slot pointed at anything speaking OpenAI protocols. That last choice is worth pulling apart, because "OpenAI-compatible API" sounds like a dependency and is not one. Point it at `llama-server` on localhost: ```bash llama-server -hf ggml-org/gemma-4-E4B-it-GGUF -np 2 -c 65536 -fa on --swa-full speech-to-speech serve \ --model_name "ggml-org/gemma-4-E4B-it-GGUF" \ --responses_api_base_url "http://127.0.0.1:8080/v1" \ --responses_api_api_key "" ``` Now the whole pipeline is local and the LLM is still reached over HTTP. Keeping the model behind a protocol boundary rather than in-process is what makes the slot genuinely swappable — the pipeline never learns which model it is talking to. For fully disconnected operation, run the exact configuration once online to warm the caches, then set `HF_HUB_OFFLINE=1`. The repo is specific that this covers STT, LLM, TTS, Silero VAD, NLTK *and* Smart Turn assets — the kind of list you only write after being caught out by one of them. There is also a `--stt none` mode that skips transcription entirely and hands each VAD-segmented audio chunk straight to an audio-input model over `/v1/chat/completions`. The README is blunt that this needs a model that actually accepts audio, and that the default `gpt-5.4-mini` does not. ## Smart Turn is the interesting part Silero VAD tells you *that* speech stopped. It cannot tell you whether the person was **finished**. That gap is why voice agents interrupt people who paused to think. [Smart Turn v3.2](https://huggingface.co/pipecat-ai/smart-turn-v3) classifies the turn using content and prosody, and speech-to-speech wires it in speculatively rather than as a gate: - **Complete turns** start STT and the LLM immediately, with `--speculative_reopen_ms` (800 ms) before output is committed. - **Incomplete turns** wait `--smart_turn_incomplete_delay_ms` (600 ms) before spending anything, and their output stays gated by `--smart_turn_max_wait_ms` (2 s). - **If speech resumes during either delay**, the turn is reopened as a newer revision, the accumulated audio is re-emitted, and work from the previous revision is discarded *before it reaches the user*. That third rule is what makes the first two safe. Speculation is only free if the wrong guesses are invisible, and revision numbering is the mechanism that makes them invisible. You spend tokens you might throw away in exchange for latency you cannot otherwise reclaim — a reasonable trade in a pipeline where every millisecond between "user stopped" and "audio starts" is audible. It ships enabled by default, as a quantized CPU ONNX checkpoint, so the cost of running it is not a GPU. ## The part that gives it weight > This pipeline runs in production as the conversation backend for thousands of [Reachy Mini](https://huggingface.co/blog/reachy-mini) robots. Voice-agent demos are cheap and voice agents that hold up in a room with background noise are not. A deployed fleet is the only evidence that separates the two, and it is the reason to read this repo rather than one of the dozens of similar cascades. ## What to be aware of **The Realtime surface is a subset**, and the repo says so. If your client depends on an event outside the tested set, it will not work, and "OpenAI-compatible" will have been true and useless simultaneously. **No latency numbers.** For a project whose headline is "low-latency," there is no published end-to-end figure — no time-to-first-audio, no comparison against hosted OpenAI Realtime, on any hardware. The architecture is clearly built for latency; how much it achieves is unmeasured in public. **Installation has sharp edges.** The Qwen3-TTS GGML backend's default PyPI wheel targets CUDA 12.8, and the README carries a table of alternate wheels for CUDA 13.x, 12.4, and CPU. That is honest documentation of a real problem, and also a sign of how much platform-specific machinery sits under `pip install speech-to-speech`. **Some components moved to `archive/`** — Moonshine STT, MeloTTS, Parler TTS. Worth knowing before you build on a backend that is on its way out. The design I would steal is the one at the boundary: implement someone else's protocol exactly enough to be tested against their client, then make everything behind it yours. It converts a hosted API from a dependency into an interface. --- # Toast 1: what happens when you stop making the frontier model do the searching > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/toast-1 > date: 2026-08-14 > tags: retrieval, agents, rag, search, cost, explainer [Toast 1](https://www.mixedbread.com/blog/toast-1) is Mixedbread's first specialised search agent, released 2026-08-13. The pitch is division of labour: instead of a frontier model burning its context window navigating a corpus, Toast 1 takes the whole search loop — decomposes the query into subqueries, gathers evidence, inspects sources, curates what matters — and hands back a package. The frontier model spends its tokens reasoning instead. It runs standalone or as a subagent, and it is backend-agnostic: co-designed with Mixedbread Search but able to run over an existing index. ## The result worth reading twice Harvey's LAB firm-knowledge benchmark, on a 33-task subset, with **one model and one evaluation** and only the retrieval stack changing between runs: | configuration | tokens | turns/task | score | |---|---|---|---| | vanilla agent | 80.6M | 21.7 | 55 | | + Mixedbread Search | 47.0M | 14.6 | 55 | | + Toast 1 subagent | 23.0M | 11.2 | 55 | The score is the finding. It does not move. If quality had gone up you would be looking at a better agent; because it is identical across all three rows, what the experiment demonstrates is that **57.6M of the vanilla agent's 80.6M tokens were not contributing to the answer**. They were the cost of looking. I recomputed the deltas and they reproduce: 47.0/80.6 is −41.7% against a stated −42%, 23.0/47.0 is −51.1% against −51%, and 80.6/23.0 is 3.50× against a stated 3.5×. Two caveats belong right next to that, and Mixedbread states the first itself in a footnote: this is a **randomly selected 33-task subset**, chosen "to make repeated comparative runs tractable." And a score that lands on exactly 55 three times is coarse enough that a small quality change would not necessarily show up in it. The token reduction is a much more precisely measured quantity than the quality preservation it is paired with. ## The headline benchmark, and whose numbers they are
On [OfficeQA Pro V2](https://www.mixedbread.com/blog/toast-1) — 90 questions on enterprise financial situations, released by Databricks — GPT-5.6 Sol running in Codex with Toast 1 as a subagent reaches **70% correctness at about $1.15 per task**. The previous best in Databricks' evaluation, Claude Fable 5 on Databricks Genie, was 60% at roughly $4. The comparison that carries the most information is the one against itself: **GPT-5.6 Sol in Codex without Toast 1 reaches 33%**. Same model, same harness, and correctness doubles when the search loop is delegated. The chart's own footnote is the thing to hold onto: "Genie and harness numbers as reported by Databricks; Codex + Toast 1 runs are ours." Half the points come from the benchmark's authors and half from the vendor being evaluated. That is a normal and disclosed arrangement, and it is still a different evidential status than a single evaluator running everything. ## As a standalone retriever
Evaluated as a retriever rather than a subagent — BrowseComp Plus, OfficeQA Pro and LongSeal, scored by NDCG@10 — Toast 1's fusion configuration lands in the same band as GPT-5.6 Sol and above Kimi K3, GLM, Opus 5 and Sonnet 5, while sitting an order of magnitude to the left on cost. The chart shows something the prose does not dwell on: every other system is drawn as a *sweep*, a short line tracing what more reasoning effort buys. Toast 1's line is short and nearly flat. Whatever it is doing, spending more on it does not move quality much — which is the expected shape for a specialised model that is already doing the one thing it was trained for. ## The economics A standard run is **$0.016–$0.023 per query at an eight-second median**; the fusion configuration is **$0.05–$0.07 at eleven seconds**. Token pricing is $0.30/M input, $0.04/M cached input with free cache writes, and $0.80/M output. Against the frontier retrieval agents in the same evaluation, Mixedbread claims 7–11× cheaper. That multiple is the only cost figure given for the comparison group, so the band in the diagram above is their claim inverted rather than a published measurement — worth flagging, because the latency comparison beside it needs no such inference: **20 seconds to four minutes**, quoted directly. An eight-second search subagent and a four-minute one are different products before price enters the discussion. If a frontier model is going to call search several times per task, the difference compounds into whether the task is interactive at all. ## What is not disclosed No architecture. No parameter count. No training details, data, or method. No information about what Toast 1 is beyond what it does and what it costs. This is a product launch, not a model release, and every number in it is a system-level measurement. That matters for one specific reason: the headline results are all **system** results. "GPT-5.6 Sol + Toast 1 in Codex reaches 70%" is a claim about a pipeline with at least three moving parts, and the contribution of each is not separable from the published data. The Harvey ladder is the closest thing to a controlled experiment on offer, and it is the one I would weight most — one model, one task set, one evaluator, one variable. Mixedbread's own footnote places this alongside SID-1 and Chroma's Context-1 as a growing category of specialised search agents. That framing is right, and it is the more interesting story than any single benchmark: the bet is that retrieval is a distinct enough skill to be worth a dedicated model, and that the frontier model's context window is too expensive to spend on navigation. The Harvey numbers are the strongest evidence for that bet I have seen stated plainly. Three quarters of a vanilla agent's tokens went to finding things, and removing that cost changed nothing about the answers. --- # WorldClaw: a 3D world generator that is really a Blender programmer > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/worldclaw > date: 2026-08-14 > tags: 3d-generation, agents, world-models, procedural-generation, paper, explainer [WorldClaw](https://arxiv.org/abs/2608.05248) (arXiv 2608.05248, 2026-08-05, Tencent Hunyuan) generates large, freely explorable 3D worlds from open-ended text. The thing that separates it from most work in this area is the output format: not a radiance field, not a mesh soup, but **explicit instance-level assets sitting on a continuous terrain**, all of it editable afterwards in Blender. The way it gets there is the interesting part. WorldClaw does not generate 3D. It writes programs that generate 3D. ## The pipeline
Formally the paper writes it as three functions and a compose: - `P = F_plan(q)` — a text prompt becomes a structured specification of regions, terrain conditions and object conditions - `T = F_terrain(P)` — the specification becomes a global terrain - `O = F_region(P, T)` — regions that need detail get objects, conditioned on the terrain already built - `S = Compose(T, O)` The ordering carries the whole design. Terrain is built first and objects are generated *conditioned on it*, so a hut sits on a slope the system already knows about rather than being placed onto a surface it has to discover.
That middle panel is the system in miniature. The agent's deliverable is a Python function that composes noise octaves and landform primitives into a height field. Not a heightmap image — a program that computes one. ## The height field Equation 6 is the core of the terrain stage: `H(x) = Σ_r m̃_r(x) · [ h_r + Σ_k w_r,k N_r,k(x) + Σ_j α_r,j G_r,j(x) ]` Each region *r* contributes a base elevation `h_r`, a weighted sum of noise octaves `N`, and a weighted sum of landform primitives `G` — and the whole contribution is gated by a **normalized** region mask `m̃_r`. The normalization is the part worth dwelling on. Because the masks sum to one everywhere, adjacent regions blend rather than abut. A beach becomes a forest without a seam, and no post-hoc stitching step is needed. That is what lets the planning agent describe regions independently — writing "coastal", "dense jungle", "volcanic ridge" as separate specifications — and still get a single continuous world out. It is also why this is a genuinely different approach from tiling. There is one field. It just happens to be authored per region. ## Placing objects is a solved geometry problem For regions that need detail, WorldClaw renders the terrain from a viewpoint, generates a **composition image** conditioned on that render, segments the objects out of it, reconstructs each as a textured mesh, and then has to work out where each mesh goes in 3D. That last step is where the paper does real work rather than prompting. Placement is recovered by solving for a similarity transform per object: a scale from the ratio of depths and focal lengths (`s_i = (Z_t/Z_o)(f_i^o/f̂_i)`), a rotation, and a translation, assembled into `T_place`. There is also a bounded contact constraint — the projected base of each object has to land within a tolerance band of a reference height, written as a two-sided inequality rather than an exact equality. This is the difference between an agent that *asks* a model where the tree goes and one that computes it. The bounded constraint in particular is doing something specific: it permits a tree to sink slightly into a slope or stand slightly proud, which is what contact looks like on real terrain, while forbidding it from floating. ## Six models in a trench coat WorldClaw uses **Claude Opus 4.8** as the agent model, with task-specific skills that wrap GPT-Image-2, SAM3, SAM3D and Hunyuan3D, executing into **Blender 5.1.1** on 4× NVIDIA H20 GPUs. The Limitations section is more candid than most, and it is the most useful part of the paper: > In our experiments, current open-source language models often struggled to generate procedural terrain and materials that were both executable and consistent with user requirements. Likewise, open-source image generation models frequently failed to produce usable semantic layout maps or to preserve object appearance and pose. And then, plainly: > Consequently, fully validating this decoupled pipeline at the current stage still requires capable models such as Claude Opus 4.8, GPT-Image-2, and Hunyuan3D. Decomposing a task into stages is supposed to make each stage easier. Here it did the opposite for the two stages whose output has to be *executable*: a plan that becomes a Blender program either runs or does not, and a layout map is either segmentable or is not. Neither degrades gracefully when the model gets weaker. The second limitation is the one anyone building on this should read twice. Several stages depend on LLM-generated programs, and: > Errors in scale estimation, numerical parameters, or node connectivity directly manifest in the resulting 3D scene as inconsistent landforms, inaccurate material effects, or object layouts that deviate from the user intent, often necessitating multiple render–inspect–refine iterations. A wrong number in a generated program is not a crash, it is a mountain in the wrong place. The render-inspect-refine loop exists because the failure mode is silent and visual. ## What the paper does not contain There are no quantitative results. Section 3.2 is "Qualitative Results" and 3.3 is "Qualitative Comparison"; there are thirteen tables in the HTML and every one of them is a display equation. No user study, no CLIP or FID-style score, no timing table, no ablation with numbers.
For a system paper this is more defensible than it would be for a model paper — the claim is "you can build worlds this way and they are editable afterwards," and a render plus an instance mask demonstrates that. But it means nothing here is measured. The comparison figures show WorldClaw's worlds looking bigger and better organized than the baselines', and that is the entire evidential basis. The third limitation is the honest counterweight: generating and reconstructing every object separately, then iterating refinement over terrain, assets and contacts, "incurs substantial inference latency and computational cost," and the pipeline "can be unnecessarily lengthy and inefficient for simpler scenes that holistic generation methods can synthesize in fewer steps." No wall-clock figure is given for either. ## Why it is still worth attention The output format is the argument. A generated radiance field is a thing you can look at; a terrain with named regions, instance-level meshes, PBR materials and solved placements is a thing you can *open* — move a building, restyle a material, swap an asset, run physics against the ground. The paper's stated next step is generating objects as executable node graphs too, which would make composition and material logic editable in the same way the terrain already is. The cost is that WorldClaw is currently less a model than an orchestration of four proprietary and open systems that its own authors could not substitute. Whether that is a stepping stone or a ceiling depends entirely on whether open models get good enough at writing Blender. --- # Liquid time constants and gated delta rules: two literatures, one recurrence > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/ltc-gated-delta > date: 2026-08-10 > tags: linear-attention, state-space-models, attention, math, explainer, open-source There are two separate research literatures about neural networks that forget at an input-dependent rate, and as far as I can tell they mostly do not read each other. One starts in continuous time. [**Liquid Time-constant Networks**](https://arxiv.org/abs/2006.04439) (Hasani, Lechner, Amini, Rus and Grosu, 2020) are ODEs, motivated by the neural dynamics of *C. elegans*, analysed with stability theorems and solved with numerical integrators. The other starts in discrete time. [**Gated DeltaNet**](https://arxiv.org/abs/2412.06464) (Yang, Kautz and Hatamizadeh, 2024) is a linear attention variant, motivated by retrieval failures in efficient Transformers, analysed through online learning and implemented with chunkwise GPU kernels. They are the same recurrence. Not analogous — the same. This piece derives that correspondence from each paper's own equations, works through what each tradition figured out that the other did not, and then reads [**LTCAttention**](https://github.com/Rikka-Botan/LTCAttention), an implementation published today that sits deliberately between them. This is a mechanism piece, so the maths is the point rather than an aside. Everything is derived from the two papers' numbered equations, and the LTCAttention section is read off its source and its checked-in result JSON. Where I run a calculation the papers do not — the discretization bridge, and a scaling estimate at the end — I say so and show the working. ## Part 1 — What "liquid" means A plain continuous-time RNN decays toward its input at a fixed rate: $dx/dt = -x/\tau + S(t)$, where $\tau$ is a learned constant. Every input, every timestep, same $\tau$. LTC's move is to let the decay rate be a function of the state and the input. Substituting $S(t) = f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)(A - \mathbf{x}(t))$ gives the paper's Equation 1: $$ \frac{d\mathbf{x}(t)}{dt} = -\left[\frac{1}{\tau} + f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)\right]\mathbf{x}(t) + f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)\,A $$ Read the bracket. The coefficient multiplying $\mathbf{x}$ is the decay rate, and it now contains $f$ — a neural network. So the **system time constant** is $$ \tau_{\text{sys}} = \frac{\tau}{1 + \tau f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)} $$ which is a number the network computes fresh at every point in time from whatever it is currently looking at. That is the whole idea, and the name: a time constant that flows. The paper's two theorems are what make this more than a reparameterization. Because $f$ is a bounded sigmoidal nonlinearity, **Theorem 1** traps the time constant: $$ \frac{\tau_i}{1 + \tau_i W_i} \le \tau_{\text{sys}_i} \le \tau_i $$ and **Theorem 2** traps the state itself between $\min(0, A_i^{\min})$ and $\max(0, A_i^{\max})$, "which guarantees that the outputs of LTCs never explode even if their inputs grow to infinity." Those are unusual guarantees. A model whose decay rate is an unconstrained network output could in principle be driven to instability by an adversarial input; LTC's cannot, by construction. This is worth flagging because the same argument recurs, unattributed, throughout modern gated linear attention. Every one of these architectures constrains its gate — Mamba2 through a softplus and a discretization, Gated DeltaNet by requiring $\alpha_t \in (0,1)$, [Kimi K3's KDA](/articles/kda-half-life) through a `gate_lower_bound` on $\log\alpha$. The reason is always the same one LTC proved in 2020: an unbounded forgetting rate is an unbounded system. ## Part 2 — The bridge Now the part neither literature states, which falls out of LTC's own **Algorithm 1**. Solving Equation 1 in closed form is not possible, so the paper introduces a *fused solver* — a semi-implicit Euler step that reads $$ \mathbf{x}(t + \Delta t) = \frac{\mathbf{x}(t) + \Delta t \, f(\mathbf{x}(t), \mathbf{I}(t), t, \theta) \odot A}{1 + \Delta t\left(\frac{1}{\tau} + f(\mathbf{x}(t), \mathbf{I}(t), t, \theta)\right)} $$ Look at the denominator. Writing $g$ for the gate output, $$ 1 + \Delta t\left(\tfrac{1}{\tau} + g\right) = 1 + \Delta t\,\frac{1 + \tau g}{\tau} = 1 + \frac{\Delta t}{\tau_{\text{sys}}} $$ because $\tau_{\text{sys}} = \tau/(1 + \tau g)$ by definition. So the entire update is $$ \mathbf{x}_{t+1} = \bar{\alpha}\,\mathbf{x}_t + \bar{\alpha}\,\Delta t\, g \odot A, \qquad \bar{\alpha} = \frac{1}{1 + \Delta t/\tau_{\text{sys}}} $$ **That is a gated linear recurrence.** Previous state times a scalar in $(0,1)$, plus a write. The scalar depends on the input, through $g$. This is structurally identical to what Mamba2, Gated DeltaNet and KDA do — the object those papers call $\alpha_t$ and describe as a "data-dependent gating term." The only difference is which approximation of the exponential you use. The exact solution of the linear ODE over a step decays by $e^{-\Delta t/\tau_{\text{sys}}}$; LTC's fused solver uses $1/(1 + \Delta t/\tau_{\text{sys}})$, which is the $[0/1]$ Padé approximant of that exponential. In the regime these models actually operate in — long memory, so $\Delta t \ll \tau_{\text{sys}}$ — the two agree to a fraction of a percent. LTC picked the Padé form because it is what makes the implicit Euler step solvable in closed form; the linear-attention literature picked the exponential because $\alpha^n$ composes cleanly across a chunk, which is what its parallel scan needs. Same recurrence, two discretizations, chosen for two different implementation reasons. Which means the three quantities have one meaning: | tradition | symbol | this article's other coverage | |---|---|---| | continuous-time / LTC | $\tau_{\text{sys}}$, a time constant in seconds | Part 1 above | | gated linear attention | $\alpha_t = e^{-\Delta t/\tau}$, retention per token | [KDA has a half-life](/articles/kda-half-life) | | what you should think in | $n_{1/2} = \ln 0.5 / \ln \alpha$, a horizon in tokens | same | I have argued the third column before, and the LTC connection strengthens it: a half-life is just $\tau_{\text{sys}}$ in units a language model can be reasoned about in. $\tau$, $\alpha$ and $n_{1/2}$ are one number in three coordinate systems. ## Part 3 — What gating alone cannot do If the story ended there, Gated DeltaNet would be LTC with better kernels. It is not, and the difference is the delta rule. A gated linear attention state is a matrix $\mathbf{S}$ holding key-value associations. Pure gating updates it as $\mathbf{S}_t = \alpha_t \mathbf{S}_{t-1} + v_t k_t^\top$: scale everything down, add the new pair. The problem the Gated DeltaNet paper identifies is that $\alpha_t$ is a single number multiplying the entire state. It can dump everything, and it can hold everything, and it has no way to express *forget this one fact, keep the rest*. DeltaNet solved that with the delta rule, which subtracts the state's existing content at the current key before writing the new one — but, as the paper puts it, "since this process only modifies a single key-value pair at a time, the model lacks the ability to rapidly clear outdated or irrelevant information, especially during context switches." One mechanism clears the table but cannot pick up a single plate; the other picks up single plates but cannot clear the table. The gated delta rule (Equation 8) is both terms in one product: $$ \mathbf{S}_t = \mathbf{S}_{t-1}\left(\alpha_t\left(\mathbf{I} - \beta_t k_t k_t^\top\right)\right) + \beta_t v_t k_t^\top $$ The two bars are the whole argument. Along any direction orthogonal to the current key, the surviving fraction is $\alpha_t$ — the global forgetting knob. Along $k_t$ itself it is $\alpha_t(1 - \beta_t)$ — global decay *and* targeted erasure. Set $\alpha_t \to 1$ and you have DeltaNet; set $\beta_t \to 0$ and you have Mamba2; the useful region is the interior. Worth keeping straight against the version of this update I covered in [KDA has a half-life](/articles/kda-half-life). Kimi K3 writes it as $S_t = (I - \beta_t k_t k_t^\top)\,\mathrm{Diag}(\alpha_t)\,S_{t-1} + \beta_t k_t v_t^\top$ — the same two factors, transposed convention, but $\alpha_t$ is a **vector** with one entry per channel rather than Gated DeltaNet's **scalar** per head. That is not cosmetic. A scalar $\alpha$ gives a head one memory horizon; a diagonal $\mathrm{Diag}(\alpha)$ gives it a whole spectrum at once, which is the difference between a head that forgets at one rate and a head that runs a filter bank. ## Part 4 — LTCAttention, and a third place to put a time constant Both traditions above put the time constant on a **recurrent state**. [LTCAttention by Rikka Botan](https://github.com/Rikka-Botan/LTCAttention), published today under MIT, puts it somewhere else: on the attention score itself.
The construction is worth following because it is genuinely clever. Each KV head carries $M$ learned directions, orthonormalized by QR so that $u_m^\top u_n = \delta_{mn}$. The first token of the causal block, $x_0$, sets every mode's time constant through one linear projection: $$ \tau_{h,m}(x_0) = \frac{\tau_{\min}}{\sigma\!\left(r_{h,m} + \delta_{h,m}\right)} > \tau_{\min} $$ That is the LTC principle exactly — a positive, input-conditioned, *bounded-below* time constant, with the sigmoid playing the role LTC's Theorem 1 played. Because $x_0$ is visible to every position in the block, reading it keeps the controller causal. For a query at $i$ and a key at $j \le i$, with key age $\Delta = i - j$, mode $m$ retains $\lambda_m(\Delta) = e^{-\Delta/\tau_m}$, and the modes assemble into $$ M_\Delta(x_0) = \prod_{m=1}^{M}\left[\mathbf{I} - (1 - \lambda_m)u_mu_m^\top\right] = \mathbf{I} + \sum_{m=1}^{M}\left(\lambda_m(\Delta, x_0) - 1\right)u_mu_m^\top $$ which drops into the score as $s_{ij} = q_i^\top M_{i-j}(x_0)\,k_j/\sqrt{d}$. The effect: the learned orthogonal complement passes through untouched, while each temporal mode is an eigenvector with eigenvalue $\lambda_m$. Since $\lambda_m(\Delta) = a_m^\Delta$ with $a_m = e^{-1/\tau_m}$, this is the same stable diagonal decay law as an SSM — just expressed as a metric on an inner product rather than a state update. ### The factorization is the load-bearing trick, and it checks out Applying a different $M_\Delta$ to every $(i,j)$ pair naively means building a $T \times T \times d$ object. LTCAttention avoids it by pushing the decay into the queries and keys separately, around a fixed center $c$: $$ q_i' = q_i + \sum_m \left(e^{-\frac{i-c}{\tau_m}} - 1\right)(q_i^\top u_m)u_m, \qquad k_j' = k_j + \sum_m \left(e^{\frac{j-c}{\tau_m}} - 1\right)(k_j^\top u_m)u_m $$ I checked the algebra rather than taking it on faith. Decompose $q_i = q_\perp + \sum_m (q_i^\top u_m)u_m$ using orthonormality; the transform replaces each modal coefficient by $e^{-(i-c)/\tau_m}(q_i^\top u_m)$ and leaves $q_\perp$ alone, and symmetrically for $k$. Their inner product is then $$ q_i'^\top k_j' = q_\perp^\top k_\perp + \sum_m e^{-\frac{i-c}{\tau_m}}e^{\frac{j-c}{\tau_m}}(q_i^\top u_m)(k_j^\top u_m) = q_\perp^\top k_\perp + \sum_m e^{-\frac{i-j}{\tau_m}}(q_i^\top u_m)(k_j^\top u_m) $$ and expanding $q_i^\top M_\Delta k_j$ directly gives the same thing. The center $c$ cancels, exactly as claimed. The modal projections cost $O(TMd)$, so **scaled dot-product attention remains the only quadratic operation** — the mechanism is free at the asymptotic level and the standard SDPA kernel is still doing the heavy lifting. There is a real numerical hazard hiding in that trick, and the code knows it. The factors $e^{-(i-c)/\tau}$ and $e^{(j-c)/\tau}$ are individually huge or tiny even though their product is bounded by 1; they cancel only algebraically. The implementation handles this two ways. It computes the exponents in FP32 or FP64 regardless of the BF16 activation dtype, with a comment saying exactly why. And it fixes $c$ at the middle of the context, `centre = 0.5 * (max_positions - 1)`, rather than recomputing it per prefix — which both keeps cached keys valid as the KV cache grows and halves the worst-case exponent. The choice of $\tau_{\min}$ then finishes the job, and this is my favourite detail in the repository. The default is `min_tau = max_positions / 12`. Combined with the centered origin, the largest exponent magnitude is $$ \frac{(T-1)/2}{T/12} = \frac{6(T-1)}{T} \approx 6 $$ **independent of context length.** Whatever $T$ you configure, the factorization's intermediate values stay inside roughly $e^{\pm 6}$. That is not a coincidence; it is a bound chosen so the trick cannot overflow. ### The experiment, and the number that worries me The repository ships a real controlled study rather than a claim: three seeds, a paired comparison, SHA-256 checksums on the tokenized data, one epoch over 287,588,352 FineWeb-Edu tokens consumed without replacement, and the full result JSON checked in.
Reading the numbers straight out of `results/fineweb_edu_fullrank_29m_6layer_half_3seeds.json`: | | validation loss | perplexity | |---|---|---| | standard | 4.13471 ± 0.02367 | 62.49 | | LTC | 4.07281 ± 0.01777 | 58.73 | | paired difference | **−0.06190 ± 0.00646** | | The per-seed differences are −0.0544, −0.0702 and −0.0610 — negative in all three, with a spread ten times smaller than the effect. As a paired result at this scale that is about as clean as three seeds get, and the README is careful to say that "three seeds and one small model scale do not establish broad scaling behavior." Two confounds are worth quantifying, and the repository reports exactly the numbers needed to do it. **Parameters.** LTC adds 345,600 of them, +1.20%. Borrowing the Chinchilla-form sensitivity $\partial L \approx \alpha\,(A/N^\alpha)\,(\partial N/N)$ with $\alpha = 0.34$, a 1.20% parameter increase at 28.8M is worth roughly **0.005 nats**. The observed effect is more than ten times that. The gain is not just parameter count. **Compute.** This is the one. LTC also runs **12.40% slower** (142,473 vs 162,638 tokens/sec, measured and reported by the author). The comparison is token-matched, not wall-clock-matched. Spend that same 12.4% on more training tokens for the baseline instead, and the same scaling form ($\beta = 0.28$ on the data term) predicts a gain of roughly **0.061 nats** — which is, to two decimal places, the entire measured effect. I want to be precise about what that estimate is and is not. The coefficients come from a scaling law fitted on a different corpus, tokenizer and budget, so the *absolute* numbers do not transfer; I am borrowing only the sensitivity, and the error bars on that are wide. The near-exact agreement between 0.061 and 0.062 is a coincidence of a rough calculation, not a measurement. But the direction is robust: at this scale, a 12% throughput penalty buys enough extra tokens to be the same order as the observed quality gain. **The missing experiment is a wall-clock-matched run**, and until someone does it the honest reading is that LTCAttention is better per token and undetermined per second. That the author reported the throughput cost at all is what makes this check possible — most releases do not. One further limitation, stated plainly in the repo: the released code is LTC-only, and the baseline artifacts are "retained only as experiment provenance." So the comparison cannot currently be re-run from this repository, only re-read. ## What each tradition knows Setting the implementations aside, the two literatures have complementary blind spots. **LTC knows about stability and it knows about time.** It has proofs that the time constant and the state stay bounded under arbitrary input. It treats $\Delta t$ as a real quantity, which means it handles irregularly sampled sequences natively — a capability the discrete-time literature mostly gave up without noticing, because tokens arrive on a uniform grid. And it thinks in a unit, seconds, that forces you to ask how long a memory is supposed to last. **Gated linear attention knows about scale and it knows about writing.** It has the chunkwise parallel algorithms that make these recurrences trainable on modern hardware at all, which is the entire reason the idea reached billion-parameter models. And it has the delta rule — a way to modify one association without disturbing the others that has no counterpart in the LTC formulation, where the "write" is just $f \cdot A$ added to a decaying state. LTCAttention is interesting mostly as evidence that the gap is crossable in either direction: it takes LTC's bounded input-conditioned $\tau$, GDN's adaptive retention, and applies them to a third substrate neither paper considered. Whether that particular hybrid pays for its 12% is, on the evidence available, not yet settled. Whether the two literatures should be reading each other seems to me much clearer. --- *Sources: [Liquid Time-constant Networks](https://arxiv.org/abs/2006.04439) (arXiv 2006.04439, Hasani, Lechner, Amini, Rus, Grosu) for Equation 1, Algorithm 1, and Theorems 1–2, read via ar5iv; [Gated Delta Networks: Improving Mamba2 with Delta Rule](https://arxiv.org/abs/2412.06464) (arXiv 2412.06464, Yang, Kautz, Hatamizadeh) for Equation 8 and the complementarity argument; and the [LTCAttention repository](https://github.com/Rikka-Botan/LTCAttention) at its 2026-08-10 state — `README.md`, `model.py`, `config/`, and `results/fineweb_edu_fullrank_29m_6layer_half_3seeds.json`. The three figures are LTCAttention's own, flattened onto white. The fused-solver-to-gated-recurrence derivation, the verification of the query-key factorization, the $\tau_{\min} = T/12$ bound, and both scaling estimates are mine and are shown in full above so they can be checked. All four interactives are mine.* --- # Muse Glimmer: an agentic model designed backwards from a 24 GB budget > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/muse-glimmer > date: 2026-08-10 > tags: llm, agents, open-weights, on-device, attention, explainer Most model releases describe an architecture and then mention, near the end, what hardware it runs on. [**Muse Glimmer**](https://huggingface.co/meta-models/Muse-Glimmer-30B) — released by Meta Superintelligence Labs on 2026-08-09 under Apache 2.0 — reads the other way round. Nearly every structural choice in it is answering the same question, which is *how do we fit a competent agent, its KV cache at 131K tokens, a vision encoder and a speculative-decoding drafter inside 24 GB at once?* That is a good question to design against, and the interesting thing is that you can check the answer. Every number below comes from `config.json`, the safetensors index, and the file sizes in the four repositories Meta actually shipped. The 24 GB claim is not a vibe; it is arithmetic, and it closes with 2.4 GB to spare. ## The budget, and what makes it work Start at the end. The card says 4-bit quantization shrinks the language model to "under 20 GB," leaving headroom for the KV cache, the perception encoder and the drafter within a 24 GB envelope. Here are the measured artifact sizes from the GGUF repository, plus a KV cache computed from the config rather than quoted: The KV arithmetic is simple enough to do in one line. With 2 KV heads at head dimension 128 in bf16, each token costs `2 × 2 × 128 × 2 = 1024` bytes per layer. Thirteen of the fifty-two layers are global and hold the whole context; the other thirty-nine are capped at a 2048-token window. At the full 131,072-token context that is **1.83 GB** — and 16.76 + 1.63 + 1.40 + 1.83 = **21.6 GB**. Now take away the sliding-window pattern and make all 52 layers global. The KV cache becomes **6.98 GB**, the total becomes 26.8 GB, and it no longer fits on the card the release is named after. That is the sentence worth keeping: the 3-local-1-global stack isn't an efficiency refinement applied to a model that already fit, it is the reason the model fits at all. Take away the 16:1 GQA as well and the KV cache alone is 112 GB. ## The attention stack, verified line by line The card describes the attention as "[Local, Local, Local, Global] repeating" with "RoPE (θ = 500,000), local layers only." Both claims are checkable per layer, because `config.json` carries two 52-element arrays. `layer_types` gives `L L L G` thirteen times exactly. `layer_rope_theta` is 500,000 on every local layer and **0 on every global one**, with zero mismatches across all 52. So the layers that see the entire context run with no positional encoding at all. That is worth pausing on, because this site has now covered three independent labs converging on it inside a year: [Kimi K3](/articles/kimi-k3) applies NoPE to its full-attention layers, and [Maple-Preview](/articles/maple-preview) sets `nope_on_global_attention: true` on the same 3:1 pattern at 24 layers. Muse Glimmer makes it explicit per layer rather than as a flag. The shared argument is length extrapolation: a layer carrying no notion of absolute distance has nothing to be surprised by when the context gets longer than anything it saw in training, and the local layers underneath have already encoded order well enough to reconstruct it. A few things the config says that the card does not: - **`final_logit_softcapping: 20.0`** — logits are squashed through a bounded function before the softmax, a Gemma-style stabilizer that the model card never mentions. - **`qk_scale_factor: 3.87`** — the attention scale is not the textbook $1/\sqrt{d_k}$. At head dimension 128 that would be 0.0884; this multiplies it by 3.87. - **`output_multiplier: 0.19611613513818404`** — which is exactly $16/\sqrt{6656}$, a residual-stream rescale tied to the hidden size. - **`post_norm_eps: 1e-08`**, separate from `rms_norm_eps: 1e-05` — implying a post-norm alongside the pre-norm rather than one or the other. None of these is exotic on its own. Together they're a reminder that the published table of hyperparameters is a summary, and the config is the document. ## It is a distillation, and the blog says so The model card describes what Muse Glimmer is. The blog post describes where it came from, and this is the part that most changes how you should read the benchmark numbers: > **Pre-Training.** We trained Muse Glimmer on Muse Spark's outputs using logit distillation, > leveraging a similar data mix as the teacher. > **Mid-Training.** We trained the model on longer-context, more agent-heavy data with richer > reasoning traces, alongside organic data. > **Post-Training.** We combined supervised fine-tuning with a mix of on-policy distillation and > reinforcement learning across general, reasoning, coding, and agentic domains. Distillation from Muse Spark at *both* ends — logit distillation during pretraining, on-policy distillation during post-training. Muse Glimmer is not a small model trained well; it is a large model compressed, twice, with RL in between. That framing explains the shape of the results better than "30B punches above its weight" does: what transferred is the teacher's *behaviour on agentic trajectories*, which is exactly where the model is strongest. It also sets up the safety argument later on, which leans on Muse Glimmer being "broadly weaker than Muse Spark 1.0" — a claim that is much easier to make about a distilled student than about an independently trained model. ## Speculative decoding, and why the same drafter is worth 3.1× or 1.5× The second optimization is a companion "drafter" based on [DFlash](https://arxiv.org/abs/2602.06036) that proposes an entire block of 16 tokens in a single forward pass, which the main model then verifies in parallel. The shipped drafter is a 5-layer model at the target's full 6656 width — 2.56B parameters, 5.1 GB in bf16, 1.63 GB quantized. Calling it "lightweight" is fair relative to 30B, but it is 8.6% of the model and it has to be resident. The headline is 3.1× on an RTX 5090 and 1.5× on an M4 Max, and the gap between those two numbers is the whole mechanism. Single-stream decoding is memory-bandwidth-bound: you read the entire 17 GB of weights to emit one token. Proposing sixteen and verifying them together amortizes that read across all sixteen, so the gain depends on how much *spare arithmetic* the device has once the weights are already moving. A 5090 has an enormous compute-to-bandwidth ratio and converts nearly the whole block size into speedup; Apple's unified memory narrows that ratio, so verification stops being nearly free.
The chart carries information the card's table drops. Those error bars span seven prompt categories, and on the 5090 the DFlash result runs from roughly 132 to 340 tok/s. So the honest statement is that speculation is worth somewhere between **1.8× and 4.5×** depending on what you ask, and 3.1× is the midpoint of a wide distribution rather than a number you should expect on your workload. ## The benchmark table Meta loses a third of
Re-tallied against the better of the two rivals in each row: Muse Glimmer leads **12 of the 22 scored rows**. Qwen3.6-27B takes 8, Gemma4-31B takes 2, and the losses are not decorative: OSWorld-Verified by 9.7 points and TerminalBench 2.1 by 9.0, both to a model three billion parameters smaller. Filter that ledger to *agentic* and the profile becomes legible. The rows Muse Glimmer wins by a distance — MCP Atlas by 21 points over Gemma, τ³-Banking by 41% relative, DeepSearch QA, Gaia2 — are the ones measuring tool schemas and multi-turn task completion inside a scaffold. The rows it loses are computer-use (OSWorld) and long-horizon terminal work (TerminalBench). For a model distilled specifically on agentic trajectories, that is exactly the shape you would predict, and it is more informative than a uniform win would have been. Publishing it in that form is the least common thing about this release. A comparison table where your own model is beaten in a third of the rows, by a competitor, in your own launch material, is not the norm. ## The table where losing is the point The chem/bio section inverts the usual reading of a benchmark, and it is worth explaining because the presentation is initially confusing. Meta bolds the *most performant* model in each row — and Muse Glimmer is deliberately not it in four of six: | | Muse Glimmer-30B | Gemma4-31B | Qwen3.6-27B | *Kimi K3* | |---|---|---|---|---| | MBCT | 41.5% | **50.6%** | 45.9% | *58.9%* | | HPCT | 52.3% | **54.0%** | 48.7% | *59.6%* | | VCT | 37.0% | **43.5%** | 33.7% | *48.0%* | | WMDP (Bio) | **86.5%** | 85.9% | 84.8% | *89.1%* | | WMDP (Chem) | 75.2% | **80.5%** | 74.8% | *84.2%* | | Lab Bench (ProtocolQA) | **80.2%** | 75.8% | 69.1% | *81.9%* | The argument is that Muse Glimmer sits "approximately in line with other models in its size class, while showing strictly lower capabilities than larger open-weight models, suggesting that it is unlikely to materially enable new threats upon release." Including [Kimi K3](/articles/kimi-k3) as an uncontested upper bound on every row is doing real work here: it establishes that whatever Muse Glimmer can tell you about wet-lab protocols, an already-open model tells you more. The same inversion runs through the safety rows of the main table, and there Muse Glimmer genuinely loses. Gemma4-31B has less than half its contextual-integrity violation rate (12.1 vs 26.4) and a lower prompt-injection attack success rate (25.6 vs 28.4). Muse Glimmer's compensation is utility — 94.2 on AgentDojo against Gemma's 90.8 — which is the familiar helpfulness/safety trade stated in numbers instead of prose. Whether 26.4% is an acceptable violation rate for a model explicitly recommended for agents with "deep access to personal context" is a judgment the card leaves to you, and it does at least give you the number to judge with. The preparedness section is unusually explicit about its own reasoning: Muse Glimmer "does not fall under the definition of 'Frontier AI' in Meta's Advanced AI Scaling Framework, since it is generally less capable than Muse Spark," and its Cyber and Loss-of-Control designations are marked as **inferred** from that comparison rather than measured directly. Naming an inference as an inference is good practice. It also means two of the three risk designations rest on the distillation relationship rather than on evaluations of this model. ## What actually shipped I checked the "Released Artifacts" table against Hugging Face, because promised artifacts and present artifacts are frequently different things. All four exist, in three sibling repositories the card does not link: | Artifact | Where | Size | |---|---|---| | BF16 weights | `Muse-Glimmer-30B` | 59.55 GB, 2 shards | | 4-bit, 24 GB target | `-GGUF` / `muse-glimmer-30B-kquant-17gb.gguf` | 16.76 GB | | 4-bit, 32 GB target | `-GGUF` / `muse-glimmer-30B-kquant-dynamic.gguf` | 19.65 GB | | DFlash drafter | `-GGUF` / `dflash-kquant.gguf`, `-assistant` | 1.63 GB / 5.11 GB bf16 | | Vision projector | `-GGUF` / `mmproj-kquant.gguf` | 1.40 GB | | ExecuTorch builds | `-ExecuTorch-PTE` | metal + sm80, text and text-image | The drafter's own `config.json` confirms the card's spec exactly — 5 layers, `block_size: 16`, sliding window 2048 on all five, 32 query heads and 8 KV heads. The ExecuTorch repository ships separate `.pte` files for Apple Metal and NVIDIA sm80, in text-only and text-plus-image variants, which is how the M4/M5 numbers were produced. One small correction to the card while I'm counting: it states total parameters as "~29.6B" twice. The safetensors index says **29,776,626,688** — 29.78B. A 0.6% understatement, and I mention it only because everything else in that table matched the config to the digit. ## The take The reason this release is worth reading closely is not the benchmark line, which is good but contested by a smaller competitor. It is that Muse Glimmer is a clean worked example of designing an architecture against a deployment constraint and then publishing enough for someone outside the lab to check the constraint was met. The three decisions that matter — 16:1 GQA, three sliding layers per global one, and 4-bit quantization validated at 1.0% degradation — are not independent optimizations. They are one budget, allocated. Remove any of them and the model stops being the thing the announcement describes. That is a more useful artifact than a leaderboard position, because the budget is the part that generalizes: the next person trying to fit an agent on a laptop has a worked example with all its numbers exposed. What is missing is the same thing that is always missing on day one. There is no third-party evaluation of any of these numbers, the methodology report is a link rather than a paper, and the quantization degradation figure — 1.0% averaged across 15 benchmarks — is exactly the kind of average that can hide a specific capability falling over. The card says the compression was validated on agentic tasks; it does not show that table. --- *Sources: the [Muse Glimmer announcement](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model) (Meta AI Research, 2026-08-10) and the [Muse-Glimmer-30B model card](https://huggingface.co/meta-models/Muse-Glimmer-30B), plus `config.json`, `model.safetensors.index.json` and the file trees of the `-GGUF`, `-assistant` and `-ExecuTorch-PTE` repositories, all as of 2026-08-10. Benchmark numbers are Meta's own, with no third-party replication. The two figures are Meta's, flattened onto white. The KV-cache arithmetic, the 24 GB budget reconciliation and its counterfactuals, the per-layer verification of the attention and RoPE arrays, and the re-tally of the comparison table are mine and are computed from the published config and file sizes. All four interactives are mine.* --- # The Skaling law: Chinchilla assumes model size and data don't interact, and they do > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/skaling-law > date: 2026-08-10 > tags: scaling-laws, llm, math, training, explainer The Chinchilla scaling law is one of the most quoted equations in the field. It says the loss of a language model decomposes into an irreducible floor plus two independent power-law terms, one for model size and one for training data: $$ L(N, D) = E + \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}} $$ [**Skaling: Chinchilla's Exponents Meet Kaplan's Coupling**](https://arxiv.org/abs/2608.07222) (arXiv 2608.07222, Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz and Kartik Ahuja, FAIR at Meta, 7 August 2026) points at the plus sign in the middle and observes that it is a very strong claim nobody ever tested. A sum of a function of $N$ and a function of $D$ has a cross-derivative of exactly zero. Not approximately zero, not small — zero, as an algebraic identity. The additive form *asserts* that how much a training token is worth does not depend on how big your model is. That assertion was never a finding. It was a modelling convenience that came along for the ride. Personal disclosure up front, because it changes how I read this paper: four days ago I published a [back-of-envelope estimate](/articles/ltc-gated-delta) that leaned on exactly this additive form to argue that a small model's reported gain might vanish under a compute-matched comparison. The last section of this piece redoes that calculation with the paper's tools. It moves. ## The saddle Fit the additive law to a dense grid of trained models and the residuals are not noise. They have structure.
The paper's description of that first panel is precise: Chinchilla "is accurate in the interior of the grid but develops large, oppositely-signed errors toward the corners, reaching several percent where $N$ and $D$ are most imbalanced. This is the saddle-shaped residual expected when the $N$–$D$ interaction is omitted." A saddle is the signature of a missing product term. If your model of a surface has no $xy$ term and the true surface has one, the errors you get are positive on one diagonal and negative on the other — which is exactly what the left panel shows. The paper backs this up with a direct measurement of the cross-derivative $\partial^2 L/\partial N \partial D$ from local quadratic fits (their Figure 3), finding it non-zero and structured. ## One exponent The fix is small enough to state in a line. Kaplan's original 2020 form did couple $N$ and $D$ — $L(N,D) = [(N_c/N)^{\alpha_N/\alpha_D} + D_c/D]^{\alpha_D}$ — but it tied the inner exponents together through the ratio $\alpha_N/\alpha_D$, so the per-axis decay rates were no longer independent. Chinchilla threw out the coupling to get the independence back. Skaling keeps both: $$ L(N, D) = \left(\frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}}\right)^{k} + E $$ Chinchilla's interpretable base terms and independent inner exponents, raised to a single free outer exponent $k$. At $k = 1$ it *is* the additive law — Chinchilla is a special case, not a rival. For any $k \neq 1$ the cross-derivative is non-zero. And because $k > 0$, the loss is still strictly decreasing in both arguments, so adding capacity or data can never be predicted to hurt. The fitted value is what makes this more than a formality. On the Farseer grid, $k = 0.41 \pm 0.01$ — nowhere near 1, and tightly determined. Differentiating, $$ \frac{\partial^2 L}{\partial \ln N\, \partial \ln D} = k(k-1)\,\alpha\beta\left(\frac{A}{N^{\alpha}}\right)\left(\frac{B}{D^{\beta}}\right)R^{\,k-2}, \qquad R = \frac{A}{N^{\alpha}} + \frac{B}{D^{\beta}} $$ With $k < 1$ the factor $k(k-1)$ is negative, so the cross-derivative is negative: since $\partial L/\partial \ln D$ is already negative, making it *more* negative means **bigger models extract more from the same token**. That is not a surprising claim — it is roughly what everyone believes — but the additive law is structurally incapable of expressing it. One honest wrinkle the paper raises itself. The coupled reducible term decays more slowly at large scale, so it absorbs curvature the additive law can only represent through a larger $E$. Skaling therefore pushes $E$ down — on Farseer almost to zero (0.03 ± 0.02, against Chinchilla's 0.45 ± 0.01). Since none of the runs reach the scale where loss saturates, the data fix the total loss but not the split between a decaying term and a constant floor. So $E$ should not be read as a measured irreducible entropy in either fit. Every other parameter is determined to within a few percent. ## Interpolation quality is not evidence This is the methodological point I most want people to take away, and it is stated bluntly in the paper: "High interpolation fit quality is not enough to validate a scaling law." Chinchilla achieves $R^2 = 0.995$ on the full Farseer grid. By the standard people usually apply, that is a solved problem. Its extrapolation error is three to four times Skaling's. The failure mode "is therefore not a poor fit to the interior, but a systematic misprediction of how the loss surface bends away from the observed region" — which is the only thing anyone ever uses a scaling law for. Nobody fits a scaling law to predict a run they already did. The compute argument in the second half of that figure is the one with budget consequences. The authors pair the coupled form with an **L-shaped sparse grid**: instead of spreading held-out points across the whole $(N, D)$ plane, restrict training runs to the low-compute edges — a row of small models across many data budgets, and a column of small data budgets across many model sizes. Skaling fitted on that L-shape, at roughly a tenth of the FLOPs, extrapolates **better than Chinchilla fitted on the entire grid** in every held-out regime. On Farseer: 0.89% vs 1.48% on larger models, 1.35% vs 1.98% on more data, 1.51% vs 2.46% far outside both. It is also worth reading the row that is not Skaling. The nine-parameter Farseer law is *worse* than three-parameter Chinchilla on several columns and carries huge fold-to-fold variance (±1.93 on far extrapolation). More parameters bought instability, not accuracy. The Skaling result is a one-parameter change that improves things, which is a different and much stronger kind of claim. ## The part that changes decisions Everything above is about fit quality. This is about where the money goes. The compute-optimal token-to-parameter ratio $D^\star/N^\star$ is the quantity that answers "should the next dollar buy a bigger model or more data?" The paper recovers it two ways without assuming any parametric law — a global Gaussian-process surrogate and a local moving-least-squares surrogate, finding the point where the log-slopes balance — and then compares against what each fitted law predicts.
The two model-free estimates of the optimum give exponents of **−0.14 and −0.15**. Skaling recovers **−0.11**. The refitted additive law gives **+0.03** — the opposite sign. Inside the observed data range all four agree, which is exactly why a good interpolation $R^2$ told you nothing. Outside it they diverge, and the paper reports that one order of magnitude beyond the data the allocations differ by more than 10×, with the additive law heading toward hundreds of tokens per parameter while the empirical fits and Skaling fall to the tens. Two caveats before anyone reallocates a training budget on this. The empirical exponent is itself an extrapolation of a fit to a surrogate of a finite grid, and the two surrogates agreeing with each other is weaker evidence than two independent measurements. And "the additive law has the wrong sign" is a claim about *these* datasets, at *these* scales, with these architectures. But the sign disagreement is not subtle, it reproduces across two independently constructed grids, and it lands on the one number the whole scaling-law enterprise exists to produce. ## What was actually measured Worth being concrete about the evidence base, because scaling-law papers vary enormously here. **Farseer** is an existing public grid; the fitting set is 302 configurations totalling ~5.0×10²² FLOPs, with held-out sets for larger models (36 points, 1.5B–6.4B), more data (66 points), and far extrapolation (7 runs at 2.3B–25B parameters on 126B–453B tokens, beyond both axes). **SK-Grid** is the authors' own: 134 configurations, 15 model sizes from 134M to 4.9B, 16 data budgets from 316M to 316B tokens, with far-extrapolation runs at ~10²² FLOPs on 5.8B–10.8B parameter models. Two more datasets appear in the appendix with the same protocol and the same ranking. All laws are fitted identically — Huber loss in log space, L-BFGS-B with 2000 basin-hopping restarts, analytic gradients — so the comparison is not confounded by one law getting a better optimizer. The paper also notes that the improvement survives holding the protocol fixed, "confirming the gain comes from the functional form rather than the protocol." What is *not* here: no runs at frontier scale, so the far-extrapolation column tops out around 25B parameters; no test of whether $k$ is stable across architecture families, tokenizers or data mixtures, which is the obvious next question given that $k$ is now carrying the entire interaction; and no mixture-of-experts models, where "model size" is ambiguous enough that it is unclear which $N$ even belongs in the formula. ## Redoing my own arithmetic Four days ago, writing about [liquid time constants and gated delta rules](/articles/ltc-gated-delta), I used the additive Chinchilla form to estimate a confound. The setup: a 29M-parameter model, LTCAttention, beat its baseline by 0.062 nats on a token-matched comparison while running 12.4% slower. I asked what the baseline would have gained if it had spent that 12.4% on extra tokens instead, got roughly 0.061 nats, and concluded that a compute-matched comparison might erase the entire result. I used Hoffmann et al.'s 2022 coefficients. This paper supplies two better options: the same additive form refitted on Farseer, and the coupled form on the same runs. The three answers span 2.4×. The estimate I published was the largest of them. Most of the movement is not the coupling — it is that Hoffmann's coefficients were fitted on a different corpus and tokenizer, and refitting the *same* functional form on Farseer roughly halves the answer, from 0.061 to 0.031 nats. The coupled form then trims it further, to 0.026. There is also a pointed detail: LTCAttention's configuration sits at $D/N = 10.0$, which falls in the band the paper reports as Chinchilla's **worst** regime — 3.47% MAPE in the optimal-ratio third, where its pooled number hides the failure. So the honest revision: a compute-matched baseline would probably have recovered somewhere around **half** of LTCAttention's measured gain, not all of it. The conclusion I actually drew still stands — the missing experiment is a wall-clock-matched run, and until someone does it the result is better per token and undetermined per second — but I stated the confound about twice as strongly as the evidence supports. That correction is now in the record here rather than only in my own notes. The broader lesson is the one I would take from this paper even if I had no stake in it. Scaling-law arithmetic is routinely used the way I used it: pull the canonical coefficients, differentiate, get a number, cite it as though it were a measurement. It is not. It is a prediction from a functional form fitted to somebody else's grid, and both the form and the grid are doing real work. When the answer matters, quote the range. ## The take The contribution is one exponent, and the reason it is a good paper rather than a small one is that the exponent is load-bearing. It removes a structural bias that was invisible in interpolation error and severe at the boundaries, it makes accurate extrapolation possible from a tenth of the compute, and it flips the sign of the trend in the single number the field uses to allocate training budgets. What it does not do is settle anything at frontier scale, where nobody has published the grid that would test it. And there is a mild irony worth naming: the paper's own argument implies that its $k = 0.41$ is a property of these datasets, and the honest way to use the Skaling law is to refit it on your own runs rather than to quote 0.41 the way people have been quoting 20 tokens per parameter for four years. --- *Sources: [Skaling: Chinchilla's Exponents Meet Kaplan's Coupling](https://arxiv.org/abs/2608.07222) (arXiv 2608.07222v1, Videau, Youbi-Idrissi, Lopez-Paz, Ahuja, FAIR at Meta, 7 August 2026, CC BY 4.0), read in full via the arXiv HTML rendering. Equation 3, the fitted coefficients in Table 2, the MAPE figures in Tables 1 and 3, and the compute-optimal exponents in Figure 6 are quoted as published. Both figures are the paper's own, flattened onto white. The cross-derivative expression, the recomputation of my earlier LTCAttention estimate, and the sensitivity arithmetic behind the last interactive are mine, computed from the paper's published coefficients at LTCAttention's reported N and D. All four interactives are mine.* --- # BTL-4: reading a model card against its own weights > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/btl-4 > date: 2026-08-06 > tags: llm, open-weights, benchmarks, evaluation, lora, explainer [BTL-4](https://huggingface.co/badtheorylabs/BTL-4) went up on Hugging Face on 2026-08-05: a 35B agentic reasoning model from Bad Theory Labs, Apache-2.0, 21 safetensors shards, and a benchmark table with **78.4% on SWE-bench Verified** in it. There is no technical report. There is no arXiv paper. There is no third-party evaluation. At the time of writing the repository has 38 likes and, according to the Hugging Face API, **zero downloads — all-time**. Nobody has run this model. That combination is common enough now that "how do I read this?" is a real question rather than a rhetorical one. The answer I want to argue for is that a model card is not the only evidence a release ships. The artifact itself — `config.json`, the tensor index, the bytes in the shards — is evidence too, it is machine-checkable, and it is often more informative than the prose. This piece is that check, run end to end on BTL-4. Some of the card holds up exactly. Some of it does not. **What this is and isn't.** I did not download 70 GB and run a benchmark; I have not reproduced or refuted any score here. Everything below comes from public metadata and HTTP range requests totalling a few megabytes. Where the card is contradicted, it is contradicted by *its own numbers* or by the artifact it ships, not by a competing measurement of mine. And Bad Theory Labs is not a drive-by account — it has seven models going back to June 2026, including [BTL-3](https://huggingface.co/badtheorylabs/BTL-3) and a BTL-4-Compact posted the day after this one. Read this as an audit of one release, not a verdict on a lab. ## What checks out Start with the parts that survive, because they are the majority and because a check that only ever finds problems isn't a check. **The base-model claim is exactly right.** The card declares `base_model: Ornith-1.0-35B`, which is an unqualified string rather than a resolvable repository id, so nothing on Hugging Face verifies it for you. But [`ornith-ai/Ornith-1.0-35B`](https://huggingface.co/ornith-ai/Ornith-1.0-35B) exists — a real MIT-licensed model with 2.67M downloads — and its `config.json` matches BTL-4's on every architectural field: | | Ornith-1.0-35B | BTL-4 | |---|---|---| | architecture | `Qwen3_5MoeForConditionalGeneration` | same | | layers · hidden | 40 · 2048 | same | | experts · active | 256 · top-8 | same | | MoE / shared intermediate | 512 · 512 | same | | vocabulary | 248,320 | same | | context | 262,144 | same | | full-attention interval | every 4th layer | same | | vision tower | 27 layers · 1152 wide | same | The tensor maps are identical too: **31,666 tensors, same names, in both**. Whatever else is true, BTL-4 is a derivative of Ornith-1.0-35B and not of something else wearing its name. **The weights are real and complete.** 21 shards, 70.21 GB, 35.11B parameters in bf16, a coherent `model.safetensors.index.json`, vision tower included. This is not an empty repository with a good README. **The lineage is worth stating**, because it puts BTL-4 next to work already covered here. Ornith-1.0-35B's architecture is `qwen3_5_moe` — the same family as [Intern-S2-Mobius](/articles/intern-s2-mobius), which is 40 layers at hidden 2048 with the same 512-wide experts and the same every-fourth-layer full attention. Both descend from Qwen's `Qwen3.5-35B-A3B`, whose parameter count (35,951,822,704) is also the exact figure I measured for [Macaron-V1-Tall's](/articles/macaron-v1) base checkpoint. Three unrelated labs, one 35B Qwen substrate. That is worth noticing on its own. **And one section of the card is genuinely useful**, which I'll come back to at the end — it isn't the benchmark table. ## The benchmark table doesn't close Here is the LiveCodeBench v6 section of the card, quoted in full. Aggregate **66.1%**, and: | | pass@1 | |---|---| | easy | 99.1% | | medium | 86.7% | | hard | 60.5% | followed by: *"The set is 45% hard problems, which is what pulls the aggregate down."* Those four numbers cannot all be true, and you do not need the benchmark to see it. An aggregate pass rate is a weighted mean of the per-difficulty rates. Fix the hard share and the aggregate is pinned inside an interval — lowest when every remaining problem is medium, highest when every remaining problem is easy. At 45% hard, the aggregate has to land between **74.9%** and **81.7%**. The card reports 66.1%, roughly nine points below the floor of what its own difficulty breakdown allows. Going the other way: to produce a 66.1% aggregate from a 60.5% hard bucket and an 86.7% medium bucket, you would need **78.6% hard problems and zero easy ones** — which would leave the 99.1% easy row reporting a score for an empty set. I want to be careful about what this does and does not establish. It does not tell you the model is bad, and it does not tell you which number is wrong. Any one of four edits reconciles it: the aggregate, the hard share, one of the bucket rates, or an unstated detail about how the aggregate was computed (a different problem set, a different pass@k, a subset that the difficulty table doesn't describe). What it does establish is that **the table was never checked against itself**, which is a fact about the release process rather than about the model. Numbers that were run, recorded, and then arithmetically verified do not do this. The same section has a smaller tension worth flagging. The card says the runs used "full splits, no subsetting," and in the next breath specifies "442 problems, 2024-08 → 2025-05." A date window is LiveCodeBench's intended usage — the whole point of the benchmark is contamination-controlled time slices — so the window is legitimate. But a date-windowed 442-problem slice is, definitionally, a subset, and "no subsetting" is the wrong way to describe it. I could not independently confirm LiveCodeBench v6's true composition for that window, so I can't say which of the card's figures the real distribution would support. ## What the config says that the card doesn't `config.json` is written by the training code, not by the person writing the README, which makes it the more candid of the two documents. BTL-4's contains this: ```json { "model_name": "/vol/merged/btl4-pilot", "transformers_version": "5.13.1", "unsloth_version": "2026.7.6" } ``` Three things leak out of five lines. The checkpoint was produced with **Unsloth**, a LoRA and QLoRA fine-tuning library. It was loaded from a directory called **`merged`**, which is what you call the output of folding an adapter back into its base. And the run was named **`btl4-pilot`**. None of that is damning — LoRA is a completely normal way to fine-tune a 35B MoE, and merging is the normal way to ship one. But the card's training section says only: *"Fine-tuned from Ornith-1.0-35B on an execution-gated reasoning corpus."* A reader deciding whether a +4.3-point BFCL gain is likely to generalize would want to know it came from a merged adapter rather than a full fine-tune, and the card does not say. The interesting question is whether the weights agree with the config. They do. ## Reading 70 GB without downloading it A safetensors file opens with 8 bytes giving a header length, followed by that many bytes of JSON describing every tensor: dtype, shape, and byte offsets into the rest of the file. That means two small HTTP range requests per shard buy you the complete layout of a 70 GB checkpoint. Once you have offsets, you can range-request *one specific tensor* out of the middle of a shard and compare it against the same tensor in another repository, having transferred a few hundred kilobytes. I ran that against BTL-4 and Ornith-1.0-35B across ten groups of tensors. The first pass was wrong, and the way it was wrong is the most useful thing in this article. Every normalization weight came back CHANGED — all forty layers' input and post-attention norms, the final norm, even norms inside the vision tower. That looked like a substantial finding. It was an artifact: BTL-4 stores norms as **F32** where Ornith stores them as **BF16**, so I was comparing 8,192 bytes of one format against 4,096 bytes of another and reading the inevitable mismatch as training. The check that settles it is arithmetic on file sizes. The two checkpoints differ in total size by **603,136 bytes**. BTL-4's metadata reports exactly **301,568 parameters stored in F32**; Ornith reports none. An F32 parameter costs two bytes more than a BF16 one, and 301,568 × 2 = 603,136. The entire size difference between the two models is the norm upcast and nothing else — which is what turns "I should exclude those rows" from a hunch into a fact. ## What the change map means With dtype-mismatched tensors excluded, the pattern is unusually clean: - **Changed**, in every layer sampled: expert `gate_proj` and `down_proj`, the shared expert's `up_proj`, the linear-attention `in_proj_qkv`, and full-attention `q_proj`. - **Unchanged**, in every window sampled: the MoE routers, token embeddings, the output head, the entire 27-layer vision tower, and the linear-attention `A_log` and `dt_bias`. That is a LoRA target set, drawn from life. Adapters go on the projection matrices; routers, embeddings, output heads and frozen encoders are left alone. Combined with `unsloth_version` and `/vol/merged/`, the artifact is telling a consistent story that the prose omits. Two of those frozen tensors deserve their own note. **`A_log` and `dt_bias` are untouched at every layer**, and unlike the big matrices these are small enough to compare in full — 64 bytes each, byte-for-byte identical. In this architecture family `A_log` is the per-head base rate of the linear-attention decay gate, the parameter whose exponential sets how fast a channel forgets. [KDA has a half-life](/articles/kda-half-life) works through what that number means: it converts directly into a memory horizon measured in tokens. So BTL-4's forgetting timescales are Ornith's, unmodified. Whatever the fine-tune taught the model about tool calling, it did not touch the mechanism that decides how long the model can hold something. **The vision tower is entirely unchanged, and entirely still there.** BTL-4 ships `processor_config.json`, an `image_token_id`, a `video_token_id`, and 27 untouched vision layers. The card sets `pipeline_tag: text-generation`, describes a text-only training corpus, and reports no multimodal evaluation whatsoever. Nothing wrong with that — you inherit a capability you didn't train and don't claim. But a reader should know that roughly a tenth of what they'd be downloading is an unexercised, unevaluated image encoder, and that the model's multimodal behaviour is entirely Ornith's. ## The number with the least behind it Of the three headline benchmarks, note which one is documented and which is not. BFCL v4 gets a full protocol sentence: official `ast_checker`, all 1240 cases, run in-house, and — best practice, this — an explicitly paired comparison against the base with "identical harness, identical decoding, only the weights differ." That is exactly how a fine-tuning claim should be stated, and the +4.3 points is the only number on the card that isolates what the training actually bought. LiveCodeBench gets a protocol sentence too, though the numbers under it don't close. **SWE-bench Verified 78.4% gets three words: "official harness."** No base comparison, so there is no way to see what the fine-tune contributed. No statement of who ran it, while the two benchmarks above it are explicitly labelled in-house. No scaffold named, which for SWE-bench is most of the result — the agent loop around the model routinely moves that score more than the model does, a point [the harness effect](/articles/harness-effect) makes at length. No trajectory logs, no leaderboard submission. It is also, by a distance, the biggest claim on the page. 78.4% would place a 35B model with roughly 3B active parameters within striking distance of the frontier systems this site has covered — the Macaron-V1 table has Claude Opus 4.8 at 88.6% on the same benchmark. Extraordinary is the wrong frame; *unverifiable* is the right one. The claim isn't refuted here. It's simply the one number with the least behind it, presented with the least detail, on a model that has been downloaded zero times. Zero downloads, all-time, is not a snark — it's a structural fact about what any reader can know. Every number on this card is unreplicated *by construction*, because nobody has yet obtained the weights to try. Likes are not replication. Until someone runs it, the honest status of all three benchmarks is "reported, unverified," and the honest status of the LiveCodeBench row specifically is "reported, internally inconsistent." ## The part of the card that's actually good None of the above touches the most useful section, and it deserves to be lifted out because it is the kind of thing most cards leave you to discover in production: > **Reasoning accumulates across agent turns.** The chat template strips prior reasoning from older > turns, but this only works if your harness separates it into `reasoning_content`. With vLLM, that > means `--reasoning-parser qwen3`. Without it, thinking lands in `content`, accumulates every turn, > and long agent runs degrade. That is correct, specific, non-obvious, and expensive to learn by yourself. A reasoning model whose chat template prunes old thinking blocks depends on the serving layer routing them to the right field; get it wrong and you don't see an error, you see an agent that quietly gets worse over a long session while your context bill climbs. Whoever wrote that paragraph has actually run this thing in a loop. The same section is candid that the model is verbose, not a chat model, and token-hungry, and the generation-settings note — LiveCodeBench moving 60.9% → 66.1% purely by raising the output budget from 16K to 32K — is a real, useful observation about evaluating reasoning models even if the endpoint number is the one that doesn't reconcile. There's a smaller inconsistency in the same neighbourhood: the card claims 262K native context and the `vllm serve` command it gives sets `--max-model-len 131072`. Both are defensible individually — you often can't fit the full window on the hardware you have — but nothing explains the gap. ## The take The reusable part here isn't the verdict on BTL-4, it's the sequence. Four checks, none of which require downloading a model or running a benchmark, in increasing order of effort: 1. **Does the table close?** Weighted means have to be consistent with their parts. This caught the LiveCodeBench contradiction in one line of arithmetic, before anything was fetched. 2. **Does the declared base exist, and does the config match it?** Field-by-field comparison confirmed BTL-4's lineage exactly, which is the strongest positive result in this whole piece. 3. **What does the config leak?** Training code writes provenance the README never mentions — library versions, working-directory paths, run names. 4. **Do the weights agree with the story?** Safetensors headers plus range requests turn "trust the training section" into a measurement, for a few hundred kilobytes of traffic. Applied to BTL-4 the result is mixed rather than damning: a real fine-tune, of the model it says it is, with complete weights, shipped with one benchmark row that contradicts itself, one number that carries the most weight and the least evidence, and a training method the artifact discloses more honestly than the prose does. What I'd want before believing the headline is a base-model row for SWE-bench and the scaffold used to get it — the same paired-comparison discipline the card already applies to BFCL, extended to the number people will actually quote. --- *Sources: the [BTL-4 model card](https://huggingface.co/badtheorylabs/BTL-4) (README, `config.json`, `model.safetensors.index.json`, safetensors headers), the [Ornith-1.0-35B](https://huggingface.co/ornith-ai/Ornith-1.0-35B) repository, and the Hugging Face models API for download, like, and parameter counts, all as of 2026-08-06. The tensor comparison was performed with HTTP range requests against both repositories' shards; the method, its dtype correction, and the limits of what a windowed comparison can prove are described in the second tab of the diff figure above. The LiveCodeBench arithmetic uses only figures printed on the card. No model was downloaded and no benchmark was re-run. Both interactives are mine; the repository ships no figures, so there are none to embed.* --- # Intern-S2-Mobius: a 35B model that separates memory from reasoning > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/intern-s2-mobius > date: 2026-08-06 > tags: llm, mixture-of-experts, architecture, linear-attention, explainer [Intern-S2](/articles/intern-s2) was Shanghai AI Lab's case for specialization: a 397B model that learns straight off the raw page of a scientific paper. **Intern-S2-Mobius** is a different experiment entirely, and much smaller — 35B parameters, continual-pretrained from Qwen3.5-35B, and the point isn't science. It's architecture. The model card's claim is that you can pull a transformer apart into two pieces — a store of learned knowledge and the computation that queries it — and get a real efficiency win from doing so: reasoning traces up to 5.0× shorter, average throughput up to 4.6× higher, at matched or better scores than the plain Transformer it's compared against. That comparison is the thing to hold onto while reading this. Mobius isn't benchmarked against GPT-5.5, Gemini, or even other 35B-class open models — every number on its card is Mobius versus its own base model, Qwen3.5-35B, continual-pretrained the same way. It's an ablation, not a leaderboard entry. That makes it a cleaner test of what the architecture buys, and a much weaker basis for "should I use this instead of X." ## What it is - **35B parameters**, bf16, five safetensors shards on Hugging Face totalling about 73 GB — not gated, apache-2.0, actually downloadable. - **`image-text-to-text`**: a vision tower (27 layers, 1152-wide, patch 16 — the same depth and width as the SigLIP-So400M family of encoders used across a lot of current VLMs) feeds a text backbone with a 256K-token context window. - **Continual-pretrained from Qwen3.5-35B**, then SFT and RL, on a new architecture the card calls **Mobius-v0**, "realized by Xtuner and LMDeploy." - Deploys single-GPU on LMDeploy, vLLM, or Transformers, with an `mtp` speculative-decoding mode (`qwen3_5_mtp`) recommended in the quickstart. The README links InternLM's [ArchSpace](https://github.com/InternLM/archspace) — a public architecture-experimentation program that turns community proposals into trained, evaluated, published results — as a related project. It doesn't say Mobius came out of that pipeline, so I'm not claiming it did; it's worth knowing the program exists, because it's the same lab publicly running exactly this kind of experiment at scale. ## What "Mobius" names Not a routing scheme, not a training recipe — an architecture. The card's own framing: > Instead of binding knowledge storage and reasoning computation layer by layer as in conventional > Transformer models, Mobius organizes knowledge into a globally shared **Memory** and lets > multiple **Reasoners** iteratively query and refine hidden states against this shared repository. Two capabilities follow from that split, per the card: **Backward Residual Connection** (a deep layer can reach knowledge a shallow layer used, not just what forward propagation handed it), and **Dynamic Latent Reasoning** (deliberation gets internalized into hidden states instead of written out as visible chain-of-thought tokens). Both are described in prose. The released code lets you check what's literally true of the shipped model versus what's evocative marketing language for the same idea — and it turns out you can, because InternLM shipped the modeling file along with the weights. ## Forty layers, four memory banks Here's what `modeling_interns2_mobius.py` actually does. `config.json` sets `num_blocks: 4`. The model builds exactly four `InternS2MobiusMetaMoeBlock` objects — each one a router plus 2560 routed experts — and holds them in one list, `meta_mlp`. Every one of the 40 decoder layers keeps its own attention and layernorms, but for its routed-expert lookup it computes `block_idx = layer_idx % num_blocks` and reads from `meta_mlp[block_idx]`. Layers 0, 4, 8 … 36 all route into the *same physical weight tensors* — not four separately-trained-but-similar banks, one set of parameters, referenced by ten different layers. A standard MoE transformer ties knowledge to depth: layer *k* owns bank *k*, and whatever it learned lives only there. Mobius reuses the same four banks across the whole stack instead, so a layer near the input and a layer near the output can draw on the identical knowledge subspace. That's the concrete mechanism behind "Backward Residual Connection" — not a literal skip connection running backward through the network, but a shared address space that any depth can query. It's also a real parameter-efficiency trade: with four banks instead of forty, the routed-expert weight mass is $$ \theta_{\text{experts}} \approx N_{\text{blocks}} \times N_{\text{experts}} \times \big(2\,d_{\text{ffn}}\,d_{\text{model}} + d_{\text{model}}\,d_{\text{ffn}}\big) = 4 \times 2560 \times 3{,}145{,}728 \approx 32.2\text{B params}, $$ roughly 90% of the model's total, and roughly consistent with the ~36.5B implied by the 73 GB of bf16 weights on disk. Per token, only one bank is queried per layer and only 8 of its 2560 experts fire — call it ~28M active FFN parameters per layer (8 routed experts plus the always-on per-layer shared expert), times 40 layers. That's a back-of-envelope estimate from `config.json`, not a number the card states; unlike [Intern-S2-Preview-397B's](/articles/intern-s2) plain top-8-of-512 math, Mobius's shared-bank routing makes a clean "active parameters" headline harder to state, and InternLM doesn't attempt one. One more thing falls out of matching two arrays in the same config: `layer_types` cycles linear-attention, linear-attention, linear-attention, full-attention every four layers (`full_attention_interval: 4`), the same period as the memory-bank assignment. Bank 3 is *always* the one full-attention layer in its group of four; banks 0–2 are always linear attention — a Gated DeltaNet variant, the same family covered in [KDA's half-life](/articles/kda-half-life) for Kimi K3's linear attention. That alignment isn't asserted anywhere in the README. It's just what the two config arrays do when you line them up. Whether "iteratively query and refine" is literally true of inference is a fair question to ask of any of this. The released `InternS2MobiusTextModel.forward()` is a single straight-through pass over 40 layers — no runtime loop, no repeated pass over the same weights within one layer. What does repeat, ten times, is the pattern: attend, then query one of four shared memory banks, at increasing depth. If that reads like an unrolled recurrence rather than free-form iteration, that's a fair description — it's a coarser, more surgical form of weight sharing than a fully [looped transformer](/articles/looped-models-done-right), which ties whole layers (attention included) across depth, or [LOTUS](/articles/lotus-latent-reasoning), which loops the same weights over a fixed latent region multiple passes per token. Mobius ties only the expert banks, once each, spread across depth rather than iterated in place. ## The benchmarks Everything on the card is Mobius vs. Qwen3.5-35B — its own continual-pretraining source, not a frontier model. On general reasoning: The average hides a mixed picture. Mobius leads on MMLU Pro (89.05 vs 85.31), IMO Bench (81.25 vs 77.50), HMMT 2026 (85.51 vs 78.50), AIME 2026, GPQA Diamond, AMO, and SimpleQA. It loses on two: UGD hard (73.02 vs Qwen's 78.02) and HLE (19.11 vs 22.40) — worth stating plainly, since the card's own bullet points don't mention either. Scientific tasks show the wider gap, and it's the same shape as the S2-Preview-397B story at a different scale: Biology-Instructions carries that average almost alone: 51.40 vs 3.77, a 13.6× gap. Mol-Instructions (45.73 vs 21.70) and MolecularIQ (59.29 vs 29.13) are more modest but still roughly double. I'd read this less as "Mobius learned multi-omics" and more as evidence that whatever mix of continual pretraining and RL Shanghai AI Lab runs across the Intern-S2 family leans hard on scientific data — consistent with, though far less extreme than, [Intern-S2-Preview-397B's](/articles/intern-s2) own scientific dominance.
## Shorter traces, faster serving The headline claim is "nearly 4x speedup reported in the technical report" — a report the model card references but never links or cites; there's no arXiv listing for Mobius as of this writing. What the card does show directly is Fig. 1: request throughput at batch sizes 16 through 256, averaged across six reasoning benchmarks, with Mobius **2.9× faster at batch 16 and 4.6× faster at batch 256**.
Zoom into the five subplots behind that average and the story isn't uniform. MMLU Pro and GPQA Diamond show a wide, cleanly growing gap in Mobius's favor — that's most of what drags the average up. The three math-competition benchmarks look nothing like it. On **AIME 2026 and HMMT 2026 the lines cross, and the Transformer baseline is the faster of the two at three of the five batch sizes plotted** — including 2⁷, where AIME's gap is widest in the baseline's favour. IMO Bench does stay in Mobius's favour at every point, but by a margin closer to 1.1× than to anything in the headline. The 2.9–4.6× number describes the boxed average panel. It doesn't describe every benchmark that average is built from, and the chart says so plainly if you look past the box. Most of the throughput gain traces back to shorter output, not cheaper per-token compute — Fig. 2 gives average trace length directly, and the "Nx shorter" figures on it are exact, not chart-estimated:
The same pattern repeats: GPQA Diamond and MMLU Pro compress the most and are also where the throughput gap is widest and cleanest; the math-competition benchmarks compress the least and are where the throughput lines cross. Shorter traces plus fewer live tokens in the KV cache is a coherent story for why throughput goes up — it just doesn't go up evenly. Card gaps worth naming plainly. There's no linked technical report or arXiv paper — "reported in the technical report" points at a document I could not find. No active-parameter figure is given (fair, given the shared-bank routing makes one less simple to state than usual). Every benchmark comparison is against Qwen3.5-35B specifically, not against any external model, so there's no frontier read and no read against comparably-sized open peers either. And one of the card's own figures — the reasoning-trace case study — labels its Mobius column **"Intern-Spin-35B"** instead of Intern-S2-Mobius-35B, an internal-codename leftover that suggests the card was assembled in a hurry. ## Licence, and whether you can run it Apache-2.0, same family as [Intern-S2-Preview-397B](/articles/intern-s2). The weights are real: five bf16 safetensors shards on Hugging Face (`internlm/Intern-S2-Mobius`), about 73 GB total, not gated, mirrored on ModelScope. That's a workstation-class footprint next to the 397B model's frontier-hardware requirement — LMDeploy's quickstart serves it on a single GPU (`--tp 1`), MTP speculative decoding recommended for the throughput numbers above. ## What I make of it - **The mechanism is real and it's in the code, not just the prose.** `block_idx = layer_idx % num_blocks` is a two-line change with a genuinely different parameter-sharing shape than a standard MoE — four memory banks instead of forty, each queried by ten layers spread across depth. That's checkable, and it checks out. - **"Dynamic Latent Reasoning" oversells what the inference code shows.** There's no runtime loop — it's a single forward pass with a repeating depth-wise pattern, which is a more modest and more precise thing than "iterative refinement" suggests. - **The efficiency win is real but uneven, and it tracks trace compression.** Where output collapses — GPQA Diamond 5.0× shorter, MMLU Pro 4.6× — throughput climbs cleanly. Where it barely moves — HMMT 1.2×, IMO Bench 1.4×, AIME 1.5× — the throughput advantage narrows to nothing or inverts. That is a coherent mechanism rather than a mystery: the speedup is mostly fewer tokens, not cheaper tokens. It also means the gain should be expected to shrink on any task where the model still needs to think at length. - **This is an ablation, not a leaderboard entry.** Every comparison on the card is Mobius against its own untouched base model. That's the right comparison for isolating what the architecture buys. It's the wrong comparison for deciding whether to run Mobius instead of anything else. --- *Sources: the [Intern-S2-Mobius model card](https://huggingface.co/internlm/Intern-S2-Mobius) (README, `config.json`, `configuration_interns2_mobius.py`, `modeling_interns2_mobius.py`) and the [Intern-S2-Preview-397B model card](https://huggingface.co/internlm/Intern-S2-Preview-397B), both InternLM / Shanghai AI Lab. Benchmark numbers and figures are quoted as reported on the Mobius model card; no independent technical report or arXiv paper could be located.* --- # Maple-Preview: a 20B reasoning model where 97% of the weights are −1, 0, or +1 > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/maple-preview > date: 2026-08-06 > tags: llm, quantization, mixture-of-experts, on-device, open-weights, explainer Most quantization is something you do *to* a model after it is trained. You take bf16 weights, find a rounding scheme that hurts least, and accept the damage. **[Maple-Preview](https://huggingface.co/deepgrove/maple-preview)**, released by DeepGrove on 2026-08-04 under MIT, is the other thing: a 20B-A1.49B mixture-of-experts reasoning model where the weights were *trained* to be ternary, so nearly every parameter in the network is one of exactly three values — −1, 0, or +1, scaled. The headline numbers are a 5.31 GB checkpoint and 218 tokens/sec on a Mac mini M4. Both are the kind of claim worth checking rather than repeating, and this time both check out — with one significant asterisk about what the released repository actually contains. ## The claim, verified from the bytes You do not need to download 40 GB to test "is it ternary." A safetensors file begins with a header giving every tensor's dtype, shape and byte offsets, so two small range requests buy you the layout, and one more pulls a single row out of the middle of a shard. Count the distinct values in that row. A normal bf16 weight row of 2048 elements has on the order of two thousand distinct values. A ternary one has three. Every row measured came back with exactly three values, perfectly symmetric — `−s`, `0`, `+s` — with `s` changing from row to row. That is ternary with a **per-output-channel scale**, the BitNet b1.58 shape. About 38–43% of the weights in each row are exactly zero, which is what absmean ternarization does to a roughly Gaussian weight distribution: everything inside the rounding threshold collapses to nothing. What stays in full precision is as interesting as what doesn't. **96.9% of parameters are ternary**; the exceptions are the two embedding tables, the norms, and — pointedly — the MoE routers. A router chooses 8 experts out of 256 based on the *margin* between logits. Crushing that margin to three levels would scramble which expert fires long before it degraded any individual expert's arithmetic, so the router is the one 12M-parameter tensor per layer that stays sharp. ## The 5.31 GB claim reconciles, and the leftover is the vocabulary A ternary weight carries $\log_2 3 \approx 1.585$ bits of information, so that is the floor for any lossless packing. Working from the actual parameter census — 19.58B ternary, 0.64B full precision — the arithmetic lands where it should. 5.31 GB implies **1.65 bits per ternary weight**: just above the entropy floor, comfortably below naive 2-bit, and about where you land packing five trits into a byte ($3^5 = 243$ fits in 256) plus the per-row scales. The claim is not merely plausible, it is consistent with the measured parameter split to within a few percent. The second-order effect is the one I did not expect. Once you have crushed 97% of the model to under two bits, the **un-quantized embedding tables are about a quarter of the entire file** — 1.27 GB of 5.31 GB, for a 151,936-token vocabulary at 2048 wide, twice over (input and output are untied). At this compression ratio the interesting problem stops being the weights and starts being the vocabulary. Anyone chasing the next factor of two on-device has to go after the embeddings. **What you download is not the 5.31 GB artifact.** The Hugging Face repository ships nine shards totalling **40.43 GB** — the ternary values stored one-per-bf16, unpacked. The 5.31 GB figure describes a packed checkpoint that is not in the repository. The README is upfront that the Apple Silicon result "uses a separate on-device runtime," and I would extend that caveat: the packed format and the kernels that make 218 tok/s possible are both part of that unreleased runtime. What is public is the weights and a reference implementation. ## The shipped code does no quantization This is worth stating plainly because `config.json` looks like it says otherwise. It contains `"quantize": true` — and nothing in the released code reads it. `MapleConfig.__init__` in `configuration_maple.py` does not declare a `quantize` parameter, so the flag lands in `**kwargs` and is stored and ignored. The only occurrence of the word in 1,052 lines of Python is a comment in `fa3.py`. The forward pass confirms it. `MapleMLP.forward` is a plain dense matmul on the dequantized bf16 tensors: ```python def forward(self, x): gate_weight, up_weight, down_weight = self.gate_proj.weight, self.up_proj.weight, self.down_proj.weight return torch.nn.functional.linear( self.act_fn(torch.clamp(torch.nn.functional.linear(x, gate_weight), max=7.0)) * torch.clamp(torch.nn.functional.linear(x, up_weight), min=-7.0, max=7.0), down_weight, ) ``` There is no packing, no unpacking, no ternary kernel. Run this and you get a correct model that occupies 40 GB and runs at ordinary dense-MoE speed, with none of the benefit that motivated the architecture. The clamps are the tell that quantization-aware training happened somewhere else. `clamp(gate, max=7.0)` and `clamp(up, min=-7.0, max=7.0)` bound the activations going into the down-projection. Activation clamping is a standard QAT ingredient — you cannot quantize weights aggressively if the activations they multiply are free to blow up — and its presence in the inference path is a residue of the training recipe, kept because removing it would change the model's behaviour. ## The architecture around the quantization The config describes a design clearly built for a memory-bound device rather than a datacenter. | | | |---|---| | layers | 24 | | hidden | 2048 · head_dim 128 · 16 heads · 4 KV heads | | experts | 256, top-8, **no shared expert**, `moe_intermediate_size` 512 | | attention | 3:1 sliding-window (512) to global | | position | `partial_rotary_factor` 0.5, `nope_on_global_attention: true` | | context | 131,072 | | vocabulary | 151,936 (Qwen tokenizer) | The `layer_types` array spells the attention pattern out exactly: `s s s G` repeated six times, with global attention at layers 3, 7, 11, 15, 19 and 23. Only a quarter of the layers hold a full-length KV cache; the rest are capped at a 512-token window. For a 131K context on a Mac mini that is not a refinement, it is the difference between fitting and not fitting. Two details are worth pulling out. **`nope_on_global_attention: true`** means the global layers get no positional encoding at all — the sliding layers carry position through RoPE (at half the head dimension, per `partial_rotary_factor: 0.5`) and the global layers are left to infer order from what the local ones already encoded. The same trick appears in [Kimi K3](/articles/kimi-k3)'s attention stack, and the argument for it is that removing RoPE from the layers that see the whole sequence is what lets length extrapolation work. And **there is no shared expert** — `num_shared_experts: 0`. Most recent MoE designs keep one or two always-on experts to absorb generic computation. Maple routes everything, which is consistent with the rest of the design: a shared expert is a dense tensor every token pays for, and this model is built to minimize exactly that. ## What the benchmarks say, and what the chart leaves out
Maple-Preview averages **78.7** across LiveCodeBench v6, AIME 2026, HMMT 2026 and GPQA-Diamond, at 1.49B active parameters. That beats GPT-OSS 20B (76.3), Qwen3 30B-A3B (76.6), Qwen3.5 9B (76.3), GLM 4.7 Flash (77.4) and the other ternary entry, Ternary Bonsai 27B (77.1). It does not beat **Qwen3.5 35B-A3B at 82.9**, and the gap is not evenly distributed. On LiveCodeBench Maple actually leads (75.1 vs 74.6); on AIME and HMMT it trails by a few points; on GPQA-Diamond it trails by **10.7 points** (73.5 vs 84.2) and is beaten even by Qwen3.5 9B (81.7). That shape — competitive on code and competition math, weak on GPQA — is the signature of a model with strong reasoning and thinner world knowledge, which is exactly what you would predict from a 1.49B active budget where the knowledge has to survive ternarization.
The frontier chart is the release's strongest visual and its most selective one. Maple sits alone in the top right, roughly 3.5× the throughput of the nearest model at comparable quality. But notice who is *not* plotted: **Qwen3.5 35B-A3B and GLM 4.7 Flash, the two models that beat or match Maple on the score table, do not appear on the speed chart at all.** In fairness that is close to the point — a 35B model in bf16 does not fit on a Mac mini, which is the whole argument for building this way — but "a new point on the Pareto frontier" is being claimed against a field that excludes the strongest competitor rather than measuring it. The honest version of the claim is narrower and still interesting: *among models that fit and run fast on consumer hardware*, nothing else is close. ## Credit where the card gives it DeepGrove's own limitations section is short and unusually candid for a launch: > This preview received minimal post-training for agentic tasks and only small-scale general > reinforcement learning. and, in the evaluation section, "this preview is focused primarily on raw reasoning and, as such, may underperform on agentic benchmarks." That is a lab telling you which axis it did not optimize before anyone can discover it. It also explains the naming — this is `maple-preview`, not `maple`, and the card says extended training is coming. ## The take The interesting claim here is not the benchmark row, it is that **quantization-aware training at 1.58 bits now produces a model that competes with bf16 models several times its active size**. That claim survives inspection: the weights really are ternary, the compression really does land near the entropy bound, and the resulting artifact really is small enough to matter on a laptop. Ternary training has been a research thread for a couple of years, mostly at scales small enough to dismiss. A 20B model scoring 78.7 average is harder to wave away. What is missing is the half that makes the numbers real. The packed checkpoint, the ternary kernels and the Apple Silicon runtime are all unreleased, and the reference implementation in the repository reproduces the model's *outputs* but none of its *economics*. Right now you can verify that DeepGrove trained what they said they trained. You cannot yet run it the way they ran it. --- *Sources: the [Maple-Preview model card](https://huggingface.co/deepgrove/maple-preview) — README, `config.json`, `configuration_maple.py`, `modeling_maple.py`, `model.safetensors.index.json` and the safetensors headers — as of 2026-08-06. Both figures are DeepGrove's own, downloaded and flattened onto white. The per-row ternary measurements and the parameter census were taken with HTTP range requests against the published shards; no checkpoint was downloaded in full and no benchmark was re-run. Benchmark numbers are DeepGrove's as printed on their table, with no third-party replication. Both interactives are mine.* --- # Pokee-Isaac 28B: 10M tokens on one GPU, and an architecture the report never explains > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/pokee-isaac-28b > date: 2026-08-06 > tags: llm, long-context, agentic, on-device, explainer Pokee AI's [technical report](https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf) makes two claims about **Pokee-Isaac 28B**. The specific one: a **28-billion-parameter, non-decoder-only** model that holds retrieval fidelity across **10 million tokens**, running on a single NVIDIA B200 at up to **137,200 tokens/s prefill** and **335 tokens/s decode**, scoring **93.3% on RULER** at that length, and priced at **$0.15 / $1.00** per million input/output tokens. The general one, stated right in the abstract: long-context agentic capability has been cloud-only because of infrastructure cost, which locks it out of regulated industries, the public sector, and anywhere data can't leave the building — and a model this small changes that. The general claim is worth taking seriously. The specific one is where a close read gets uncomfortable: the report names its architecture "non-decoder-only" twice, in the abstract and the introduction, and then never says what that means anywhere in nineteen pages. Take both in turn. **What this is, precisely.** A technical report posted to Pokee AI's own product console (`console.pokee.ai`) and stamped `arXiv:submit/7908231 [cs.AI]` — a submission-tracking number, not a public arXiv identifier. At the time of writing it has not been assigned one, has not been peer reviewed, and ships no weights, no config file, and no independent replication. Every number below is Pokee's own, on Pokee's own infrastructure, unless marked otherwise. That doesn't make the numbers wrong. It means nobody outside Pokee has checked them yet. ## "Non-decoder-only": the claim the report doesn't explain Here is the entirety of what the report says about its own architecture. From the abstract: "Pokee-Isaac 28B is a **non-decoder-only** foundation model that reasons, plans, and uses tools." From the introduction: "we introduce Pokee-Isaac 28B, a 10M-token context window, **non-decoder-only** model engineered to operate within a compact compute budget." That's it. Those two sentences are the complete architectural disclosure. There is no architecture section, no diagram, no equation, no named mechanism, no ablation isolating what the non-decoder component contributes. The word "decoder" does not appear again anywhere else in the document. That absence is conspicuous by comparison. [Kimi K3's technical report](/articles/kimi-k3) spends its first several sections laying out Kimi Delta Attention's exact recurrence, ships the kernel that implements it, and publishes a `config.json` that pins down which of its 93 layers are which. [KDA's decay mechanism](/articles/kda-half-life) is something a reader can derive and check independently precisely because Moonshot wrote the state-update equation down. Pokee's report has the same page budget and spends none of it there — it moves directly from the abstract to evaluation tables. One passage nearby is the closest thing to a hint, and it complicates the claim rather than supporting it: "Some weights in Pokee-Isaac are fine-tuned from **Qwen3.6-27B** (Apache 2.0)." Qwen3.6-27B is, per its own citation in the same report, a conventional **27B dense** model — ordinary decoder-only transformer. Isaac is a 28B dense model. The arithmetic sits right there: a ~1B-parameter gap between a decoder-only base and a model described as non-decoder-only. That is not evidence of anything specific — it could be a new module bolted onto an inherited decoder backbone, a retrieval or state component, a modified embedding or head, or something else — and the report gives no way to distinguish between those. I'm flagging the arithmetic because it's the one concrete data point available, not because it resolves the question. It doesn't. So: what actually gives Isaac its 10M-token window and makes it "non-decoder-only" is not something this report lets a reader verify. That is a real gap, not a stylistic one — it's the single most technically interesting claim in the paper, and it's asserted rather than shown. Treat everything below as evaluation of *what Isaac does*, because that's what the report actually lets you check; *how* it does it stays a closed question. ## What the retrieval numbers say Whatever the mechanism, the report does back the context claim with two established long-context benchmarks, run against five named baselines: **GPT-5.6 Luna**, **Gemini 3.5 Flash Lite**, and **Claude Haiku 4.5** (the cost-optimized tier of the three big cloud providers), plus **Nemotron 3 Super 120B** and **Qwen 3.5 122B** (open-weight models an organization can self-host). Frontier flagships — GPT-5.6 Sol, Claude Opus 5, Gemini 3.1 Pro — are explicitly excluded as costing roughly an order of magnitude more per token and addressing a different deployment envelope. That's a defensible exclusion, but worth naming: Isaac isn't compared against the actual frontier, only against the cheap tier of it.
The RULER protocol is NVIDIA's own official pipeline: 10 samples per task configuration, 13 configurations at 256K and 512K, and — since common-words extraction needs a small fixed vocabulary that stops being meaningful past 1M tokens — 12 configurations from 1M onward. Isaac's own scores across the sweep are **96.9 / 96.7 / 95.0 / 95.8 / 96.7 / 93.3** at 256K/512K/1M/2M/4M/10M. Read that sequence closely and it isn't a smooth decay curve — it dips at 1M, recovers at 4M, then drops again at 10M. At 10 samples per configuration that's within the noise you'd expect, not a story about Isaac getting worse and then better with more context, but it does mean 93.3% at 10M is a fairly small-sample number, not a tight measurement. Every other baseline falls off a cliff. GPT-5.6 Luna and Gemini 3.5 Flash Lite track Isaac closely through 512K, then hit **context-overflow errors** at 1M and score zero from there — they simply can't be run at that length, which the report scores as a failure rather than excusing. Claude Haiku 4.5 and Qwen 3.5 122B score zero across the entire sweep, consistent with their native windows (200K and 262K) being smaller than even the first column tested. Nemotron 3 Super 120B is the one row worth reading carefully: the 256K–512K–1M figures in the table are **NVIDIA's own self-reported numbers, not Pokee's measurement** — marked with a superscript `s` and dashed in the chart — while the 2M-onward zeros are Pokee's direct measurement. Mixing a vendor's self-reported numbers into one row of your own comparison table, clearly labeled, is honest; it's still worth noticing when you're reading the row, since it isn't measured the same way as the rest of the table. On **MRCR v2** — a harder multi-needle variant that distributes several targets through a long synthetic conversation rather than one — Isaac leads throughout: **0.607 / 0.743 / 0.500** at 256K/512K/1M (again non-monotonic — it peaks at 512K, not 256K). Gemini 3.5 Flash Lite is the closest competitor and the gap widens with length, from a 0.133 margin at 256K to 0.295 at 1M. GPT-5.6 Luna collapses to 0.050 at 1M despite scoring 95.0% on RULER at 256K — a reminder that single-needle retrieval and multi-needle disambiguation measure genuinely different failure modes, and a model can be strong at one and weak at the other. ### The comparison this site already has an anchor for The most natural comparison for a 10M-token claim is [Kimi K3](/articles/kimi-k3), the largest context window documented on this site until now: **1M tokens** on a **2.78-trillion-parameter** open model. Isaac's framing — implicit in the numbers, not stated by Pokee this directly — is 10× the context at roughly 1% of the parameters. That ratio is real arithmetic. It is not, however, a like-for-like measurement: K3's 1M is what Moonshot trained it up to; Isaac's 10M is a RULER score Pokee measured at that length. Those are different kinds of number — one a training-curriculum endpoint, the other a benchmark result — and neither the report nor this piece can turn them into a single fair ratio. What the comparison can support is narrower and still notable: a 28B dense model holding measured retrieval accuracy at a context length ten times past where a 2.8T model's *training* stopped. Whether Isaac would still say 93% if someone ran RULER on it at 20M or 50M tokens is not something either report answers. ## Why 137,200 tokens/s and 335 tokens/s are both true The efficiency section (Table 8, on a single B200-class GPU under the RULER workload) reports **time-to-first-token** directly and derives **prefill throughput** from it — context length divided by TTFT. Decode throughput is reported separately and holds close to flat regardless of context length: The two numbers describe [different bottlenecks](/articles/how-llm-inference-works): prefill processes the whole prompt as one large matrix-matrix multiply and is compute-bound, so throughput scales with how much parallel work is available — which is why it actually *rises* with context length, from ~42K tokens/s at 1M to 137K at 10M. Decode generates one token at a time against an already-populated cache, a matrix-vector operation gated by memory bandwidth rather than arithmetic, so it doesn't get faster no matter how much context is resident — 335 tokens/s at 1M, 337 at 10M, 322 under four-way concurrency. The report draws out the one number worth remembering: a 10× jump in context costs about 3× the time-to-first-token (23.6s → 72.9s), not 10× — which is the behavior that makes a 10M window usable rather than merely addressable. A full 10M-token prefill landing its first output token at 72.9 seconds is a real number to plan around, not an abstraction. ## The agentic benchmarks: where "matches or exceeds" holds and where it doesn't The abstract's claim is specific: Isaac "matches or exceeds the strongest cost-optimized cloud systems on **function calling, multi-turn interactive execution, tool orchestration, and terminal work**." Four categories, four benchmarks. Worth checking each against the report's own tables, because they don't all say the same thing. **Function calling — BFCL v4.** The Berkeley Function-Calling Leaderboard, programmatically scored throughout (no LLM judge), combining five components under fixed weights (0.40 agentic + 0.30 multi-turn + 0.10 live + 0.10 non-live + 0.10 hallucination) over 5,106 scored entries. Isaac leads the panel at **70.94**, just ahead of GPT-5.6 Luna's **70.61**. The report itself calls this "parity rather than a decisive lead," which is the right read of a 0.33-point gap — and it's a fair characterization to give credit for. This category holds up. **Multi-turn interactive execution — τ³-bench.** Sierra's benchmark runs an agent against an LLM-simulated user across four domains, verified by a five-criteria rubric rather than an LLM judge's opinion of fluency. Isaac leads the four-domain average at **0.662**, but that average hides two domains where it doesn't win: | Domain | Isaac | Luna | Gemini | Haiku | Nemotron | Qwen | |---|---|---|---|---|---|---| | Retail | **0.789** | 0.623 | 0.719 | 0.667 | 0.614 | 0.693 | | Airline | **0.760** | 0.720 | 0.700 | 0.500 | 0.688 | 0.660 | | Telecom | 0.912 | 0.579 | 0.904 | 0.404 | 0.368 | **0.947** | | Banking | 0.186 | 0.186 | **0.203** | 0.062 | 0.033 | 0.144 | | Average | **0.662** | 0.527 | 0.631 | 0.408 | 0.426 | 0.611 | Qwen 3.5 122B edges Isaac on telecom (0.947 vs. 0.912), and Gemini edges it on banking (0.203 vs. 0.186) — the domain the report itself calls "by a wide margin the hardest," where the policy an agent needs lives across 698 documents rather than the prompt. Isaac's own banking score, 18.6%, sits below the 25.5% pass@1 the report cites as the strongest *previously reported* result on that domain. This category holds up on average, not on every domain. **Tool orchestration — MCP-Atlas.** This is the one built specifically to avoid mock tool surfaces: 500 tasks against a live 36-server sandbox of real production MCP servers (GitHub, Slack, Google Workspace, Notion, and more), scored by mean claim coverage under a shared judge. It's also the one category where the abstract's claim doesn't survive contact with the table: Isaac places **third of six**, behind both GPT-5.6 Luna and Gemini 3.5 Flash Lite — two of the three cost-optimized cloud systems named in the abstract's own comparison panel. The report's honest mitigating point is efficiency, not score: Isaac reaches within 2.1 points of Gemini using 9.10 tool-call turns against Gemini's 14.99, about 60% of the trajectory length for comparable coverage. That's a genuine and worth-stating efficiency result. It is a different claim from "matches or exceeds," and on the benchmark built to be hardest to game, the report's own number doesn't back the headline phrase. **Terminal work — Terminal-Bench 2.1.** An agent at a bare command line with no enumerated action space, every model driven by the same harness (Harbor 0.20.0 with Terminus-2), evaluated on the 86 text-compatible tasks of the 89-task suite: Second of six, four tasks behind GPT-5.6 Luna, well clear of everyone else including two open-weight models an order of magnitude larger. The report is direct about this one: "Terminal-Bench is the one benchmark in this report where a cloud baseline finishes ahead of Isaac, and we report it as measured" — which is honest as far as it goes, but reads oddly next to MCP-Atlas two sections earlier, where Isaac trails not one but two cloud baselines. Calling Terminal-Bench "the one" undersells what MCP-Atlas already showed. So: of the four categories in the headline claim, function calling holds up as genuine parity, multi-turn execution holds up on average but not on every domain, and tool orchestration and terminal work both show Isaac behind at least one of the named cost-optimized cloud systems — behind two of them on the benchmark specifically designed to be hardest to inflate. "Matches or exceeds" is a fair summary of roughly half the evidence and an optimistic gloss on the rest. ## Security: safest on attacks, third on capability Isaac is evaluated on **DTAP**, a red-teaming benchmark measuring attack success rate (ASR, lower is safer) and benign task success rate (BSR, higher is better) across 12 Linux-Docker domains and 6,195 judged tasks. Isaac is the safest of the six models on both direct and indirect attack rates and their combination — 35.6% combined ASR against a range of 37.9% to 66.3% for the rest — and shows the tightest balance between direct and indirect attacks (0.8 points), where models with less refusal training swing 13 to 38 points toward direct attacks specifically. On capability (BSR), though, Isaac places third: 82.5%, against 85.1% for GPT-5.6 Luna and 83.3% for Gemini 3.5 Flash Lite — thin margins, but not a win. And there's a footnote worth reading rather than skipping: "the five baselines were run under the benchmark's stock runner; Isaac was run under the Pokee harness, which is the one condition that still differs across rows." A safety benchmark is exactly the place where the evaluation harness itself matters, and this one isn't held constant. ## The deployment argument, taken on its own terms Strip away the specific benchmark rows and there's a real argument underneath this report, and it's the one I'd give the most weight to. Long-context agentic capability today is delivered almost entirely from the cloud, because the infrastructure to serve it any other way has been expensive. That forecloses the option entirely for organizations that can't send data across a boundary at all — regulated industries, public-sector deployments, on-device applications — not because of price, but because the data isn't allowed to leave. A model that holds long-context agentic capability at 28B dense parameters changes what's *possible* to deploy inside that boundary, independent of whether it's the best model available outside it. The pricing table backs a narrower, more concrete version of this point better than the headline "$0.15/$1.00 beats everyone" framing does. Of the five baselines, only two — GPT-5.6 Luna and Gemini 3.5 Flash Lite — can actually be *bought* at the context lengths this report tests. Claude Haiku 4.5 caps at 200K, Qwen 3.5 122B at 262K, and Nemotron 3 Super 120B's public endpoints all cap at 262K despite a 1M native window — so three of five baselines simply aren't commercially available at long context, at any price. Against the two that are, Isaac is cheaper on both meters ($0.25 and $0.80 below Luna on input/output; $0.15 and $1.50 below Gemini) while covering an order of magnitude more context. That's a real and checkable comparison, distinct from the sovereignty argument, and it holds up on its own — though it's list pricing marked **provisional and subject to confirmation at launch**, so treat the exact numbers as directional rather than final. The product page at `console.pokee.ai` fills in what the report doesn't need to say: an OpenAI-compatible endpoint at `api.pokee.ai/v1/chat/completions`, streaming over SSE with a background mode that survives a disconnect, and three concrete deployment tiers — a single B200-class GPU for datacenter serving, a single consumer **RTX 4090 or 5090** for a private workstation, and Qualcomm or Intel Panther Lake NPU silicon for on-device edge inference. None of that page explains the architecture either — it repeats "purpose-built agentic architecture" without elaborating, which is consistent with the report rather than a missed opportunity to say more. Whether this specific model earns that framing is exactly what the benchmark section above complicates. But the underlying argument — that long-context agentic capability has had no in-boundary path at all, not merely an expensive one — is real, underserved by the current market, and worth taking seriously as a category even while staying skeptical of any one vendor's report about their own entry into it. ## Portability, and what's still thin Beyond the B200 numbers, Pokee reports adaptation to client and edge silicon. On an **Intel Arc Pro B70**, their own serving stack reaches 1,087–1,500 tokens/s prefill against 305 tokens/s for stock `llama.cpp` on the same hardware — a 3.6–5× gain — and 58.8 tokens/s decode against 25.7, a 2.3× gain (with a note that a 90 tokens/s decode path is still "in development," meaning the shipped number is below their own internal target). On a 12-core Xe3 **Panther Lake** SoC, fully on-device with no discrete GPU: 150.7 tokens/s prefill, 22.84 decode. On a **Snapdragon X2 Elite**: 124.95 prefill, 23.54 decode. AMD support is listed as in progress. One more benchmark result is worth flagging for provenance rather than dismissing: on something called the "Pinchbench 116-task SuperClaw suite," Isaac scores 0.9567 against 0.929 for a cloud-hosted 744B model and 0.866 for an 80B local model. Unlike RULER, MRCR, BFCL, τ³-bench, MCP-Atlas, and Terminal-Bench — every one of which is independently authored and citable — this suite has no citation, no public description, and no other appearance I could find outside this report. That doesn't make the number false. It means it can't be checked the way the rest of this report's benchmarks can, and it shouldn't carry the same weight in your own read of the model. **The report states three limitations itself, plainly, and they're worth repeating rather than summarizing.** (1) **Text-only** — no image, audio, or video input, which is also why the evaluation runs 86 of Terminal-Bench's 89 tasks and only the text track of τ³-bench. (2) **Coding was not a training priority** for this release, and the report contains no code-authoring benchmark; Terminal-Bench measures shell-agent execution, not code generation, and the report explicitly says Isaac's placement there is "neither evidence of coding strength nor evidence of its absence." (3) **Hardware adaptation is partial** — AMD support is still in progress, and most silicon families aren't covered yet. ## What I'd want before trusting this further Take the report's own framing at face value on one thing: everything here is Pokee's measurement, on Pokee's infrastructure, reported by Pokee, with no third-party replication and no released weights or config to check independently — a gap the report is upfront about, but a gap all the same. Specific to this piece, five things I couldn't verify or that need a second source: - **The architecture.** Nothing in the report or the product page says what makes Isaac "non-decoder-only" or how the 10M window is achieved. This is the biggest open question in the whole document, and it stays open here too — I'd rather say that plainly than guess at a mechanism the source doesn't support. - **The MCP-Atlas and Terminal-Bench placements**, where Isaac trails one or two of the exact cost-optimized cloud systems the abstract claims it matches or exceeds. - **The DTAP harness difference** — Isaac run under Pokee's own harness against five baselines run under the benchmark's stock runner, on the one evaluation most directly about safety. - **The "Pinchbench" result**, which has no independent citation or description to check it against. - **Pricing and general availability.** Rates are explicitly provisional, and neither the report nor the product page states a launch date or confirms the model is generally available yet. ## The take The mechanism claim doesn't survive scrutiny, because there's nothing offered to scrutinize — "non-decoder-only" is asserted twice and explained zero times, in a report that had the exact same page budget Moonshot used to write down KDA's recurrence in full. The benchmark claim survives partially: real parity on function calling, a real average lead on multi-turn tasks that hides two domain losses, and two categories — tool orchestration and terminal work, including the one benchmark built specifically to resist gaming — where Isaac trails cost-optimized cloud systems by the report's own numbers, not a critic's. What does survive, and what I'd actually flag as the interesting part of this report, is the deployment argument underneath all of it. Long-context agentic capability being cloud-only is a real constraint today, and it genuinely does foreclose entire categories of deployment — not because of price, but because data can't leave a boundary at all. A 28B dense model that holds measured retrieval accuracy for even a fraction of a 10M-token claim, running on hardware small enough to sit in a private workstation, is a meaningfully different option than what existed before — regardless of whether this particular model, from this particular report, is the one that delivers it best. That argument deserves to be taken on its merits. This report, on its own, isn't yet the evidence that settles it. --- *Sources: the [Pokee-Isaac 28B technical report](https://console.pokee.ai/pokee-isaac-28b-v0-technical-report.pdf) (all benchmark tables, the efficiency and pricing profile, and the two architecture sentences quoted in full above), and [console.pokee.ai/model](https://console.pokee.ai/model) (API details, deployment tiers, pricing display). Figure 1 here is the report's own Figure 1, cropped from the source PDF and flattened onto white for legibility in both themes — no relabeling. RULER, MRCR v2, BFCL v4, τ³-bench, MCP-Atlas, Terminal-Bench 2.1, and DTAP are each independently authored benchmarks cited in the report; none of the scores above are this site's own measurement. The context-length and prefill/decode interactives are mine, built entirely from numbers in the report's own tables — no extrapolated or simulated figures appear in either.* --- # Prime Agent: the interface is a Python REPL, not a tool-call schema > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/prime-agent > date: 2026-08-06 > tags: agents, coding-agent, harness, open-source, prime-intellect, explainer Most agent harnesses give the model a menu. You define `grep(pattern, path)`, `read_file(path)`, `run_tests()`, each with a JSON schema, and the model picks one, the harness parses the call, runs it, and hands the result back as another message in the transcript. Composition — loop over these results, retry that one, spawn three of these in parallel — happens in the *conversation*, one role-tagged message at a time, because the schema has no concept of control flow. [Prime Agent](https://github.com/PrimeIntellect-ai/prime-agent), Prime Intellect's open-source coding and research agent, makes a different bet. It gives the model one tool: a persistent Python interpreter. Composition is not a transcript pattern the harness has to support — it is just Python. This piece is a tour of that decision, built from reading the actual TypeScript host and Python runtime in the repo (not just the README), plus the second idea Prime Agent ships alongside it: a harness that edits its own supplemental state through `/refine`, with one part of itself — the base system prompt — locked out of the edit path in code, not just in the prompt. **On the clone.** A first pass at this piece was written against a *shallow* clone, which turns out to be a much worse instrument than I gave it credit for — it shows one commit and tells you nothing about a project's real age. Everything below is built from a full clone instead: 4,473 commits of history, the complete TypeScript host, and the Python runtime. The maturity section near the end is where that difference bites hardest. **Read this before the rest.** Prime Agent is open source (MIT) but it is a vendor's product for its own stack — there is no third-party evaluation of it anywhere. I read the whole README, the docs under `packages/coding-agent/docs/`, and the source, and there is not one accuracy, pass-rate, or SWE-bench-style number in any of it. Not a "we're competitive with X," not a chart, nothing. So treat everything below as an architecture and a set of design decisions — genuinely well-documented ones — not a measured result. I say more about what maturity signals *do* exist near the end. This continues two threads already on this site. [Agent harnesses](/articles/agent-harness) argued the loop wrapped around a model matters as much as the model itself, and [the harness effect](/articles/harness-effect) showed orchestration — not the model — sets an agent's token bill. Prime Agent is a concrete instance of both claims taken further: the orchestration layer here is not a fixed loop around a fixed tool menu, it is a programming environment, and the harness state that shapes behavior is itself something the agent is allowed to edit. ## One tool, not a menu Here is the entire tool surface Prime Agent gives the model, from `packages/coding-agent/src/core/tools/ipython.ts`: ```typescript const ipythonSchema = Type.Object({ code: Type.String({ description: "Python scratchpad code or `%%bash` shell cells to execute in the agent kernel. Use the target project's own environment for project imports, tests, scripts, CLIs, and dependency checks instead of direct kernel imports.", }), }) ``` That is one parameter: `code`, a string. Compare that to a typical schema-based agent, which carries a dozen or more tool definitions — `read_file`, `write_file`, `bash`, `grep`, `glob`, maybe a bespoke one per integration — each with its own JSON schema, each repeated in every request the provider sees. Prime Agent's own `CHANGELOG.md` records the direction of travel: an early entry reads "Removed the interactive `!` / `!!` bash shortcuts; use IPython for shell commands." Shell access isn't a separate tool bolted on next to Python. It is a magic cell (`%%bash`) inside the same interpreter. The kernel is genuinely persistent. Variables, imports, and open file handles survive across tool calls — and across context compaction — because they live in the interpreter's process, not in the token transcript the model rereads every turn: ```python from pathlib import Path config_files = list(Path(".").rglob("*.toml")) large_files = [path for path in config_files if path.stat().st_size > 10_000] ``` `config_files` is still there three turns later. Nothing re-lists the directory, and nothing re-sends the file list back through the model's context to remind it what it found — the model holds a *reference* to the data, not a copy of it in its own working memory. That is the "prompt-as-a-variable" half of the [Recursive Language Model](https://www.primeintellect.ai/blog/rlm) idea Prime Agent is built on: context becomes something you slice with Python, not a transcript you re-read. ## Context as variables, made concrete Take a task like "grep across 60 files for a pattern and summarize the hits." A schema-based agent issues one `grep` call, gets a list of matches back, and then — if it wants to actually look at what it found rather than trust the grep output blind — issues one more round trip per file it wants to inspect, and a final call to write the summary. The loop lives in the conversation: every iteration is a full model turn, with the tool schemas and message envelope repeated each time. An RLM turn writes the loop instead of living inside one: ```python import subprocess hits = subprocess.run( ["grep", "-rl", "TODO(perf)", "src/"], capture_output=True, text=True ).stdout.splitlines() summaries = [] for path in hits: text = Path(path).read_text() summaries.append(f"{path}: {text.count('TODO(perf)')} occurrences") print("\n".join(summaries[:10])) print(f"... {len(summaries)} files total") ``` The `for` loop, the file reads, and the counting all happen inside one `ipython` call. The model sees one printed summary, not sixty round trips of tool call and tool result. `summaries` stays a Python list the model can filter, sort, or hand to another cell — it does not have to be re-stated in the transcript to stay usable. Drag the slider above. The gap is not a fixed multiplier — it is linear-versus-flat, so it gets more dramatic exactly where it matters most: large fan-out tasks. The numbers there are a cost model built to make the *shape* of the tradeoff visible, not a benchmark; Prime Agent doesn't publish one, so neither do I. This is also where the honesty has to cut both ways. Working in a persistent kernel does not make the model's own attention free — if `summaries` is genuinely large, someone (or something) still has to look at it, and dumping ten thousand lines of `print()` output into the transcript defeats the entire point. The advantage is that the *decision* about how much of the data to surface is a line of Python (`summaries[:10]`) instead of a constraint baked into the tool schema. It is a better failure mode, not an absent one. ## The schema didn't disappear — it moved Here is the thing I got least precise about the first time, and it is the most interesting mechanism in the codebase. "One tool" is true of what the *provider* sees. It is not true of what the model can reach. When Python in the kernel calls `rlm(...)`, or `goal.get()`, or `agent_message.send(...)`, it is not doing the work locally. It opens a Jupyter comm target named `host.request` and sends a typed request back across the ZeroMQ boundary to the TypeScript `AgentSession`, which does the work and replies. `rlm-runtime.md` is blunt about the division: the Python `rlm` package "is a model-facing shim; the TypeScript host owns child execution, persistence, usage accounting, and lifecycle," and "the Python side does not call providers or implement an agent loop." The dispatch table is built in `_createKernelHostHandlers()` in `agent-session.ts`. I counted twenty-one entries: Two separate things fall out of that, and they are worth not conflating. The cheap one is token economics. A conventional harness pays for its tool surface in every single request — twenty-one JSON schemas re-serialized into the prompt on every turn, forever. Prime Agent pays for its surface once, in the kernel bootstrap and the skills' `SKILL.md` files, and the per-turn cost of the whole bridge is zero. That is the same argument as the turn-count one above, pointed at a different axis. The load-bearing one is that most of these handlers are registered **conditionally**. Goals are wired up only `if (this._includeGoals)`. Compaction only `if (this._includeCompactSkill)`. Refinement only `if (this._autoRefineAllowedForSession())`. Messaging requires both a controller *and* that the `agent-message` skill be in the model-visible set. In a session where those flags are off, the handler is not in the map at all — the model can write the exact right Python and get an error from the host, because there is nothing on the other end of the comm to answer it. That reframes the security story in a way the README's "not a security sandbox" warning does not, and it is a genuine tension inside the design rather than a resolution of it. Arbitrary Python against the filesystem really is unbounded: the kernel runs with the user's permissions and can do whatever Python can do. But the *agent-control* surface — spawn a child, open a goal, edit the harness, message a sibling, read another session's transcript — is not open. It is a typed, argument-validated, conditionally-registered table in the host, which is exactly the property a tool schema is supposed to give you. Prime Agent kept the schema and moved it somewhere the model cannot see or enumerate, then gave the model a general-purpose language for calling into it. ## Subagents are function calls The same move applies to delegation. `rlm` is preloaded in the kernel as a callable: ```python handle = await rlm("Review the authentication flow for security issues", name="auth-reviewer") print(handle.rlm_child_id, handle.name, handle.session_dir, handle.model) ``` `rlm(...)` is admission, not completion — it returns as soon as the TypeScript host has created a real child `AgentSession` with its own context and session directory, and it never blocks waiting for the child's answer. That is a real API design decision, not an implementation detail: the `CHANGELOG.md` for `0.6.0` records changing `rlm(...)` from waiting for the child to finish to returning a spawn handle at admission, specifically because treating `asyncio.gather()` over several `rlm()` calls as fan-in was the wrong mental model — spawning three reviewers is three independent calls, not a scatter-gather: ```python api_review = await rlm("Review the public API", name="api-reviewer") test_review = await rlm("Review the test coverage", name="test-reviewer") integration_audit = await rlm("Run the slow integration audit", name="integration-audit") ``` A child reports back only through an explicit message, never through the `rlm()` return value: ```python await agent_message.send(message, receiver_role="parent") ``` the same daemon-routed messaging `prime-agent send "..."` uses from the shell. Reach is deliberately narrow — an agent may message or observe only its parent, siblings, and direct children (the `0.6.0` changelog calls this "the nuclear family"); reaching a grandchild means relaying through the intermediate child. That is a real constraint on the "agents can message each other and orchestrate without routing through the user" claim: it is true, but bounded, not an open mesh. ## Skills are Python you can call, not prompts you re-paste Skills follow the standard [Agent Skills](https://agentskills.io/specification) markdown format (a `SKILL.md` with frontmatter Prime Agent loads lazily), extended with a Python-backed variant: a skill directory with a `pyproject.toml` gets installed into the kernel's virtualenv and exposed by import name. ```python report = await release_audit(repository=".", target_version="0.4.0") ``` That's a real callable, not a re-explained prompt — Prime Agent's built-in `skill-creator` skill turns a described workflow into exactly this shape: `SKILL.md` plus `src//__init__.py` plus a documented `run()`. Worth being precise about a distinction the docs themselves flag: an *installed* Python skill is a package on disk; a continual-harness *skill entry* (below) is a persisted description of a reusable call. `/refine` can create or update the description after it sees a repeated pattern, but it never packages the executable capability itself — that stays `skill-creator`'s job. ## MCP, without adding a tool The clearest test of whether a "one tool" design is a real commitment or a slogan is what happens the first time someone wants Linear and Notion in the agent. The default answer everywhere else is to mount the MCP server's tools into the model's tool list, which is how a clean six-tool agent becomes a forty-tool agent nobody planned. Prime Agent's docs refuse the move in the first paragraph: "Consistent with Prime Agent's single-tool design, MCP integrations are **not** exposed as new agent tools." An integration is a Python skill whose module subclasses `McpIntegration`, and the MCP connection runs *inside the kernel* on the official `mcp` Python SDK. The host's only jobs are browser OAuth and keeping a token fresh in `auth.json`. ```python import linear for tool in await linear.list_tools(): print(tool["name"], "-", tool["description"]) help(linear.list_issues) # schema, after list_tools() has run issues = await linear.list_issues(team="Engineering") ``` Every discovered tool is bound as an async method on the integration object; results come back as parsed Python rather than JSON to unpack; a tool whose name isn't a valid identifier (Notion's `notion-search`) falls back to `await notion.call_tool("notion-search", {...})`. Authoring your own is a `pyproject.toml`, an `mcpServers` entry in settings, and roughly ten lines subclassing the base. Two details are more interesting than the API itself. The first is that **discovery moved into the turn**. In a schema-mounted MCP integration, tool definitions are resolved when the harness connects and then frozen into the prompt; the model gets the server's surface whether it needs it or not, and a server that changes its tools mid-session is a stale-schema bug. Here the docs tell the model to `list_tools()` and `help()` before calling rather than hardcoding, because "tool names and argument schemas come from the server and can change." The model pays for the schema only in the turns where it actually looks it up. The second is a small landmine that says a lot about how the kernel works. The reference integration's module-level `__getattr__` forwards unknown attributes to the instance, but keeps a reserved list: ```python _RESERVED = {"run", "__wrapped__", "__call__"} ``` Forwarding `run` would make the kernel bootstrap, which probes modules for a callable entrypoint, mistake the whole integration module for a callable skill and break dispatch. That is the flavour of bug you only get when your tool boundary is Python's attribute protocol instead of a JSON schema — more expressive, and with sharper edges. The auth-gating is asymmetric in a way worth knowing before you write one. **Built-in** integrations (Linear, Notion) ship installed but disabled, are excluded from the prompt, and only get imported into the kernel once credentials exist. **User-authored** ones are not gated that way at all: drop a skill into a skills directory and it is visible and imported immediately, failing at call time with `NotEnabled` until you log in. So your `SKILL.md` has to tell the model how to connect — and tell it the right way, since `/mcp login` works only for OAuth servers and reports "Unknown MCP integration" for a bearer-token one. ## The Continual Harness: durable state the agent is allowed to edit Everything so far is inside one turn. The [Continual Harness](https://arxiv.org/abs/2605.09998) (arXiv 2605.09998) is about state that outlives the turn — and the session, if you ask for it to. It has four editable kinds, defined in `refinement.ts`: `prompt` (supplemental behavioral notes), `memory` (durable facts and decisions), `skill` (a description of a reusable Python call), and `subagent` (a reusable delegation role). Each entry lives in one of two scopes — `local`, written to the current session's own `harness/harness_state.json` and gone with the session unless promoted, or `global`, written to `~/.prime/agent/harness/` and available to every future session. `/refine` is the mechanism that writes to this state. It reviews the current trajectory and, when it finds something worth persisting, emits small Create/Update/Delete edits — never a full rewrite. From the actual system prompt the host sends to the refiner model: ```text Use the trajectory, current continual harness state, and prior refinement history. Prefer small evidence-backed edits. If prior refinements caused issues, rollback or replace the faulty editable entries. Never edit source files directly. ``` ## The one thing `/refine` cannot touch The interesting design choice is not that the harness can improve — plenty of systems do prompt optimization. It's what's carved out of the edit surface, and how that carve-out is enforced. `validateEdit()` in `refinement.ts` runs before any edit is applied: ```typescript if (edit.kind === "prompt" && (edit.id === "base_system_prompt" || computedId === "base_system_prompt")) { return "base system prompt is not editable"; } ``` That is not a prompt instruction the model could talk itself out of — it's a function that runs on every proposed edit, in the host, outside the model's control. The base system prompt is compiled once from the harness's own instructions, and any attempt to create, update, or delete an entry with that id is rejected before the edit ever lands. Everything the harness learns goes into one of the four editable kinds instead, injected at the top of the compiled prompt as clearly subordinate material: "Use these continual harness prompt notes, memories, skills, and subagent specs when they are relevant. The base system prompt is immutable; prompt entries below are supplemental notes only." That matters because it draws a hard line between two very different kinds of self-modification. The model can accumulate memories, refine delegation roles, and tighten behavioral notes — real, compounding change to how it behaves — but it can never touch the instructions that define what counts as a legitimate edit in the first place. Nothing in the four editable kinds can rewrite the rule that keeps them editable-only. It's the same shape as a constitution that can be amended but whose amendment procedure is (by design) not itself amendable through the ordinary amendment process. Every applied edit is versioned and every refinement pass is appended to `refinements.jsonl` with before/after entry state, which is what makes rollback possible: `refineHarness()` accepts a `rollbackId` and, instead of running the LLM proposal pass again, replays a target refinement's prior state as the new edit. If a `/refine` pass turns out to have been wrong, the fix is pointing the entry back at an earlier recorded version — not trusting a second LLM call to undo the first one's mistake correctly. Read next to [Recursive Harness Self-Improvement](/articles/recursive-harness-self-improvement), published today, the contrast is worth stating plainly. Sakana and Berkeley's method compares a harness against its own immediately-previous version and keeps the winner — a research method with a real information-theoretic argument for why pairwise beats population search, but no product around it. `/refine` is the shipped, product-side sibling of that same instinct: also self-vs-self in spirit (evidence from *this* trajectory, checked against *this* harness's own history), but with no comparison objective, no accept/reject criterion beyond "small and evidence-backed," and a rollback button instead of a formal proof. One is a method with a Bradley-Terry argument behind it; the other is a feature with a JSONL log behind it. Neither is a lesser idea for that — they're answering different questions — but they shouldn't be mistaken for the same rigor. [MemHarness](/articles/memharness), also published today, is a useful contrast in the other direction. MemHarness's argument is that retrieved memory should be *reconstructed* — critiqued and rewritten against the current state — every time it's used, because verbatim replay of a stale memory can hurt more than having none. The Continual Harness takes the opposite bet on when the work happens: refinement is a deliberate, evidence-gated event ("prefer small evidence-backed edits") that happens rarely, and once written, an entry is trusted and injected verbatim into every future compiled prompt until the next refinement touches it. MemHarness spends compute at *read* time, on every retrieval; Prime Agent spends it at *write* time, once, on `/refine`. Neither is obviously right — cheap reads with occasional expensive writes versus expensive reads with cheap storage — but it's worth knowing you're choosing between them, and Prime Agent has made the choice, not left it implicit. ## Two memories, and only one of them forgets A persistent kernel gives an agent two independent memory systems, and I don't think that gets said plainly enough. The transcript is one: bounded by the context window, and periodically summarized away. The kernel namespace is the other: bounded by RAM, and *never* summarized. Compaction only touches the first. Auto-compaction fires when `contextTokens > contextWindow - reserveTokens` — 16,384 reserved by default — walks backwards from the newest message accumulating tokens until it has kept `keepRecentTokens` (20k by default), summarizes everything before that cut into a structured document, and reloads the session as summary-plus-recent. `long-running-agents.md` states the kernel's exemption directly: "The IPython kernel persists through compaction, so variables, imports, helper functions, and task state remain available." So the same object can be simultaneously forgotten and present. The model may no longer have the message where it built `summaries`, but `summaries` is still bound in the interpreter. Both halves of that are useful and both can bite: the good case is that fifteen minutes of expensive analysis survives a compaction intact; the bad case is that the model retains a variable whose *provenance* was summarized away, and has to re-derive what it means. The structured summary format is clearly designed against this — it carries explicit `## Critical Context`, `` and `` blocks precisely so the pointers outlive the prose. Three implementation details reveal where the pressure actually is: - **Tool results are truncated to 2,000 characters** during the serialization that feeds the summarizer, with a marker recording how much was dropped — because, in the docs' own words, tool results "especially from `ipython` and optional `bash`, are typically the largest contributors to context size." The REPL design makes compaction *harder*, and this is the mitigation. - **Split turns get two summaries.** Normally the cut lands on a turn boundary. When one turn is itself bigger than `keepRecentTokens` — which is exactly what a long autonomous stretch of kernel work produces — the cut lands mid-turn on an assistant message, and Prime Agent summarizes the history and the turn prefix separately, then merges them. Never at a tool result: those must stay attached to their call. - **Compaction is explicitly not a stopping condition.** It "does not stop goals, autonomous continuations, heartbeats, or existing child sessions." A harness that treated a full context as the end of a task would quietly cap every long-running job at one context window. Both compaction and the `/tree` branch summarizer accumulate file operations *cumulatively* across passes, so the record of what was read and modified survives repeated compactions rather than being re-derived from a summary of a summary. ## What runs when nobody is attached The last piece, and the one easiest to miss from the README alone: Prime Agent is built for sessions with no human in front of them. Sessions live in resident daemon worker processes, so closing the terminal detaches a client rather than stopping the work, and there are four separate mechanisms for producing a prompt when no user is typing. - **Heartbeats**, in two flavours. `/heartbeat every 10m ...` is the user's single visible recurring instruction; `rlm_heartbeat.create(...)` is the agent's own, plural and programmatic. The Python skill deliberately cannot clear or replace the user-owned one. - **Schedules** — `prime-agent schedule add worker "0 9 * * 1-5" -- "Review open work"` — persisted per session, surviving detach. The reliability detail is good: due ticks are claimed before delivery so a crash cannot replay an uncertain prompt, and missed ticks are coalesced rather than accumulating into a backlog. - **Goals**, which store a durable objective plus token usage, elapsed time, continuation count and an optional budget. Only `await goal.complete()` marks one done. The docs are careful that a goal is "an explicit user or host action, not something the agent should infer from every task." - **Autonomous mode**, which is the policy that decides whether to inject another continuation. It is bounded on four axes at once — continuations, assistant turns, tokens, wall clock — and gated on shell commands (`--autonomous-gate "npm run check"`) that must pass before the session may finish, with a failed gate's bounded output returned to the agent for another attempt. The division there is sharper than most agent products bother with: the goal holds *what* and *how far along*, autonomous mode decides *whether to continue*. And one line in the gate policy is worth stealing outright — Prime Agent "avoids rerunning the same failed gate when the workspace has not changed." An agent that reruns a two-minute test suite against a byte-identical tree is not verifying anything, it is billing you for a cached failure. All of it funnels into the same queue. From the session queue onward, a prompt from a heartbeat, a cron schedule, a goal continuation, autonomous mode, or another agent takes exactly the same path as one typed by a person, which is why none of these features needed a parallel execution mode. One consequence of that design shows up in `--mode acp`, added in `0.6.0`, which runs Prime Agent as an [Agent Client Protocol](https://agentclientprotocol.com) agent so arbitrary clients can drive it. IPython maps cleanly onto ACP's `execute` tool call. Everything above does not: subagents, autonomous gate state, compaction, goals, heartbeats and continual-harness refinement have no native ACP concept, so they ride in a namespaced `ai.primeintellect.prime-agent` `_meta` envelope that vanilla clients ignore. That is a fair summary of where this design sits relative to the emerging standards — the single-tool core is portable, and most of what makes it interesting is an extension. ## What programmatic execution costs The tradeoff the whole design rests on: a persistent interpreter is more capable and much harder to bound than a fixed tool schema. The README says this plainly, not buried in a docs page: > Prime Agent executes model-generated Python and project commands with your user permissions. > Its worker and kernel processes improve lifecycle isolation and recovery; they are **not** a > security sandbox. Review changes and use trusted repositories, instructions, skills, and > extensions only. `rlm.md`'s trust-model section says the same thing about the kernel specifically: it "runs model-generated Python and project commands with the worker's operating-system permissions. It is a durable control environment, not a security sandbox." A fixed tool schema at least gives you an enumerable attack surface — every action the model can take is one of N defined functions, each individually auditable and individually deniable. A REPL's attack surface is "anything Python (and `%%bash`) can do," which is a much larger set to reason about, and a much easier one for a malicious skill or a compromised MCP integration to abuse. The host bridge splits that claim in two, and the split is the honest version. Against the *machine* — files, network, processes, credentials on disk — the REPL really is unbounded, and no amount of typed dispatch changes that. Against the *agent system* — spawning children, opening goals, editing harness state, steering a sibling, reading another session's transcript — the surface is exactly as enumerable as a tool schema, because it *is* one: twenty-one named handlers, arguments validated in TypeScript, most of them absent unless a session flag turned them on. When you read "not a security sandbox," read it as a statement about the filesystem, not about the agent graph. Persistence has an operational cost too, separate from the security one: a kernel that gets stuck stays stuck. The host's own busy-kernel handling spells out the tradeoff directly — interrupting a runaway cell and it still hasn't stopped, the choices are "wait" (preserve state, keep waiting) or "kill" (lose every in-memory variable, import, and running task and restart clean). There is no third option where you get both a responsive kernel and the state back. A schema-based tool call that hangs just times out; a wedged interpreter is holding real, valuable state hostage to its own unresponsiveness. The daemon-backed background sessions and inter-agent messaging compound this rather than replace it. Sessions run in resident worker processes that survive a detached terminal — genuinely useful for long tasks, and `long-running-agents.md` is honest that this is a lifecycle property, not a security one: "Daemon workers are process-isolated for lifecycle and failure containment, not security-sandboxed. They normally run with the same operating-system permissions as the client." Add agents that can message and steer each other's active work and the blast radius of one compromised or badly-instructed agent is no longer just its own kernel — it's whatever its parent, siblings, and children will act on without a human turn in between. The nuclear-family reach limit added in `0.6.0` is a real mitigation, but it bounds propagation, it doesn't remove the surface. And `0.7.0` moved in the other direction on the same axis: agent messages now *always* steer, injecting into a running turn, with the option to queue politely behind the current work removed from every API. That is almost certainly the right default for responsiveness, and it does mean an inbound message from a sibling always interrupts. None of this makes Prime Agent unusual among agent products with shell and code-execution access — it makes the tradeoff explicit and names it in the docs instead of marketing around it, which is more than most. ## Maturity, honestly This is the section the shallow clone got wrong, and the fix is more interesting than the error. On that first pass I could not audit the commit history at all — the clone showed one commit — so I declined to claim anything about the project's age. That was the right call given the instrument, and the wrong instrument. A full clone answers the question completely, and the answer is not what either a shallow clone or the changelog suggests. **4,473 commits, 231 distinct authors, 48 release tags**, first commit 2025-08-09, latest 2026-08-06, with pull request numbers past #660. That is a year-old project with a real contributor base, not a three-month-old repo — and it is worth saying that my earlier hedge, read charitably, was still an *underestimate* by an order of magnitude. But the interesting number is the split. **Mario Zechner authored 3,099 of those commits — 69%** — and his last one is 2026-05-08. Prime Intellect's first commit lands 2026-05-21. Everything before the gap is [`pi-mono`](https://github.com/badlogic/pi-mono) under its own name; the `clean up legacy pi artifacts` commit lands 2026-05-19. In the three months since, **482 commits by 17 authors** have turned it into Prime Agent. So "how mature is this?" has two honest answers depending on what you're asking. The *codebase* is a year old, heavily iterated (December 2025 and January 2026 alone account for 2,096 commits), and was production software before Prime Intellect touched it. The *product* — the RLM framing, the continual harness, `/refine`, the daemon-backed agent tree, the MCP-as-skill design — is three months old and mostly the work of a small team. If you are evaluating engineering quality, use the first number. If you are evaluating how settled the agent architecture is, use the second, and note that `0.6.0` and `0.7.0` both shipped breaking API changes inside 24 hours of each other. The changelog backs that up rather than contradicting it. `0.6.1` and `0.7.0` both landed in the three days before this was written, and `0.7.0`'s single breaking change is instructive: agent messages now "always use steering delivery," and the `mode` parameter is gone from the Python, CLI, RPC, and connection APIs. The three-mode design (`auto`, `steer`, `follow_up`) I would have described as a feature two days ago has been collapsed into one behaviour. `long-running-agents.md` still documents all three and still shows `mode="auto"` in its example; the actual `send()` in `agent-message/src/agent_message/__init__.py` no longer accepts it. Docs lagging source by one release is a normal cost of moving this fast, and a reason to read the Python, not the markdown. The rest of what I can check from the repository holds up: version `0.7.0` across all four TypeScript workspaces (`ai`, `agent`, `tui`, `coding-agent`), a `CHANGELOG.md` per package that tracks breaking changes deliberately (a house rule bans touching already-released version sections), GitHub Actions for CI and for building versioned release binaries with SHA-256 checksums, and 414 TypeScript test files plus 4 Python test files under `prime-agent-runtime/test/` across roughly 341,000 lines of TypeScript. The daemon protocol is explicitly versioned (`DAEMON_PROTOCOL_VERSION`, a schema revision — now at 13 — and compatibility maps for old-client/ new-daemon and new-client/old-daemon pairs) — the kind of care you only add after being burned by version-skew bugs. One provenance correction while I'm here. I called Prime Agent "an acknowledged hard fork" of pi-mono, which is how the docs describe it and is fair as a statement about lineage. The git history says something more literal: this is not a copy of pi-mono's code, it *is* pi-mono's repository, history unbroken from Zechner's first commit through to today's. The license reflects it — MIT, copyright jointly held by Mario Zechner (2025) and Prime Intellect (2026) — and the README credits pi-mono in its header links. Second: `assets/` in the repo has a brand logo (an SVG butterfly mark) and nothing else — I checked specifically for architecture or benchmark figures to embed as this site's house style asks for, and there aren't any. The only other images in the repository are TUI screenshots under `packages/coding-agent/docs/images/`, and they carry the pre-rename `pi-mono` branding from before the fork was productized, which would misrepresent the current product if reproduced here. So this article ships no cover image and no embedded repo figures — the four interactives above are original, built from reading the code and measuring the repository, not redrawn from anything Prime Intellect published. ## Where this sits [Scaling agentic RL](/articles/scaling-agentic-rl) covered Prime Intellect's environments side — 23 agentic tasksets, roughly 365,000 tasks behind one taskset API, each with a graded, reproducible reward. Prime Agent is the natural agent-side counterpart to that stack: the same company building the environments an RL loop trains against is also shipping the agent architecture that would run inside them. I want to be careful about what that observation is and isn't — I did not find any published result training or evaluating Prime Agent against that taskset catalog, so this is a structural connection (same company, complementary halves of an agentic-RL stack), not a reported one. If that pairing produces a number, it belongs in a different article than this one. What Prime Agent actually is, stripped of both the marketing framing and my own enthusiasm for the design: an open-source harness that replaces a tool-call schema with a programming environment, and a harness state that can accumulate evidence-backed edits without ever being allowed to rewrite the rule that makes those edits legitimate. Both are real, checkable design decisions. Neither comes with a number attached. The revision changed my read on one of them. "Replaces a tool-call schema with a programming environment" is the marketing line and it is half right. What Prime Agent actually did is *demote* the schema — out of the provider payload, where it is re-billed every turn and every entry competes for the model's attention, and into a private typed bridge the model reaches through a general purpose language. Twenty-one operations, argument-validated, conditionally registered. That is a better idea than abolishing the schema would have been, and it is the part I would steal. --- *Sources: the [prime-agent repository](https://github.com/PrimeIntellect-ai/prime-agent) at commit `fix(coding-agent): isolate kernel state tests (#661)`, 2026-08-06 — specifically the docs under `packages/coding-agent/docs/` (`architecture.md`, `rlm-runtime.md`, `rlm.md`, `compaction.md`, `mcp-integrations.md`, `long-running-agents.md`, `acp.md`), the TypeScript host (`agent-session.ts`, `refinement.ts`, `agent-messages.ts`, `tools/ipython.ts`), the Python runtime under `prime-agent-runtime/src/rlm/` and `packages/coding-agent/skills/`, and the per-package `CHANGELOG.md` files. All repository statistics — commit counts, author counts, per-month distribution, tag count, line counts — were measured with `git` against a full clone and are reproducible from the commands in the source of the commit-history figure. The Continual Harness paper is [arXiv 2605.09998](https://arxiv.org/abs/2605.09998); the RLM framing is Prime Intellect's [RLM post](https://www.primeintellect.ai/blog/rlm). The four interactives are mine.* --- # Recursive Language Models: context as a variable, recursion as a function call > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/recursive-language-models > date: 2026-08-06 > tags: agents, llm, recursion, context-management, prime-intellect, explainer The standard agent loop has one data structure at its center: the transcript. The model reads a growing conversation, emits a tool call that matches a JSON schema, a harness parses it, runs it, and appends the result back as another message. Everything the model can act on has to live in that transcript. Everything it does has to fit the schema. [Agent harnesses](/articles/agent-harness) and [the harness effect](/articles/harness-effect) are both, in different ways, about how much that loop shape costs — in design effort and in tokens. A **Recursive Language Model** (RLM) doesn't optimize the loop. It replaces the data structure. An RLM gives the model a persistent Python REPL instead of a transcript. Two things follow from that, and they are the spine of this piece: 1. **Prompt-as-a-variable.** The input — a document, a codebase, a 500 MB corpus — becomes a value bound to a name in the REPL, not text sitting in the context window. The model writes code to slice, filter, and search it, and only what it chooses to print ever enters its own context. 2. **Programmatic recursion.** Calling another language model is a function call — `rlm(...)` — that returns a value like any other call. It composes with `for` loops, `if` statements, `map`, and error handling, because it *is* one of those, not a special harness verb bolted on next to them. **Idea, implementation, product — three different things, one name.** Alex Zhang (MIT CSAIL) coined "Recursive Language Model" for a specific inference-time technique: one long prompt held as a REPL variable, decomposed by recursive sub-LM calls, terminated by an explicit `FINAL()`/`FINAL_VAR()`. Prime Intellect built two things that carry the same name and the same two mechanisms but are not the same system as each other or as Zhang's original design: `RLMEnv`, a research-eval reproduction inside their `verifiers` library, and Prime Agent, a general coding harness that generalizes both mechanisms — context-as-data, recursion-as-a-call — to its entire working environment, not just one input string. [Prime Agent](/articles/prime-agent) covers that second one in full, including its separate Continual Harness feature. This piece is about the idea and how the paper measures it; read that one for the shipped product. ## Who actually built this Prime Intellect's own post is straightforward about credit, and I'll be too: "the Recursive Language Model (RLM), introduced by Alex Zhang in October 2025 as a blog post, and now available as a full paper," with an acknowledgment thanking him "for his original work on recursive language models." Zhang's post frames the mechanism plainly — an RLM is "a thin wrapper around a LM that can spawn (recursive) LM calls for intermediate computation," with an API meant to be a drop-in replacement for an ordinary completion call: `rlm.completion(messages)` where you'd otherwise write `gpt5.completion(messages)`. The motivating problem is what he calls **context rot**: model recall degrades as context grows, independent of whether the context still technically fits the window. The idea was formalized two months later in [*Recursive Language Models*](https://arxiv.org/abs/2512.24601) (arXiv 2512.24601, submitted 2025-12-31), authored by Alex L. Zhang, Tim Kraska, and Omar Khattab, all MIT CSAIL. Their own framing in the abstract: RLMs are "a general inference paradigm that treats long prompts as part of an external environment and allows the LLM to programmatically examine, decompose, and recursively call itself over snippets of the prompt." The paper is explicit about what it's reacting against, and credits the right ancestors rather than claiming recursion or code-as-tool-use as new: - **CodeAct** established writing code as the tool-call format, but in a standard coding agent that code still executes inside the same context-constrained loop — sooner or later the harness has to compact. RLMs offload the entire prompt as an external variable instead, so the REPL's addressable state isn't bounded by the model's window at all. - **MemGPT** manages context explicitly, paging things in and out of a single model's working memory. RLMs don't build a memory hierarchy; they let the model itself decide what to look at, programmatically, each time. - **ReAct**-style sub-calls are verbalized autoregressively — described in natural-language turns inside one transcript. RLM sub-calls are constructed programmatically and their results are stored as REPL variables, which is what lets a `for` loop over sub-calls do real accumulated work instead of restating each result back into the same window. So the general idea — treat a long input as an external, programmatically addressable environment, and let recursive delegation happen through control flow instead of prose — has a real, credited origin, and it is not Prime Intellect. What Prime Intellect has done is build two different things on top of it. Their research post says plainly: "we at Prime Intellect have implemented our version of the RLM in [verifiers](https://github.com/PrimeIntellect-ai/verifiers/) so that it is ready to be used in any environment," landing as the experimental `RLMEnv` — a reasonably faithful reproduction of Zhang's design, built for running controlled evaluations. Separately, `prime-agent-runtime`'s `rlm` package — the one this site's [Prime Agent](/articles/prime-agent) piece covers — takes the same two mechanisms and applies them to an entire general-purpose coding agent: files, shell commands, skills, and subagents all go through the same persistent kernel, not just one oversized input prompt. Zhang's design restricts the root model to *metadata* about the prompt (its length, a prefix) until it explicitly decides to look closer; Prime Agent's root model just has an ordinary working context plus a kernel, because it's built to be a general agent, not a single-prompt inference technique. Related, useful, and worth keeping straight — not the same artifact. ## The prompt becomes a variable
In the paper's own algorithm, the root model never receives the prompt as tokens in its context. It receives metadata — length, a prefix, how to access it — and a REPL where that prompt already sits as a variable. The loop is: the model writes code, the REPL executes it, truncated stdout comes back, and this repeats until the model calls `FINAL(answer)` to return a string directly or `FINAL_VAR(name)` to return whatever a REPL variable currently holds. Nothing about the prompt's actual content is ever force-fed into the root model's window; the model decides what to look at, a slice at a time. Prime Agent's version of this same bet is less specialized but the mechanism is identical in spirit: a persistent IPython kernel that survives across turns, with `rlm` preloaded in the namespace. Here is the actual shim that puts it there, from `prime-agent-runtime/src/rlm/__init__.py`: ```python class _RLMCallable: async def run(self, prompt: str, **kwargs: Any) -> RLMSpawnHandle: return await run(prompt, **kwargs) async def __call__(self, prompt: str, **kwargs: Any) -> RLMSpawnHandle: return await run(prompt, **kwargs) rlm = _RLMCallable() ``` `rlm` is not a tool the model selects from a menu. It is a plain Python object with a `__call__` method, sitting in the kernel's global namespace the same way any import would. Calling it is calling a function, full stop — which is the whole point: nothing about `await rlm(...)` needs a harness to specially recognize the string `"rlm"` and route it through a different code path than any other line of Python. Here's the same underlying claim made concrete with a task. Say the question is "which of these log files mentions an out-of-memory kill, and which one is worst." A schema-based harness does this as a sequence of round trips — each one a full model turn, a parsed tool call, and an appended result: ```text # illustrative — the general shape of a schema-based harness, not quoted from a specific product assistant: tool_call grep(pattern="Out of memory", path="logs/") tool_result: {"matches": ["logs/worker-014.log:88231", "logs/worker-014.log:88245", ...340 more]} assistant: tool_call read_file(path="logs/worker-014.log", offset=88200, limit=100) tool_result: "<100 lines of log text>" # …and one more round trip per file the model wants to actually look inside ``` The RLM version, written in the same idiom as the docs' own `config_files` example (`packages/coding-agent/docs/rlm.md`): ```python from pathlib import Path hits = [p for p in Path("logs").rglob("*.log") if "Out of memory" in p.read_text(errors="ignore")] worst = max(hits, key=lambda p: p.stat().st_size) print(f"{len(hits)} files mention OOM; worst by size: {worst}") ``` `hits` is a real Python list, still there next turn if the model wants to `map` something else over it. The grep, the read, and the size comparison happen inside one cell. The model's context grows by one printed line, not by one message per file. That's the mechanism at data-scale, not code-scale: the top bar is what has to happen when the only way to look at something is to read it into the window — the window caps out at 272K tokens for GPT-5 regardless of how the corpus is chunked, so a 500 MB corpus needs on the order of hundreds of read-and-compact rounds just to scan once, and any single round can only ever see a 272K-token slice. The bottom bar is the REPL: the corpus is bounded by machine memory, not context budget, and only a found, printed slice ever reaches the model. This is the same shape as the paper's own S-NIAH and OOLONG scaling runs — hold the task fixed and grow the input, and one line stays flat while the other falls off past the window boundary. ## Recursion is just a call The second inversion is about delegation. In a schema-based harness, spawning a subagent is a distinct, specially-recognized action — usually literally called `Task` or `subagent` in the tool list, with its own parsing path in the harness. In an RLM, `rlm(...)` is not a different *kind* of call from anything else in the REPL. It's an `async` function that happens to start another agent instead of, say, reading a file. Here's the actual implementation, trimmed from the same file: ```python async def run(prompt: str, **kwargs: Any) -> RLMSpawnHandle: """Spawn a recursive Prime Agent child and return once its task is admitted.""" if not isinstance(prompt, str): raise TypeError(f"prompt must be str, got {type(prompt).__name__}") payload = await host_request("rlm.run", {"prompt": prompt, "kwargs": kwargs}) return _spawn_handle_from_payload(payload) ``` `host_request` opens a Jupyter comm to the TypeScript host, which creates a real child `AgentSession` and returns as soon as the task is *admitted* — not when it's *done*. That admission- not-completion design is a deliberate choice recorded in Prime Agent's own changelog (covered in more depth in [Prime Agent](/articles/prime-agent)), and it's what makes fan-out compose with ordinary control flow instead of blocking on it: ```python # real — packages/coding-agent/docs/rlm.md api_review = await rlm("Review the public API", name="api-reviewer") test_review = await rlm("Review the test coverage", name="test-reviewer") integration_audit = await rlm("Run the slow integration audit", name="integration-audit") ``` Three lines, three independent children, one turn. In a schema-based harness the equivalent is three separate structured messages, each requiring the model to emit a full tool call and the harness to parse and dispatch it — and, if the harness's subagent tool is synchronous, a wait on each before the next line can even be written. Here it's `for child in reviewers: await rlm(child)` if you want a loop, or three independent statements if you don't. Recursion composes with the rest of the language because it's written in the rest of the language. The tree above is the concrete version of "a task over 200 files is a `for` loop with 200 model calls made by the program, not 200 round trips through the model's own context." Each spawned session's own context holds exactly one task — the child at `RLM_DEPTH=1` never sees its siblings, never sees the root's other work. What holds the shape of the whole job is the root session's Python namespace: the list of handles, the loop that produced them, the code that will eventually read their replies. That's a different place for "the state of the whole task" to live than any single model's context window, and it's why the unit of work stops being "how much fits in 200K tokens" and starts being "how many child sessions can the host actually run." ## What's actually measured, and by whom Two different evaluations exist, and they're worth keeping apart because they measure different things with different rigor. **The paper's own numbers** (arXiv 2512.24601) are the more citable evidence, because they're in a reviewable artifact with a stated protocol: GPT-5 and Qwen3-Coder-480B-A35B-Instruct, compared against RLM wrappers of themselves, across S-NIAH, OOLONG, OOLONG-Pairs, BrowseComp-Plus, and CodeQA. The cleanest figures to quote are the ones in the abstract, because they are stated exactly rather than read off a chart, and because they are relative to the right baselines. Wrapping GPT-5 in an RLM beats — by a median across the evaluated benchmarks — **26% against compaction, 130% against CodeAct with sub-calls, and 13% against Claude Code**, "while having comparable cost." That middle number is the one that matters most for the argument here: CodeAct *also* gives the model code execution and sub-calls. The gap between it and an RLM isn't code-versus-schema, it's whether the input lives outside the model's context or inside it. The paper also claims inputs "up to two orders of magnitude beyond model context windows." And one contribution the write-ups mostly skip: they didn't only wrap existing models, they **post-trained one for the paradigm**. RLM-Qwen3-8B beats plain Qwen3-8B by 28.3% on average and, per the abstract, "approaches the quality of vanilla GPT-5 on three long-context tasks." An 8B model approaching a frontier model on long-context work by being trained to drive a REPL rather than to read further is the most interesting claim in the paper, and the one I'd most want replicated. The headline figure scales input length from roughly 8K to over 1M tokens on S-NIAH, OOLONG, and OOLONG-Pairs: GPT-5 degrades sharply as input grows, especially past its own 272K-token window where it structurally cannot see the rest of the input at all, while RLM(GPT-5, depth=1) stays roughly flat across the same range. In the paper's tables — read from its figures, so treat the exact decimal as approximate rather than a number I recomputed myself — GPT-5 alone scores around 44% on OOLONG versus roughly 56–58% for the RLM wrapper at recursion depths 1 and 3; on BrowseComp-Plus (1,000 documents, 6–11M tokens total) GPT-5 alone scores 0% because the input doesn't fit at all, versus roughly 91–92% for RLM(GPT-5); on CodeQA, GPT-5 scores around 24% versus roughly 62–66% for the RLM wrapper. These are real, single-paper, not-yet-independently-replicated numbers — but they come with a stated model, a stated task, and a stated context length, which is more than most of what gets cited as evidence for an agent architecture. **Prime Intellect's own post** runs a separate, smaller evaluation: GPT-5-mini through their `RLMEnv` implementation, across four `verifiers` environments — DeepDive (web research), Math-python, Oolong, and Verbatim-copy — at 50 rollouts each. This is where the honesty has to cut both ways. RLM helps on DeepDive (with explicit strategy tips pushing sub-LLM calls further), helps on Oolong (the plain LLM gets close to zero reward on the longest real-data contexts; RLM keeps working out to roughly 1.5M characters), and helps on Verbatim-copy across most content types. On Math-python it does not: the post reports RLM performing *worse* than the plain LLM, and ablating the REPL's timeout up to 600 seconds doesn't close the gap. That's a genuine negative result from the people building the thing, reported plainly rather than left out — a useful data point on where "run more code" isn't automatically the right move for a task that's mostly reasoning, not search. Charts, not tables: the post shows relative comparisons rather than a numeric results table, and says so itself — "this is not a measurement of any model's absolute performance on any benchmark." A day before this piece, Prime Intellect published a separate launch post for Prime Agent reporting Opus 5 scoring 95.5% Best@1 on ARC-AGI-3 (183 test levels), against a self-reported 30.2% baseline for Opus 5 on its own harness, and a table comparing Prime Agent across models against Pi-mono, Claude Code, and Codex on OOLONG, OBLIQ-Bench, and LongBenchv2. That's Prime Intellect's own first-party number, for the product, not in the open-source repository — consistent with what [Prime Agent](/articles/prime-agent) already found reading the repo itself: no benchmark numbers ship in the code, only in the marketing post. It belongs in a product-level discussion, not this one; I mention it only so this piece doesn't read as unaware of it. ## What this costs None of the above is free, and the paper and the Prime Agent docs are both honest about the price. **Arbitrary code execution is the primary interface, not a fallback.** A fixed tool schema gives you an enumerable, individually-auditable set of actions. A REPL's action space is "anything Python (and a shell cell) can do." Prime Agent's own trust-model documentation says this plainly: the kernel "is a durable control environment, not a security sandbox." Zhang's design narrows the blast radius somewhat by keeping the root model restricted to prompt metadata until it asks for more — but the sub-LM calls and any tool access still execute inside the same interpreter. **A stateful interpreter can wedge.** Persistent state across turns is the entire point of the design, but it means a hung cell doesn't just time out cleanly the way a stuck tool call does — Prime Agent's own busy-kernel handling frames the choice as "wait" (preserve state, keep waiting indefinitely) or "kill" (lose every in-memory variable and restart clean). There's no option that gets you both a responsive kernel and the state back. **Recursion needs an enforced budget, not a polite one.** The recursion tree above has a real cap behind it: `RLM_MAX_DEPTH` defaults to 1, and the check — ```typescript // packages/coding-agent/src/core/agent-session.ts if (this._rlmDepth >= this._rlmMaxDepth) { throw new Error( `RLM recursion depth limit reached (RLM_DEPTH=${this._rlmDepth}, RLM_MAX_DEPTH=${this._rlmMaxDepth})`, ); } ``` — runs in the host before a comm channel even opens, not as a prompt instruction the model could argue its way past. That matters because the cost of recursion is exponential in depth if fan-out is uncapped: `fanout^depth` sessions, each billed and each capable of spawning more, versus a `for` loop's cost growing linearly in the number of iterations. A depth cap enforced in code is the difference between "a program that does 200 things" and "a program that can, in principle, spawn without bound." **Harder to sandbox and audit than a fixed schema.** Every action a schema-based agent can take is one of N defined functions — individually reviewable, individually denyable. "Any code the model writes" is a much larger surface for a malicious skill, a compromised MCP integration, or a bad instruction to abuse, and it's a correspondingly harder surface to review after the fact. **Debugging shifts from reading a transcript to debugging a program.** A schema-based agent's failure mode is usually legible from the transcript alone: read the tool calls and results in order. An RLM's failure mode can be a bug in generated code, a REPL state that's subtly wrong three cells after the mistake that caused it, or a child session that never replies because nothing in the parent's code checks for it. That's a different, and for most engineers a more familiar, debugging discipline — but it is a different one, and treating it like transcript-reading will miss real bugs. ## Where this sits RLM is an inversion of the model [Lilian Weng's harness framing](/articles/agent-harness) and [the harness effect](/articles/harness-effect) both describe: a fixed loop around a fixed tool schema, where the transcript is the only place state can live. RLM doesn't optimize that loop's token economics — it removes the transcript as the place state lives at all, replacing it with an interpreter. It's also a different axis from two other pieces published alongside it. [Recursive Harness Self-Improvement](/articles/recursive-harness-self-improvement) treats the harness as a single text prompt and improves it by comparing it against its own immediately-previous version — harness as a string being optimized. RLM treats the harness as a program the model writes fresh each turn — harness (or at least the working state) as code being executed. And the Continual Harness [paper](https://arxiv.org/abs/2605.09998) — *Continual Harness: Online Adaptation for Self-Improving Foundation Agents*, by Seth Karten, Joel Zhang, Tersoo Upaa Jr, Ruirong Feng, Wenzhe Li, Chengshuai Shi, Chi Jin and Kiran Vodrahalli — is a third axis again: an online loop that alternates acting with refining the agent's *own* prompts, skills, memory, and subagent specs *during* a run, without resetting. (Its Zhang is Joel Zhang, not RLM's Alex L. Zhang — same surname, different author.) Prime Agent's own Continual Harness feature, covered in [Prime Agent](/articles/prime-agent), draws on this and composes with its RLM runtime. The connection is more than citational: the paper's lead author, Seth Karten, is an active committer to the prime-agent repository, with 39 commits in its history as of 2026-08-06. So the two ideas arrive in the same product from the same people — but they answer different questions. RLM is about how one task executes; Continual Harness is about how the harness's own configuration evolves across tasks. [MemHarness](/articles/memharness) is a useful contrast in the opposite direction from RLM's whole bet. MemHarness's argument is that a retrieved memory should be reconstructed — critiqued and rewritten — against the current state every time it's used, because stale verbatim replay can hurt more than no memory at all. RLM doesn't reconstruct anything: the corpus sits in a variable exactly as it was written, and the model's job is to write code that finds the right slice of it, not to have that slice handed to it pre-digested. Reconstruction spends compute making memory trustworthy before use; RLM spends compute letting the model decide what's worth looking at, each time, from an unmodified source. Different failure modes follow from each: a MemHarness-style system can misreconstruct; an RLM-style system can simply fail to look at the part that mattered. The idea itself is Zhang, Kraska, and Khattab's, credited honestly by the company building on it. What Prime Intellect has actually built is two separate systems that take the same two mechanisms — context as a variable, recursion as a call — and apply them at different scopes: one faithful to the original single-prompt design, one generalized into an entire agent. Both are real, checkable architecture decisions. The paper's numbers are the most rigorous evidence either has; Prime Intellect's own eval is smaller, honestly mixed, and shown as charts rather than a table; and the product-level ARC-AGI-3 number belongs to a different piece than this one. --- # TencentDB Agent Memory: the four-tier pyramid is really a cache-stability hierarchy > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/tencentdb-agent-memory > date: 2026-08-06 > tags: agents, memory, context-management, open-source, retrieval, explainer [TencentDB Agent Memory](https://github.com/TencentCloud/TencentDB-Agent-Memory) is Tencent Cloud's open-source memory layer for coding agents — MIT, TypeScript, about 148,000 lines across three services, with adapters for OpenClaw, Hermes, Claude Code and CodeBuddy. The pitch is the one every agent-memory product makes: stop re-explaining your project to every new session. The pitch is not the interesting part. The interesting part is a decision buried in a source comment, which reframes the whole design and is worth stealing whether or not you ever run this software. ## The pyramid everyone builds Start with what the README shows you. Conversations are captured raw and refined by an async pipeline into four levels:
So far this is the standard picture, and on its own it does not tell you much — "summarize old things into shorter things" describes most memory systems ever built. The obvious reading is that the pyramid is about *abstraction*: L1 is more general than L0, L3 more general than L2. Read the code and a different axis appears. ## Sorted by volatility, not abstraction `MemoryProxy/src/injection/injectors/tdai-profile-memory-injector.ts` opens with a comment explaining how each tier reaches the model, and the three answers are all different: - **L3 persona** — injected in full. "稳定且通常较短" — stable and usually short. - **L2 scenarios** — **only the scene-navigation index** goes in: a list of paths plus a one-line summary each. The stated reason is that L2 full text "经常上千 chars × N 个" — often thousands of characters times N blocks. The agent reads the full scene through a tool once it has decided the scene matters. - **L0 and L1** — not injected at all on this path. The comment is blunt: "不再自动召回" — no longer auto-recalled. They are exposed as read-only search tools. The proxy README gives the reason in one clause, and it is the whole thesis of the system: > injects Skills, Knowledge and Memory L2/L3 into the system prompt on demand; L0/L1 are exposed as > read-only tools for the model to query proactively, **avoiding upstream KV-cache invalidation**. That is a token-economics argument, not a knowledge-management one. Anything you put in the system prompt that differs from last turn breaks the provider's prompt cache and makes you re-pay for the entire prefix — so the question "which tier goes where" is really "which tier is stable enough to sit in a cached prefix." Sorted that way, the pyramid falls out for a completely different reason than the abstraction story suggests. L3 is at the top not because it is the most abstract but because it is the most *stable*. L0 is at the bottom because it grows every single turn. The discipline shows up again in the deployment guide, which is where you can tell someone has run this in production: > For multi-node deployments you must use `storage.backend=cos` and explicitly set > `injection.externalGatewayUrl`, otherwise each instance caches independently and causes upstream > KV-cache misses. Treating a prompt-cache miss as a documented operational failure mode — in an *installation* doc — is not something most agent-memory projects think to do. It rhymes with what [the harness effect](/articles/harness-effect) argued from the other direction: the orchestration layer, not the model, sets the bill. The same repository makes the opposite choice on its other integration path. `MemoryCore`'s `auto-recall` hook — used by the OpenClaw plugin — *does* automatically retrieve L1 and inject it into context before the agent runs. So the product ships two integration surfaces with two different injection policies, and only the proxy one is cache-preserving. Neither doc mentions the divergence. If you are evaluating this, which path you take changes the token behaviour substantially. ## The retrieval underneath is real It would be easy to ship "hybrid search" as a marketing phrase. This is not that. `auto-recall.ts` runs FTS5 BM25 for the keyword side and cosine similarity over a vector store for the dense side, then merges the two rank lists with reciprocal rank fusion, with the constant spelled out and attributed: ```typescript // RRF merge: k=60 is a standard constant from the RRF paper const RRF_K = 60; ``` Scoring a record at rank *r* as $1/(k + r)$ and summing across lists is the standard formulation, and k = 60 is the value from Cormack et al. If the backing store can do dense-plus-sparse-plus-RRF server-side it short-circuits to one API call; on the SQLite path it runs both sides in parallel and fuses client-side. There is a graceful degradation if FTS5 is unavailable — the keyword list comes back empty and RRF operates on the dense side alone. This is the same fusion idea this site's own search uses, and it is the right default: BM25 finds the document that says the exact identifier you typed, embeddings find the one that means what you meant, and RRF combines them without needing calibrated scores from either. ## What actually bounds an injection The README says results are "further capped by item count, character budget, and timeout limits to prevent memory from overwhelming the context window." Checking the shipped defaults in `MemoryCore/src/config.ts`, that is two-thirds true. `maxResults` defaults to 5, `scoreThreshold` to 0.3, `timeoutMs` to 5000 — all binding. But `maxCharsPerMemory` and `maxTotalRecallChars` both default to **0**, and the budgeting function short-circuits when they are: ```typescript if (!maxCharsPerMemory && !maxTotalRecallChars) { return lines; } ``` So the character budget exists, is properly implemented with truncation markers and drop counts, and ships turned off. Five results still bounds things, but five results of unbounded length is a different guarantee than the sentence implies — and a single sprawling L1 memory is exactly the case a character budget exists to catch. It is a one-line config fix, not a design flaw, but you have to know to make it. A smaller drift in the same file: `l1IdleTimeoutSeconds` is documented in its own doc comment as "default: 30" and initialized to `600`. Twenty times the documented value. ## The guide it injects into your agent One more thing the code shows that no doc mentions. Alongside the memories, MemoryCore injects a usage guide telling the model how to retrieve more — and it is hardcoded **in Chinese**, in a repository whose README, install guide and contributing guide are all bilingual: ```text ### ⚠️ 调用次数限制 每轮对话中,tdai_memory_search 和 tdai_conversation_search 合计最多调用 3 次。 ``` *"Per conversation turn, `tdai_memory_search` and `tdai_conversation_search` may be called at most 3 times combined."* The guide goes on to instruct the model that if three searches turn up nothing, the information is not in memory and it should answer from what it has rather than keep searching. Two observations. The **3-call ceiling is a good idea** — an agent that can search its own memory without limit will, and each miss costs a round trip. Naming the budget in the prompt and telling the model what to do when it is exhausted is more thoughtful than most retrieval integrations manage. And the **language is a real deployment consideration**: a fixed Chinese-language instruction block enters the context of every agent this wraps, including English ones. Models handle it, but it consumes tokens in a tokenizer that is not optimized for it and it sets the instruction language for that portion of the prompt. ## Permissions, checked against the code The visibility model is the part I expected to be thinnest and it is the most carefully built. The README promises `private` means private "not even team admins," and `permission-checker.ts` backs it with a dated comment explaining the choice: ```typescript case "private": // 私密语义(2026-07 变更):严格私密,只有 owner_user_id 能访问。 // 团队 admin 也不放行 —— 因为第 2 步 owner 判定已优先返回 ALLOW, // 走到这里说明当前 user 不是 owner,即使是 admin 也一律拒绝。 return { allowed: false, reason: "visibility_restricted" }; ``` A July 2026 semantics change, the reasoning preserved in the source, and the consequences enumerated underneath it — including that admin `list-accessible` calls must not return other people's private assets. That is a team that had the "should admins see everything?" argument and wrote down how it ended. One gap worth naming, because the README's framing does not survive it. `restricted` is described as "precise access via User / Role / Agent ACLs," and for ordinary members that is exactly what the code does — an explicit ACL match is the only way in. But the check is gated on `membership.role !== "admin"`, so **team admins skip the ACL entirely** and fall through to role defaults. Defensible — someone has to administer the thing — but "strict ACL whitelist" is true for members and not for admins, and the docs do not say so. ## The architecture, briefly
Three services. **MemoryCore** owns storage and the L0→L3 pipeline. **MemoryKnowledge** builds the Wiki and CodeGraph assets. **MemoryProxy** is the clever piece: a transparent LLM proxy that forwards OpenAI `/v1/chat/completions` and Anthropic `/v1/messages` verbatim, doing session setup, injection and write-back on the way past. Point your coding agent's base URL at it and you get team memory "without changing a single line of code." That is a genuine integration strategy rather than a shortcut. It also means the proxy sits in the path of every request and every response, holding your model credentials, which is a trust decision worth making deliberately rather than by following a quickstart. The unification is the real product claim: Chat Memory, Skills, Wiki and CodeGraph are all registered as **Memory Assets** with owner, version, status, visibility and agent bindings, retrieved through one permission-scoped surface. The README's comparison table puts it well — RAG answers "what can be found?", and this also answers "who can use it, which version is valid, and which agent should receive it." Whether that ontology is worth its complexity depends entirely on whether you have a team; for one person with one agent it is overhead. ## The number There is exactly one benchmark in the repository, and it is in the README: | Benchmark | Without | With | Relative | |---|---|---|---| | PersonaMem | 48% | 76% | +59% | That is the entire evaluation. No harness, no model named, no agent configuration, no seed count, no link to a run. I searched the repository for any other mention of PersonaMem and found two — the same table in the Chinese README. So there is no reproduction script here, and the claim is first-party and unreplicated. To be fair on two counts: a memory layer improving a *memory* benchmark is not a surprising result, and the repository's own Notes section is refreshingly frank about what is unfinished — CodeGraph "currently prioritizes public HTTPS repositories," the Hub supports manual binding while "fully automated memory routing is still under iteration," and Team Memory is labelled Beta. A project that tells you which parts are not done yet has earned some patience about the parts it has not measured. The provenance is also handled properly. The acknowledgements credit [CodeGraph](https://github.com/colbymchenry/codegraph) for code the CodeGraph module "uses," Nous Research's Hermes Agent for part of the Skill management code, and Karpathy's LLM-wiki gist for the Wiki design — specific about what was borrowed rather than a generic thank-you list. ## The take Most agent-memory projects are a retrieval index with an ontology bolted on, and the ontology is where the marketing lives. This one has a real idea underneath it, and the idea is not the pyramid. It is that **memory has to be sorted by how often it changes, because the cost of memory is not storage, it is the prompt prefix you invalidate by updating it.** Once you see the four tiers as a cache-stability ordering rather than an abstraction ordering, the delivery mechanism for each one stops being arbitrary: stable things get injected, semi-stable things get injected as an index, volatile things become tools with a call budget. That principle is portable to any agent you are building, with or without this software. What comes with the software is a competent hybrid retriever, a genuinely careful permission model, a transparent proxy that is a real integration story and a real trust decision, one unreplicated benchmark number, two integration paths that disagree about injection policy, and a character budget you should turn on before you rely on it. --- *Sources: the [TencentDB-Agent-Memory repository](https://github.com/TencentCloud/TencentDB-Agent-Memory) at its 2026-08-06 state — `README.md`, `INSTALL.md`, `MemoryProxy/README.md`, and the TypeScript in `MemoryCore/src/config.ts`, `MemoryCore/src/core/hooks/auto-recall.ts`, `MemoryCore/src/metadata/service/permission-checker.ts` and `MemoryProxy/src/injection/injectors/`. Both figures are the project's own, flattened onto white; the pyramid is its English-language variant. Chinese source comments are quoted verbatim with my translations. The PersonaMem figure is the project's own and is not independently replicated. Both interactives are mine.* --- # ABot-World-0: a 5B world model that wins on efficiency, not the leaderboard > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/abot-world > date: 2026-08-03 > tags: world-models, video-generation, diffusion, open-weights Most interactive-world demos are a video model you watch. ABot-World-0, from Alibaba's AMAP CV Lab, is one you steer: upload a starting image, hold WASD, and a 5B-parameter model streams a continuously-controllable 720p world back at up to 16 frames per second, from a single RTX 5090, with 1.2 seconds between an action and its first frame on screen. Code, weights (Apache-2.0), and a 500-hour action-annotated dataset are all released. It does not top the benchmark it's measured on. That's the more useful fact to lead with, not bury.
## Bidirectional teacher, causal student, then fix the drift The model is a distilled video-diffusion world model, trained in three stages on top of the `Wan2.2-TI2V-5B` backbone. First, a high-quality **bidirectional** action-conditioned teacher learns good dynamics over full temporal context -- accurate, but it has to see the whole clip at once, so it can't stream. Second, teacher forcing plus causal ODE distillation compress that into a **causal**, few-step student that approximates the teacher's denoising trajectory turn by turn, which is what makes low-latency interaction possible at all. Third, and this is the paper's actual technical contribution: **LongForcing**.
Plain causal distillation only ever supervises short, clean trajectories -- it never sees what its own errors compound into over a minute of rollout. LongForcing closes that gap directly: let the causal student generate a long self-rollout, then correct the *distribution* of that self-rollout against an extended-horizon teacher, rather than only imitating clean short clips. It is, in effect, training the model against its own failure mode instead of only against ground truth. Measured over a 60-second rollout against a plain causal-forcing baseline, the effect shows up as less accumulated visual damage, not a benchmark score:
The baseline doesn't fail suddenly -- it drifts. Causal Forcing looks comparable to LongForcing for the first 15-20 seconds on all four curves, then peels away: color saturation creeps up, the image blurs, patches start repeating. That's exactly the accumulated-error problem autoregressive video generation is known for, and LongForcing's fix is to train against long rollouts directly rather than assume short-horizon quality generalizes. ## Real-time is five separate wins, not one "Few-step generation does not automatically translate into real-time interaction" is the paper's own line, and Table 2 backs it up in a way that's genuinely counterintuitive: adding a faster attention kernel by itself does nothing, because the model doesn't fit in memory to begin with. Every one of those five changes is load-bearing. Skip the VAE swap and the faster attention kernel just gets you a faster out-of-memory error. That's a more honest way to read "single desktop GPU" than treating it as one clever optimization -- it's a full-stack co-design where the first fix is the one that makes the rest of the stack possible to even measure. ## WorldRoamBench: a real third-party number, and it doesn't sweep WorldRoamBench is not Alibaba's benchmark -- that independence is worth stating plainly, because it means ABot-World-0's score wasn't set by the people reporting it. Against Genie 3, HappyOyster, LingBot-World (14B), and HY-World 1.5 (8.3B): ABot-World-0 sits second, not first, on Strict Accuracy -- and that pattern holds across the rest of the benchmark's sub-metrics too: Two honest qualifications on top of what the chart above already shows. First, neither Genie 3 nor HappyOyster has a disclosed parameter count, so the efficiency claim is only verifiable against the two comparators whose sizes are public -- LingBot-World and HY-World 1.5 -- not against the benchmark's actual leader. Second, ABot-World-0 running on a single consumer GPU is not a property this benchmark measures at all; WorldRoamBench scores output quality and controllability, not deployment cost. The efficiency story and the benchmark score are two separate claims, and only one of them is what WorldRoamBench actually tested. Reproducibility here is unusually complete for this space: code, weights, and a 500-hour action-annotated dataset (`ABot-World-Explorer-500h`) are all released under Apache-2.0. The release README documents a staged rollout from 2026-07-09 through 2026-08-03 -- today, by this piece's own dateline. That's a meaningfully higher bar than a paper with numbers and no artifacts. ## What's missing The paper's qualitative claims -- physically plausible responses despite no explicit physics training, coherent hour- and day-scale rollouts, generalization to out-of-domain controls -- are demonstrated with cherry-picked keyframe strips, not a systematic user study. That's standard for this genre of paper, not a special flaw of this one, but it means "plausible physical responses" is an illustration, not a measured claim the way WorldRoamBench's numbers are. The data-collection system behind all of this, WorldExplorer, is also worth a sentence on its own: it's closed-loop and distribution-aware, meaning it uses the current model's own failure modes to decide where to collect more data next, rather than collecting blind -- a genuinely different approach from scraping video and hoping coverage works out, though the paper's evidence for how well that targeting works is qualitative too. ## The take ABot-World-0 is not the best model on WorldRoamBench. HappyOyster beats it on six of seven reported sub-metrics, and the benchmark's own leader has no disclosed size to compare against. What ABot-World-0 actually demonstrates is that a 5B model, with the right three-stage distillation and a genuinely load-bearing systems stack, beats two larger open rivals on every metric measured while being the only one of the group that runs interactively on one desktop GPU. That's a real, checkable claim, and it's a more interesting one than a clean sweep would have been -- a paper that only won everywhere would have less to say about where the wins actually come from. --- *Built on Alibaba AMAP CV Lab's [ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU](https://arxiv.org/abs/2607.19191) (Jiang et al., 2026) and the [amap-cvlab/ABot-World](https://github.com/amap-cvlab/ABot-World) repository (Apache-2.0). Figures 1, 3, and 10 are reproduced from the paper for commentary, flattened onto white; the systems-ablation and WorldRoamBench explorers are original visualizations of the paper's Table 2 and Table 3 data, not measured traces. Benchmark numbers are as reported in the paper and on WorldRoamBench.* --- # AngelSpec: specialize the drafter, share the verification budget > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/angelspec > date: 2026-08-03 > tags: inference, speculative-decoding, llm, systems, explainer Speculative decoding's basic trade is well covered on this site already: a cheap draft model proposes tokens, the target model verifies them all in one pass, and the target only ever emits what it would have sampled anyway. [EAGLE-3](/articles/eagle-3-speculative-decoding) fixed the draft model's own scaling law; [DSpark](/articles/deepseek-dspark) paired a semi-autoregressive drafter with a load-aware verifier. [AngelSpec](https://arxiv.org/abs/2607.25852) (Liu, Cen, Shi, et al., Tencent) is a different kind of paper: it is not one new trick, it is what a production team building on top of that whole lineage — [multi-token prediction](/articles/multi-token-prediction), DFlash, DFlare, DSpark, EAGLE-3's Training-Time Test — actually ships for [Hunyuan Hy3](/articles/hunyuan-hy3). Two ideas carry the paper. First: don't train one universal drafter, train a different architecture for each workload's entropy. Second: don't give every request a fixed verification budget, treat verification depth as a resource the whole batch shares. I am reading this as **engineering**, not a new algorithm. DFly extends DFlash's target-conditioning and DFlare's layer fusion; the Training-Time Test principle is EAGLE-3's. AngelSpec's own contribution is combining them into one training framework, specializing them per workload, and running the combination on real serving traffic. That last part is the rare thing here — most speculative-decoding papers stop at static benchmarks. ## One drafter doesn't fit both workloads Chat is high-entropy and open-ended; the next few tokens are genuinely hard to guess. Code and math are the opposite — once you're two lines into a `for` loop or three steps into an algebraic simplification, a lot of what comes next is close to deterministic. A single drafter trained on a uniform mixture of both has to compromise. AngelSpec's answer is to stop compromising: train an autoregressive multi-token-prediction (MTP) drafter on conversation-heavy data for chat, and a block-parallel diffusion drafter — **DFly** — strengthened with code and math data, for everything else. ## MTP: fix position 2 and 3, not position 1 The MTP drafter reuses one physical Transformer block recurrently at increasing logical depth, each depth predicting one more token ahead. The problem EAGLE-3 already diagnosed for feature prediction shows up again here for direct token prediction: depth 1 is trained against a clean prefix, but at inference depth 2 has to condition on depth 1's own (possibly wrong) output — a distribution it never saw in training. AngelSpec's fix is the same principle EAGLE-3 calls Training-Time Test: unroll the drafter over its own predictions during training, so what it practices on matches what it sees at inference. The loss stack that gets it there is a genuine progression, not a single choice: hard-label cross-entropy, then forward-KL, then an adaptively-blended KL/total-variation objective ("LK loss"), finally switching to an end-to-end objective that directly optimizes expected accepted length — $$ \mathcal{L}_{e2e} = 1 - \frac{1}{|I| D} \sum_{i \in I} \sum_{m=0}^{D-1} \prod_{k=0}^{m} \alpha_{i,k} $$ — a product, not a sum, because one rejection anywhere in the prefix invalidates every position after it. Training directly on raw total variation from a cold start is worse than plain KL (the TV gradient is too weak far from alignment); the paper's own ablation shows the cold-start-then-switch recipe is necessary, not decorative. The payoff shows up exactly where the theory predicts: the first drafted token barely moves, positions two and three — the ones with no defense against drift before TTT — improve the most. ## DFly: a hybrid backbone, then a cheap causal patch DFly starts from DFlash's move (feed every draft layer the same shared cross-layer-projected target context) and adds DFlare's move (a layer-specific weighted fusion of target features), combined rather than chosen between: $$ g^{(i)}_t = \text{RMSNorm}\big(c_t + f^{(i)}_t\big) $$ $c_t$ is DFlash's shared basis, $f^{(i)}_t$ is DFlare's depth-dependent refinement — the hybrid adds only $D \times T$ scalar fusion weights over DFlash alone, precomputable once training finishes.
Block-parallel diffusion drafts $B$ tokens in one shot, which is where the latency amortization comes from — but a one-shot draft has no mechanism to make token $t{+}2$ aware that token $t{+}1$ was just chosen. DFly's **hidden-correction head** patches that in afterward, cheaply: a small SwiGLU pass folds the previous position's embedding into each hidden state before the LM head runs, turning independent marginals into a causal chain $$ q\big(X_{t+1:t+B} \mid x_{\le t}\big) = \prod_{i} q_i\big(x_{t+i} \mid x_{\le t}, x_{t+1:t+i-1}\big) $$ while the expensive backbone stays fully parallel — only this small head runs sequentially. Tested against a Markov-style low-rank correction (DSpark's approach), hidden-correction wins on both accepted length and, notably, on the break-even latency it needs to beat MTP — the more accurate head is also the cheaper one to run. ## The numbers On Hy3-A21B, cumulative ablation (backbone → AR head → domain data) takes mean accepted length from 3.77 to 4.75; against the other drafters on the same target: That's +59.7% over MTP and +29.8% over DFlash on this target — the paper is upfront that DSpark isn't in this row (it's only benchmarked against Hy3 as MTP/DFlash/DFly; DSpark's own comparison runs on Qwen3-8B, where it still wins MT-Bench, consistent with AngelSpec's own framing that DFly targets code and math, not chat). Production throughput on Hy3-295B-A21B, 8×TP, tells the concurrency story: DFly wins the average speedup at every tested concurrency, 4 through 64. The more interesting detail is what happens at the high end: at concurrency 64, DFlash's own speedup actually **drops** below MTP-3's (1.89× vs 2.08×), while DFly stays ahead at 2.11×. DFly isn't just faster — it's the one that degrades least gracefully into the regime where the GPU is already saturated with verification work. ## D-cut: verification depth is a shared resource, not a per-request setting Here's the fact that makes D-cut make sense: median target-model verification (`execute_model`) latency runs 19.77–64.16ms; drafting and sampling (`sample_tokens`) runs 0.89–4.49ms. Verification dominates decode-step cost by roughly an order of magnitude. So the thing worth optimizing at serving time isn't the drafter — it's how much of that expensive verification you spend, and where.
The mechanism is a genuine reallocation, not a threshold. Per request $i$, expected progress from keeping $n_i$ drafted positions is estimated from the drafter's own prefix-confidence product, $\hat A_i(n_i) = \sum_{k=0}^{n_i} s_{i,k}$. D-cut doesn't pick $n_i$ per request — it flattens every position across the **whole batch**, ranks by that same confidence score, and takes a global top-K: $$ K_\rho(B) = \max\big(B,\ \lceil \rho\, B (D{+}1) \rceil\big) $$ restricted to four ratios, $\rho \in \{0.25, 0.5, 0.75, 1.0\}$, chosen each step by a pre-profiled runtime latency table that picks whichever $\rho$ maximizes projected throughput, not just kept length. It only ever discards drafts — verification stays exact, so the target distribution is untouched. ## Live traffic: the validation that actually matters Static benchmarks are where DFly's story ends for most papers in this space. AngelSpec adds one more figure, replaying real Hunyuan production traffic on 8×H20 at concurrency 2 through 64 — and this is the evidence that made me want to write the piece up.
DFly alone saturates past concurrency 48 (~848–860 tok/s, flat). D-cut keeps rising: +3.0% at c48, +9.2% at c56, **+15.7% at c64**. At matched per-user decode speed (~15.3 tok/s), D-cut sustains 981 tok/s at c64 versus DFly's 858 tok/s at c56 — 14% more aggregate throughput at the same latency. Against plain autoregressive decoding, DFly's own speedup peaks at 1.33× (c24–c40) and then falls back to 1.25× at c64; D-cut keeps climbing to 1.45× at c56 and 1.44× at c64. All of that for a pruning cost of just **1.5%** average reduction in accepted length (2.50 → 2.46), rising to only **2.8%** even at the most contended concurrency tested (2.50 → 2.43). Read the caption on the paper's own Figure 4 carefully — it says the comparison **understates D-cut**: "DFly uses full-and-piecewise CUDA graph capture and D-cut piecewise capture only." D-cut is running with a documented implementation disadvantage relative to DFly and still wins. That is an unusually candid thing for a paper to put in its own headline figure's caption. ## What's honest here, and what isn't new Two disclosures matter more than most papers in this space bother to make. All production numbers — throughput Tables 7–8, the live-traffic Figure 4 — run on **NVIDIA H20**, the export-compliant part, not a flagship H100 or B200. The paper doesn't claim these numbers generalize to other accelerators; it just tells you what it actually ran on. And the "this comparison understates D-cut" line above is the second: a paper flagging that its own reported advantage is a conservative lower bound is rarer than papers that quietly let an asymmetric comparison flatter them. What isn't new: DFly is DFlash plus DFlare plus a hidden-correction head borrowed from TreeFlash's idea; MTP's training recipe leans on EAGLE-3's Training-Time Test and an external LK-loss paper; D-cut's "verification as a shared resource" framing groups itself explicitly with DSpark as the other method doing this, rather than claiming to invent the idea. None of that is a knock — production systems earn their keep by combining existing pieces well, not by mandating a new algorithm — but it means the honest read of AngelSpec is "well-executed systems integration with real production validation," not "a new speculative-decoding algorithm." A few more gaps worth knowing before you cite this: DSpark is compared on Qwen3-8B, not on the paper's own Hy3-A21B target — so the strongest same-class competitor is missing from the main Hy3 table. The released DFly is mode-specific (Table 6: no-think and high-think drafters don't transfer across each other), which roughly doubles the drafters you maintain for a model family serving both modes. And every reported baseline number — including DFlash and DSpark's — was retrained and measured by the authors inside their own stack; there's no independent third party re-running any of it. ## The take The interesting move in AngelSpec isn't a new speculative-decoding trick — it's refusing to ship one universal answer. Chat and code/math have different entropy profiles, so they get different drafters. Verification cost and drafter confidence vary across requests and load, so verification depth becomes a batch-level resource instead of a fixed setting. Neither idea is exotic on its own; what makes the paper worth reading is that both survive contact with real Hunyuan traffic on honestly-disclosed hardware, with the one place it could have inflated its own result — the CUDA-graph asymmetry — disclosed instead of hidden. --- *Source: [AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding](https://arxiv.org/abs/2607.25852) (Hong Liu, Rui Cen, Junhan Shi, Guangshuo Qin, Jiebin Zhang, Tianyu Liu, Runzhi Fan, Guoliang Zhao, Ruobing Xie, Kai Zhang, Song Liu, Guanghua Yu, Jianchen Zhu — Tencent), arXiv:2607.25852. Figures 2, 3, and 4 are reproduced from the paper for commentary; the interactives are mine, built on the paper's own reported numbers and formulas.* --- # AutoCompact: teaching an agent to decide when to forget > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/autocompact > date: 2026-08-03 > tags: agents, llm, context-management, explainer Get the caveat out of the way first, because it changes how to read everything below: [AutoCompact](https://autocompact.github.io/) has **no paper, no arXiv listing, and no released code**. I checked the arXiv API by title and by author — zero hits. The project page is the only artifact, and the two charts that carry its headline claims (Figures 4–6) are unlabeled-axis line plots of "pass rate vs. inference-cost budget," not a results table. The one hard number in the whole post is "+10.6% RL gain on average on SWE-bench Verified," stated in prose, not shown as a table row you could check against a baseline. So this is a research blog post, and I'm treating it as one. I still think the idea is worth explaining, because the mechanism — training a model to make a decision it doesn't naturally make, using an LLM judge to correct its trajectory rather than hand-labeling it — is a genuinely interesting instance of a pattern that shows up across post-training right now. Just don't read the rest of this as a verified result. ## The problem: compaction is usually a clock, not a decision Long-horizon coding agents run out of context. The standard fix — what the post says OpenAI uses for ChatGPT and Codex — is a **fixed-threshold** compaction: once the trajectory crosses some token count, everything before it gets replaced by a summary. It doesn't matter whether the agent is mid-hypothesis or about to write the fix; the clock doesn't know the difference. AutoCompact's bet is that *when* to compact is a decision the agent should make about its own task state, not a number a harness enforces on it. The model gets a `compact()` tool call it can invoke at any point. When it does, the trajectory so far gets replaced by a generated summary — objective, the localized issue, files touched, what's verified, what's next — while the original task description and the most recent turns stay verbatim. Everything else (dead-end hypotheses, reverted edits, redundant file reads) gets dropped. ## Training it: a judge corrects the trajectory, not the label The harder problem is that models don't call `compact()` well without training — they don't reliably notice when a phase has ended. AutoCompact's answer is judge-guided correction rather than hand-authored demonstrations. At each step of a rollout, a judge (GPT-5.5-Codex) sees only the history visible so far, plus the fact that `compact()` exists, and makes one of three calls: leave the model's proposed action alone, **replace** it with a `compact()` call if this is a good moment to summarize, or **repair** a summary/continuation that's missing state or drifting off-track. That turns "teach the model to self-manage its own context" into "step-level correction under an annotation protocol" — something you can run at scale with a well-prompted LLM instead of a small army of human raters. From 379 SWE-rebench tasks, after filtering malformed and off-track examples, this produces **1,052 SFT examples** — a genuinely small cold-start set, split roughly 24% teaching *when* to trigger, 53% teaching *what to preserve*, 23% teaching *how to continue* after compaction. From that SFT checkpoint, GRPO reinforcement learning on SWE-Gym with a binary pass/fail reward pushes further: active compaction rate rises from **44.3%** of tasks (SFT) to **58.5%** (SFT+RL) — the RL stage doesn't just improve quality, it makes the model reach for `compact()` more often, because doing so is apparently what correlates with solving the task. Two more self-reported quality numbers from the post: generated summaries retain relevant state **99.8%** of the time and specify a concrete next action **97.8%** of the time. None of the pass-rate comparisons on the project page come with an exact number. All three head-to-head evaluations — no forced compaction vs. adaptive compaction, SFT vs. SFT+RL, and a forced 16k-token regime testing whether adaptive timing still beats a fixed threshold at the *same* limit — are qualitative line charts ("AutoCompact solves more tasks at every budget shown"), not tables. The "+10.6%" figure is the only exact number stated for RL over SFT, and even that is a TL;DR-level claim without a breakdown by task or budget. This is also a context-management idea, which puts it in conversation with [how a harness manages context more generally](/articles/agent-harness) — Lilian Weng's point that durable state belongs on disk, not in an ever-growing prompt. AutoCompact is one layer higher: it's not asking *where* state should live, it's asking *when* the model itself should decide to shed it. Both are betting that context is the scarce resource and the harness (or the model, here) needs an explicit policy for spending it. ## The take The training method — judge-corrected trajectories teaching a behavior the base model doesn't do on its own — is a pattern worth knowing regardless of whether AutoCompact's specific numbers hold up: it's a cheap way to get supervision for a decision (when to compact, when to stop, when to ask) that's hard to hand-label at scale but easy for a strong model to critique step by step. What I can't tell you is how good the resulting agent actually is, because there's nothing outside one team's own charts to check it against. If code or a paper ships later, the interesting question is whether the qualitative "wins at every budget" story survives being reduced to a table. --- *Source: the [AutoCompact project page](https://autocompact.github.io/) (Xuan Zhang, Longtao Zheng, Cunxiao Du, Bo An, Xin Dong; July 30, 2026). No code or paper release exists at time of writing; all figures on the source page are JS-rendered widgets, not downloadable images — the timeline above is my own illustration of the mechanism, using an invented trajectory and thresholds, not a reproduction of anything on the page.* --- # A.X-K2: a sparse-attention upgrade that costs nothing, trained natively in FP8 > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/ax-k2 > date: 2026-08-03 > tags: llm, mixture-of-experts, attention, sparse-attention, quantization, fp8, long-context SK Telecom's **A.X-K2** is a 688B-parameter, 33B-active Mixture-of-Experts model, trained from scratch as part of Korea's Sovereign AI foundation-model project and shipped under Apache 2.0. It is the successor to A.X-K1, and the tech report ([SKT, 2026](https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf)) is unusually specific about what changed: a full architecture table, a full FP8 training recipe, full RL hyperparameters, and — the detail worth building an article around — an ablation that reports the *quality cost* of its own efficiency trick, and the cost is close to zero. Two things carry this piece. First, **Sparse Gated Attention (SGA)**: a lightweight top-k indexer bolted onto gated Multi-head Latent Attention, where SKT states the LongBench score before and after adding sparsity — 62.80 to 62.99 — rather than only the speedup. Second, A.X-K2 is **trained natively in FP8**, forward and backward, from the first optimizer step. Not quantized after the fact. Those two facts turn out to be connected: the same architectural choices that make SGA cheap (GatedNorm suppressing outlier activations) are the choices that make FP8 training survivable at this scale. ## What the weights say Total tokens: about 8.5T (8.2T pre-training, the remainder post-training) — fewer than A.X-K1's roughly 10T, and SKT is direct about why that's the headline, not the parameter count: *"despite training on fewer tokens than A.X K1 (∼10T), A.X K2 shows substantial improvements across the board — over 30 percentage points on some benchmarks — reflecting substantial gains in token efficiency."* The architecture, read straight off `config.json` and Table 1 of the report: | | | |---|---| | Total / active parameters | **688B / 33B** | | Layers | **61** (1 dense, 60 MoE) | | Hidden size | 7,168 · 64 attention heads (Q = KV) | | Routed / shared experts | **256 / 1**, 8 active + 1 shared per token | | Expert routing | sigmoid score, `noaux_tc`, 8 groups, group top-k 4, routing scale 2.5 | | Attention | MLA + head-specific output gate + QK-norm, plus a top-*k* sparse indexer (k = 2,048) | | KV-lora / Q-lora rank | 512 / 1,536 | | Context | 128K native (ABF), 256K via YaRN (factor 2.0), zero-shot to 512K (factor 4.0) | | Vocabulary | 163,840 (unchanged from A.X-K1), 5 languages | | Training precision | native **FP8** (MXFP8, E4M3, block 32) forward + backward | | Checkpoint | block-scaled FP8 (E4M3, 128×128), ~646 GB on disk | Relative to A.X-K1 (519B-A33B), the entire parameter growth is in expert count: 192 to 256, chosen because it's a power of two and a multiple of 128 for sharding, targeting a compute-derived total-parameter budget under a fixed 70-day, 512-GPU schedule. Active parameters didn't move. That's a scale-up in *capacity*, not in per-token compute — worth holding onto before the next section. ## Sparse Gated Attention Start from what A.X-K1 already had: Multi-head Latent Attention (MLA) with a **head-specific output gate** — a learned, input-dependent gate applied to the attention output before the `Wo` projection, present in every layer throughout pretraining, not something bolted on for long context. The report's framing for why the gate matters: it "introduces non-linearity into the attention output, mitigates attention sinks, and improves loss convergence." Attention sinks are the few uninformative token positions (often the first token) that vanilla softmax attention dumps disproportionate probability mass onto — a gate that suppresses that mass gives every downstream consumer of the attention output a cleaner signal. **SGA** is the second half of the name: a lightweight indexer, adopted from DeepSeek-AI's sparse-attention design, that scores every cached key and keeps only the **top 2,048 tokens per query** — selected at individual-token granularity, not in fixed blocks. (That's a genuine design fork from [MiniMax Sparse Attention](/articles/minimax-sparse-attention), which scores and selects in 128-token blocks specifically so memory access stays contiguous. A.X-K2 trades that contiguity for finer-grained selection.) MLA then runs exactly over the selected set, `KV[I_topk]`, instead of the full cache. Here is the paper's own architecture figure for the resulting block:
Read it as a data path: GatedNorm output splits into the Indexer→Selector pair (which decides *what* MLA gets to read) and the Output Gate (which decides how much of what MLA computes gets through). The mechanism I built to walk through the dynamics — how the fixed 2,048-token budget shrinks as a fraction of a growing context, and what changes when the Selector is switched off entirely: The report calls the gate and the indexer **mutually reinforcing**, and the causal story runs one direction: because the output gate already suppresses attention-sink mass throughout pretraining, the attention distribution the indexer is trained to imitate is better-calibrated before the indexer ever sees it — so its top-*k* budget goes to genuinely relevant positions instead of partly being spent re-discovering which tokens are sinks. The indexer itself is trained with a KL-divergence loss against the (already gated) attention distribution, introduced in a dedicated Stage 3C after the model is natively trained to 128K context. The number that makes this worth an article: on LongBench, A.X-K2 scores **62.80 before** the sparse adaptation and **62.99 after** — sparsity made it very slightly *better*, not worse. A lab publishing the comparison that shows its own efficiency trick is nearly free — rather than only the speedup — is worth crediting on its own. And the adaptation recipe has a second, smaller honest claim attached: unlike DeepSeek-V3.2 and GLM-5, which warm the indexer up against the *dense* (full) attention distribution before switching on sparse selection, SKT trains the indexer against the **sparse** top-k selection from the outset — a "sparse warmup" — because it's cheaper (sparse attention is less compute per step than dense) and, in their experiments, cost no measurable downstream quality relative to the dense-warmup alternative used elsewhere. The efficiency payoff shows up exactly where you'd expect — long-context serving. Reading the report's own inference sweep at 120K input tokens (concurrency 32, dp8/ep8): total-token throughput goes from roughly 9,100 tok/s for A.X-K1 to roughly 12,200 for A.X-K2 with a bf16 KV cache, and roughly 14,600 with an FP8 KV cache — with per-token latency and time-to-first-token moving the same direction, each a few tens of percent lower for K2 than K1 at that length. (Approximate, read off the report's chart, not a published table — but the direction and rough magnitude are unambiguous.) One more piece worth naming here because it recurs in the next section: **GatedNorm** replaces A.X-K1's dual-normalization scheme entirely — a single input-dependent gate applied right after RMSNorm, instead of stacking normalization layers around attention and the MLP the way A.X-K1 (and Gemma-style designs) did. SKT ran the ablation at 20B-A3B scale and found GatedNorm alone matches the loss curve of the full dual-norm design; stacking a second post-MLP norm on top added nothing. The reason GatedNorm matters beyond training stability: it suppresses **massive activations** — the small number of hidden units that run orders of magnitude larger than the rest and persist across layers — which is exactly the failure mode that wrecks narrow low-precision formats. Which is where the second half of this piece starts. ## The scale, next to what else is disclosed A.X-K2 sits in the middle of this range by total parameters and at the small end by active parameters. [Kimi K3](/articles/kimi-k3) — 2.8T total, 104B active — took the opposite bet on the same axis: where A.X-K2 grew total capacity 519B → 688B while holding active compute flat at 33B (a pure expert-count expansion, 192 → 256), K3 tripled its active parameters alongside its total, spending its extra headroom on a bigger per-token forward pass rather than more parked capacity. Both are legitimate ways to spend a training budget; they're just different bets about where the marginal FLOP is worth spending. ## Trained natively in FP8 Everything above assumes a working low-precision model. A.X-K2 gets there by training natively in FP8 from the start rather than quantizing a full-precision model afterward — MXFP8, E4M3, block size 32, forward *and* backward pass, with FP32 master weights and BF16 optimizer state as the only higher-precision parts of the recipe. I made the same argument at a different precision two weeks ago in [Neutrino-1](/articles/neutrino-1): quantization is a decision you make before training starts, not a knob you turn on a finished checkpoint. Neutrino-1 showed the cliff that decision avoids — ternary weights rounded post-hoc land at 24.2–24.7 on 5-shot MMLU, against a 25.0 chance line, while the same ternary format trained in from scratch reaches 72.1. A.X-K2 is the same principle, replayed at FP8 instead of ternary, at 688B instead of 8B. The practical consequence: A.X-K2 has no BF16 form to compare itself against, because none was ever trained. Its FP8 checkpoint *is* the master weights, not a rounded-down copy of something else. Serving it in NVFP4 — a further post-hoc step, applied only to expert weights (W4A4) — is a much smaller step down than Neutrino-1's ternary rounding, because the base it's stepping down from was already trained natively in a narrow format: The report's own robustness table backs this up on eleven benchmarks: NVFP4 tracks FP8 within about a point on most of them — CLIcK 84.21 → 84.06, MMLU 82.27 → 82.00, KoBEST-BoolQ 96.72 → 96.01 — with GSM8K (−2.50) and MATH (−2.42) as the honest outliers, and HumanEval and KoBEST-COPA actually improving slightly. Compare that spread to Neutrino-1's cliff and the shape of the difference is the whole argument: rounding *into* a format a model never trained in collapses to chance; stepping *further down* from a format it was already native in costs a couple of points at most.
This same commitment shows up again, more sharply, inside RL post-training — and it's the cleanest evidence in the whole report that "native FP8" is an infrastructure discipline, not just a training-time flag. RL needs the trainer (Transformer Engine, on Blackwell) and the rollout engine (vLLM) to agree numerically. But Blackwell defaults to MXFP8 while vLLM's mature MoE FP8 path targets the older *blockwise* FP8 recipe built for Hopper — so if you leave each side on its native default, they diverge. Figure 7 above shows what that divergence does: a trainer running MXFP8 against a blockwise-FP8 rollout looks fine early, then the reward curve stalls and drifts down. SKT's own diagnosis is the sentence worth keeping: *"Applying TIS does not prevent this collapse, indicating that token-level intervention alone cannot remove the underlying trainer–rollout precision mismatch."* Truncated Importance Sampling is a standard token-level correction for exactly this kind of train/inference distribution drift, and it doesn't work here — the fix has to be architectural (a patched Transformer Engine branch that forces blockwise FP8 on Blackwell, matching vLLM's format end to end), not a loss-side patch. That's a small, honest, specific admission: a common trick from the RL toolbox failed, and they said so instead of quietly switching methods without comment. One more low-precision data point, smaller but concrete: on Rebellions' ATOM-Max NPU, A.X-K2 reports **107% performance-per-watt** relative to a comparable NVIDIA L40S GPU — a real deployment-hardware number, not a simulation. ## The benchmarks
A.X-K2 leads this five-model, five-benchmark slice outright, and the gap on **Apex** is the most striking: 45.8 against a next-best of 28.1 (DeepSeek-V4 Flash) — more than double the third-place score. Beyond what's in that chart, SKT reports two non-benchmark math results worth noting because they aren't self-scored evals: **35/42 on IMO 2025** (the gold-medal threshold is 35, with a perfect 7/7 on each of the first five problems), and correct proofs for all eight KMO26 second-round problems, using an iterative proof-refinement method borrowed from DeepSeekMath-V2's approach. Long-context quality holds up on RULER, staying above 92 out to 128K and only easing to 86.6 at the full 256K: | Context | 4K | 8K | 16K | 32K | 64K | 128K | 256K | Overall | |---|---:|---:|---:|---:|---:|---:|---:|---:| | RULER | 97.5 | 97.2 | 97.5 | 96.5 | 94.3 | 92.7 | 86.6 | **94.6** | Needle-in-a-haystack retrieval is a clean 100 at every position tested at both 256K (YaRN factor 2) and a zero-shot 512K (YaRN factor 4) — including after NVFP4 quantization, which is the same "the base format survives further compression" story as the precision ladder above, applied to retrieval instead of MMLU. ## Where it's honest about losing Every number above is self-reported by SK Telecom, on their own harness, with no independent reproduction I could find. The eval protocol is disciplined — all baseline open-weight models run through OpenRouter fixed to the model's own publisher, `xhigh` reasoning effort, pass@1 averaged over multiple generations (8 for math) rather than single-shot — but it's still one lab grading a comparison it designed. The clearest weak spot, and SKT names the reason itself: 9.3 is worst of all seven models with a reported score — GLM-5.1 leads at 29.1, and even the next-worst (Nemotron 3 Ultra at 13.4) beats A.X-K2 by 44%. The model card's own explanation: *"Agentic performance is moderate — A.X K2 trails the strongest compared models on BrowseComp — reflecting limited agentic RL during post-training."* That's the right way to publish a weak number — attribute it to a specific, checkable cause (the RL data mixture allocates only 18% of SFT tokens and a modest RL slice to agentic tool use, against much heavier agentic investment in a model like [Kimi K3](/articles/kimi-k3)) rather than burying it. One methodology caveat the report itself surfaces: A.X-K2's only tool on this benchmark was Brave Search's API, capped at ≤10 searches per problem — a real constraint, not full open browsing — and the report doesn't state whether every compared model ran under the same cap. That could shift the absolute number somewhat; it's very unlikely to explain a 3× gap to the next-worst model. BrowseComp isn't the only place A.X-K2 comes second. On **GPQA Diamond** it's mid-pack, and the surprise is which model beats it on a Korean-language benchmark: DeepSeek-V4 Flash — not a Korean-focused lab — beats SK Telecom's own Korean-sovereign model on a Korean benchmark, by 2.3 points. A.X-K2 still wins the *other* two Korean benchmarks in the comparison (KMMLU-Pro, CLIcK), so this is one loss inside a category it otherwise leads, not a category-wide miss — but it's exactly the kind of specific, checkable number a self-reported table should surface rather than smooth over. Rounding out the mid-pack results: LiveCodeBench v6 at 84.0 (DeepSeek-V4 Flash leads at 89.4), SciCode at 41.0 (near the bottom of the field; Kimi-K2.6 leads at 53.5), and IFBench at 75.9 (DeepSeek-V4 Flash leads at 81.2). None of these are collapses — they're a model that wins decisively on math and most of Korean, and trails on strict-instruction-following, code-execution benchmarks, and — sharply — on open-web agentic search. Two more limitations the model card states plainly, worth repeating because they're easy to omit: A.X-K2 is text-only (no native multimodality, listed as future work), and SKT explicitly did not run a dedicated quantitative bias or fairness evaluation. ## The take Two disclosures make A.X-K2 worth writing about on their own, independent of where it lands on any single leaderboard. It's one of the only sparse-attention releases I've seen that reports the ablation showing its sparsity is nearly free (62.80 → 62.99 on LongBench) instead of only the speedup — crediting the reader with the question "what did this cost?" instead of hoping nobody asks. And it's trained natively in FP8 end to end, with the RL infrastructure section going out of its way to show a standard fix (TIS) failing against a real precision mismatch rather than quietly working around it off-page. Set against [Neutrino-1](/articles/neutrino-1)'s ternary cliff and [MiniMax Sparse Attention](/articles/minimax-sparse-attention)'s block-granularity bet, A.X-K2 reads as the same 2026 pattern — quantization and sparsity are training-time commitments now, not deployment-time knobs — applied at a scale and with a level of self-disclosure that makes the whole argument checkable, including the parts (BrowseComp, KoBALT) where the honest answer is that it lost. --- *Sources: the [A.X K2 Technical Report](https://github.com/SKT-AI/A.X-K2/blob/main/A_X_K2_Tech_Report.pdf) (SK Telecom, dated 2026-07-28 — architecture, training recipe, RL infrastructure, evaluation tables) and the [model card and config](https://huggingface.co/skt/A.X-K2). Figures 1 and 3 here are the report's Figures 2 and 7, reproduced for commentary; the benchmark comparison figure is the report's Figure 1. All benchmark numbers are SK Telecom's own, on their own harness, with no independent reproduction found. The inference-efficiency numbers in the Sparse Gated Attention section are approximate, read off the report's chart rather than a published table. Interactive diagrams are mine. Related: [Neutrino-1](/articles/neutrino-1) on training-native vs. post-hoc quantization, [Kimi K3](/articles/kimi-k3) on the other end of the MoE sparsity-ratio spectrum, and [MiniMax Sparse Attention](/articles/minimax-sparse-attention) on block- vs. token-granularity top-k selection.* --- # Chimera: unbundling RoPE into a diffusion Transformer that extrapolates 6× on video > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/chimera-diffusion > date: 2026-08-03 > tags: diffusion, linear-attention, video-generation, scaling-laws, positional-encoding, explainer Visual generation is hitting the same wall language models hit a few years back: the tokens keep multiplying. A high-resolution image is thousands of tokens, a video clip is tens of thousands, and once you want text, image, and video sharing one context, full attention's quadratic cost stops being a rounding error and starts being the budget. Language models solved their version of this with linear and hybrid attention. The catch, as [Chimera](https://arxiv.org/abs/2607.28611) — Adobe Research's new hybrid visual diffusion Transformer — points out, is that those solutions don't transfer directly: a diffusion backbone has to preserve spatiotemporal locality and support genuinely bidirectional interaction across modalities, neither of which a causal language model has to worry about. Chimera's answer is a single-stream backbone that processes text, image, and video tokens together, mixing them with **Kimi Delta Attention (KDA)** for cheap $O(N)$ state tracking, periodic **Multi-head Latent Attention (MLA)** for exact global interaction, and **modality-aware short convolutions** for local structure — with no positional embeddings anywhere in the stack. The paper backs this with **HeteroP**, a hyperparameter-transfer scheme built for a backbone that is not one uniform shape, and fits genuine Chinchilla-style scaling laws on top of it. The headline results: a real zero-shot **6× video-length extrapolation** (5-second training clips generalizing to 30 seconds with only 6.5% FID degradation, versus 50%+ for two full-attention baselines), and a compute-efficiency claim over a matched Wan2.1 baseline that the abstract prints as **7.3×** but the paper's own arithmetic, two pages later, computes as **6.8×**. Both numbers are worth seeing, and the mechanism behind the extrapolation result is the more interesting story.
## One stream, three mechanisms, no positions Text tokens (from a frozen T5-style encoder) and visual tokens (from a frozen Wan2.1 VAE, patchified) are concatenated into a single sequence and pushed through the same stack. Visual tokens are flattened in **temporal-major raster order** — position $(i,j,k)$ in a $(T, H, W)$ grid maps to sequence index $m = iHW + jW + k$, with a single image just a one-frame video — so there is one token order for every modality, not a per-modality scheme bolted on afterward. Inside a block, the attention sublayer is either KDA or MLA on a fixed **3:1 KDA-to-MLA schedule**: three linear layers, then one global layer, repeating. The first and last blocks use a dense SwiGLU feed-forward; every other block routes through **sparse MoE** — 56 experts, top-8 active, no shared experts, balanced by an auxiliary-loss-free bias added only at top-K selection (not to the mixture weights). With that bias, batch-level `MaxVio` — the metric tracking how far the busiest expert's load sits above the average — settles near 0.5; strip the bias out and it blows past 5, close to the theoretical collapse bound of 6. Every sublayer — attention and FFN alike — is wrapped in **Identity Hyper-Connections (iHC)**: the residual stream is duplicated into $M{=}4$ parallel copies with token-dependent read/write gates, a simplified version of hyper-connections that fixes the residual-mixing matrix to the identity instead of learning a doubly-stochastic one via Sinkhorn iterations — cheaper, at the cost of losing learned cross-stream mixing. None of KDA, MLA, or NoPE is new by itself — they're the same three ideas [Kimi K3](/articles/kimi-k3) uses to run a 1M-token language model, adopted here essentially as-is and pointed at diffusion instead of next-token prediction. **KDA** keeps a fixed-size recurrent state per head instead of a growing KV cache: $$ S_t = \big(I - \beta_t k_t k_t^{\top}\big)\,\mathrm{Diag}(\alpha_t)\,S_{t-1} + \beta_t k_t v_t^{\top} $$ with $\alpha_t \in (0,1)^{d_k}$ a **per-channel** forget gate and $\beta_t$ a scalar write strength — the exact recurrence [KDA has a half-life](/articles/kda-half-life) walks through: each channel forgets a fixed fraction of its state per step, so it has a half-life $n_{1/2} = \ln(0.5)/\ln(\alpha)$ measured in tokens. That article works the math out for language tokens; Chimera is the same forget-gate law now doing memory management for pixels and frames instead of words. **MLA** restores exact bidirectional interaction on top, compressing keys and values through a low-rank projection, with one twist: the direct key path is left **unrotated** — no RoPE at all, on either mechanism. That's the part worth slowing down on. ## What RoPE was actually doing Every visual diffusion Transformer before this one, and most language models, lean on RoPE to inject position. Chimera's authors run a mechanistic audit of what that rotation is actually buying, on Qwen3-4B first and then on the visual diffusion models FLUX.2 and Wan2.2, and decompose each attention logit into its per-frequency-pair cosine contributions (the pieces sum to the real logit to within $3 \times 10^{-13}$, so this isn't an approximation of the mechanism — it's an exact accounting of it). The clearest case is layer 0, head 1 in Qwen3-4B, a strong "previous-token" head: 98.5% of its queries attend most strongly to the immediately preceding token. Trace *why*, and it's one channel pair — the fastest-rotating one, turning roughly one radian per token — whose cosine happens to peak exactly at offset 1. The slow, high-amplitude pairs sitting alongside it contribute a flat, content-driven background that doesn't move the peak at all. Search across every head for the best "attend exactly $n$ tokens back" detector and the same pattern holds: the strongest head hits a 0.56 fraction at $n{=}2$, but only 0.07 at $n{=}10$ — RoPE-based position selection is a short-range tool, not a general one, in both language and visual models alike. From that audit the paper pulls out three things RoPE is doing at once, inside the same rotated channels: 1. **Position selection** — a canonical previous-token/induction-head trick, and (per the numbers above) reliable only at short offsets. 2. **Recency decay** — an *implicit* average property of the summed rotations, not an explicit mechanism, and one a trained head can learn to bypass. 3. **Layout encoding** — the channel partition across positional axes (time, height, width, text index) is set by hand at design time; every additional axis divides the available channels further and never adapts. Chimera's move is to stop asking one set of rotated channels to do all three jobs and give each its own dedicated module instead: Token order for KDA comes for free from the recurrence itself — a scan is inherently order-aware, no rotation required. Position selection moves to the **modality-aware short convolution**: a depthwise kernel mixing tokens at explicit index offsets, which is a cheap, parameter-light way to do exactly the short-range job the audit found RoPE was actually good at. Recency decay moves to **KDA's own forget gate** $\alpha_t$ — explicit and content-adaptive, rather than an emergent side effect. And layout encoding falls out of the convolution's native shape: **causal 1D** along the token index for text, **causal-in-time 3D** over the $(T,H,W)$ grid for video — each modality's structure is encoded by the operator itself, with no channels spent partitioning anything. Text and visual tokens each get their own kernel here, implemented as one fused Triton pass that measures 2.2–2.3× faster forward, 1.5–1.8× faster forward-plus-backward, and up to 4× less peak activation memory than the naive gather-convolve-scatter version. With all three biases reassigned, MLA is left to do only content matching — the paper's framing is that MLA without RoPE is the limiting case where every channel has zero rotary frequency, so its logits depend purely on content. KDA's queries and keys carry no positional phase either. Nothing in the stack is tied to how long the training sequences were, which is the actual, mechanistic reason extrapolation works — not a property tacked on after the fact, but the direct consequence of where each inductive bias now lives. It's the same bet [Kimi K3](/articles/kimi-k3) makes for a 1M-token language model: no RoPE means nothing to rescale when the context grows past training length. Chimera is that bet, replayed in a diffusion Transformer over image and video tokens instead of a causal LM over text. ## HeteroP: a scaling ratio per tensor, not per model Fitting a Chinchilla-style law needs a family of models at different sizes, each trained with hyperparameters as good as they'd be at full scale — otherwise you're not comparing model sizes, you're comparing tuning quality. The standard fix is µP-style hyperparameter transfer: tune a small proxy model, then derive the large model's learning rate, init, and weight decay from a single width ratio. Chimera's backbone breaks that assumption, because it isn't one uniform width. Widening the model changes the KDA head width, the MLA compression rank, the MoE expert width, the router width, and the timestep-conditioning MLP width all differently — a single global ratio, tuned for the backbone, is the wrong ratio for the rest. **HeteroP**'s fix is to stop pretending there's one ratio. For each parameter group $W$, it computes its own width ratio from that group's own **functional fan-in**, plus one shared depth ratio from the block-count ratio: $$ (m_W, m_L) = \left(\frac{\mathrm{fan\text{-}in}(W)}{\mathrm{fan\text{-}in}(W^{(0)})},\ \frac{n_{blk}}{n_{blk}^{(0)}}\right) $$ Concretely: hidden weights get init variance and learning rate scaled by $m_W^{-1}$ and weight decay scaled by $m_W$ (keeping the LR-times-decay product invariant); attention and FFN residual branches get an additional $m_L^{-1}$ output scaling, a depth correction borrowed from CompleteP; input adapters, norms, and the readout keep their base LR and init (standard µP convention), with the readout's forward pass separately rescaled by $m_W^{-1}$. The proxy model is width 512, depth 4 (22M activated, 59M total parameters); the largest is width 2048, depth 32. The validation is direct: under HeteroP, the optimal base learning rate sits at about $10^{-3}$ across a 56× range in activated parameters (20M to 1.12B) and an 8× range in depth (4 to 32 layers). Under standard parameterization — one global ratio for everything — the optimum drifts sixfold, from $10^{-4}$ to $6\times 10^{-4}$, and several of the high-learning-rate runs at large scale diverge outright. That drift isn't just an inconvenience for the scaling-law fit, it actively biases it: trained without HeteroP, the same image model family gives a fitted exponent of $N_{opt}\propto C^{0.588}$ (envelope) or $C^{0.581}$ (isoFLOP) — inflated by 0.08–0.10 over HeteroP's 0.505/0.481 — which the paper's own extrapolation shows prescribes a compute-optimal model roughly **2× oversized (and correspondingly undertrained)** three orders of magnitude of compute out from where it was fit. Get the transfer wrong, and the law tells you to build the wrong-shaped model at scale. ## What the scaling law says With HeteroP holding hyperparameter quality constant across scale, the paper fits $\hat L(N,D) = E + AN^{-a} + BD^{-b}$ — activated parameters $N$, visual-latent-token count $D$ — using three independent estimators (a training-loss envelope, an isoFLOP profile, and a direct parametric fit) for image and video pretraining separately. All three estimators agree, and they disagree with each other by modality: | Modality | $N_{opt}$ exponent (across 3 estimators) | Split | |---|---|---| | Image (256²) | 0.48–0.52 | Nearly balanced between model size and data | | Video (180p) | 0.53–0.56 | Modestly favors model size at higher budgets | The parametric fits — $\hat L_{image} = 0.126 + 5.28N^{-0.315} + 33.8D^{-0.336}$ ($R^2{=}0.993$) and $\hat L_{video} = 0.124 + 8.07N^{-0.330} + 145.2D^{-0.394}$ — land on nearly identical irreducible-loss terms (0.126 vs 0.124), consistent with both modalities sharing the same denoiser and VAE. The paper adds an axis prior scaling-law work doesn't have: the compute-optimal **image-to-video data ratio**. It drifts from roughly 4:1 to 3:1 as compute grows from $10^{18}$ to $10^{19}$ FLOPs (image gets relatively cheaper to learn from per token as budget grows), while the video-loss-optimal ratio stays pinned at 1:1 — the video-heaviest mixture the authors actually tested, so that half of the result is a boundary effect, not a discovered optimum, and the paper is upfront that it didn't search past it. ## The number that doesn't quite add up Guided by those laws, the paper trains an 11B-total / 2B-activated Chimera and compares it against matched 2B full-attention baselines — Wan2.1 and Z-Image — all four models trained in-house to the same $5\times10^{20}$-FLOP budget on identical data. At a shared training loss of 0.149, Wan2.1 needs $4.29\times10^{20}$ FLOPs and Z-Image needs $3.75\times10^{20}$; Chimera-dense (no MoE, no iHC, no HeteroP) reaches it in $2.55\times10^{20}$ — a clean **1.7×**. The complete configuration (MoE + iHC + HeteroP) reaches it in $6.27\times10^{19}$ FLOPs.
Do the division on the two numbers printed a page earlier and you get $4.29\times10^{20} / 6.27\times10^{19} = 6.84$ — which is exactly what Section 5.5's own sentence says: "a 6.8× compute-efficiency gain over Wan." The abstract, the introduction, and the plotted label in Figure 12 above all instead say **7.3×**, from the same pair of FLOPs figures. I'm not accusing anyone of anything here — I re-fetched the paper's own HTML and confirmed both numbers appear verbatim: "6.8 × compute-efficiency gain over Wan" in the Section 5.5 prose, right next to the $4.29\times10^{20}$ and $6.27\times10^{19}$ FLOPs figures it's computed from, and "7.3 ×" in the abstract, the introduction, and baked into Figure 12's own plotted label. $4.29 / 0.627 = 6.84$, not $7.3$, using the numbers exactly as printed. It's possible the true, unrounded internal FLOPs values reconcile to 7.3× and the 3-significant-figure numbers printed in the text are what drifted — the paper doesn't say either way, and I found no footnote reconciling the two. What I can say: **using the checkable numbers, the arithmetic supports 6.8×, not 7.3×.** If you're going to cite Chimera's headline efficiency gain, cite 6.8×, or go verify the underlying FLOPs yourself. Worth separating, too: the component ablation (dense → +MoE → +iHC → +HeteroP) reports a **4.1×** cumulative gain at loss 0.149 — MoE alone gets to 1.5×, +iHC to 1.7×, +HeteroP the rest of the way. That 4.1× is measured *relative to Chimera-dense*, not to Wan2.1, so it isn't a third candidate for the headline number — it's answering a different question (how much of the complete system's win comes from which piece), and it's consistent with either the 6.8× or 7.3× reading of the Wan-relative number, since $1.7 \times 4.1 \approx 7.0$, which lands between the two and settles nothing on its own. ## Zero-shot length extrapolation: NoPE's actual payoff This is where the RoPE audit cashes out. Chimera is trained only on 5-second, 81-frame clips, then asked — with **no length-specific fine-tuning at all** — to generate 30 seconds, 6× its training length. Every metric below is computed only on the final 5 seconds of each generated clip, isolating the extrapolated region, over 512 generated vs. 512 reference videos at matched prompts, seeds, resolution, and fps:
The numbers: Chimera's FID goes from 77.1 at 5 seconds to 82.1 at 30 — a **6.5%** degradation. FVD moves from 685.8 to 829.5, up 20.9%. Wan2.1-T2V-1.3B's FID degrades **50.5%** over the same stretch, HunyuanVideo-1.5's **53.6%** — both well past the point where a video model's later seconds are visibly falling apart. Chimera also posts the lowest *absolute* FID and FVD of the three at 30 seconds, not merely the smallest percentage move — it isn't winning by having started worse and degrading less, it's ahead the whole way. This is the sibling result to [SANA-Video 2.0](/articles/sana-video2), the other linear-attention video approach covered here, and the two make an interesting contrast. SANA-Video 2.0 keeps the same 3:1 linear-to-global attention idea and Block Attention Residuals for cross-depth flow, but keeps RoPE and optimizes for raw single-GPU latency at a fixed, modest clip length. Chimera keeps RoPE out entirely and stakes the design on exactly the axis SANA-Video 2.0 doesn't test: generalizing far past the lengths it was trained on. Different bets, same underlying conviction that softmax attention over every token pair was never the part of video generation worth paying full price for. ## What $O(N)$ buys in memory and latency A softmax KV cache costs $O(N \cdot H \cdot d_h)$ — grows with sequence length. KDA's recurrent state costs $O(H \cdot d_h^2)$ — fixed, independent of $N$. Measured directly: a matched KDA/MLA and MHA/MLA backbone, both around 2B activated parameters at the same 3:1 ratio, batch size 1, BF16, 512 text tokens plus 18×28 visual tokens per frame, on one NVIDIA A100-SXM4-80GB — — the linear variant supports 1.68× longer sequences before it runs out of memory, and runs 2.14× faster at the 255k-token point both backbones can reach. The paper is careful about what this comparison actually shows: FlashAttention removes the quadratic attention *workspace*, but not the quadratic *arithmetic* — so this isn't a straw-man comparison against an un-optimized baseline, it's the honest gap that remains after the standard fix. ## Benchmarks, and what 600 H100-days buys Trained for only about 600 H100-days, Chimera is competitive on text-to-image quality with models that cost far more to build: On GenEval, Chimera lands at 0.82 overall — tied with Z-Image-Turbo, matching FLUX.1-dev, beaten only by Seedream 3.0's 0.84 — and it beats both FLUX.1-dev and Z-Image-Turbo on DPG-Bench specifically. The paper also quotes Z-Image-Turbo's *own* reported training budget, about 12.4K H100-days, as roughly 20× Chimera's — worth reading as a cross-lab comparison rather than a controlled one: different codebases, different clusters, and a number each lab measured on its own infrastructure, not a shared benchmark. The GenEval/DPG-Bench baseline rows themselves are the field's normal practice, too — each competitor's own published number, not re-run under Chimera's exact sampling protocol. Worth knowing before quoting either comparison as settled. ## Honest limits The paper is unusually direct about what it hasn't shown yet: - **MoE underperforms its LM-scaling expectation.** Sparsity buys only about **1.5×** compute efficiency here, well short of the commonly cited $\sqrt{\text{sparsity}}$ heuristic — the authors attribute this to weak expert specialization (routing stays close to uniform across tokens and timesteps) and call it out as an architectural ceiling, not a training bug. A negative result reported plainly rather than smoothed over. - **Muon underperforms AdamW throughout** their tests. They hypothesize Adam's implicit low-rank bias matters for diffusion training specifically, but say so as a hypothesis, backed by a brief spectral-analysis follow-up, not a systematic sweep. - Several structural ratios are **held fixed across the entire scaling study** and never made scale-dependent: iHC's stream count, MoE's expert count and top-K, and MLA's KV-compression ratio. The paper flags directly that whether the optimal compression ratio is itself scale-dependent is left to future work. - Only **text-to-image and text-to-video generation** are evaluated — no multimodal *understanding* task, despite the single-stream design being a natural fit for one (the paper name-drops this as a target direction, not a result). - The timestep-conditioning MLP's width is scale-sensitive enough to destabilize training if mismatched to the backbone — a 4096-dim MLP paired with a 1024-wide backbone went unstable — and it's patched ad hoc rather than covered by the main HeteroP table. Set against that, the scaling-law and compute-efficiency measurements themselves are unusually rigorous for the genre: Wan2.1 and Z-Image aren't cited from their own papers here, they're re-implemented and trained in-house at matched 2B scale on identical data, specifically so the efficiency comparison isn't citing someone else's number under someone else's conditions. ## The take The RoPE audit is the part of this paper worth remembering past the benchmark tables. It's not "we removed positional embeddings and it worked" — it's a demonstration, with an exact per-frequency accounting, that RoPE was quietly doing three separate jobs through the same rotated channels, that one of those jobs (position selection) only works at short range anyway, and that giving each job its own dedicated, non-attention mechanism is what actually buys extrapolation — not a side effect of going linear, a direct consequence of where position now lives in the model. HeteroP is the less flashy but equally load-bearing half: none of the scaling-law numbers mean anything if the hyperparameters drift as you change scale, and a heterogeneous backbone needs a heterogeneous transfer scheme to keep them from drifting. The 6.8-versus-7.3 gap doesn't undercut any of that — it's a rounding-sized discrepancy in one headline multiplier, not in the mechanism. But it's exactly the kind of thing worth checking yourself before repeating a number, which is the whole reason to read the arithmetic instead of just the abstract. --- *Source: [Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers](https://arxiv.org/abs/2607.28611) (Ge, Jiang, Wang et al., Adobe Research, 2026), read from the arXiv HTML render. Figures 1–3 here are the paper's Figures 2, 12, and 16b, reproduced for commentary; all benchmark and scaling-law numbers are the paper's own, self-measured against in-house-trained baselines except where marked as cited. The RoPE re-assignment, HeteroP drift, and length-extrapolation diagrams are mine — the first is a schematic of the paper's own finding, the second reproduces the real numbers from its Figure 7 ablation on an illustrative loss curve, the third traces the actual tested points from its Figure 16b. Related: [Kimi K3](/articles/kimi-k3) for where KDA, MLA, and NoPE come from; [KDA has a half-life](/articles/kda-half-life) for the forget-gate math this piece leans on; and [SANA-Video 2.0](/articles/sana-video2) for the other linear-attention video architecture on this site.* --- # DCFormer: an ICML oral that let attention heads borrow each other's circuits, and quietly shipped anyway > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/dcformer > date: 2026-08-03 > tags: attention, transformers, architecture, scaling, llm **Dynamically Composable Multi-Head Attention** (DCMHA) is a two-year-old idea: May 2024, an ICML oral, a real mechanism, a real headline number — DCPythia-6.9B beats open Pythia-12B on Pile validation perplexity (5.95 vs 6.01) with close to half the parameters. By any normal measure that should have traveled. Semantic Scholar currently shows it at roughly a dozen citations, one flagged "influential." That's a modest footprint for an ICML oral two years out. What it does have is a real deployment: Caiyun Technology, the authors' own industry affiliation, put it into a production language model and an AI-RPG platform, and kept extending the training code through early 2025. The honest frame is not "field-changing" — it's a technically solid piece of work that quietly shipped without the citation graph to match, and is worth understanding on its own terms. There is a **second, unrelated** paper that reuses the same name: "DCFormer: Efficient 3D Vision-Language Modeling with Decomposed Convolutions" (arXiv 2502.05091, 2025) is about decomposed convolutions for 3D vision-language models — nothing to do with attention heads. Everything below is arXiv 2405.08553, Xiao, Meng, Li, and Yuan, "Improving Transformers with Dynamically Composable Multi-Head Attention." Check the ID if you go looking for it. ## Two problems with heads that never talk to each other Standard multi-head attention runs $h$ heads in parallel, each on its own learned $Q$/$K$/$V$ projection, and concatenates the results. The heads never see each other's work. That independence is also the design's two weaknesses. First, the **low-rank bottleneck**: each head's attention score matrix is a rank-limited function of a head-dimension-sized projection, and Bhojanapalli et al. (2020) showed that widening the per-head QK dimension relieves it — but widening every head's projection is expensive. Second, **head redundancy**: with nothing coupling them, heads are free to learn overlapping, partially duplicated functions instead of covering the space of useful attention patterns efficiently. DCMHA's answer isn't to make heads bigger. It's to let them **compose** — combine their scores and weights across heads, per token — which the paper shows buys the same kind of expressivity gain that a wider QK projection would, without actually widening anything. ## What Compose does Two calls to a function named `Compose` are inserted into ordinary multi-head attention: one right after the scores are computed (pre-softmax), one right after softmax turns them into weights (post-softmax).
For a fixed query/key pair, stack the $H$ heads' scores (or weights) into one vector $A_{:ij} \in \mathbb{R}^H$ — "the attention vector." `Compose` turns that into a new vector $A'_{:ij}$ by summing five branches: - **B1**, a static base projection (in practice, DCMHA drops this in favor of a plain skip connection, with no measurable loss); - **B2/B3**, a query-wise dynamic low-rank projection and a query-wise dynamic gate, both generated from $Q_i$ by a small FFN; - **B4/B5**, the same pair generated from $K_j$ instead. Every one of B2 through B5 is **input-dependent** — the weights that decide how much of head $h'$ leaks into head $h$'s new value are computed fresh from the actual query or key vector at that position, not fixed at training time. ## Why it has to be dynamic, not just wider The paper proves something specific about the *static* version of this idea first. Compose one head's score with a fixed matrix $C \in \mathbb{R}^{H \times H}$, and that is provably identical to concatenating an $H$-fold expanded QK projection (Theorem 2.1); do the same to the post-softmax weights, and it's identical to an expanded V/O projection (Theorem 2.2). In other words: a **static** composition matrix buys you exactly what a wider head dimension buys you — the fix for the low-rank bottleneck — and nothing more. **Talking-Heads Attention** (Shazeer et al., 2020) is that static case: it already composes both scores and weights, just with one fixed matrix reused for every token, every input, forever. DCMHA's own ablation measures the gap this leaves on the table — adding the static projection alone gets Pile validation perplexity from 11.68 down to 11.17, but the full dynamic Compose reaches 10.79. The static version is doing real work; the dynamic version is doing about 60% more of it, by this measure. Query-wise and key-wise branches contribute nearly as well on their own as together, and post-compose (on the weights) alone beats pre-compose (on the scores) alone, 11.05 vs 11.54 — the low- rank projection branches (B2/B4) matter more than the gates (B3/B5). ## Why it's cheap The reason DCMHA doesn't cost what a full head-to-head transform would is a decomposition, not a shortcut. Conceptually, composing every head with every other head needs an $H \times H$ transform per query/key pair — quadratic in the number of heads. DCMHA factors that tensor into a query-wise term plus a key-wise term (row + column), and factors each of *those* into a rank-$R$ product plus a diagonal gate (low-rank + diagonal decomposition). The cost drops from $H^2$ to $2HR + H$, and — a nice side effect — the key-wise half can be computed once and cached alongside K/V, which is exactly what a serving stack needs. At the paper's own 6.9B-scale example ($D_h = 128$, $R = 2$): roughly 1.3% extra parameters and 1.9–3.3% extra FLOPs, depending on sequence length. Rank $R = 2$ turns out to be close to a sweet spot in the ablation ($R{=}1$: 10.87 ppl, $R{=}2$: 10.83, $R{=}4$: 10.89 — non-monotonic, and not worth pushing higher). ## The headline number Trained on The Pile, matched Chinchilla-style token budgets, three model families: the scaling curves show **DCFormer-834M matches a plain Transformer trained with roughly 1.87 times the compute**, and DCFormer++ (RoPE + SwiGLU added to both sides) matches its own baseline at roughly 1.67 times. That gap doesn't shrink with scale — DCMHA's relative improvement decays more slowly than the RoPE+SwiGLU improvement does, which is the favorable direction. The result that carries the abstract is the 300B-token run against the actual Pythia suite:
DCPythia-6.9B's 5.95 beats Pythia-12B's 6.01 — a model with close to half the parameters, ahead on the metric that matters for pretraining. It also edges out on average 0-shot downstream accuracy (56.7 vs 56.5) and 5-shot (57.7 vs 57.2). The gap over its own size class is not close: Pythia-6.9B sits at 6.29, meaningfully behind. | Model | Pile ppl | Flan ppl | Avg 0-shot acc | |---|---|---|---| | Pythia-2.8B | 6.63 | 8.16 | 53.1 | | DCPythia-2.8B | 6.36 | 7.68 | 54.5 | | Pythia-6.9B | 6.29 | 7.85 | 55.1 | | **DCPythia-6.9B** | **5.95** | **7.13** | **56.7** | | Pythia-12B | 6.01 | — | 56.5 | The gap is largest on the Flan Collection (instruction-following/few-shot/CoT data) and grows with scale, which the authors read as DCMHA disproportionately helping the harder, more compositional end of the task distribution — a reading the paper backs up with a purpose-built test. ## A synthetic test built to need composition The authors built a 74-task, 888-example diagnostic where getting the right answer requires *simultaneously* attending to the right source token and applying the right output transformation (e.g., mapping an object to its superclass) — precisely the combination a head that only ever reads its own fixed QK/OV circuit should struggle with: Perplexity on this set drops from 10.05 to 7.36 alongside the accuracy jump — a much bigger swing than on Pile or Flan, and the paper's own explanation is the one you'd expect: this task rewards recombining an existing head's QK circuit with a different head's OV circuit on the fly, which is the one thing static heads structurally cannot do. The head-diversity analysis (captured variance of concatenated QK and OV projection matrices, lower meaning more diverse heads) backs this qualitatively too — DCPythia shows markedly more QK-circuit diversity than Pythia, and moderately more OV-circuit diversity. ## The honest costs None of this is free, and the paper says so plainly. Composition is I/O-bound, not compute-bound, and the reference implementation has no fused kernel — plain JAX for training, plain PyTorch for inference: | Size | Training throughput (DCFM++ / TFM++) | Inference throughput (DCFM++ / TFM++) | |---|---|---| | 2.8B | 74.5% | 81–88% | | 6.9B | 83.1% | 89–95% | | 13B | 84.4% | 90–95% | | 33B | 89.2% | 90–95% | The overhead shrinks as models scale up, and the authors are explicit that a fused kernel — FlashAttention-style — is headroom they haven't taken. A separate lever recovers most of it directly: raising the local-to-global attention ratio and composing only query-wise (dropping the key-wise branches) pushes DCFormer++-6.9B's training throughput from 83.1% back up to 92.5% of baseline, at a small, still-net-favorable cost in perplexity. Two more honest limits, stated in the paper's own words. First: **DCMHA doesn't transplant onto a pretrained model.** Continual-pretraining a 1.4B LLaMA-style checkpoint into a DCFormer for a tenth of its original training steps produced no real improvement — the composition that matters most happens in early layers, and early-layer gradients are too small during fine-tuning to move already-settled MHA weights. DCFormer has to be trained from scratch. Second: the paper is explicit that **matching SOTA was never the goal** — DCPythia deliberately keeps every other Pythia hyperparameter fixed, to isolate what DCMHA alone contributes, rather than stacking it with every other efficiency trick to chase a leaderboard number. The mechanism transfers outside language too: on ImageNet-1K, DCViT-S/16 at 1.03× the baseline's parameters (68.0 top-1 at epoch 90) matches ViT-M/16 at 1.72× the parameters (67.1 top-1) — the same roughly 1.7× parameter-efficiency story, in a different domain, on one held-out test. ## A different axis from the memory-side attention variants If you've read the [field guide to attention mechanisms](/articles/attention-mechanisms) on this site, it's worth being precise about where DCMHA sits relative to that map. MQA, GQA, and MLA all operate on what that piece calls the **memory axis** — they *share* or *compress* the K/V heads to shrink the KV cache, trading some quality for less memory bandwidth at decode time. DCMHA doesn't touch the cache at all: the number of physical K/V heads is unchanged, nothing shrinks. It operates on an orthogonal axis entirely — not how many heads you cache, but what each head is allowed to compute, by letting it borrow another head's QK or OV circuit, per token. A model could in principle combine GQA's cache savings with DCMHA's composition; the paper doesn't test that combination, so read it as plausible, not demonstrated. (For the general design question of specializing attention below the layer level, [HydraHead](/articles/hydrahead) is the other piece on this site working that seam, from a different angle.) The broader landscape of architecture choices — attention, position encoding, MoE, diffusion — is mapped at [/architectures](/architectures). ## How much of this actually caught on Being fair to the number in the abstract requires two caveats the paper itself doesn't hide. The Pythia baselines are **re-run by the authors** under matched settings, not copied from the original paper — a genuinely controlled comparison, and the authors say so directly ("our aim is not to obtain SoTA results, but to clearly quantify the gain"). And the compute-equivalence multipliers (1.87×, 1.67×, 1.85×, 1.97×) come from fitting scaling-law lines to **three data points per curve** — reasonable given the cost of training each point, but a thinner fit than, say, Chinchilla's own study. Two years after an ICML oral, roughly a dozen citations and one flagged "influential" is a modest academic footprint — the kind of number that would normally suggest an idea that didn't pan out. What actually happened looks different: Caiyun Technology, the paper's own industry co-author's employer, shipped DCFormer into a production language model and upgraded an AI-RPG platform to run on it, and the GitHub repository was still being extended — DeepSpeed ZeRO support, Hugging Face Trainer integration — well into 2025. No successor paper has benchmarked against DCFormer as a state-of-the-art baseline to beat; the adoption signal here is industrial, not academic. That's a real but different kind of validation than a citation count measures, and it's the honest way to read this one: not a paper that changed the field's direction, but a working piece of architecture that one production system actually adopted, sitting quietly under-cited. ## The take Fixed, independent attention heads leave two things on the table: a low-rank bottleneck that a wider head dimension fixes at a real cost, and redundancy nothing forces heads to avoid. DCMHA's Compose function fixes both by recombining heads' scores and weights per token, through a decomposition cheap enough that a 6.9B model pays about 1.3% more parameters for it. The result — DCPythia-6.9B beating Pythia-12B on perplexity at roughly half the parameters — is real, reproduced by the authors under controlled settings, and backed by a synthetic test built specifically to need what static heads can't do. It just hasn't been the paper everyone cites. Production adoption at one company and a quiet GitHub repository are what two years actually bought it — which is a fine outcome for a piece of architecture, even if it isn't the one the citation count would lead you to expect. --- *Built on [Improving Transformers with Dynamically Composable Multi-Head Attention](https://arxiv.org/abs/2405.08553) (Da Xiao, Qingye Meng, Shengping Li, Xingyuan Yuan; Beijing University of Posts and Telecommunications / Caiyun AI; ICML 2024, Oral) and its [code release](https://github.com/Caiyun-AI/DCFormer). Figures are the paper's own Figures 2 and 4, reproduced for commentary. Tables and numbers are the authors' except where marked as this site's own illustrative simplification (the Compose bar demo, the head-cost demo); interactive diagrams are mine.* --- # Dream-Cubed: diffusion directly on Minecraft's block IDs, where inpainting comes free > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/dream-cubed > date: 2026-08-03 > tags: diffusion, generative-models, minecraft, world-models, 3d Most generative work on 3D game worlds goes through a pixel renderer or a learned latent space before it touches anything the game itself understands. **Dream-Cubed** skips both. It trains a diffusion model directly on Minecraft's own vocabulary — integer block IDs, the same representation the game engine uses — and gets a specific, useful property for free as a result: place a block by hand anywhere in a chunk, and the model will build around it *exactly*, with no special training and no extra inference-time machinery. The paper is a three-month-old preprint from a team at NYU and Sakana AI, posted with no peer review and, as of this writing, no citations. It also lands in a niche that's suddenly crowded — the paper names two concurrent Minecraft/voxel generators of its own accord — and its own limitations section is unusually candid about how little its evaluation actually proves. All of that is worth knowing before the mechanism, which is genuinely worth understanding. ## Blocks as tokens, not pixels Each training sample is a $32 \times 32 \times 32$ tensor of integer block IDs — dirt, water, stone, whatever the game placed there — over a vocabulary of 117 block types in the core procedurally-generated set, extended to 177 once six professionally human- authored maps are folded in. The dataset totals 1,667,781 procedural chunks plus 358,762 human-authored ones: **2,026,543 chunks**, tens of billions of tokens at $32^3 = 32{,}768$ voxels per chunk, spanning fifteen biomes from ocean to village to cave. One backbone serves both diffusion families the paper compares: a 280M-parameter 3D Diffusion Transformer, 25 blocks, hidden dimension 768, 8 attention heads. A 3D convolution patchifies each chunk into non-overlapping voxel patches; fixed 3D sine-cosine position embeddings and a biome-label embedding condition every block through AdaLN modulation, the same conditioning mechanism DiT uses for timestep and class in image diffusion. ## Two diffusion families, one backbone **Masked discrete diffusion (MD4).** Add a `[MASK]` token to the block vocabulary. The forward process independently masks each voxel with probability $p_{\text{mask}}(t) = \sin(\pi t / 2)$ for $t \sim \mathrm{Uniform}(0,1)$ — at $t{=}0$ nothing is masked, at $t{=}1$ everything is. The network sees the corrupted chunk and predicts the original block ID at every masked position, trained with cross-entropy on masked positions only. Sampling starts from an all-`[MASK]` chunk and iteratively unmasks positions over a fixed number of steps. **Continuous diffusion (DDPM), in an embedding space.** Every block name — "dirt", "sand" — is embedded once via OpenAI's `text-embedding-3-small`, giving a frozen 16- dimensional lookup table with a semantic prior baked in for free. A standard cosine noise schedule runs $x_t = a_t x_0 + b_t \varepsilon$ over 1000 steps, trained with v- prediction. At the end of sampling, the continuous output is decoded back to discrete block IDs by nearest-neighbor lookup against the embedding table. Both are trained on the identical backbone, the identical data, the identical compute budget — the paper's stated goal is a controlled, apples-to-apples comparison of the two diffusion formulations, not a fight either one is rigged to win. ## Why inpainting is free under one formulation and not the other This is the mechanism worth sitting with. MD4's forward process decides, **independently, per voxel**, whether that voxel gets masked. A voxel the user has placed by hand simply isn't a candidate for that decision — it's excluded from the start, at every step, all the way through sampling. There is no moment where the model has to reconcile "what I was trained to expect here" with "what's actually here," because the fixed voxel was never part of the corruption process to begin with. A continuous DDPM cannot get the same thing for the same reason. Its forward process adds Gaussian noise to *every* position, at every step, following one global schedule — clean voxels aren't a special case the network was ever trained to see. If you clamp a user- placed block to its clean embedding partway through sampling — the natural thing to try — the model is still conditioned on that position carrying noise amplitude $b(t)$ at timestep $t$, and it sees zero instead. That's a real mismatch between what training taught the network to expect and what inference is handing it, not a cosmetic one. The paper is upfront that it doesn't solve this: closing that gap needs extra machinery — RePaint-style repeated re-noising and re-sampling — which it flags in an appendix as an unresolved comparison point, not a capability it demonstrates for the DDPM side. Exact conditioning is what falls out of the masked formulation for nothing; it's what the continuous one would have to be re-engineered to approximate.
If you've read the piece on [iLLaDA](/articles/illada-diffusion-language-model) on this site, the mechanism will look familiar: masked discrete diffusion over text tokens is the same "absorbing-state" idea MD4 applies here to voxels — a masking probability schedule, a network trained to fill in exactly the masked positions, bidirectional context by construction. iLLaDA's own masking ratio is closer to a straight linear schedule ($t$ itself, roughly); Dream-Cubed's MD4 uses the $\sin(\pi t/2)$ reparameterization from Shi et al.'s original MD4 paper — a detail, not a different mechanism. What changes here is the alphabet the diffusion runs over: block IDs instead of vocabulary tokens, arranged on a 3D grid instead of a 1D sequence. Same masking idea, different token space — see also the [masked-diffusion-lm entry](/architectures) in the architecture map for where this sits relative to the wider non-autoregressive-LM family. ## Outpainting is the same trick, tiled Generating a world larger than one $32^3$ chunk uses a sliding window: partition the larger canvas into overlapping cells, generate them in sequence, and for every cell after the first, treat the already-generated overlap with its neighbors as more fixed context — recursively the same "these voxels are excluded from masking" trick, just applied at world scale instead of one seeded pattern.
The cost of this is real and disclosed: a single 5×5 outpainted world takes over an hour of H100 inference time, generated cell by cell, sequentially — the paper calls inference speed "a practical barrier to all envisioned applications," not a solved problem. ## What the numbers actually say **MD4 and DDPM land in a statistical tie on the paper's own metric.** Adjusted FID (generated minus a reference FID from held-out chunks) averages 59.26 for MD4 at patch size 2 versus 59.29 for DDPM at the same patch size — indistinguishable overall, with MD4 winning 9 of 15 biomes and DDPM winning 6. **Patch size is where the two formulations actually separate.** MD4 holds up at patch sizes 2 (4,096 tokens per chunk) and 4 (512 tokens), with visible artifacts only at patch 8; DDPM works at patch 2 but **fails outright at patch 4** under the identical configuration. That's the one place in the paper where discrete and continuous diffusion give clearly different answers, and it favors the discrete side. **Naive frequency matching doesn't work for rare, structured content.** Three data mixtures were compared: a balanced split, natural biome frequency, and a village-boosted split. Natural frequency wins on average FID — but ocean chunks, over-represented 5.3× relative to balanced, and forest, at 1.7×, improve, while village and cave, both rare and structurally complex, get worse. Boosting village samples specifically recovers the village-biome losses. The honest reading: matching real-world frequency is not automatically the right training mixture once some categories are both rare and hard. The human study is small and its own authors say so: 19 Minecraft-experienced participants (all from the authors' own institution), roughly 1,000 two-alternative forced-choice trials, free pan/zoom/rotate: Both MD4 configurations beat real chunks at statistical significance (patch 2: p less than 0.001, patch 4: p = 0.042) — a result the authors attribute candidly to classifier-free guidance pushing generated samples toward a "prototypical" idealized biome, more uniform than messy real terrain, rather than claiming their model has somehow out-built reality. MD4 (patch 2) and DDPM (patch 2) tie head-to-head at 49.4%, consistent with the FID result above. | Min FID gap between two models | Agreement with human preference | n | p | |---|---|---|---| | 0 | 54.3% | 512 | 0.029 | | 5 | 61.7% | 227 | less than 0.001 | | 10 | 62.9% | 159 | less than 0.001 | | 15 | 66.1% | 109 | less than 0.001 | Agreement between FID and human raters rises with the size of the FID gap being compared, but tops out at 66.1% even at the largest gap tested — a coin flip with a thumb on the scale, offered by the authors themselves as evidence that FID is only weakly informative here, and only at large gaps. The evaluation the whole paper hangs its numbers on has real, self-acknowledged holes. FID is a render-based metric: it cannot see building interiors, cave systems, or whether a door is actually reachable — none of the things that make a Minecraft structure *functional* rather than merely picturesque. It's computed from only 1,500 rendered images per model, and costs roughly 60 GPU-hours to run — expensive enough that the authors say it can't be used for model selection during training, only for a final after-the-fact score. And the human study, small as it is, evaluated only biome- conditioned generation. It never tested the inpainting or outpainting capability the paper actually leads with — the figures above are demonstrations, not measured results. The dataset itself is drawn from Minecraft version 1.12.2 (2017), for tooling compatibility, so newer blocks and biomes aren't represented at all. Training cost is disclosed cleanly: 4×H100 GPUs, classifier-free guidance with 20% label dropout during training and a guidance scale of 4.0 at inference, patch-2 models run 20 epochs and patch-4 models 160 (matched for equal token exposure), roughly 192 GPU-hours total across every model in the paper. Inference is the bottleneck end of the system: about 2.5 minutes per chunk at patch 2, 25 seconds at patch 4. ## A crowded moment, honestly disclosed Dream-Cubed cites its own competition directly rather than presenting itself as singular: **Scaffold Diffusion** (a NeurIPS 2025 workshop paper that conditions on an input occupancy scaffold instead of generating from nothing) and **PERSIST** (arXiv 2603.03482, roughly a month earlier, which uses a 3D DiT with rectified flow matching but as one component inside a video-generation system, not a standalone generator) are both named as concurrent work on the same general problem, in the same few months of 2026. **Solaris** (arXiv 2602.22208) is cited as another concurrent voxel/world-modeling effort in the same window. None of WorldGAN, Scaffold Diffusion, or XCube — the paper's narrative comparison points — are actually benchmarked against Dream-Cubed's FID or human-preference numbers on shared ground; the positioning against prior work is qualitative throughout, and there is no table anywhere in the paper showing Dream-Cubed beating a previously published Minecraft or voxel generator on a metric both were scored on. For a reader wanting a settled state-of-the-art claim, that table doesn't exist yet. Code, data, and all pretrained models are released ([github.com/SakanaAI/DreamCubed](https://github.com/SakanaAI/DreamCubed)), which is worth crediting on its own — a preprint this new, this openly scored against its own limitations, and this fully released, is a reasonable way to publish work you don't yet have citations to back up. ## The take The technical point is narrow and real: masked discrete diffusion turns user-block- conditioning, inpainting, and outpainting into a structural guarantee — unmasked voxels were never part of the corruption process, so they can't drift — while the equivalent constraint on a continuous DDPM has to be bolted on after the fact, and the paper is explicit that it doesn't fully solve that side. Working directly in block-ID space, skipping pixels and learned latents entirely, is what makes that guarantee possible in the first place. Everything past that point is evidence you should discount appropriately: FID and DDPM come out statistically tied on the paper's own numbers, the human study that exists didn't test the paper's headline capability, and the field around this exact problem got crowded within the same few months this was written. Read Dream-Cubed for the mechanism and the pictures it produces — both hold up on inspection — and treat the quantitative claims as a first data point from one preprint, not a result that's been through the wringer yet. --- *Built on [Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes](https://arxiv.org/abs/2604.22847) (Tim Merino, Sam Earle, Ryunosuke Iwai, Julian Togelius, Edoardo Cetin; NYU / Sakana AI, preprint, April 2026) and its [code and data release](https://github.com/SakanaAI/DreamCubed). Figures are the paper's own Figures 6 and 7, reproduced for commentary. Tables and numbers are the authors' except where marked as this site's own illustrative simplification (the block-grid and schedule-mismatch demos use hand-picked, not trained, values); interactive diagrams are mine.* --- # Explorative Modeling: factor the training loop, not the generation loop > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/explorative-modeling > date: 2026-08-03 > tags: generative-models, diffusion, training-methods, robotics, explainer Every generative model has to solve the same problem: a prompt like "generate a dog" has billions of valid answers, not one. [Diffusion](/articles/set-diffusion) handles that by denoising in dozens of steps; autoregression handles it by predicting one token at a time. Both are the same trick wearing different clothes — break *generation* into enough small steps that no single prediction has to average across the billions of valid dogs. **Explorative Modeling** (Gladstone, Ji, and Du — UIUC and Harvard) asks a different question: what if you left generation alone and factored the *training loop* instead? Generate $K$ candidate outputs per example, score them all against the target, and only backpropagate through the winner. They call this **exploration**, argue it is a third axis for scaling generative models — next to parameters and data — and show it can *substitute* for step-factored generation entirely, which is what lets them train models that generate in a single forward pass and still match diffusion.
## Why generation gets factored in the first place Squared-error regression is maximum likelihood — under a fixed-variance Gaussian. That sounds like a technicality, but it has teeth: the maximum-likelihood-optimal output under a multimodal target is the *mean* of every valid answer, and the mean of "every plausible dog photo" is not a photo of a dog. It is a brown blur. Same failure in language: force a next-token model to average over "the cat sat on the _____" and you get a smear of probability mass spread across every plausible word, not a committed answer. The paper gives this capacity a name, **generative expressivity**: the number of distinct modes a training objective's loss minimizer is *allowed* to capture. A direct regressor has expressivity $E=1$ — one output, always the average — and no amount of extra parameters or data raises it, because expressivity is a property of the *objective*, not the model. That is the reason diffusion and autoregression look the way they do. Autoregression conditions each token on everything already generated, so by the time it predicts token $i$, most of the multimodality is already resolved by the tokens before it — each individual prediction is closer to unimodal. Diffusion does the analogous thing over noise levels: each denoising step only has to move a slightly noisy sample a little cleaner, not solve the whole distribution at once. **Factoring generation is a device for keeping generative expressivity high**, one small step at a time. It is also why training and inference stop matching: a diffusion model trained on isolated denoising steps gets unrolled over hundreds of them at test time, and even single-step distillations still anchor their *training* targets to the multi-step trajectory. Sampling and training are never the same procedure, so exposure bias never fully goes away. ## Forward XM: buying expressivity with candidates, not steps If factoring generation is one way to raise expressivity, exploring more candidates is another. Fix a data target $x$, draw $K$ generations $\hat y_1, \dots, \hat y_K \sim G_\theta$ from the model, and train only on the closest one: $$ \mathcal{L}_{\text{Forward}}(\theta) \;=\; \min_{i \in \{1,\dots,K\}} J(\hat y_i, x) \tag{1} $$ This is the entire mechanism — no new architecture, no new loss family, just a `for` loop around generation and a `min` before `.backward()`. Explore $K$ candidates with a plain regressor and its expressivity rises to at least $K$: with enough candidates, one of them lands near enough to any given mode that the model can commit to it instead of averaging. Play with $K$ above and the mechanics of equation (1) are exactly what's on screen: every candidate is scored against the same target, the closest one gets the gradient, and the rest are discarded for this step. Pulling $K$ up can only tighten the best-of-$K$ error — never loosen it, since you are taking a minimum over a strictly larger set — but each extra candidate is a full extra generation, so the compute for a training step scales linearly with $K$. That is the whole cost of exploration, and it is paid once, during training. The same de-blurring shows up in the model's actual outputs, not just the training loss:
XM-1 is not a badly-trained model — it is the theoretical best a direct regressor *can* do, the blurred mean equation (1) predicts. XM-50 is the same architecture, same data, same loss, with one difference: 50 candidates scored per step instead of one. ## Forward and Reverse XM, and what they actually optimize Forward XM (fix the data, explore the model's generations) is one direction. **Reverse XM** flips it: fix one generated sample $\hat y \sim G_\theta$ and explore over $K$ data targets $x_1, \dots, x_K \sim \mathcal D$, training toward whichever is closest: $$ \mathcal{L}_{\text{Reverse}}(\theta) \;=\; \min_{i \in \{1,\dots,K\}} J(\hat y, x_i) \tag{2} $$
These are not cosmetically different. The paper works out what each one minimizes in the limit. Write $p_\theta$ for the model's output distribution blurred by the reconstruction kernel (Gaussian, for squared error) and $p^*_\sigma$ for the data blurred the same way. Then, in their smooth relaxations: $$ \underbrace{\mathrm{KL}(p^* \,\|\, p_\theta) + H(p^*)}_{\text{Forward XM}} \qquad\text{and}\qquad \underbrace{\mathrm{KL}(g_\theta \,\|\, p^*_\sigma) + H(g_\theta)}_{\text{Reverse XM}} $$ Forward XM's entropy term, $H(p^*)$, belongs to the *data* — a constant the model can't touch — so Forward XM is just maximum likelihood over its $K$-candidate mixture, for every $K$. Mass-covering, never collapsing, but the recall comes at the price of running $K$ full generations per step, so it struggles to scale to very high-multimodality targets. Reverse XM's entropy term, $H(g_\theta)$, belongs to the *model* — something it can shrink by narrowing its own spread — so Reverse XM drifts toward collapse on its own and needs an explicit entropy bonus to stay at the true reverse-KL optimum instead. The paper is candid that Reverse XM's fix is "largely left for future work"; Forward XM is what every downstream result in the paper actually runs. ## Substitutable, not just additive Here is the move that turns this from "a training trick" into "a scaling axis." Factoring generation exists only to supply expressivity. Exploration supplies the same quantity a different way. If that's right, the two should be interchangeable — you should be able to trade generation steps for exploration and land in the same place. The paper tests this directly with **Jumpy** models, a family that interpolates between direct regression (one jump) and full continuous-time flow (infinite jumps) by varying the number of steps. Take two Jumpy models, one with fewer jumps and one with more, and add exploration to both: the model with *fewer* jumps — the more end-to-end one — gains more from exploration than the one that already had step-factorization doing the work. That is the substitution effect, measured rather than asserted: the less a model already leans on factored generation, the more it has to gain from factoring training instead. Push that trade all the way and you get **XM's other headline**: a model that samples exactly the way it trained, in one forward pass, with no separate multi-step inference procedure to keep in sync. The paper calls a model "end-to-end" when it never faces inputs at inference it wasn't trained on — no denoising schedule to unroll, no exposure bias from a mismatched sampling procedure. This is the same argument that ended hand-designed feature pipelines after AlexNet, aimed now at the one corner of deep learning that never fully got the memo: [diffusion language models](/articles/illada-diffusion-language-model) and [autoregressive decoding](/articles/how-llm-inference-works) both still train on one procedure and sample with another; exploration is what lets a model close that gap without giving up quality. The trade is exactly compute, moved to a different place in the pipeline: [MrFlow](/articles/mrflow-diffusion-acceleration) and [Set Diffusion](/articles/set-diffusion) both attack the *inference* side of this same step-factorization: reshuffle where a fixed step budget gets spent, or change which tokens get decoded together, but the sample is still built from many forward passes. Explorative Modeling is a different lever entirely — it doesn't make the multi-step generator cheaper, it removes the requirement to be multi-step in the first place, by paying for expressivity up front instead of on every draw. ## Results: three modalities, two robot benchmarks **Image generation.** Added to RAE, a near-state-of-the-art ImageNet latent-diffusion recipe, exploration (XRAE, using XM-2) reaches a near-SOTA **1.43 FID** without classifier-free guidance: | Method | FID (no CFG) ↓ | |---|---| | DiT | 9.62 | | SiT | 8.61 | | VA-VAE | 2.17 | | REPA-E | 1.70 | | Latent Diffusion + RAE | 1.55 | | **XRAE (RAE + XM-2)** | **1.43** | That 47% parameter-efficiency figure is unrelated to the next number, which happens to share a digit: RAE itself converges 47x faster than SiT (a separate, prior result the paper is building on), and stacking XRAE's 6.2x sample efficiency on top of *that* puts the whole recipe at roughly **300x faster to converge than plain SiT** — the paper's arithmetic, not an independent measurement. One negative result worth keeping: minibatch optimal-transport coupling, an alternative de-blurring trick, made FID *worse* (46.3 → 54.5 at the Small scale) — exploration wins here specifically, not "adding any anti-blur trick" generically. **Scale doesn't dilute the gain — it grows it.** Going from XM-5 to no exploration, the improvement climbs from 13% to 23% as model size scales up, and from 7% to 36% as data scales up. That is the opposite of what you'd expect from a scaling axis that's about to run out of room — the paper's reading is that generative expressivity becomes a *larger* bottleneck as the other two axes get pushed harder, because parameters and data stop being the limiting factor first. **Video (Something-Something V2).** FID/FVD improve monotonically with more explored modes. The more interesting number is generalization, not fit: best achievable FVD is **30.0 with exploration versus 37.5 without** — less overfitting on a fixed dataset, which the paper frames as a compute-generalization tradeoff: extra training compute spent on exploration buys generalization the way more data usually does. **Robot policies (Behavior Cloning, Robomimic).** This is where "single forward pass, matches diffusion" gets tested against a real baseline: Explorative Policy matches Diffusion Policy on Lift and Can (both 100%), and beats it on Square (96% vs. 94%), Transport (74% vs. 72%), and ties on Tool Hang (86%) — at **1 forward pass instead of 100**. **Goal-conditioned world models (Maze2D), vs. Diffuser:** Average score edges up too (130.0 vs. 127.2), at 16-256x fewer denoising steps depending on the maze size (4 vs. 64 on U-Maze, 1 vs. 256 on Medium). This pairing — matching or slightly beating a strong multi-step baseline, at two orders of magnitude less inference compute per sample — is the article's headline for a reason: it is the plainest demonstration that the compute the paper claims you save at inference is compute it actually spent, once, at training. ## What I make of it - **The conceptual reframe is the real contribution.** "Factor the training loop instead of the generation loop" is a genuinely different axis, not a repackaging of an existing trick — best-of-$K$ training has appeared before, but treating it as *substitutable* for diffusion/AR step-factorization, and confirming that substitution empirically with the Jumpy-model ablation, is new. - **"Third scaling axis" is the authors' framing, argued from one paper's worth of experiments** — real equations, a real KL derivation, and empirical scaling curves that trend the right way, but not yet a claim anyone outside this group has stress-tested. Treat it as a strong hypothesis with supporting evidence, not an established fact. - **The evaluation is real but narrow at the edges that matter most.** The robotics results — the ones carrying the "matches diffusion at 100-256x less inference compute" headline — are Robomimic proficient-human/state-observation behavior cloning and Maze2D planning: small, well-studied benchmarks, and the paper says outright the control experiments got "barely any tuning." That's stated as a limitation working in XM's favor (untuned and already competitive), but it also means these are not yet frontier-scale robot-learning results, and there's no evidence here about vision-based control, long-horizon manipulation, or real hardware. Gains on autoregressive language models are the weakest reported of any modality — the paper's own explanation is that next-token prediction is already close to unimodal, so there's less blur left for exploration to fix. - **Code is Apache-2.0 and real, which counts in its favor** — `--xm_best_of_k K` is a genuine flag in a runnable repo, not a promise. But as of this writing the repository explicitly marks the code behind the headline results as not yet released: the RAE image-generation runs use a separate codebase "to be released separately," and masked-diffusion-language-model and control-task (robot policy / world model) code are both marked "coming soon." What's public today is the general XM training scaffolding, not a drop-in reproduction of the paper's own numbers. - **No third-party replication yet** — the paper is a July 2026 preprint. The authors are transparent about a related limit: Diffusion Policy's numbers had to be reproduced under their own setup because they used a newer Robomimic version than the original paper, so the baseline is a good-faith re-run, not a quoted number — a small but real point of honesty worth crediting. The clean way to hold all of this: exploration is a training-time payment for a capability generation factorization normally buys at inference-time, over and over. That's a real trade, mechanically well argued, and it works on real benchmarks. Whether it holds at frontier model scale, on harder control tasks, or once other labs have run the numbers, is still open. --- *Built on [Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation](https://arxiv.org/abs/2607.27372) (Gladstone, Ji & Du — UIUC and Harvard, 2026). Code: [github.com/alexiglad/XM](https://github.com/alexiglad/XM) (Apache-2.0). Figures 1-3 are reproduced from the paper; all numbers are from its Tables 1-3 and Sections 4.1-4.2.* --- # Fara 1.5: an open, vision-only browser agent at 4B, 9B, and 27B > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/fara-1-5 > date: 2026-08-03 > tags: agents, computer-use, open-weights, vision-language-models [Qwen-CUA](/articles/qwen-cua), published earlier today, controls a full desktop from screenshots alone with a 397B-A17B mixture-of-experts model. Fara1.5, from Microsoft Research, does the narrower version of the same job -- browser only, no desktop -- at 4B, 9B, and 27B parameters, and ships all three sizes under the MIT license, weights included. That is two orders of magnitude smaller than Qwen-CUA's backbone, scoped to one surface instead of an entire OS, and open in a way Qwen-CUA specifically isn't: Qwen-CUA's repository is Apache-2.0, but its own README states plainly that "model weights are not included" -- paper and demo only. Fara1.5 publishes the checkpoints. The interesting question isn't whether the biggest Fara model is good. It's what the 4B-to-27B ladder says about how much capability a vision-only agent actually needs.
## The same narrow interface, a much smaller model Fara1.5 is a multimodal decoder-only model built on Qwen3.5, at three sizes, all trained the same way. Given a goal, the current screenshot, and the last three steps of history, it reasons in text and then emits exactly one action from a fixed vocabulary: click, type, scroll, drag, `visit_url`, `web_search`, `go_back`, plus meta-actions for longer horizons -- `memorize` to persist a fact past the three-screenshot window, `ask_user` to pause on a critical point, `finish` to stop. Coordinates are predicted directly from pixels, the same design choice Qwen-CUA makes for an entire desktop: skip the DOM and accessibility tree, and bet that whatever a human can operate with a screen and two input devices is a general enough interface. Fara1.5 just makes that bet at a fraction of the parameter count, and only for the browser. The safety mechanism is worth naming precisely because it's a real constraint, not a suggestion: the model is trained to trigger `ask_user` at eight defined "critical point" types -- across three dimensions (permission granted or not, task fully specified or not, action reversible or not) -- covering things like entering personal information, submitting payment or shipping details, or sending a message on the user's behalf. Microsoft's recommended deployment wrapper, MagenticLite, is sandboxed and pausable specifically so a human is in the loop at those points. `ask_user` and `memorize` are context-management tools in exactly Lilian Weng's sense from [Agent harnesses](/articles/agent-harness) -- deciding what to carry forward and when to hand control back -- just built into the action vocabulary itself rather than the surrounding scaffold. ## Trained on trajectories it generated for itself Nearly all of Fara1.5's training data comes from FaraGen1.5, its own synthetic-data pipeline. A solver -- GPT-5.4, paired with a user simulator that withholds task details the way a real user would -- attempts tasks in two kinds of environments: the live, open web, and six sandboxed synthetic apps (Mail, Calendar, Stream, ML, Stay, Scheduler) whose functional code was itself generated by a coding agent (GitHub Copilot CLI) rather than scraped or mocked. Every resulting trajectory then has to clear three independent verifiers before it counts as training data. That data becomes roughly 2 million training samples, 60% still ordinary open-web trajectories, the rest split across synthetic environments, deliberately ambiguous form-filling, grounding, and a small slice of VQA and drag gestures. It's a real answer to the standard complaint about computer-use training data -- human demonstrations are slow and expensive to collect -- but it's worth being precise about what "generated" means here: the solver, the user simulator, and the verifiers are themselves LLM judgments, not ground truth. A verifier checking "did this trajectory ask before an irreversible action" is exactly as reliable as the model doing the checking. ## What the ladder buys Fara1.5-27B reaches 72.3% on Online-Mind2Web, ahead of Gemini 2.5 Computer Use (57.3%), OpenAI Operator (58.3%), and Yutori Navigator n1 (64.7%) -- three proprietary systems, all evaluated on an independently maintained academic benchmark, not one Microsoft built. That's a genuine result: an open, MIT-licensed family beating closed competitors on a benchmark none of them control. But it only holds at the top of the ladder. WebTailBench v1.5 -- Microsoft's own 609-task eval set, worth flagging as self-authored rather than independent -- shows the same monotonic climb, no crossover to check it against: Read against the predecessor, Fara-7B, the jump looks even sharper: Fara1.5-9B improves +29.3 points on Online-Mind2Web, +13.1 on WebVoyager, +8.3 on WebTailBench, +18.1 on ScreenSpot-Pro grounding, +8.9 on OSWorld-G Refined. That comparison is real but not clean -- it conflates a parameter increase (7B to 9B) with a full training-pipeline change (FaraGen1.5 replacing whatever generated the original Fara's data). Stated as a training-pipeline improvement, it overclaims; stated as "the current generation beats the last one," it's exactly as strong as it sounds and no stronger. The Online-Mind2Web and WebVoyager comparisons against Operator, Gemini 2.5 CU, and Navigator n1 are self-reported by Microsoft on benchmarks those three systems don't control -- a meaningfully better setup than grading your own exam, but still not an independently run leaderboard. No third-party replication of these specific numbers was found for this piece. ## Where the vision-only bet costs something The model card is direct about the downsides of skipping the DOM: English-only, vulnerable to visual deception and prompt injection embedded in page content, error accumulation over long multi-step trajectories, and explicitly **not suitable** for legal, health, or financial use. None of that is unique to Fara1.5 -- Qwen-CUA's paper documents the same shape of limitation for the same underlying reason -- but a 4B vision-only model has less capacity to notice something is wrong mid-trajectory than a 397B-A17B one, and the model card doesn't pretend otherwise. ## The take Two orders of magnitude smaller than Qwen-CUA, scoped to a browser instead of a desktop, and shipping actual weights under MIT where Qwen-CUA ships code and a paper but withholds the checkpoints: Fara1.5 is a genuinely different point in the design space, not a smaller copy of the same idea. The headline -- 27B beats three proprietary computer-use agents on a benchmark none of them own -- is real. The more useful reading of the paper is the ladder underneath it: at 4B, Fara1.5 merely ties the weakest of those three baselines; the win only fully arrives at 27B. Vision-only browser control is not a capability a small model gets for free. It's bought, roughly a third of it per step up the ladder, exactly as parameter-scaling laws would predict. --- *Built on Microsoft Research's [Fara1.5: Scalable Learning Environments for Computer Use Agents](https://arxiv.org/abs/2606.20785) (Awadallah et al., 2026) and the [microsoft/fara](https://github.com/microsoft/fara) repository (MIT license). Figures 5 and 7 are reproduced from the paper for commentary, flattened onto white; the FaraGen1.5 pipeline diagram and scaling-vs-baseline chart are original illustrations of the paper's Figure 2 and Table 3 / Figure 7 data, not measured traces. Benchmark numbers are as reported in the paper.* --- # Language models are injective, so a KV-cache is not a summary — it's the prompt, in another basis > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/injective-language-models > date: 2026-08-03 > tags: llm, interpretability, privacy, theory, kv-cache Every individual piece of a transformer is lossy. LayerNorm throws away a per-token scale and shift. Softmax attention collapses many key-value pairs into one weighted average. Low-rank projections shrink dimensions on purpose. It would be reasonable to conclude that the hidden state a transformer produces for a prompt is a compressed, irreversible summary of that prompt — the way a hash or a JPEG is. [Nikolaou, Mencattini, Crisostomi, Santilli, Panagakis, and Rodolà](https://arxiv.org/abs/2510.15511) prove that conclusion is wrong. Decoder-only transformer language models, as a whole, are **almost surely injective**: two different prompts essentially never produce the same last-token hidden state. Not "usually don't" — provably, with probability one, for any model trained by gradient descent for any finite number of steps. And they don't stop at the proof. **SipIt** is an algorithm that exploits injectivity to reconstruct a prompt *exactly* from its hidden states, in time linear in the prompt's length, with 100% accuracy in their tests. The paper is ten months old, already accepted at ICLR 2026, and has picked up roughly 30 citations (7 flagged "influential") in that time — fast uptake for a theory paper. The result has an immediate consequence worth sitting with: a hidden state, or a cache built from one, is not a lossy fingerprint of what a user typed. It is that text, in a different basis, and it can be read back out. ## Non-injective parts, injective whole The paper's move is to stop looking at individual components and instead look at the whole map. Write $f: \mathcal{V}^{\le K} \times \mathbb{R}^p \to \mathbb{R}^d$ for the model: a vocabulary $\mathcal{V}$, a context bound $K$, parameters $\theta \in \mathbb{R}^p$, and $r(s;\theta)$ for the last-token hidden state of prompt $s$. The claim is that for any two distinct prompts $s \ne s'$, $$ \Pr_\theta\big[\, r(s;\theta) = r(s';\theta) \,\big] = 0. $$ The argument runs through **real-analyticity**. Embedding lookups, affine projections, softmax, LayerNorm with $\varepsilon > 0$, and every real-analytic activation in common use (GELU, SiLU, SwiGLU, GeGLU) are all real-analytic functions of their inputs and parameters. Real-analytic functions are closed under addition, multiplication, division away from poles, and composition — so the entire network, prompt fixed, is a real-analytic function of $\theta$. That matters because of a classical dichotomy. Fix two prompts $s \ne s'$ and define $h(\theta) = \lVert r(s;\theta) - r(s';\theta) \rVert^2$. Because $h$ is real-analytic, exactly one of two things is true: either $h \equiv 0$ everywhere, or the zero set $\{\theta : h(\theta) = 0\}$ has Lebesgue measure zero — a thin, lower-dimensional slice of parameter space, not a region with any volume. The paper rules out $h \equiv 0$ by hand: it exhibits one concrete $\theta$ where $s$ and $s'$ provably map to different states (for instance, freeze the network down to embeddings plus positions and point at two distinct rows). So the collision set for that pair is measure zero. That picture is the whole argument. Any initializer with a continuous density — Gaussian, uniform, Xavier — places exactly zero probability mass on a measure-zero set, so at initialization the odds of landing on the collision curve are zero. Training doesn't change that: a gradient step $\varphi(\theta) = \theta - \eta \nabla L(\theta)$ is itself real-analytic, so its Jacobian determinant $\det D\varphi(\theta)$ is real-analytic and not identically zero, which makes $\{\theta : \det D\varphi = 0\}$ measure zero too. Away from that set, the inverse function theorem says $\varphi$ is a local diffeomorphism, and a diffeomorphism cannot squash a positive-volume region down onto a lower-dimensional set. Push an absolutely-continuous parameter distribution through enough of these steps — full-batch, mini-batch, or even adversarially chosen batches — and it stays absolutely continuous. Induct over any finite training horizon $T$ and injectivity holds with probability one after training, not only at init. The same argument extends to any finite set of prompts being pairwise distinct simultaneously, not just one pair at a time. The theorem is explicit about how a collision *would* have to happen: two vocabulary items given exactly identical embedding rows, or two positional encodings set exactly equal by hand while everything else is tuned to suppress positional signal. Both are measure-zero, hand-engineered pathologies — never reached by continuous initialization plus gradient descent, but not physically impossible if someone builds them on purpose. That is exactly what "hand-engineered" does in the picture above. ## What "almost surely" does and doesn't buy you This is worth being precise about, because it is the obvious objection. "Probability one" is a statement about a *distribution* over parameters — the set of continuous initializers, pushed through finitely many gradient steps. It is not a certificate stamped on any one specific, already-trained, already-quantized checkpoint sitting on a GPU. A model built by deliberately engineering a collision (identical embedding rows, say) would sit exactly on the measure-zero set and would not be covered by the "almost surely" — the theorem says that model is vanishingly unlikely to arise by chance, not that no such model can exist. The other half of the objection is floating point. The proof is a statement about real numbers; a GPU computes in fp16, bf16, or fp32 with rounding at every step. The paper's own practical collision test uses `torch.allclose` with `rtol=1e-5, atol=1e-8` — a floating-point *tolerance* check, not exact real-number equality. So what gets verified empirically is "no near-collisions above this threshold," which is a weaker, computational claim standing in for the idealized real-valued one. The paper doesn't claim its proof covers the discretized forward pass directly — only the empirical tests do, and only up to that tolerance. The empirical side is where finite precision gets tested directly. Table 2 measures the minimum pairwise distance at the final layer under FP4, INT8, and FP32 for three models — for Llama-3.1-8B: **2.281 (FP4) · 6.597 (INT8) · 1.274 (FP32)**. Quantization didn't shrink the separation margin in these tests; if anything the coarser formats measured *larger* minimum distances. That's reassuring, but it's a different computational object than the real-analytic map the theorem is stated over, checked empirically rather than derived from the proof. ## Zero collisions, at a scale that matters
A separate run — 100,000 prompts sampled from Wikipedia, C4, The Pile, and GitHub Python, roughly 5 billion pairwise comparisons — measured the minimum pairwise distance across four different models at three depths (layer 1, the middle layer, the last layer), against a collision threshold of $10^{-6}$: | Model | Layer 1 | Layer L/2 | Layer L | |---|---|---|---| | Llama-3.1-8B | 0.001 | 0.129 | 0.620 | | Mistral-7B-v0.1 | 0.002 | 0.187 | 1.274 | | Phi-4-mini-instruct | 0.014 | 1.336 | 9.020 | | TinyStories-33M | 0.029 | 1.434 | 2.793 | Zero collisions, and the separation grows by roughly two to three orders of magnitude from the first layer to the last — consistent with the boxplot above. The closest pairs the authors found anywhere, on manual inspection, were near-duplicate code and documentation snippets differing only by a trailing newline token — still far above the threshold. Then the authors went looking on purpose. They took the ten closest prompts in their sample and appended *every* vocabulary token as a one-token continuation to each, an exhaustive collision hunt rather than a random sample: **over 343 billion prompt pairs per model**. Zero collisions, on both GPT-2 Small and Gemma3-1B. That is the number in the abstract, and it is worth being precise about its scope: it is the most exhaustive test in the paper, and it ran on two of the *smaller* models tested. The bigger models — Phi-4 (14B) and Llama-3.1-70B — were checked under the sampled protocol above (Table 3), not the exhaustive one; injectivity at 70B+ scale rests on the same theorem plus a smaller, sampled empirical check, not the 343-billion-pair stress test.
Two more things the tests deliberately probed, both reported candidly rather than cherry-picked: separation does not shrink as prompts get longer (above), and — a genuinely counter-intuitive result the authors report without softening it — inverting **random, out-of-distribution token sequences is faster than inverting natural language** (146s vs. 107s mean, GPT-2, 100-token prompts). Their read: natural-language hidden states sit on a more structured, clustered manifold, which is flatter and worse-conditioned for the gradient-guided search described next; OOD states are more dispersed, giving sharper gradients to follow. One distinction worth making explicit, because it's easy to blur: injectivity is a claim about the **hidden state**, not about what a model eventually says. Two prompts can produce the identical next-token answer — "the sum is 12", a translation landing on the same word, a completion ending in "dog" — while their hidden states remain measurably distinct underneath. The paper stress-tests exactly this: translation pairs, arithmetic pairs, and ten thousand different Wikipedia prefixes all forced to the same fixed suffix and the same output token all still show real, measurable separation at the hidden-state level, growing with depth just like everything else. Collapsing to the same output is common and expected; collapsing to the same internal state is what the theorem rules out. ## SipIt: turning the proof into an algorithm Injectivity is a static fact about the map. **SipIt** (Sequential Inverse Prompt via ITerative updates) is what you get when you notice the map is also *causal*: the hidden state at position $t$ depends only on the prefix already fixed and the token at $t$. That means the one-step map $v_j \mapsto h_t(\pi \oplus v_j)$, for a fixed correct prefix $\pi$ and candidate $v_j$ ranging over the vocabulary, is itself almost-surely injective — the same argument, run one position at a time. So the algorithm doesn't need to solve the whole sequence at once: ``` for t = 1..T: for each candidate v_j (policy: gradient-guided, or random): if the candidate's predicted hidden state matches h_t within tolerance ε: append v_j to the reconstructed prefix; move to position t+1 ``` **Correctness (Theorem 3.1):** this recovers the true sequence with probability 1, in at most $T \cdot |\mathcal{V}|$ candidate checks in the worst case — linear in the prompt length $T$ for a vocabulary of fixed size, which is the "linear time" the abstract promises. **Robustness (Theorem 3.2):** it still recovers the exact sequence under bounded perturbation of the observed state, as long as the perturbation stays under half the minimum pairwise distance among candidate continuations at that step — which is exactly the separation margin measured above, and exactly why that margin *growing* with depth matters practically, not just theoretically. In practice SipIt doesn't try candidates in vocabulary order. It uses a **gradient-guided policy** — clip the gradient norm to 1, periodically re-project the running estimate back to the nearest true token embedding every 50 proposals — rather than the brute-force random order its own ablation uses as a baseline. And it is explicit about its threat model: it assumes an attacker who already holds the **full per-position hidden-state sequence at some layer** $\ell$ — the paper's own examples are "a leaked KV-cache, a shared-inference pipeline, or an API exposing intermediate activations." Recovering a prompt from *only* the final embedding is asserted to be theoretically possible under the same theorem, but no efficient algorithm for it is demonstrated — that's left as future work. Everything below is about the case SipIt actually solves: someone already has the hidden states. Also worth naming: Thomas et al. (2025), the paper's own "most closely related" citation, recovers prompts from hidden states with a similar sequential structure but without an injectivity guarantee behind it — so it has to score close to the entire vocabulary at each step before committing to a token. SipIt's early exit, explored in under a quarter of one percent of the vocabulary in these experiments, is what the guarantee buys on top of the same basic idea. ## How fast, and how much of the vocabulary On 100 prompts (90% real sentences, 10% random tokens), 20 tokens each, GPT-2 Small: | Method | Mean time (s) | Accuracy | |---|---|---| | HardPrompts (gradient prompt search) | 6132.59 ± 104.61 | 0% | | BruteForce (SipIt, random-order ablation) | 3889.61 ± 691.17 | 100% | | **SipIt** (gradient-guided) | **28.01 ± 35.87** | **100%** | HardPrompts — the standard gradient-based *approximate* prompt-search baseline, adapted by the authors from its original vision-language objective to a text-only one — never lands on the exact sequence: it optimizes toward *a* prompt, not *the* prompt. Brute-force random search gets there eventually, at over two minutes an average token. SipIt matches brute force's accuracy at roughly **1/140th the time**, purely by trying candidates in a smarter order. That gap holds up against real vocabularies, not just GPT-2's ~50K tokens. Under FP4 quantization, 50 prompts, 10 tokens each: | Model | Vocab size | Accuracy | Time (s) | Vocabulary explored | |---|---|---|---|---| | Mistral-7B-v0.1 | 32,000 | 100% | 111.78 ± 46.50 | 0.19 ± 0.08% | | Llama-3.1-8B | 128,255 | 100% | 549.48 ± 265.75 | 0.21 ± 0.10% | The unquantized appendix ablation lands within noise of the same numbers (0.21% and 0.22% explored respectively) — quantizing the model barely moves how much of the vocabulary SipIt has to touch. Every measurement here is a single NVIDIA A100-SXM (64GB), no custom kernels for SipIt itself — the ~28-second figure is an unoptimized, single-GPU number, not a lower bound on how fast this can go. ## The consequence: a KV-cache is the prompt, in another basis Put the two halves together. The hidden state is provably (almost surely) a lossless encoding of the prompt that produced it, and there is a linear-time algorithm that decodes it back exactly, needing only a sliver of the vocabulary and no training data of its own. That means the sentence "the model doesn't store your prompt, it just computes with it" is not true in the way people mean it. The computation *is* the storage. A hidden state is not a fingerprint or a hash of the input — it's the input, run through an invertible function. This is precisely what a [KV-cache](/articles/how-llm-inference-works) is built from. Every serving stack that skips recomputing attention for tokens it has already seen — prefix caching in vLLM and SGLang, multi-tenant inference sharing a cache across requests, cache offload from GPU to CPU DRAM or disk when memory is tight, activation logging for debugging or evals — is, under this paper's result, holding recoverable prompt text, not an opaque compressed artifact. One nuance worth being exact about: SipIt's proven target is the residual-stream hidden state $r(s;\theta)$ itself, not the $K$ and $V$ tensors a serving stack actually caches — those are per-head linear projections of that hidden state, generically lower-dimensional per head. But stacked across every layer and every head, a full KV-cache is a far *higher*-dimensional linear image of exactly the same per-position hidden-state sequence the paper's own threat model names as its motivating example: "a leaked KV-cache, a shared-inference pipeline, or an API exposing intermediate activations." The paper doesn't run SipIt against raw $K$/$V$ tensors — it inverts hidden states directly — so read "the cache is invertible" as the natural extension the authors themselves point at, not a number they measured. A concrete, current example: [Kimi K3](/articles/kimi-k3)'s reinforcement-learning infrastructure writes idle KV prefixes out to an external CPU DRAM pool between rollouts, so paused sandboxes stay cheap. Under this paper's result, that pool isn't holding compressed activations — it's holding recoverable prompt text, sitting outside the GPU's usual trust boundary, on a different piece of hardware entirely. That's not a criticism specific to K3 — it's the same design every prefix-cache and cache-offload system makes for the same performance reasons — it's just a live instance to point at. And compression doesn't obviously buy you out of this: quantizing the cache, as [TurboQuant](/articles/turboquant-kv-cache) does for entirely different (memory and throughput) reasons, doesn't collapse distinct prompts into a shared entry either — Table 2 above shows minimum pairwise distances at FP4/INT8 holding up or growing relative to FP32. Shrinking the cache for efficiency and erasing what's recoverable from it are different problems, and solving the first doesn't solve the second. There's a regulatory angle here too, which the paper raises directly. The Hamburg Data Protection Commissioner argued in 2024 that a model's trained *parameters* aren't personal data, because training folds the data into abstract, non-retrievable representations. The authors' point is narrower and, on their result, correct as far as it goes: that argument is about parameters at rest, not about **inference-time hidden states**, which this paper shows are lossless, recoverable encodings of whatever a specific user typed, right now. Any system that stores or transmits those states — including as a cache — is storing something closer to the original text than "abstract representation" suggests. (If you've read the [Jacobian lens](/articles/jacobian-lens) piece: that method also reads information out of a hidden state, but by linearizing the model around a corpus average — an approximation. SipIt's guarantee is exact, because it has an injectivity theorem underneath it instead of a linear approximation.) ## How new is this, and what to weigh The paper is genuinely recent — submitted October 2025 — but it isn't a fringe preprint sitting uncited. It's accepted at **ICLR 2026**, and by the time of writing has around **30 citations**, **7** of them flagged "influential" by Semantic Scholar, which is a fast citation trajectory for a ten-month-old theory paper. Code (`SIPIT`) is public. A few things to weigh before taking every number at face value. **Baselines are re-implemented, not re-run verbatim**: HardPrompts is the authors' own adaptation of a gradient prompt-search method originally built for vision-language models, ported to a text-only $\ell_2$ objective — its 0% accuracy reflects that adaptation, not a hostile misreading of someone else's code. **All numbers are self-reported**; no independent third-party reproduction exists yet at ten months old. **Every timing number is single-GPU, no custom kernels** — treat "28 seconds" as this implementation's number, not a hardware-independent constant. And the theorem's scope is decoder-only transformers with real-analytic activations — the paper surveys 18 widely-used LLMs and finds all 18 use real-analytic FFN activations (SwiGLU, SiLU, GeGLU, GELU), but a classic ReLU network sits outside the theorem's direct coverage, since ReLU isn't real-analytic. None of that undermines the core claims — billions of comparisons across multiple independent experimental setups, a working algorithm with two proven theorems behind it, and honest reporting of the results that don't flatter the paper (OOD prompts inverting faster than natural language; the biggest models tested under the less exhaustive protocol). It's the right amount of scrutiny for a result this consequential, not a reason to discount it. ## The take The intuition that hidden states are lossy comes from staring at individual layers — LayerNorm, softmax, low-rank projections — each of which really is lossy on its own. The paper's point is that lossiness doesn't compose the way that intuition assumes: the full map from prompt to last-token state is, almost surely, injective, and an algorithm exists that inverts it exactly, in linear time, using a sliver of the vocabulary. The finite-precision and hand-engineered caveats are real, and worth stating precisely rather than waving away — but they don't touch the core result, which is that a hidden state is not a summary of a prompt. It's the prompt. Anything built to store, cache, offload, or ship hidden states around — for speed, for multi-tenancy, for debugging — is, whether it says so or not, in the business of storing exact user text. --- *Built on [Language Models are Injective and Hence Invertible](https://arxiv.org/abs/2510.15511) (Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà; EPFL / Sapienza University of Rome / University of Athens / Archimedes, Athena RC; accepted ICLR 2026), and its [SipIt code release](https://github.com/giorgosnikolaou/SIPIT). Figures are the paper's own Figures 3 and 9, reproduced for commentary. Tables and numbers are the authors' except where marked as this site's own illustrative simplification (the SipIt walker's compressed vocabulary); interactive diagrams are mine.* --- # Inkling-Small: what on-policy distillation actually buys a reasoning model > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/inkling-small > date: 2026-08-03 > tags: llm, mixture-of-experts, distillation, reinforcement-learning, explainer Thinking Machines' [Inkling](/articles/inkling) shipped with an unusually candid pitch: *"not the strongest overall model,"* a broad multimodal base meant for fine-tuning. **Inkling-Small** is the smaller sibling promised in that release, and its own pitch is just as specific: *"an efficient open-weights model that achieves comparable performance to Inkling at a quarter of its size."* Same architecture family, same 256-expert MoE backbone, same 1M-token context — but a materially different post-training story. Inkling-Small was **post-trained from an earlier checkpoint using on-policy distillation with Inkling as the teacher**, then pushed through **two weeks of scaled agentic-coding RL**. The result, per Thinking Machines: Inkling-Small now **surpasses Inkling** on reasoning and agentic coding benchmarks, while Inkling **keeps the edge on knowledge and factuality**. That is the interesting part of this release — not the size, the recipe. On-policy distillation is a genuinely different training signal from ordinary distillation, and it is worth being precise about why. Along the way there is also a smaller, checkable finding: the parameter counts both Thinking Machines channels quote for Inkling-Small and Inkling do not match what the released weights actually contain. ## Same backbone, one size down Inkling-Small shares its architecture with Inkling almost feature-for-feature — the [full mixture-of-experts and hybrid-attention design is covered in the Inkling piece](/articles/inkling), so here is just the shape, read from each model's `config.json`: | | Inkling-Small | Inkling | |---|---|---| | Hidden size | 4096 | 6144 | | Layers | 42 | 66 | | Attention heads / KV heads | 32 / 8 | 64 / 8 | | Sliding-window heads / KV heads | 32 / 8 | 64 / 16 | | Sliding-window size | 512 | 512 | | Routed experts | 256 | 256 | | Active experts / shared experts | 6 / 2 | 6 / 2 | | MoE expert intermediate size | 2048 | 3072 | | Dense-layer intermediate size | 16384 | 24576 | | Context length | 1,048,576 | 1,048,576 | | Multi-token-prediction heads | 8 | 8 | Same routing scheme (256 routed experts, 6 active, 2 always-on shared experts, sigmoid router with post-top-k norm), same [relative position bias](/articles/how-llm-inference-works) and short-conv mixing, same encoder-free image/audio path, same million-token context. Inkling-Small is a narrower, shallower cut of the identical design — fewer layers, a smaller residual stream, and a tighter expert width. What changes is everything downstream of pretraining. ## Off-policy vs on-policy: who generates, who scores Ordinary distillation — call it off-policy — has the teacher generate the training data. The teacher produces a sequence of tokens (an answer, a reasoning trace, a full trajectory), and the student is trained by cross-entropy to reproduce those exact tokens. It is imitation: match the teacher's output distribution on the teacher's own text. That works fine for short, single-step outputs. It runs into a specific problem for anything autoregressive and long — like a chain of reasoning. At inference time the student has no teacher transcript to fall back on; it has to sample its own next token from its own distribution, condition on that, sample the next one, and so on. The moment a sampled token differs even slightly from what the teacher would have produced at that step, the student is in a state its training never covered — and every token after that is generated conditioned on an increasingly unfamiliar prefix. This is **exposure bias**: errors compound because the training signal only ever showed the model teacher-generated prefixes, never its own. **On-policy distillation** removes the reference trajectory entirely. The *student* generates its own rollout, token by token, under its own policy — and the teacher's only job is to score the tokens the student actually produced (as a per-token reward or a log-probability target, depending on the recipe). There is no teacher transcript to drift away from, because training never showed the student one. Whatever state the student's own sampling puts it in, that is exactly the state it gets graded and corrected in. Drag through the two modes below and scrub the rollout step to see the difference concretely: This is precisely why on-policy distillation matters more for a reasoning model than for a plain chat model: a reasoning trace is long, autoregressive, and self-referential — later steps depend directly on the model's own earlier steps. A student trained only to imitate a teacher's specific path is fragile exactly where it counts, the moment its own sampling wanders off that path. A student whose own rollouts are the only thing ever scored has no such cliff to fall off. Thinking Machines describes the Inkling-Small recipe directly: *"we post-trained an earlier checkpoint, Inkling-Small (preview), in part using on-policy distillation with Inkling as the teacher. Starting from that checkpoint, we continued scaling agentic coding RL for two weeks."* Inkling — the larger, already-trained sibling — is the sole teacher; the smaller model generates, Inkling grades. This site has covered two other takes on the same idea, and the contrast is worth naming. [Kimi K3's post-training](/articles/kimi-k3#post-training-nine-experts-then-one) trains **nine** separate RL specialists (three domains times three effort levels) and then uses **Multi-Teacher On-Policy Distillation** to collapse all nine back into one shipped checkpoint — many teachers, all of them versions of the model itself. [Agents-A1](/articles/agents-a1) does something similar with six domain specialists, routing each training trajectory to the one teacher that owns its domain. Inkling-Small's version is the simplest point in that space: **one** teacher, and it is not a specialist expert of the student — it is a wholly separate, larger, already-shipped model. Same underlying mechanism (student generates, teacher scores the student's own tokens), different teacher cardinality and a different relationship between student and teacher. ## Two weeks of RL — read the disclosure level honestly After the on-policy distillation stage, Thinking Machines says it *"continued scaling agentic coding RL for two weeks."* Read that number for what it actually specifies and what it does not. It tells you the wall-clock duration of one training phase. It tells you nothing about cluster size, GPU count, rollout throughput, number of environments, or total compute — so "two weeks" from a 64-GPU pod and "two weeks" from a full GB300 NVL72 cluster are the same sentence describing wildly different amounts of work. [Scaling agentic RL](/articles/scaling-agentic-rl) is mostly an environments-and-infrastructure problem — verified, reproducible task environments at scale is usually the actual bottleneck, not algorithm novelty — and none of that infrastructure detail is disclosed here either: no environment count, no rollout count, no reward model description beyond "agentic coding." Compare that to Kimi K3's post-training write-up, which at least names concrete infrastructure numbers (sandbox counts, checkpoint latencies) for its agentic RL stage. Thinking Machines' own blog names the training hardware for the base models (NVIDIA GB300 NVL72) but not specifically for this RL phase. Two weeks is a real number and a real signal that the recipe kept running rather than stopping early — it is just not, by itself, a compute disclosure. ## The parameter count: stated vs measured Both Thinking Machines channels — the announcement blog and the Hugging Face model card — quote the same rounded parameter counts for both models: > "Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, > trained on NVIDIA GB300 NVL72 systems." — Thinking Machines blog > Params (B) (activated/total): Inkling-Small "12/276", Inkling "41/975" — HF model card, > evaluations table Fetching each repository's `safetensors` metadata directly from the Hugging Face API (`api/models/thinkingmachines/{Inkling-Small,Inkling}`, checked 2026-08-03) gives a different number — the literal count of parameters in the released weight files: | | Stated (blog + HF card) | Measured (HF `safetensors.total`) | Difference | |---|---|---|---| | Inkling-Small | 276B total | **265,956,439,090** (≈265.96B) | +10.04B, ≈3.8% above measured | | Inkling | 975B total | **952,377,623,626** (≈952.38B) | +22.62B, ≈2.4% above measured | No accusation implied here — both numbers come from official Thinking Machines channels, and this is simply what the weight files measure against what both channels quote. It is consistent across both models and both channels, so it reads as a rounding-and-carry-forward convention rather than a one-off typo. The active-parameter figures (12B / 41B) cannot be checked the same way — they describe how many parameters fire per token, which depends on live MoE routing at inference and cannot be read off static weight metadata. Take those as self-reported. The **"a quarter of its size"** framing is worth checking on its own terms too. A literal quarter means Inkling should be 4x Inkling-Small. On the stated numbers, 975 ÷ 276 ≈ 3.53x; on the measured numbers, 952.38 ÷ 265.96 ≈ 3.58x. Either way, Inkling-Small is closer to 28% of Inkling's size than 25% — "a quarter" is a round-down of a real but smaller ratio, not a precise figure. One more data point that tracks the same rough ratio: Tinker's stated output pricing is $1.20 per million tokens for Inkling-Small against $4.05 for Inkling — Inkling-Small at about 30% of Inkling's price, in the same neighborhood as the ≈28% size ratio. ## Benchmarks — where it wins, and where it doesn't Thinking Machines' own evaluation suite backs the headline claim: Inkling-Small beats its own larger sibling on most reasoning, coding, and agentic benchmarks. That pattern — Small ahead of its own larger sibling, and ahead of every similarly sized open peer Thinking Machines tested — holds cleanly on SciCode, GPQA Diamond, ARC-AGI-1/2, CritPt, and Toolathlon Verified. It is not universal: on **SWE-Bench Pro** Inkling-Small (55.9%) sits in a three-way near-tie, edged out slightly by MiMo V2.5 (56.1%) and Minimax M2.7 (56.2%) even as it still beats its own sibling Inkling (54.3%). And it does not hold at all on knowledge-recall tasks. Thinking Machines states that exception directly: *"Inkling maintains an advantage on knowledge coverage and factuality."* SimpleQA Verified is the sharpest case: Inkling-Small loses to its larger sibling by more than 23 points here, and a similar gap shows up on **AA Omniscience** (Inkling-Small −9.0 vs. Inkling +2.1). Both are pure knowledge-recall benchmarks, not reasoning or agentic ones — exactly where the blog's caveat says to expect the loss, and the evaluation table backs it up cleanly. The next-largest gaps are **Tau³ Banking** (15.5% vs. 23.7% — an 8-point, roughly one-third relative deficit) and **FORTRESS adversarial** safety (71.6% vs. 78.0%); smaller, single-digit-point trails also show up on **AIME 2026**, **Global-MMLU-Lite**, and the multimodal audio/voice suite. None of these are reasoning or agentic benchmarks either — they cluster around knowledge, safety-adversarial robustness, and multimodal recall, consistent with "knowledge and factuality" being the one axis the bigger model still owns. **Scope the evaluation methodology, not just the scores.** All of this is Thinking Machines grading its own model against a provider-selected comparison set (Qwen3.5, MiMo V2.5, Minimax M2.7, DeepSeek V4 Flash, Nemotron 3 Ultra, plus closed models Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna) — a self-report, not a neutral third-party leaderboard. The card itself discloses several protocol caveats worth carrying forward: SWE-Bench Verified and Terminal-Bench 2.1 use an **internal, bash-only harness**, and external models' numbers on those two are self-reported by their own vendors, not run in-house. Terminal-Bench 2.1 also zeroed out "a small number of solutions... found to be contaminated from web search." HLE-with-tools numbers for MiniMax M2.7, Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna were run in-house by Thinking Machines, not vendor-reported. None of this invalidates the results, but it means the comparison set and the harness are both chosen by the same lab whose model is winning most of the charts. ## The take Inkling-Small is a useful data point for a specific question: what does distillation from a bigger sibling actually buy you, mechanically? The answer here is not "compress the teacher's knowledge into a smaller container" — Inkling-Small is clearly *worse* at raw factual recall than Inkling, which is exactly what you would expect if the distillation target was never "know what the teacher knows." The target was "generate reasoning and agentic trajectories the teacher scores well" — and on-policy distillation is the mechanism that makes that the actual training signal, because it grades the student's own rollouts instead of teaching it to imitate someone else's. Layer two weeks of agentic-coding RL on top of that checkpoint and the result tracks: gains concentrate exactly in reasoning and agentic coding, and the one place the recipe doesn't touch — static factual knowledge — is the one place the bigger sibling keeps its lead. The parameter-count gap is a smaller story, but it is the kind of thing worth checking rather than repeating: two official channels, one consistent 2.4-3.8% overstatement, verifiable in about two API calls. None of it changes what Inkling-Small actually is — an Apache-2.0, genuinely open-weights model that beats its own much larger sibling on most reasoning and coding benchmarks. It is just a reminder that "check the primary source" is worth doing even when the primary source is the model card itself. --- *Sources: the [Inkling-Small announcement](https://thinkingmachines.ai/news/inkling-small/) and the [Hugging Face model card](https://huggingface.co/thinkingmachines/Inkling-Small) (architecture, training recipe, evaluations, pricing), cross-checked against the [Inkling flagship card](https://huggingface.co/thinkingmachines/Inkling) and this site's [Inkling piece](/articles/inkling). Parameter counts were independently verified via the Hugging Face API's `safetensors.total` field for both repositories on 2026-08-03, not taken from either card. All benchmark numbers are Thinking Machines' own, on their own evaluation suite, with the harness caveats noted inline. Related reading: [Kimi K3's Multi-Teacher On-Policy Distillation](/articles/kimi-k3#post-training-nine-experts-then-one), [Agents-A1's domain-routed on-policy distillation](/articles/agents-a1), [mixture-of-experts from scratch](/articles/mixture-of-experts-from-scratch), and [scaling agentic RL](/articles/scaling-agentic-rl). Neither Thinking Machines source publishes a static architecture or benchmark figure for this release — the blog's charts are rendered client-side from inline data, not static images — so the diagrams here are original illustrations of the mechanism, not reproductions of a paper figure.* --- # Instella-MoE: a 16B MoE that never touches an NVIDIA GPU > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/instella-moe > date: 2026-08-03 > tags: llm, mixture-of-experts, amd, rocm, explainer Every big open MoE release from the last two years shares one unstated assumption: it was trained on NVIDIA GPUs. [AMD's Instella-MoE](https://github.com/AMD-AGI/Instella-MoE) breaks that assumption on purpose. It is a **16B-total, 2.8B-active** Mixture-of-Experts model, and every stage of it — pretraining, mid-training, long-context extension, SFT, DPO, RL — ran on **AMD Instinct MI300X and MI325X** GPUs under ROCm, with nothing borrowed from a CUDA cluster. AMD [released six checkpoints](https://huggingface.co/collections/amd/instella-moe), one per stage, plus the training, inference, and RL codebases, under the tagline "fully open" — a word this piece is going to hold them to.
Two things make this worth a full read rather than a spec-sheet skim. First, the architecture has a real new idea in it: **Gated MLA**, a per-token gate on the attention output, which turns out to be the same move [Kimi K3](/articles/kimi-k3) landed on independently. Second, the systems story — **FarSkip-Collective** — is a genuinely interesting trick: it makes MoE training and serving faster by *deliberately* feeding the model stale, outdated activations. That sounds like a bug. It is the entire point. ## Six checkpoints, all the way down Most "open" model releases mean one thing: a final `safetensors` file and a model card. Instella-MoE releases the **entire pipeline** — a checkpoint at every stage, not just the one you'd chat with: `Instella-MoE-16B-A3B-Pretrain` → `-Midtrain` → `-Base` → `-SFT` → `-DPO` → `-Think`. Every one of those six ships with its own weights, training config, and a named, token-counted data recipe. That is a materially different claim from "we released the weights." Pick a stage below and see exactly what shipped for it: AMD draws its own line here, and it is a useful one: in the blog post the fully-open bucket is **OLMo-3, SmolLM3, and OLMoE** — full data, full code, checkpoints along the way — while **Moonlight-16B-A3B, Qwen3.5, and Gemma-4** get filed as "open-weight": you get the final weights and usually a report, not the recipe. Instella-MoE places itself in the first group, and the checkpoint list above is the receipt. What is genuinely missing, and it matters: **no training compute figure, no cluster size, no wall-clock duration** anywhere in the repo or the blog. "Fully open" usually implies you could, in principle, reproduce the run — and the one number every reproduction attempt needs first is exactly the one AMD didn't publish. ## The shape of the model Strip away the training story and here is what actually runs at inference time: | | | |---|---| | Total / active parameters | **16B / 2.8B** | | Layers | 27 decoder layers | | Hidden dimension | 2048 | | Attention | Gated Multi-head Latent Attention (Gated MLA) | | MoE routing | 2 shared experts (always on) + 6 of 64 routed experts (top-6) | | Pretraining objective | next-token + Multi-Token Prediction | | Tokenizer / vocabulary | DeepSeek-V3 tokenizer · 128,896 tokens | | Context | 4K pretrained → 64K via YaRN + document masking | | Training frameworks | Primus (Megatron-LM based) · Miles (RL, SGLang + Slime) | | Inference | SGLang v0.5.9 with FarSkip-Collective overlays | | Hardware | AMD Instinct MI300X + MI325X, ROCm | The MoE layer is a shared-plus-routed design: 2 experts run on every token no matter what (the model's general-purpose knowledge), and a router picks 6 more out of 64 candidates per token — the same [top-*k* routing](/articles/mixture-of-experts-from-scratch) and shared-expert idea DeepSeek-MoE popularized, applied at a 16B/2.8B ratio. The pretraining objective adds [Multi-Token Prediction](/articles/multi-token-prediction) on top of ordinary next-token loss, DeepSeek-V3 style — training the model to predict a short run of future tokens, not just the next one, which both improves the base model and hands you a natural draft head for speculative decoding later. ### Gated MLA: a per-token filter on attention Standard Multi-head Latent Attention compresses the KV cache into a small latent vector, then reconstructs keys and values from it — cheap to store, same attention math otherwise. Gated MLA adds one more piece: after attention produces its output, a **dedicated linear projection reads the input token and produces a gate**, one value per channel, and that gate multiplies the attention output **before** it goes through the final output projection. Concretely: attention answers "what did this token look up." The gate answers a second, separate question — "how much of what it found is actually worth keeping" — and answers it per channel, per token, learned from data. AMD's own framing is direct: the gate lets the model "selectively attenuate low-utility attention responses for each token." Attention decides what to look at; the gate decides how much of the answer to trust. The reason this is worth pausing on: **the same idea shows up independently in [Kimi K3](/articles/kimi-k3)**, which also augments MLA with an input-dependent output gate — Moonshot's version is a heavier, full-rank gate; AMD's is a single lightweight linear projection. Two labs, thousands of miles apart, training on different hardware stacks, converged on "put a learned gate after MLA's output" as a cheap way to buy expressivity. When two independent teams reach for the same fix, that is usually a sign the fix is addressing something real in the base mechanism, not a one-off trick. ## FarSkip-Collective: paying with staleness to buy overlap Here is the problem FarSkip-Collective solves. In expert-parallel MoE training, each MoE layer's routing decision depends on that layer's own, freshly-computed attention output. Once the router picks experts, the tokens have to physically move across GPUs to wherever their chosen experts live — an **all-to-all** collective. That communication cannot start until the fresh activation exists, and the expert compute that follows cannot start until the communication finishes. Compute and communication are chained, not parallel, and on a large expert-parallel cluster that chain is expensive: the GPUs sit idle every time the network is busy, and vice versa. FarSkip-Collective's fix is to break the dependency that causes the chain. Instead of routing on the fresh activation, it deliberately routes the MoE (and attention) sub-blocks on an **outdated, partial activation** — a slightly stale copy of the signal that was already available earlier. Stale data has one property fresh data doesn't: it's already sitting there, so the communication that depends on it doesn't have to wait for this layer's compute to finish. It can start **alongside** that compute instead of after it. That is the whole trick, and it generalizes past this one model: the separate FarSkip-Collective paper (Dukler et al., MLSys 2026) reports **97.3%** prefill communication-computation overlap and **88.9%** training all-to-all overlap, validated on models from 16B up to 109B parameters — including converting Llama 4 Scout to the FarSkip architecture via self-distillation and landing within **1%** of the original's accuracy. For Instella-MoE specifically, AMD reports the trade paid off exactly as advertised:
**+12.7%** pretraining throughput from overlapping expert-parallel communication, and **up to 39.2% lower Time to First Token** when serving with expert parallelism — a systems win that costs nothing in serial correctness, because the model is trained end-to-end to expect stale inputs at those points rather than having staleness bolted on after the fact at serving time. ## Where it lands AMD ran its evaluations through **OLMES**, Allen AI's open evaluation harness — a real third-party framework, even though AMD is the one running it. On standard benchmarks, the base checkpoint lands second among six comparably-sized models, ahead of every "fully open" peer: Instella-MoE-Base runs at **2.8B active parameters** — less than every model above it except Moonlight, and well under OLMo-3-7B's 7B dense parameters. It also leads all six on `WinoGrande` at **86.5**, and posts `HumanEval+` **65.7**, a solid coding number for a base checkpoint that hasn't seen SFT yet. Long context is where the honest counterexample lives. At 64K tokens on HELMET and RULER, the **dense** 7B OLMo-3 actually wins: | Model | HELMET avg | RULER avg | |---|---|---| | OLMo-3-7B (dense) | **43.1** | **80.2** | | Instella-MoE-Base | 41.5 | 79.4 | | SmolLM3-3B-Base | 37.6 | 78.6 | AMD reports this without burying it. A sparse 2.8B-active model narrowly losing to a dense 7B on long-range retrieval is a plausible, checkable result, not a suspicious one — and it's a useful reminder that "active parameters" isn't the only variable that determines long-context strength. After SFT, the post-training funnel adds up:
SFT → DPO → Think is a steady climb, not a single post-training jump, and the RL stage's gain is concentrated exactly where you'd expect from an instruction-following-focused reward: `IFEval` moves from **77.08** after DPO to **83.70** after RL, the single largest jump on the sheet. The RL recipe itself is worth a note for anyone who followed [Rollout Routing Replay](/articles/rollout-routing-replay): Instella-MoE's IF-RL stage uses R3 alongside GRPO/DAPO-style tricks (zero-gradient filtering, active sampling, token-level loss, no KL term) — the same fix for MoE-RL's rollout/training routing mismatch, here as one ingredient in a larger recipe rather than the whole story. A second RL stage, Multi-Teacher On-Policy Distillation, then anchors the model back to both an IF-specialized teacher and the DPO checkpoint, so the instruction-following gain doesn't come at the cost of the math and code ability DPO already had. ## What's still closed "Fully open" is doing real work as a claim here, so it deserves the same scrutiny as any benchmark number. **What's released, precisely:** all six stage checkpoints on Hugging Face; the full training codebase (`Primus`-based, Megatron-LM lineage) under **MIT**; the inference overlays and RL codebase (`Miles`, built on SGLang); per-stage YAML configs; and named, token-counted data mixtures for every stage, down to the individual sub-corpus (`Nemotron-CC-Math-v1`, `Dolma3 Dolmino`, `cranecode`, `instella-gsm8k-synthetic`, and dozens more). **What isn't released:** the model **weights** carry a **Research RAIL license** — academic and research use only, not the permissive MIT the code ships under. So the license itself is split: the code is as open as it gets, the weights are open-but-restricted. Beyond licensing, three things are just missing: training compute and cluster size (undisclosed anywhere), wall-clock training duration, and a peer-reviewed technical report — the citation AMD gives today is for a *different, earlier* dense 3B Instella model plus the separate FarSkip-Collective systems paper, not an Instella-MoE-specific report. The blog itself says the report is "coming soon." There's a methodology gap worth naming too: it isn't stated whether the comparison models' scores (Qwen3.5, Gemma-4, Moonlight, SmolLM3) were re-run by AMD under the same OLMES harness, or taken from those models' own published numbers. Either is defensible, but the blog doesn't say which, so treat cross-model comparisons as directionally trustworthy rather than exactly apples-to-apples. AMD is candid about the rest: the models are released "for research purposes only," explicitly not intended for "safety-critical applications" or "health and medical applications," shipped "without any safety promises," and multilingual ability "has not been tested." That's an unusually direct limitations section for a benchmark-forward launch blog, and it's worth taking at face value rather than reading past it. ## What training off NVIDIA proves, and what it doesn't The part of this release that will get repeated the most is also the simplest to state: a competitive, 16B-parameter MoE, trained through every post-training stage including RL, ran entirely on AMD Instinct hardware. That is a real data point. It says the ROCm software stack — Primus for pretraining, Miles and SGLang for RL and serving, FarSkip-Collective's overlays making expert parallelism efficient on this hardware specifically — can carry a full modern LLM pipeline, not just a pretraining demo. It does not say AMD hardware is cheaper, faster, or as mature to develop against as the CUDA ecosystem for this workload — none of the numbers that would let you compute a cost-per-FLOP or a wall-clock comparison are published. It does not say the result would look the same at 10× the scale. And it's one vendor's own benchmark of its own model on its own hardware, run through a credible third-party harness but not independently reproduced elsewhere yet. What it *does* rule out is the null hypothesis that this can't be done at all outside NVIDIA — six checkpoints and a working RL pipeline are hard to argue with on that specific point, even while the cost question stays open. ## The take Two real ideas, evaluated honestly, and a rare complete-pipeline release: Gated MLA is a cheap, convergent fix (the same one Kimi K3 found independently) for getting more out of MLA's compressed attention; FarSkip-Collective is the more interesting systems idea, because it's not "make the network faster" — it's "make the network's timing not matter" by feeding the model activations that are already a step behind, on purpose. Together with a genuinely complete checkpoint trail — pretrain through RL, not just a final drop — Instella-MoE beats every other "fully open" peer AMD names and trails only a larger, open-weight-only Qwen3.5-4B. The unresolved part is exactly the part AMD chose not to publish: what the whole run cost, on how many GPUs, for how long. Until the technical report lands, that number stays the reader's to estimate, not AMD's to claim. --- *Sources: the [Instella-MoE GitHub repository](https://github.com/AMD-AGI/Instella-MoE) (architecture, training stages, license, data preparation), the [ROCm technical blog](https://rocm.blogs.amd.com/artificial-intelligence/instella-moe/README.html) (benchmarks, FarSkip-Collective and Gated MLA framing, figures), the [Hugging Face model collection](https://huggingface.co/collections/amd/instella-moe) (six checkpoints), and the [FarSkip-Collective paper](https://arxiv.org/abs/2511.11505) (Dukler et al., MLSys 2026 — overlap percentages, cross-scale validation, Llama 4 Scout conversion). Figures reproduced here are the blog's own Figures 1–3. Benchmark numbers are AMD's, via OLMES; the training-cost and cluster-size figures this piece flags as missing are missing because AMD has not published them, not because they were left out here. Interactive diagrams are mine; the FarSkip timeline and gate values are illustrative, not measured traces.* --- # JOSIE-2: a 4M-token fine-tune, and a labeling bug in its own benchmark table > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/josie-2 > date: 2026-08-03 > tags: open-weights, llm, fine-tuning, explainer [JOSIE-2](https://huggingface.co/collections/Goekdeniz-Guelmez/josie-2) is Gökdeniz Gülmez's third-generation personality-tuned model family: three sizes — 2B, 4B, 9B — each fine-tuned from the matching Qwen3.5 base, MIT-licensed, trained entirely on Apple Silicon. The release note makes a specific, checkable claim: *"JOSIE-2-2B-OSS consistently outperforms its 4B base model. JOSIE-2-4B-OSS consistently outperforms its 9B base model."* A 2B model beating a 4B, and a 4B beating a 9B, would be a genuinely notable result. I went and checked it against the model cards and `config.json` directly, because that's the kind of claim worth verifying before repeating. ## What the cards actually say Each `JOSIE-2--OSS` repo carries a `base_model` field in its frontmatter and a `model_name` field in `config.json`. Both agree, on all three repos: **2B is fine-tuned from Qwen/Qwen3.5-2B, 4B from Qwen/Qwen3.5-4B, 9B from Qwen/Qwen3.5-9B.** Same size, every time. There is no size-up training anywhere in the released weights — the release note's "outperforms its 4B/9B base" framing describes a comparison that, per the models' own configs, never happened.
That mismatched "4B" badge on a row that says "Qwen3.5-2B" is the same bug in miniature. The bigger version of it lives in the part of each card that Hugging Face's benchmark widget actually reads — the `model-index` YAML. There, **every one of the three cards labels its baseline row `Qwen/Qwen3.5-4B (base)`, including the 2B and 9B cards**, even though the numbers next to that label differ card to card (82.5/49.1 on the 2B card, 83.4/48.9 on the 4B card, 92.6/69.5 on the 9B card) in a way that only makes sense if each card's numbers really are its own base model's, mislabeled: Toggle between what's published and what `config.json` verifies, and the numbers don't move — only the caption does. That's what makes this read as a copy-paste templating bug rather than a fabricated result: the underlying scores look real and internally consistent per-card, but the machine-readable label attached to them is wrong on two of three cards. The 9B card has a second, unrelated gap: its own reasoning-mode benchmark row is simply blank, published as "comming soon" in the card's own chart — the one number that would most directly support "the 9B model in its best mode," and it doesn't exist yet. I want to be fair to the release here: this reads as a labeling artifact, not a fabricated claim. The scores are plausible and self-consistent within each card. The problem is narrower and more mundane — a shared table template where the label field wasn't updated per model, which happens to be exactly the kind of error you'd only catch by checking `config.json` against what the benchmark table says, which almost nobody does. ## What's actually worth taking seriously None of this makes the underlying work uninteresting. The whole JOSIE-2 family — three sizes — was fine-tuned on **the same dataset**: roughly 3,500 samples, about 4 million tokens total, generated by a pipeline that itself leaned on larger models (Gemma 4 31B, Qwen3.5-9B-Base, GPT-5.4, and an unreleased JOSIE-2-35B-A3B-RTG model) to synthesize training data far more capable than the 4M-token set alone would suggest. That's a striking ratio: a few thousand curated examples, reused across three model sizes, apparently doing real work — the ARC-C and TruthfulQA gains over each model's *actual, same-size* base are consistent and positive across all three sizes, even once you correct the label. The more genuinely interesting finding in the cards is emergent, not benchmarked: JOSIE-2's reasoning traces sometimes "roast" or insult the user mid-thought, and the cards are specific that *"no reward was introduced to enforce a uniformly polite, corporate, or sanitized internal monologue"* — the behavior wasn't trained in, it showed up during RL because nothing trained it out. It's confined to the hidden reasoning trace, not the user-facing reply, and the cards are candid that they can't yet separate "reasoning-first supervision improves the policy" from "the training data's honesty framing drives the gain" — an open question, stated as one. ## The take Check the claim, not just the headline: a 2B model plausibly outperforming its own 2B base and a 9B model plausibly outperforming its own 9B base is still a real, useful result from a tiny dataset — it just isn't the size-up story the release note tells, and the cards' own tables currently say something they don't mean to say. That's worth fixing on Gülmez's end, and worth checking on everyone else's before the "2B beats a 4B" framing gets repeated as fact. --- *Sources: the [JOSIE-2 collection](https://huggingface.co/collections/Goekdeniz-Guelmez/josie-2) and individual model cards (`Goekdeniz-Guelmez/JOSIE-2-2B-OSS`, `-4B-OSS`, `-9B-OSS`), their `config.json` files, and each repo's `benchmarks.png`, cross-checked directly for this piece. The figure is the 2B card's own chart, unedited aside from flattening onto a white background; the interactive reproduces the model-index labels and config-verified bases as published.* --- # KDA has a half-life: linear attention forgets like a radioactive isotope > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/kda-half-life > date: 2026-08-03 > tags: linear-attention, kimi, attention, explainer, math Here is a small thing I keep turning over. The forgetting mechanism inside **Kimi Delta Attention** — the linear attention in [Kimi K3](/articles/kimi-k3) — is *the same mathematics as radioactive decay*. Not "reminiscent of", not "a useful analogy". The same two-line derivation, with tokens where a physicist writes seconds. ## Two laws that are one law A radioactive sample loses a fixed *fraction* of its remaining atoms per unit time. That gives the exponential law everyone meets in school: $$ N(t) = N_0 e^{-\lambda t} $$ A KDA channel loses a fixed *fraction* of its remaining state per token. Take the recurrence and strip it to the decay term — set the write strength to zero and watch a stored value with no new input arriving: $$ S_n = \alpha^{\,n} S_0 $$ Those are the same function. Since $\alpha = e^{\ln \alpha}$, $$ \alpha^{\,n} = e^{n \ln \alpha} = e^{-\lambda n}, \qquad \lambda = -\ln \alpha $$ The retention factor $\alpha$ and the decay constant $\lambda$ are two spellings of one number. A channel with $\alpha$ close to 1 is a long-lived isotope; a channel with small $\alpha$ is one that barely outlives its own creation. ## So it has a half-life Once you accept that, the half-life comes for free. Ask for the $n$ where half the signal is gone: $$ \alpha^{\,n_{1/2}} = \tfrac{1}{2} \quad\Longrightarrow\quad n_{1/2} \, \ln \alpha = \ln \tfrac{1}{2} \quad\Longrightarrow\quad n_{1/2} = \frac{\ln 0.5}{\ln \alpha} $$ That is the whole result, and it is worth internalizing because it converts an opaque hyperparameter into a number with units you can reason about. **α = 0.99 gives a half-life of about 69 tokens.** Not "some decay" — sixty-nine tokens, roughly a long sentence. Drag it: The lever is brutally nonlinear near 1, which is the part worth feeling rather than reading. Going from α = 0.99 to α = 0.999 does not extend memory by a tenth of a percent; it multiplies the half-life by ten, from about 69 tokens to about 693. Each additional nine buys another factor of ten. That is why linear-attention gates are usually parameterized in log space — the useful resolution all lives in the last few decimal places, and a linear parameterization would spend nearly all its range on channels that forget immediately. A useful sanity check: half-life is a property of the *ratio*, not the magnitude. A channel at α = 0.99 has lost half its signal after 69 tokens, three quarters after 138, and about a thousandth of it survives to 690 — ten half-lives, the same "ten half-lives and it's gone" rule of thumb used for isotopes. ## The interesting part: α is per channel If KDA had one global α this would be a cute observation and nothing more. It doesn't. In K3, α is a **channel-wise** vector — the report writes the state update as $$ S_t = \left(I - \beta_t k_t k_t^{\top}\right) \mathrm{Diag}(\alpha_t)\, S_{t-1} + \beta_t k_t v_t^{\top} $$ where $\alpha_t \in (0,1)^{d_k}$ is a **per-channel** one-step retention factor and $\beta_t$ is the delta-rule write strength. `Diag(αₜ)` is the load-bearing notation: every one of the $d_k$ channels gets its own decay constant, so a single head carries a whole spectrum of half-lives simultaneously. This is what makes a fixed-size state genuinely useful rather than merely cheap. The head is not choosing between "remember recent things sharply" and "remember old things vaguely" — it runs both at once, on different channels. The fast channels behave like a local window: they hold the current clause and dump it. The slow channels are closer to a running summary that survives the entire context. Attention over a KV cache gets its long-range recall by *storing everything*; KDA gets a version of it by storing a small number of things at deliberately different rates. It also reframes what "training the gate" means. The model is not learning *whether* to forget. It is learning a distribution of timescales — effectively allocating channels across memory horizons, the way a filter bank allocates across frequencies. ## What K3's config actually pins down The released weights make a couple of things concrete. K3 runs **69 KDA layers out of 93**, three of every four, with a Gated MLA layer as the fourth — so most of the model's sequence mixing is this decay process, and the full-attention layers are the periodic exact-recall anchor. Head dimension is 128, and the gate is full-rank (`use_full_rank_gate: true`) rather than a low-rank approximation — though see the update below for exactly how the per-channel variation is produced. The config also carries `gate_lower_bound: -5.0`. Read as a floor on log-α, that bounds the fastest a channel is allowed to forget: $\alpha \ge e^{-5} \approx 0.0067$, which is a half-life of about **0.14 tokens** — a channel that has essentially dumped its state by the very next step. The ceiling is the interesting end and it is open: as α approaches 1 the half-life grows without bound. To keep half your signal across a full 1M-token context you need α ≈ 0.99999931. That number has seven leading nines, which is exactly why the bound is expressed in log space. **Update, 2026-08-03: the inference above is confirmed.** I originally flagged the log-α reading of `gate_lower_bound: -5.0` as an assumption — the config does not state the functional form, and the report does not either. [kimi-k3-in-c](/articles/kimi-k3-in-c), an independent C99 reimplementation, computes the gate exactly that way: ```c const float a = expf(A_log[h]); /* per HEAD */ const float u = a * (z[i] + dt_bias[i]); const float gi = lb * sigmoidf_(u); /* in (lb, 0] -> this is log alpha */ alpha[i] = expf(gi); /* in (e^lb, 1] */ ``` With `lb = -5.0`, α is bounded to $(e^{-5}, 1] \approx (0.0067, 1]$ — the 0.14-token floor holds. One refinement the code makes that the config alone did not: `A_log` is stored **per head**, and the per-channel variation comes from the `z + dt_bias` term inside the sigmoid. So a channel's decay is a per-head base rate modulated per channel, rather than a fully independent per-channel parameter. The implementation carries a pointed warning about this — the checkpoint stores `head_dim` floats but only the first `H` are nonzero, so indexing `A_log` per channel is *"a silent, fatal error"*. The `Diag(αₜ)` structure and the resulting spread of timescales are unaffected. ## Why this is more than a nice analogy Two things fall out of it that are practically useful. **It gives you a unit.** "The gate decays the state" is unfalsifiable prose. "This channel has a half-life of 69 tokens" is a claim you can check against a model's behaviour — and it tells you immediately that a channel with a 7-token half-life cannot be the thing carrying a fact across a document, no matter what the attribution heatmap suggests. **It explains the parameterization.** Every design choice around these gates — log-space parameterization, bounded gates, careful initialization near 1 — follows from the shape of $n_{1/2} = \ln 0.5 / \ln \alpha$. The function is nearly flat for most of $(0,1)$ and then explodes in the last sliver. Any scheme that samples α uniformly wastes almost all of its capacity on channels that forget within a few tokens. The same algebra runs through every gated linear-attention variant, not just KDA — Mamba's $\bar{A}$, the decay in RetNet and RWKV, the forget gate of an LSTM. They differ in how α is produced and whether it depends on the input. They agree on the underlying law, which has been sitting in physics textbooks the whole time. --- *Sources: the [Kimi K3 technical report](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf) for the KDA recurrence and the hybrid layer composition, and the released [Kimi K3 `config.json`](https://huggingface.co/moonshotai/Kimi-K3) for the layer split, head dimension, `use_full_rank_gate` and `gate_lower_bound`. The half-life framing and the derivation are mine; the channel α values in the spectrum widget are illustrative, chosen to span the range, while every half-life shown is computed exactly from them.* --- # kimi-k3-in-c: 2.78 trillion parameters, one CPU, 8 GB of RAM > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/kimi-k3-in-c > date: 2026-08-03 > tags: llm, inference, quantization, kimi, systems, c, explainer ```console $ ./bin/k3 ~/k3model --trunk ~/k3trunk --preset laptop \ --tok ~/k3model --prompt "The capital of France is" --gen 8 --incremental --- generated text --- Paris.", + "The Eiffel ---------------------- 8 tokens in 261.5 s, 32.69 s/token average PEAK RSS for the whole run: 8.24 GB ``` That is [Kimi K3](/articles/kimi-k3) — 2.78 trillion parameters, 1.56 terabytes on disk — answering a question correctly from a laptop-sized memory budget, on one CPU, with zero GPUs. The engine that did it is [kimi-k3-in-c](https://github.com/FareedKhan-dev/kimi-k3-in-c): 4,779 lines of portable C99, a single author, Apache-2.0. It is slow — 32.69 seconds for that one token — and it is a base model, so `" Paris."` is a continuation of a sentence, not a chat reply. Neither of those facts changes what the run demonstrates: a frontier-scale mixture-of-experts model, read and multiplied straight off an NVMe drive, fits in memory most people already own. I went through the source rather than just the README — `src/core/k3_ops.c`, `src/cache/k3_cache.c`, `src/io/k3_st.c` — because the interesting claims in a project like this live in code comments, not marketing copy, and this codebase's comments are unusually good. This piece is about what they say.
## The four reductions Every parameter at bfloat16 is **5,560 GB**. That number is where the project starts and it is not close to fitting anywhere. Four decisions, each one shipped in the checkpoint or written into this engine, bring it down to a measured **8.24 GB** — a 675× reduction, with no weight dropped and no approximation: 1. The routed experts already ship at **half a byte** — MXFP4 — multiplied straight out of their packed form, never expanded to floats first. 2. **KDA** gives 69 of the 93 layers a recurrent state that does not grow with context. 3. **MLA** caches one 576-wide latent per position instead of ninety-six heads of key and value. 4. The dense **trunk streams** a layer at a time instead of sitting resident, which turns the last floor into a dial. The first three are architectural decisions Moonshot made when training Kimi K3; I covered why they exist — the routing, the gating math, the attention-residual stack — in [the K3 architecture piece](/articles/kimi-k3). What is new here is the fourth one, which belongs entirely to this engine, and the fact that all four survive being reimplemented from scratch in C and checked against the released weights.
## Reduction one: the experts already ship at half a byte The routed experts are not quantised by this engine — they arrive from Moonshot already in **MXFP4**, a microscaling 4-bit float. Every weight is a 4-bit code indexing a 16-entry table, and every 32 consecutive weights share one 8-bit exponent. One routed expert is exactly 33,030,144 parameters. At half a byte plus a shared scale, that is **17,547,264 bytes** — 17.55 MB. Dequantised to fp32 first, the same expert is 132 MB. A token touches 16 experts in each of 92 MoE layers — 1,472 experts. Multiply that out and the difference stops being an abstraction: - Dequantise everything first: **194 GB** of pure format conversion, per token, before one multiply happens. - Read the nibbles directly: **25.83 GB**. The comment above the kernel that does this is the best sentence in the codebase: *"This is not an optimisation; it is what makes streaming experts possible at all."* And the mechanism is worth sitting with, because it inverts an intuition every ML engineer carries: quantisation is supposed to trade compute time for memory. Here it does the opposite. A matrix-vector product is memory bound — the arithmetic is cheap, the wait is for bytes to arrive — so reading 7.5× fewer bytes makes the packed kernel **faster** than dequantise-then-multiply, not slower. ```c /* y[rows] = W[rows][in] . x[in], with W read straight out of packed MXFP4 and never * materialised as floats. This is not an optimisation; it is what makes streaming * experts possible at all. */ void k3_matmul_mxfp4(float *y, const float *x, const unsigned char *packed, const unsigned char *scales, int in, int rows, int group) { ... for (int r = 0; r < rows; r++) { for (int g = 0; g < ngrp; g++) { const unsigned char sb = sr[g]; if (sb == 255) continue; /* NaN scale: contribute nothing */ /* expand the group of 32 to floats, dot product, then apply ONE scale */ ... acc += sub * (double)K3_E8M0[sb]; } y[r] = (float)acc; } } ``` The reason the inner loop is fast is a small lookup table: `K3_E2M1_PAIR[256][2]` maps a whole byte to both of its decoded values, so the loop does one 8-byte load instead of masking and shifting each nibble out separately. Groups of 32 elements are exactly 16 packed bytes, which is why the group size was chosen there — and the scale factors out of the inner sum entirely, applied once per group instead of once per weight. ## A floating-point contract, not a convention Kimi K3's headline claim is stronger than "it runs small": *"the same model runs in 8 GB and in 224 GB and produces byte-identical output at every budget between."* Not close. Identical. That does not happen by accident in floating point, because addition is not associative — `(a + b) + c` and `a + (b + c)` round differently once you are past a handful of terms, and this engine sums thousands of them per output, across a scalar path, an OpenMP path, and an AVX2 path, at any thread count. `k3_matmul` fixes the order rather than trusting the compiler with it: ```c void k3_matmul(float *y, const float *x, const float *W, int in, int out) { for (int o = 0; o < out; o++) { const float *row = W + (size_t)o * in; double a0 = 0.0, a1 = 0.0, a2 = 0.0, a3 = 0.0; int i = 0; for (; i + 3 < in; i += 4) { a0 += (double)row[i ] * (double)x[i ]; a1 += (double)row[i + 1] * (double)x[i + 1]; a2 += (double)row[i + 2] * (double)x[i + 2]; a3 += (double)row[i + 3] * (double)x[i + 3]; } double acc = (a0 + a1) + (a2 + a3); for (; i < in; i++) acc += (double)row[i] * (double)x[i]; y[o] = (float)acc; } } ``` Four accumulators, partitioned by `i % 4`, reduced as `(a0 + a1) + (a2 + a3)` — written out by hand rather than left for the compiler to vectorise however it likes, because *that specific split is the summation order the AVX2 path must reproduce exactly*. A `__m256d` register holds four doubles; loading four elements per iteration places element `i` in lane `i % 4`, the same partition as the scalar accumulators, and reducing with the same parenthesisation gets the same bits back. Two more details carry the guarantee: the accumulators are **double**, because a float accumulator loses precision the comparisons can see at hidden size 7168; and the AVX2 code uses a separate multiply and add, never a fused multiply-add, because `-ffp-contract=off` means the scalar path rounds twice and an FMA rounds once — a hardware capability that would otherwise quietly change the output. `k3_matmul_bf16` mirrors the same layout for the bf16 trunk, so all three paths — scalar, OpenMP, AVX2 — agree to the bit. And because output rows never depend on each other, threading the outer loop changes nothing about the arithmetic either: every row is still summed by exactly one thread in exactly this order, so the result is identical at any thread count. That determinism is also honest about its one exception. `k3_matmul_mxfp4` is *not* bit-identical to dequantise-then-multiply, and the comment says so without hedging: it sums each group of 32 under its own accumulator and applies that group's scale before combining groups, while a plain matmul sums the whole row under one accumulator — a different order. But the bound on that difference is derived, not asserted. Every individual product inside a group is **exact** in double: an E2M1 value carries 3 mantissa bits, `x` carries 24, the product needs 27 of the 53 bits double has to spend, so only the additions round at all. Reassociating exact terms moves the result by about one unit in the last place of a double — roughly `1e-16` relative. The test gate requires agreement to `1e-6`. The margin between what the reordering actually costs and what the test demands is nine orders of magnitude. A codebase that states plainly where it is *not* exact, and then bounds how far off, is doing something most numerical code does not bother to do. ## KDA in the code I wrote about [KDA's decay as a half-life](/articles/kda-half-life) from the technical report alone, and had to *infer* that the forget gate is parameterised in log-alpha space — the report gives the mechanism but not the exact functional form, so I flagged the parameterisation as unverified. This C source confirms it outright. `k3_kda_decay` computes the gate per head, then folds it per channel: ```c void k3_kda_decay(float *g, float *alpha, const float *z, const float *A_log, const float *dt_bias, int H, int D, float lb) { for (int h = 0; h < H; h++) { /* PER HEAD. The checkpoint stores head_dim floats but only the first H are * nonzero. Indexing this per channel is a silent, fatal error. */ const float a = expf(A_log[h]); for (int d = 0; d < D; d++) { const int i = h * D + d; const float u = a * (z[i] + dt_bias[i]); const float gi = lb * sigmoidf_(u); /* in (lb, 0] */ g[i] = gi; alpha[i] = expf(gi); /* in (e^lb, 1] */ } } } ``` `gi = lb * sigmoid(u)` is log-alpha directly, and `alpha[i] = expf(gi)` is exactly the exponential I had to guess at from the outside. With the checkpoint's `gate_lower_bound` of −5, alpha lands in `(e^-5, 1]`, about `(0.0067, 1]` — a per-channel retention factor, with a per-*head* base rate (`A_log` is indexed by `h`, not by the channel index `i`) modulating it. An independent reimplementation confirming an inference from the outside is a satisfying result on its own, and it is the reason these two pieces belong read together. The comment on `A_log` is worth pausing on for a second reason: it is one of five *invariants* the codebase states up front as places a plausible-looking implementation silently produces the wrong model — no crash, no NaN, just a different function that still writes fluent English. `A_log` being per-head rather than per-channel is invariant one. The recurrence itself is the delta rule, in four stages that the comments number: ```c void k3_kda_step(float *S, float *o, const float *q, const float *k, const float *v, const float *alpha, float beta, int dk, int dv) { /* 1. channel-wise decay: scale ROW i of S by alpha[i] */ for (int i = 0; i < dk; i++) { ... } /* 2. read the state along k: u = S^T k */ ... /* 3. rank-one delta write. (v - u) is the prediction error: this is what makes * it a DELTA rule rather than plain accumulation. */ for (int i = 0; i < dk; i++) { const float ki = k[i]; float *row = S + (size_t)i * dv; for (int j = 0; j < dv; j++) row[j] += ki * beta * (v[j] - u[j]); } /* 4. output from the ALREADY UPDATED state: o = S^T q */ ... } ``` Decay the state, read from it along the key, write back the *error* between the value and what the state already predicted — not the value itself — then read the output from the state that write just produced. That third stage is what turns a running sum into a rule that corrects itself: writing `v` directly would just accumulate; writing `v - u` writes only what the state did not already know. Step 4 reading from the post-write state, not the pre-write one, is the second place a plausible-looking bug hides with no visible symptom. ## The cache the project exists for The dense trunk is 108.81 GB and every layer of it runs on every token — nothing to skip there, so it streams from a packed file with a pinned prefix and one rotating ring slot. The routed experts are the opposite kind of problem: 1.45 TB of the 1.56 TB checkpoint, and only 1,472 of the 82,432 experts fire per token. The header comment on the cache that handles them does not undersell its importance: *"This is the part the project exists for."* Left uncached, one decode step reads 25.83 GB of experts. At the roughly 1.2 GB/s a commodity NVMe device sustains on cold random reads of that size, that alone is about **21 seconds per token** from storage. The cache holds those experts in the same MXFP4 bytes the matmul consumes directly — caching dequantised floats would cut the number of experts that fit by 7.5× for nothing, since nothing downstream ever wants the expanded form. The replacement policy is LRU with pinning, and the victim search is a plain linear scan, on purpose: ```c /* Least recently used unpinned slot. Linear, deliberately: a few hundred comparisons * against a 17.55 MB read is not where the time goes. */ static int pick_victim(K3Cache *c) { ... } ``` A few hundred integer comparisons next to a 17.55 MB disk read is not a place worth a heap. And the cache keeps a request histogram — 82,432 counters, 330 KB — purely so a hot set can be identified and pinned, because, as the comment puts it, *"which experts are hot is not knowable in advance"*: without measuring it, pinning is guesswork. ### The bug the code confesses to The slot table has three states, not two, and the comment explains why with a candour I have rarely seen in a repository: ```c /* >= 0 holds that key * K3_SLOT_EMPTY holds nothing, free to take * K3_SLOT_INFLIGHT reserved by a batch prefetch whose read has not finished * * The third state exists because of a real bug. The batch prefetch marks a slot empty * before reading into it ... But the empty test below is a FAST PATH that returns * immediately, ahead of the pinned check and the LRU scan -- so the next expert in the * same batch was handed the SAME slot, several parallel reads wrote into one buffer, and * the MoE multiplied garbage. It cost one wrong token (65 instead of 2494) on the real * model and nothing at all in the fixtures, because no fixture exercises the streaming * cache. */ ``` Read that last clause again. The bug was invisible to the entire test suite, because the fixtures test kernels and the fault lived in the cache. It surfaced as **exactly one token** — `65` where the model should have emitted `2494` — in a run that otherwise produced fluent, plausible text. That is the failure mode that should worry anyone building inference infrastructure: not a crash, not a NaN, but one silently wrong token inside an output that reads perfectly well. That component reuses the measured, steady-state numbers, and they are less flattering to the cache than a quick simulation suggested. K3's training process uses a technique called Quantile Balancing specifically to keep expert usage flat across the pool — good for training, and exactly what defeats an LRU cache, which needs a hot subset to be worth anything. Below about 36 GB of cache arena, the bytes read per token do not move at all, a fact the engine's own measurements caught and reported rather than smoothing over: a full-recompute trace predicted a 36% hit rate at 8 GB; steady-state incremental decode measured 0%. ## Why the trunk stays at 16 bits There is an obvious asymmetry in all of the above. The experts are 4-bit. The trunk — 108.81 GB of it — is bfloat16, and the engine has no bit-width knob for it at all. If quantisation is what made the experts streamable, why not quantise the part that has to stay resident? Because they measured it. A sensitivity study over 31 real attention tensors, quantised symmetrically per row, gives **about 1% mean relative weight error at int8 and about 17% at int4** — a ratio of roughly 18 that holds across every tensor type. The worst individual rows at int4 reach 45%, 56% and **65%**. So the trunk's precision is not an oversight or a to-do; it is a decision with a number behind it, and the absent knob is the decision being enforced rather than left to a flag. The study is honest about its own limit: it measures **weight error**, not output quality. No downstream logit or token comparison was run at int4, so the cost is bounded rather than observed. That is a narrower claim than "int4 would break the model", and the repo makes the narrower one. A second measurement worth stealing: on their hardware, `O_DIRECT` cold reads run at 3.2 GB/s and are *faster* than buffered ones. The repo flags this as **"the opposite of the usual expectation, and it is why the engine opens the trunk `O_DIRECT`"**. When you are streaming a terabyte past a model, the page cache is not helping you — it is another copy. ## The memory dial Put the streaming trunk and the streaming cache together and memory stops being a wall and becomes a knob. The engine ran the identical prompt through twelve cgroup-enforced budgets, from 8 GB to 224 GB, with `MemorySwapMax=0` so an over-budget rung fails outright instead of quietly swapping: Every one of those twelve runs produced the same token ids. Not similar — identical, at a budget span of 28×. Going from 8 GB to 224 GB buys 1.70× the speed, and the paper trail behind that number is worth respecting: three back-to-back runs of one identical configuration on a quiet machine spanned 33% just from device timing noise, so the engine's own docs treat anything under that as unproven. The 28×-memory-for-1.70×-speed result clears the noise floor with room to spare; a great many of the smaller steps in between do not, and the source says so rather than reporting every row as significant. One more result from the same measurement campaign is worth stating because it runs against instinct: at a fixed 128 GB total, giving memory to the trunk before the expert cache is 1.69× faster, even though the winning split reads *79% more* expert bytes from disk than the losing one. Optimising the number that looks obviously important — cache hit rate — actively picks the slower configuration, because the trunk is re-read in full on every single token while the experts are only ever sampled. ## What this is not This is a hobby project, version 0.1.0, one author, Linux x86-64 only. 32.69 seconds per token is not usable for anything interactive, and the project does not claim otherwise — the README's own words are "slow, and answering correctly." Every measurement in this piece is the author's own, taken on one workstation; nothing here is independently replicated the way a benchmark suite would be. What sets it apart from most projects making similar claims is that it ships the receipts: raw TSVs and JSON traces under `docs/data/`, a replicated-noise-floor study that argues against several of its own smaller results, and a fixture ladder that gates every kernel against a PyTorch reference before the released checkpoint is ever touched. None of that makes this a serving solution — nobody should run a chatbot on it. What it demonstrates is narrower and, I think, more interesting: that a 2.78 trillion parameter model can be read, multiplied, and audited on hardware someone already owns, by one person, in under 5,000 lines of a language older than most of the engineers who trained the model. That is a pedagogical and archival result, not a production one, and it is worth having regardless. --- *This engine is a C implementation of [Kimi K3](/articles/kimi-k3) — see that piece for why K3 uses KDA, MLA, and Stable LatentMoE in the first place. Its MXFP4 kernel is the concrete, memory-bound case behind the general argument in [how LLM inference actually works](/articles/how-llm-inference-works): decode is bound by bytes moved, not by arithmetic, which is exactly why reading fewer bytes wins even when it means reading them in an awkward packed format. And its confirmation of KDA's log-alpha gate closes a loop from [the half-life piece](/articles/kda-half-life) I published the same day. If you're comparing this to training-time low-precision work like [Neutrino-1](/articles/neutrino-1), the distinction is the direction: that piece is about training a model to tolerate ternary weights from the first gradient step. This one is about multiplying weights a much larger model already shipped in 4-bit form, without ever training anything.* --- # LFM2.5-Encoder: classification in one forward pass, zero completion tokens > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/lfm2-5-encoders > date: 2026-08-03 > tags: explainer, encoders, classification, nlp, open-weights The default way to classify text with an LLM today is to prompt a decoder and parse whatever comes back. Write instructions, describe the labels, ask for JSON, decode the answer one token at a time, then hand the string to a parser and hope it's valid. It works, and it's also the slow way to do something that has a much cheaper shape: a fixed-size answer picked from a fixed set of labels doesn't need to be *generated* at all. That's the pitch behind Liquid AI's **LFM2.5-Encoder** models, released a week ago alongside a fine-tuning tutorial in their [cookbook](https://github.com/Liquid4All/cookbook) repo: **one forward pass, zero completion tokens.** I read the code behind that line, not just the slide it's on, and it holds up exactly as stated. ## Two new encoders, built from a decoder **LFM2.5-Encoder-230M** and **LFM2.5-Encoder-350M** are bidirectional encoders — the BERT shape, not the chat-model shape. Their parameter counts, read straight from the safetensors metadata rather than the rounded name: **229,693,184** and **354,483,968**. Both use the LFM2 hybrid backbone (interleaved gated short convolutions and grouped-query attention), hidden size **1024**, vocabulary **65,536**, context length **8,192 tokens**, and cover 15 languages. They ship under the LFM Open License v1.0, and both were created on Hugging Face on **2026-07-27** — about a week old as I write this. The interesting part is how they were built: each size starts from the corresponding causal **LFM2.5 decoder** checkpoint and is converted into an encoder with three changes. 1. **Bidirectional attention** replaces the causal mask, so every position can see the whole sequence instead of only what came before it. 2. The short-convolution layers switch from causal padding to **symmetric center padding**, so a kernel mixes information from both neighbors instead of only the left one. 3. Pretraining continues with **masked-language-modeling at 30% token masking** — twice BERT's original 15%, which is a real methodological choice, not a rounding difference, in how hard the denoising task is made.
This is the same bidirectional-vs-causal distinction the [architectures gallery](/architectures) draws for BERT: full attention lets every position condition on the entire input, which is what you want when the job is producing *a representation* rather than continuing a sequence. Generation needs causal masking so the model can't peek at the future it's about to write; classification has no future to hide — the whole document is already there, and every token should get to see all of it before you summarize what the document means. ## The mechanism: mean-pool, one Linear, sigmoid Liquid's fine-tuning tutorial (`examples/lfm-encoder-classification/` in the cookbook) is a complete, runnable project — `train.py`, `predict.py`, a YAML config, sample data — and the model class inside it, `DocumentClassifier`, is short enough to read in full. It loads the pretrained encoder, throws away the masked-token prediction head it was pretrained with, and keeps only the backbone: ```python outputs = self.backbone(input_ids=input_ids, attention_mask=attention_mask) hidden = outputs.last_hidden_state # [batch, seq_len, 1024] mask = attention_mask.unsqueeze(-1).to(hidden.dtype) pooled = (hidden * mask).sum(dim=1) / mask.sum(dim=1).clamp_min(1.0) logits = self.classifier(self.dropout(pooled)) # one nn.Linear(1024, num_labels) loss = nn.functional.binary_cross_entropy_with_logits(logits, labels.float()) ``` That's the entire model on top of the backbone: an attention-masked mean over the last hidden state, $$ \text{pooled} = \frac{\sum_i m_i \, h_i}{\sum_i m_i} $$ (padding tokens carry $m_i = 0$ so they don't dilute the average), then one `nn.Linear` sized `[hidden, num_labels]`, trained with `binary_cross_entropy_with_logits` — the loss for **multi-label** classification, where a document can carry zero, one, or several labels, as opposed to cross-entropy's "exactly one correct class." At inference, `predict.py` does the other half in three lines: ```python with torch.inference_mode(): probabilities = torch.sigmoid(model(**inputs).logits)[0].cpu().tolist() predicted = [label for label, p, t in zip(labels, probabilities, thresholds) if p >= t] ``` No `generate()`, no sampling, no stop tokens, no text to hand to a parser. One forward pass in, a fixed-length array of probabilities out, compared against per-label thresholds tuned on a validation split. That is what "zero completion tokens" means in the code, not just the slide. ## Sigmoid, not softmax — and why it has to be The reason this needs its own head, rather than reusing whatever classifier head ships with a decoder fine-tune, is the difference between picking one thing and scoring several independently. **Softmax** turns a set of logits into a probability distribution that sums to 1 — it is built to choose exactly one winner, which is correct for single-label problems (a document *is* sports, politics, or tech, never more than one). **Sigmoid** applied per label makes each label its own independent yes/no question: $\sigma(z_i) = 1/(1+e^{-z_i})$, with no normalization across labels, so two labels can both clear the threshold, or none can. A support ticket about a double charge is legitimately both a billing issue and a technical one; softmax would be forced to pick a single "real" answer and quietly discard the other. ## Where it lands: 4th of 14, honestly Liquid's own eval — a 17-task suite spanning GLUE, SuperGLUE, and five multilingual tasks, averaged over 5 seeds with standard deviations reported — puts LFM2.5-Encoder-350M **4th of 14** models at a mean score of **81.02**, and LFM2.5-Encoder-230M **6th** at **79.29**. That ranking is worth sitting with rather than rounding up. LFM2.5-Encoder-350M sits a point below ModernBERT-large and two points below XLM-R XL — a model **ten times its size**. It is not the best encoder on this benchmark, and Liquid doesn't present it as one. It's the fourth-best, at a fraction of the parameters of the model above it, next to a considerably larger one. The 230M model separately beats ModernBERT-base despite being the larger of the two by parameter count — a real, checkable comparison the raw numbers support either way you read them.
The chart also settles a question Liquid answers candidly in the same blog post: why build a new general-purpose encoder instead of reusing their existing retrieval models? Because those retrieval-tuned siblings — **LFM2.5-ColBERT-350M** and **LFM2.5-Embedding-350M** — score **76.18** and **75.68** on this same suite, both below the general-purpose LFM2.5-Encoder-350M's 81.02. In their own words: *"Because retrieval is only a subset of what encoders enable, we chose to build a general-purpose encoder rather than adapt the existing retrievers."* A model tuned to make embeddings cluster well for search is not automatically a good classification backbone, and Liquid's own numbers show the gap rather than hiding it. On raw speed, Liquid also reports the 350M encoder running about **3.3× faster than ModernBERT-base at 8,192 tokens on CPU**; a separate blog claim puts the 230M model specifically at roughly **3.7×** faster than ModernBERT-base at the same length (about 28 seconds versus over a minute and a half). Those are two distinct comparisons, not one number restated — worth keeping straight if you quote either. ## The tutorial's own result The cookbook ships a second, harder example beyond the 4-label sample data: fine-tuning LFM2.5-Encoder-350M on **ECtHR-A** (`coastalcph/lex_glue`), European Court of Human Rights cases labeled by which of 10 Convention articles they violate — real long documents, real multi-label targets, 9,000 / 1,000 / 1,000 train/validation/test examples, CC BY 4.0. Trained at the full 8,192-token context, one seed, 3 epochs: | Split | Metric | Score | |---|---|---| | Validation | micro-F1 (after per-label threshold tuning) | 0.8060 | | Test | micro-F1 | 0.7913 | | Test | macro-F1 | 0.7062 | | Test | micro average precision | 0.8400 | The README says outright that these numbers are from one seed — no variance reported, unlike the 17-task pretraining eval above. Read it as "this recipe works on a real long-document benchmark," not as a tuned, reproducible leaderboard number. Thresholds are tuned only on the validation split and never touch test until a separate, explicit `--evaluate-test` flag is passed — a small detail, but the right one for anyone checking the tutorial's methodology. ## What you give up An encoder with a classification head cannot do several things a decoder can, and it's worth naming them plainly rather than only listing what it's good at: It cannot generate. There's no explanation, no rationale, no free-text answer — only probabilities over a label set that's fixed at training time. Adding a new label means retraining the head (a small, cheap step, but a step), not writing a new prompt. And it needs supervised examples per task: this is a fine-tuning recipe, not a zero-shot classifier out of the box, even though Liquid's own HF Spaces (prompt routing, PII detection, policy linting) show the same base encoder fine-tuned across several different classification tasks. The honesty gaps worth naming too: these models are about a week old, with limited independent adoption to point to yet. The 17-task eval is Liquid's own compilation — methodologically solid (5 seeds, reported std, a full 14-model field rather than a curated subset), but not yet replicated by anyone outside Liquid. And the cookbook repo carries no LICENSE file at its root as of this writing, which matters if you plan to reuse the tutorial code itself, distinct from the separately-licensed model weights. ## The take The argument here isn't that encoders are back or that decoders are wrong for classification — it's narrower and more useful than that: match the tool to the shape of the answer. If you already know the output is a choice from a fixed set of labels, a bidirectional encoder can produce that choice as a probability vector in one forward pass, with nothing to decode and nothing to parse. If you don't know the shape of the answer in advance — you need explanation, planning, or free text — that's what the decode loop and [its own cost structure](/articles/how-llm-inference-works) are for. LFM2.5-Encoder is a clean, current example of the first case done right: a small model, an honestly-reported 4th-of-14 ranking against models many times its size, and a fine-tuning recipe short enough to read start to finish in one sitting. --- *Built on Liquid AI's [LFM2.5-Encoders blog post](https://www.liquid.ai/blog/lfm2-5-encoders), the [LiquidAI/LFM2.5-Encoder-230M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-230M) and [LiquidAI/LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) model cards, and the [`examples/lfm-encoder-classification`](https://github.com/Liquid4All/cookbook/tree/main/examples/lfm-encoder-classification) tutorial in Liquid4All/cookbook. Parameter counts are read from HF safetensors metadata, not the rounded model names. The 17-task benchmark and both embedded figures are Liquid AI's own, reproduced for commentary; the two interactive diagrams are original illustrations of the mechanism using illustrative example data, not measured traces.* --- # Towards looped models done right: what actually separates Ouro from Huginn > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/looped-models-done-right > date: 2026-08-03 > tags: llm, looped-transformers, recurrent-depth, ablation-study, architecture, mixture-of-experts, explainer Ask why Huginn-style looped transformers tend to beat Ouro-style ones and most people reach for the same answer: random state initialization. It is the assumption inherited wholesale from deep equilibrium models — randomize the recurrent state so the loop cannot just memorize a fixed point, and the model is forced to learn a genuinely path-independent computation. It sounds right. It is also, according to a new controlled-ablation report from the Institute of Foundation Models (IFM), mostly wrong. **Towards Looped Models Done Right** does not propose a new looped architecture. It audits the two existing lineages this site has already covered — the padded-latent loop in [LOTUS](/articles/lotus-latent-reasoning) and the shipped two-pass loop in [Nanbeige4.2-3B](/articles/nanbeige-4-2-3b) both descend from this same family of [looped / recurrent-depth transformers](/architectures) — and asks a narrower, more useful question: Ouro and Huginn differ along *three* design axes at once, so which one is actually doing the work? The answer, walked through below, is not the one folk wisdom would guess. **Read this as a first look, not a finished paper.** This report lives only as a Notion "living blog," explicitly billed as Part I of a series and continuously updated — there is no arXiv listing. Code is marked "Release Soon," meaning nobody outside IFM can rerun these numbers yet, and there has been no third-party replication. Every result below comes from ablations run by one group, at 730M dense / 8B-resident MoE scale — a real, carefully controlled experiment, but not yet an independently checked one, and not yet evidence about frontier scale. ## One formalism, two lineages, three tangled axes Both Ouro and Huginn are instances of the same tied-iterative model. Take token embeddings, run them through a prelude, loop a shared recurrent body $T$ times, then run a coda: $$ \mathbf e = P_\theta(\mathbf x_0),\quad \mathbf z_0 = \phi_\theta(\mathbf e, \boldsymbol\xi),\quad \tilde{\mathbf z}_t = W_\theta(\mathbf z_t, \mathbf e),\quad \mathbf z_{t+1} = R_\theta(\tilde{\mathbf z}_t),\quad \mathbf h = C_\theta(\mathbf z_T) $$ $P_\theta$ is the prelude, $R_\theta$ the tied recurrent body (the only part that repeats), $C_\theta$ the coda, $\phi_\theta$ the state initializer (optionally seeded with noise $\boldsymbol\xi$), and $W_\theta$ the per-step write that decides how much of $\mathbf e$ gets re-injected at each pass. Set every one of $P_\theta$, $W_\theta$, $\phi_\theta$ to the identity and you get **Ouro-style**: the entire network is the recurrent body, looped over the full sequence, with $\mathbf z_0 = \mathbf e$ directly. Untie a real prelude and coda from the loop, persistently re-inject $\mathbf e$ into the core at every pass, and optionally randomize the initial state, and you get **Huginn-style**. Three independent knobs — iteration envelope, input interface, latent-state design — get flipped between the two architectures simultaneously in the existing literature. Nobody had isolated which flip mattered.
## The controls that make this an ablation, not a vibe check A myth-busting result is only as good as what it holds constant. IFM's models are matched on logical depth — Ouro-style loops a 28-block stack four times, Huginn-style runs an 8-block prelude, a 12-block core eight times, and an 8-block coda; both total 112 block executions — and on parameters (730M stored / 2.9B unrolled-equivalent dense; 8B-resident / 0.8B-active, 32B/3.2B unrolled-equivalent for the MoE runs), tokens (TxT360, swept across 58/115/230/460 tokens-per-parameter), and a ten-benchmark suite spanning knowledge (ARC-C, HellaSwag, MMLU, TriviaQA), context reasoning (BBH-CoT, DROP), math (GSM8K, MATH500), and code (HumanEval+, MBPP+). Whichever topology wins a given comparison, it isn't winning on a hidden depth or size advantage: With the controls fixed, IFM walks the transformation path from Ouro to Huginn one axis at a time — sandwich envelope, then input injection, then random state init — and measures what each addition actually buys. ## Q1: does untying the prelude and coda matter? Yes, and only for a specific kind of task. Adding a sandwich envelope — untying a prelude and coda from the loop so only the middle core repeats — lifts **MATH500 by 12.00 points** and **DROP by 2.61 points** at the 460 tokens-per-parameter budget, and the gain persists across all four token budgets tested. But knowledge-heavy benchmarks and strict-output-contract tasks like code show no consistent gain, sometimes a small decline. Read plainly: the envelope helps *instance-conditioned, multi-step reasoning*. It does not help stored-knowledge recall or spec-compliant code generation, and the report is upfront that it shouldn't be expected to. ## Q2: does persistent input injection matter? Also yes, and it's the widest-reaching single change in the whole report. Writing the prelude's representation $\mathbf e$ into the core at every pass — not just once at the start — uses a learned, per-channel gate: $$ D(\mathbf z_t, \mathbf v) = \boldsymbol\alpha \odot \mathbf z_t + \boldsymbol\delta \odot \mathbf W_{\text{in}}\mathbf v, \qquad \boldsymbol\delta = \operatorname{softplus}(\mathbf b_\delta), \qquad \boldsymbol\alpha = \exp\{-\boldsymbol\delta \odot \exp(\mathbf a)\} $$ On the middle-loop (sandwich) topology, adding this write lifts **MMLU +2.53, BBH-CoT +6.63, DROP +1.39, HumanEval+ +5.49, MBPP+ +4.23** — five benchmarks, all up, some by a lot. Bolt the same write onto the full-stack Ouro topology (injecting raw token embeddings instead of a prelude-encoded $\mathbf e$) and the same-direction gains show up smaller (MMLU +1.79, BBH-CoT +1.80, DROP +0.91, HumanEval+ +2.44, MBPP+ +4.50) — evidence the effect is real and not an artifact of one topology. It is not a free lunch. The same write **hurts** quantitative reasoning: middle-loop MATH500 drops **3.60 points**, GSM8K drops **2.51** (the full-stack Ouro version is less damaged: MATH500 −1.60, GSM8K +0.53). Persistently re-showing the model its own input, it turns out, competes with letting the loop's state evolve freely enough to carry a multi-step derivation.
Put the envelope and the injection together and the combined model beats full-stack Ouro on **8 of the 10 benchmarks** (losing only ARC-Challenge and HellaSwag). Step through both changes yourself — the diagram below is my own redrawing of the same construction path, with each stage's measured delta attached so you can see exactly which wire produced which number, before the third, more surprising change gets added: ## Q3: the myth — random state init and shared H/L hierarchies This is the report's contrarian core. Swap the direct initial state $\mathbf z_0 = \mathbf e$ for a randomly sampled one, $\mathbf z_0 \sim \mathcal N(0, I/d)$ — the equilibrium-model-inherited move everyone assumes is load-bearing — and two benchmarks improve (**ARC-C +3.34, GSM8K +1.22**) while four get worse by more than a point (**MMLU, MATH500, HumanEval+, MBPP+**). Net: **direct init wins 6 of 10 benchmarks**, and it's cheaper, since it skips sampling noise at every forward pass. The report's own words: "random initialization is not a necessary ingredient for loop language models... [it] should instead be viewed as a task- and objective-dependent inductive bias, rather than as a universally beneficial design choice." A second candidate "obviously helps" ingredient fares no better. HRM/TRM-style hierarchies split the loop into a slow high-level state and a fast low-level state cycling underneath it; IFM tests a version that **shares one recurrent body** across both states (isolating the state-hierarchy idea from the separate-modules idea) and finds gains over a point on three benchmarks, losses over a point on three more — MATH500 hit hardest — and roughly flat on the rest. Their conclusion: "a shared-module H/L hierarchy provides no consistent benefit." (They flag, honestly, that this doesn't rule out a separately-parameterized HRM-Text-style version — that variant is "still under evaluation.") Here is the believed-important story against what got measured, for both: ## The net effect: two wires did almost all the work Chain every change together — Ouro, plus envelope, plus injection, plus random init — and you land on full Huginn, which does beat Ouro on all ten dense benchmarks at the 730M/336B-token setting. But laid out per-benchmark across the whole construction path, the shape of the win is obvious: most of the climb happens in the first two steps, and the last step (random init) barely moves several benchmarks and actively costs a few.
## Does the story change at MoE scale? The two levers that mattered — envelope and injection — hold up when the recurrent body becomes a mixture-of-experts. At 8B-resident / 793.9M-active parameters (500B tokens, top-2 routing, 25 experts), Huginn-MoE beats Ouro-MoE on 8 of 10 benchmarks, with the largest gains on **GSM8K (+4.70)** and **MATH500 (+3.60)**. It also routes more evenly — a normalized load-balancing loss where lower means more balanced: A causal check backs up that the routing difference is meaningful, not noise: force loop iterations 2 through 8 to reuse iteration 1's expert *identities* (keeping the iteration-specific mixture weights) and accuracy drops on all six evaluated tasks. Whatever the loop is learning to route each pass, it matters. Against a **112-layer feedforward MoE** reference (32B resident parameters), the feedforward model still wins overall — 7 of 10 benchmarks — but Huginn-MoE beats it on DROP and GSM8K and matches it on MATH500, while using **75% fewer resident parameters**. The mean gap to the feedforward reference shrinks from 4.96 points in the dense setting to 1.71 points in the MoE setting. Looping doesn't close the gap to a much bigger feedforward model outright, but MoE narrows it substantially — the same shape of result as Nanbeige's own parameter-efficiency argument, below. ## Set against a shipped model: Nanbeige4.2-3B The most useful check on any ablation study is an independent result that wasn't trying to test the same hypothesis. [Nanbeige4.2-3B's technical report](/articles/nanbeige-4-2-3b) is exactly that: a production model that made its own looping decisions under deployment pressure, not a controlled academic sweep. Nanbeige's architecture is closer to Ouro-style — a homogeneous loop over the full stack, run twice — and its report reached three conclusions of its own: **two passes is the sweet spot** (more loop count bought little and made training less stable), **training the looped architecture from scratch beats upcycling** a pretrained feedforward model into one, and **sharing the KV cache across passes underperformed**, so they paid full attention cost at every pass rather than take the cheaper shortcut. None of Nanbeige's three findings directly tests IFM's three axes — Nanbeige never tried an untied prelude/coda, persistent injection, or random state init — so this isn't a replication in either direction. But the two reports rhyme in an interesting way: every place Nanbeige tested a cheap shortcut inside the loop (share the KV cache, upcycle instead of retraining, add more passes without changing the topology), the shortcut lost. Every place IFM tested a richer per-pass mechanism (untie the envelope, inject persistently), it won — and the one change that added complexity without adding real per-pass information (random init) was the one that didn't clearly help. Read together, the two reports point at the same underlying rule: what a looped model does *each pass* — how much fresh computation and fresh input it gets — matters more than how many times it loops or how its state gets seeded. Where they don't overlap at all is loop count itself: Nanbeige's "two is enough" is a statement about a plain full-stack loop; IFM's Ouro-style baseline already runs four passes and Huginn-style runs eight core iterations inside a smaller envelope, so the two reports are sweeping different variables and shouldn't be read as agreeing or disagreeing on "how many loops." There's a second, more mechanical echo. [LOTUS](/articles/lotus-latent-reasoning) — the site's other looped-transformer piece — already does something IFM's ablation independently flags as one of the two levers that matter: every LOTUS iteration recomputes $\mathbf e + \mathbf h^{(t-1)}$, persistently feeding the fixed input embeddings back into the loop rather than only conditioning on them once. That's a specific instance of the same principle behind IFM's write operator $D(\mathbf z_t, \mathbf e)$: keep re-showing the loop its input. LOTUS applies that idea inside a frozen-backbone latent-reasoning setup at inference time rather than IFM's from-scratch pretraining setup, so the two aren't directly comparable — but it's a second, independent place persistent input injection shows up as doing real work. ## What to trust, and what to hold loosely **The scope, precisely.** Every result above is at 730M dense / 8B-resident MoE scale — there is no evidence yet these findings hold at frontier (100B+) scale, and the report says so. The random-init result is explicitly narrower than "random init never helps": IFM's own caveat is that these evaluations "do not directly measure multi-start path independence or extrapolation to recurrent depths beyond those used during training" — the property random init is classically supposed to buy. The H/L null result is scoped to a *shared-module* hierarchy only; a closer HRM-Text replica with separate modules is still under evaluation and not included here. And this whole report has no arXiv listing, lives on a continuously-updated Notion page billed as "Part I" of a series, and ships with no code yet — treat every number as provisional until an arXiv version, released code, or a third-party rerun shows up. ## The take The intuitive story about looped transformers has always centered on the recurrent state itself — randomize it, and the model is forced to learn something more general. IFM's controlled ablations say that story is backwards for at least these two designs at this scale: the state's initialization is close to a wash, sometimes a net loss, and a fashionable H/L hierarchy adds nothing consistent once you control for everything else. What actually separates a strong looped model from a weak one is much less exotic — untie a prelude and coda from the loop so the recurrent core can specialize, and keep showing that core its input at every pass instead of just once. Neither idea needs a random number generator. That's a genuinely useful result for anyone building a looped model today, and it lines up with what Nanbeige found the hard way in production: cheap shortcuts inside the loop (shared KV cache, more passes without restructuring, upcycling instead of retraining) tend to cost you, while spending real compute and real information on each pass tends to pay. The caveat that matters most is the one IFM states themselves — this is one group's ablation, at sub-billion-to-8B scale, with no code out yet and no arXiv paper. "Part I" means there's more coming. Whether the random-init myth holds at 100B+ parameters, and whether a properly separate-module HRM-Text hierarchy fares better than the shared one tested here, are open questions the authors name as future work, not settled ones. --- *Source: [Towards Looped Models Done Right — Part I: Topology, Input Injection, Recurrent-State Design](https://ifm-research.notion.site/Towards-Looped-Models-Done-Right-3ade511912ec8128987dfeb7a5580043) (Huang, Shi, Chen, Wen, Liu, Xing, Ma; Institute of Foundation Models, 2026). All benchmark deltas, equations, and figures are quoted or reproduced from the report; the matched-depth, construction-path, and believed-vs-measured diagrams are my own illustrations of the same numbers.* --- # MemHarness: agent memory should be reconstructed, not replayed > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/memharness > date: 2026-08-03 > tags: agents, memory, reinforcement-learning, llm, explainer Most memory-augmented agents treat a retrieved experience the way a tape recorder treats a cassette: press play, get back exactly what was stored. **MemHarness**, from a Zhejiang University / Shanghai AI Lab team, argues that's the wrong model of memory entirely — and points at cognitive science to say so. Human recall isn't playback; it's reconstruction, rebuilt each time from fragments and reshaped to fit the moment. Their agent does the same: before acting on a retrieved memory, it first critiques that memory against what it's looking at *right now*, and rewrites it if the two don't match. Single paper, one lab, no third-party replication found. Everything below — every percentage, every table — is MemHarness's own reported numbers on two benchmarks (ALFWorld, WebShop), one 7B backbone (Qwen2.5-7B-Instruct). The paper reports no hardware, no wall-clock latency, and no variance across seeds for any of its evaluation numbers — see **Honesty check** near the end before you take any number as a settled fact. The interactive diagrams below are clearly labeled: two use the paper's own measured table values, one is an illustrative toy walkthrough of the mechanism. ## The failure mode: negative transfer from a memory that no longer fits Retrieval-augmented agents work like this: finish a task, distill what happened into a short natural-language "experience," store it, and next time a similar task comes up, pull the closest matches back into context. The problem is *closest* is doing a lot of work. An experience learned from one kitchen layout, one inventory state, one shopping page, gets pasted into a context where the fridge is already full or the shelf has something else on it — and the agent, having no reason to doubt its own memory, follows the stale instruction anyway. MemHarness calls this the "replay" paradigm, and its central claim is that replay's failures are systematic, not occasional: the retrieved experience is abstract and general by construction, while the state at decision time is concrete and constantly changing, and nothing in a replay pipeline reconciles the two.
This isn't a new observation for this site — [Agent harnesses: engineering the loop around the model](/articles/agent-harness) already named "context and memory lifecycle" as one of the open problems in agent engineering, and pointed out that a file-backed harness effectively turns the file system into the agent's long-term memory. MemHarness is a concrete answer to a sharper version of that problem: it's not enough to *store* memory durably, the harness also has to decide, at read time, whether a piece of stored memory still applies — and if not, what to do about it. MemHarness's answer is to make that decision itself a trained, learned skill rather than a fixed retrieval-and-paste rule. ## Three stages: retrieve, reconstruct, act Formally, the agent holds a memory bank $\mathcal{B} = \{m_i\}_{i=1}^N$ of entries $m_i = (e_i, o_i^{src})$ — an abstracted experience $e_i$ paired with the **source observation** $o_i^{src}$ it was distilled from. At each step $t$, retrieval returns the top-$k$ closest entries, $\mathcal{E}_t = \mathcal{R}(q_t, \mathcal{B})$. A pure replay policy conditions the next action directly on whatever comes back: $$ a_t \sim \pi_\theta(\cdot \mid \mathcal{T}, h_t, \mathcal{E}_t) $$ MemHarness inserts one step in between. The same policy first produces guidance $g_t$ by critiquing the retrieved experiences against the current history $h_t$, then maps that into final guidance $\tilde{g}_t$, and only then generates the action: $$ g_t \sim \pi_\theta(\cdot \mid \mathcal{T}, h_t, \mathcal{E}_t), \qquad \tilde{g}_t = f(g_t), \qquad a_t \sim \pi_\theta(\cdot \mid \mathcal{T}, h_t, \tilde{g}_t) $$ The reconstruction input concatenates the task, the recent history, and every retrieved (experience, source-state) pair: $$ x_\text{recon} = \mathcal{T} \oplus h_t \oplus \bigcup_{i=1}^{k} (e_{t,i},\, o_{t,i}^{src}) $$ and $f$ is a simple conditional: if the policy decides nothing retrieved applies, it emits the literal token ``, and $\tilde{g}_t$ falls back to a fixed self-reasoning prompt $p_\text{self}$ instead of forcing a bad match into the action context.
The mechanics of that middle stage — comparing a retrieved memory's source state against the live one, and deciding pass-through / adapt / reject — are easiest to see with a toy example. Toggle through the three cases below: Note what each branch does differently from plain replay. When the state genuinely **matches** the memory's source, reconstruction is a no-op and replay would have been fine anyway — the interesting cases are the other two. When the state has **drifted**, reconstruction rewrites the *target*, not just the wording, while replay keeps repeating an instruction the environment has already invalidated. And when retrieval turns up nothing usable, MemHarness can say so explicitly and fall back to the agent's own reasoning — a replay pipeline has no equivalent move; it either injects a weak match or injects nothing silently. ## Training: one policy, three roles, GRPO end to end The same weights play all three parts — retriever-decider, reconstructor, actor — and the whole thing is trained with **GRPO** on a sparse outcome reward plus a small format bonus: $$ R(\tau_i) = R_\text{outcome} + 0.1 \cdot R_\text{format} $$ $R_\text{outcome}$ is 10 for a successful episode and 0 otherwise; $R_\text{format}$ checks that every step emits exactly one valid `` block, one valid `` block, that memory is retrieved through valid `` blocks (one to five times per episode), and that everything is in English. Rewards are group-normalized — sample $G=8$ rollouts per prompt and standardize against the group: $$ A_i = \frac{R(\tau_i) - \text{mean}(\{R(\tau_k)\}_{k=1}^{G})}{\text{std}(\{R(\tau_k)\}_{k=1}^{G})} $$ and the policy update is the standard clipped-surrogate-plus-KL objective, with clip range $\varepsilon = 0.2$ and KL coefficient $\beta = 0.01$: $$ \mathcal{J}(\theta) = \mathbb{E}\!\left[\frac{1}{\sum_i |\tau_i|}\sum_{i=1}^{G}\sum_{j=1}^{|\tau_i|} \Big(\mathcal{L}^{\text{CLIP}}_{i,j}(\theta) - \beta\, \mathbb{D}_{\text{KL}}[\pi_\theta \| \pi_\text{ref}]\Big)\right] $$ None of this is a new RL recipe — it's the same token-level, group-relative machinery covered in [Token-level RL is a first-order approximation to the reward you actually want](/articles/first-order-rl). What's specific to MemHarness is that the *reconstruction step itself* is inside the RL loop and gets credit for the same sparse outcome reward as the final action, rather than being a fixed prompt template bolted on the side. That's also why an ablation later in this piece — replacing the trained reconstruction with a generic, untrained LLM doing the same rewriting job — measurably underperforms: rewriting text is not the same skill as rewriting text so that it wins the episode. Before RL, there's a short cold-start SFT stage — 200 trajectories with GPT-5.1-generated retrieval and reconstruction turns, plus 200 trajectory-to-memory summarization examples per benchmark — whose only job is to teach the interaction protocol (when to emit ``, how to format guidance). The paper is explicit that this stage is about "protocol and format alignment rather than task-skill acquisition," and the numbers back that up: the cold-start model alone scores a *worse* 7.6% on ALFWorld than the untrained base model's 14.5%, because it has learned to follow a longer protocol without yet having learned to solve the task. The memory bank itself lives in **Milvus**, embedded with **BGE-M3**, retrieved by cosine similarity at $k=3$. It isn't hand-curated — during training, the policy distills roughly half of its own generated trajectories (balanced between successes and failures where possible) into new memory entries, so the bank grows out of the same policy that reads from it. Write-time deduplication skips a new entry if it's too similar (cosine $> 0.85$) to something already stored — enabled for WebShop, disabled for ALFWorld — and retrieval-time deduplication thins a larger candidate pool before truncating to the top-$k$. ## Does it beat the baselines On the headline numbers: MemHarness reaches **85.2%** average success on ALFWorld's six task categories and **75.6%** on WebShop, ahead of every baseline the paper reports — including foundation models an order of magnitude larger: A few things worth being precise about here. The 16-row full table (not all shown above) mixes closed-source frontier models (GPT-4o, Gemini-2.5-Pro), prompt-only memory agents (ReAct, Reflexion, Mem0, ExpeL, MemP, SimpleMem), and RL-trained agents (RLOO, GRPO, MemRL, EvolveR, and two "+GRPO" memory hybrids) — and with one exception (**EvolveR**, explicitly marked "reproduced"), the paper doesn't say whether the other baseline numbers are copied from those methods' original papers or re-run by the authors under this setup. Given every RL-based and prompt-based baseline shares the same Qwen2.5-7B-Instruct backbone as MemHarness, it reads as an in-house re-implementation for a controlled, like-for-like comparison — which is the right thing to do for fairness, but it also means there's no independent number to check any of them against. **Mem0** in particular scores worse than the untrained base model on WebShop (2.0% vs. 7.8%) — plausible for a general-purpose memory library not tuned to this task, but a reminder that "memory system" is not automatically an improvement. ## Why raw memory can hurt: the ablation The paper's most useful table isn't the leaderboard, it's the ablation, because it isolates *why* MemHarness wins rather than just *that* it wins. Same policy, same GRPO recipe throughout — only the memory wiring changes: Two results here are worth sitting with. First, **RL + Raw Memory** — verbatim replay, grafted onto the same trained policy — actually *loses* to having no memory at all on ALFWorld (70.1% vs. 76.4%), which is the paper's sharpest evidence that unreconstructed memory is not a free win; it can be actively confusing. Second, **w/o memory** — the fully-trained MemHarness policy with retrieval switched off at test time — still beats the no-memory-ever RL baseline on both benchmarks (83.0% vs. 76.4% on ALFWorld, 73.6% vs. 66.1% on WebShop). The paper reads this as evidence that training the policy to reconstruct memories also sharpens its general reasoning, independent of whether memory is available at inference — the reconstruction objective works partly as a training-time signal, not only a run-time lookup. The training curves back this up with a second, independent kind of evidence — not an end-of-training snapshot, but what happens over the run:
Trajectories where the policy accepted a reconstructed memory track the overall success-rate curve closely; trajectories where it rejected one lag behind and stay noisier throughout training. That's a consistency check on the whole framework: if "accept vs. reject" were an arbitrary or miscalibrated signal, there'd be no reason for it to correlate with which trajectories actually succeed. MemHarness also holds up when the environment itself is unfamiliar. On ALFWorld's out-of-distribution split — unseen room layouts and object placements — it scores **85.9%**, while stripping reconstruction back out (raw memory injected, same OOD environments) drops to **76.3%**, and disabling reconstruction only at test time (same trained policy) lands at **82.4%**. The direction of every result here matches the in-distribution ablation: verbatim replay is the worst way to use memory precisely when the environment has changed most. ## The mechanism, under a microscope Everything so far shows *that* reconstruction helps. The paper also runs two controlled probes asking a narrower question: does the policy's reconstruction step actually compare the current state against the memory's recorded source state, or is it just producing plausible-sounding rewrites without really checking anything? The **source-state ablation** answers this directly: strip $o_i^{src}$ out of the reconstruction prompt entirely, and rejection rate barely moves — but success rate drops, because the policy now accepts guidance it has no way to judge as stale. Swap in a *random* memory's source state instead — a state that's guaranteed not to match — and rejection rate jumps sharply (8.7%→13.3% on ALFWorld, 56.0%→63.3% on WebShop). That asymmetry is the tell: removing the comparison signal doesn't change behavior much because the policy simply can't tell anymore, while corrupting it with a wrong-but-present signal actively triggers more rejections. The **counterfactual probe** — asking a strong LLM to make a minimal edit to 1,000 real states so a previously-applicable memory should no longer apply, then scoring only the reconstruction output — shows the same pattern from the other direction: minimal edits shift outputs measurably away from "unchanged" and toward "adapted" or "rejected" on both benchmarks, with WebShop rejecting far more often than ALFWorld in both the matched and edited conditions (72–79% vs. 0–6%), which the paper attributes to WebShop's longer, more heterogeneous page observations making a fuzzy accept riskier than a clean reject. ## Honesty check - **Self-reported, single lab, no replication.** Every number above is from this one paper. I found no independent reproduction, and the community-discussion page on alphaXiv had nothing beyond the paper's own abstract and tables at the time of writing. - **Baselines are the paper's own reruns, not cited published numbers**, as far as the text discloses — with the single exception of EvolveR, marked "(reproduced)" in the table. That's a reasonable design for a fair, same-backbone comparison, but it also means none of the sixteen rows in Table 1 have an outside number to be checked against. - **No hardware, no latency, no cost.** The paper never states what GPUs it trained or evaluated on, never reports wall-clock time, tokens/sec, or dollars, and never measures the added inference cost of the reconstruction step itself. Reconstruction is a second full decode pass through the same 7B policy on every step where memory is retrieved (retrieval decision, then reconstruction, then action) — at minimum one extra generation versus a direct-replay or no-memory baseline — and that overhead is not quantified anywhere in the paper. - **No variance, no seeds.** Every success-rate and rejection-rate number is reported as a single figure with no standard deviation, confidence interval, or multi-seed spread disclosed for the evaluation runs. - **Two benchmarks, one model scale.** ALFWorld and WebShop are both well-worn, relatively short-horizon (15–50 step) simulated environments; the paper does not test a larger backbone, a real-world tool-using agent, or a benchmark with a genuinely different observation modality. The conclusion names this directly: "future work will explore scaling to larger models and open-ended environments" — which is the authors' own way of saying this hasn't been tried yet. - **No explicit Limitations section.** The paper has no dedicated limitations discussion; what's above is reconstructed from the ablations, the conclusion, and what the method section does and doesn't measure. - **What is solid:** the mechanism probes (source-state ablation, counterfactual editing) are a genuine attempt to falsify the "it's just fluent rewriting" explanation, and they point the same direction from two independent angles. That's better methodological care than a bare leaderboard table, even without outside replication. ## The take The idea underneath MemHarness is simple enough to state in a sentence — compare the memory's source state to the current one before you trust it — and the paper's real contribution is making that comparison a *trained* skill inside the same policy, credited by the same sparse outcome reward as the action itself, rather than a hand-written heuristic bolted onto retrieval. The ablations back the framing better than the leaderboard does: raw memory replay measurably *loses* to no memory at all on one benchmark, and the reconstruction-trained policy keeps a chunk of its advantage even with memory switched off entirely, which says the training signal is doing more than teaching better lookups. What it hasn't shown yet is whether any of this survives outside two small, well-studied simulators at 7B scale, and whether the extra reconstruction pass is worth its unmeasured latency cost in a setting where that matters. "Reconstruct, don't replay" is a good design principle for any agent harness that reads back its own memory. Whether this specific recipe for teaching it — GRPO, a `` escape hatch, a Milvus bank refreshed by the policy's own trajectories — is the way to get there past ALFWorld and WebShop is still an open question the paper itself doesn't claim to answer. --- *Built on [MemHarness: Memory Is Reconstructed, Not Replayed](https://arxiv.org/abs/2607.28272) (Wu et al., 2026; arXiv:2607.28272). Figures are reproduced from the paper for commentary. The interactive diagrams use the paper's measured table values except where marked illustrative; see the Honesty check above for what is and isn't independently verified.* --- # MiniMax H3: open weights, four excluded countries, zero benchmarks > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/minimax-h3 > date: 2026-08-03 > tags: licensing, open-source, multimodal, video-generation, explainer MiniMax [announced H3](https://www.minimax.io/blog/minimax-h3) on July 31, 2026 and put weights on Hugging Face two days later — an omni-modal system that takes text, image, video, or audio in and produces video with native stereo audio out, up to 2K resolution, 15 seconds, 24 FPS. The architecture underneath is a real, disclosed piece of engineering: an encoder built on the **full pretrained weights of Qwen3-VL-32B**, sampled from its 50th layer, feeding a **33B-parameter dense Omni-Transformer** with no modality-specific attention or feed-forward blocks — only the input/output layers and a set of AdaLN branches (about 13B of the 33B) are modality-specific. I'm not writing about the capability, though. I'm writing about what shipped alongside it, because the license and the evidence base are the actual story here. ## The license MiniMax H3's weights carry the **MiniMax H3 Community License Agreement**, and its territorial scope is precise enough to quote directly. The license grants use across the "Applicable Territory," defined as worldwide **excluding** the "Excluded Territories" — and the Excluded Territories are named explicitly: **the European Union, the United Kingdom, the Republic of Korea, and the United States of America.** That is a different thing from a normal open-weight release. Apache-2.0 and MIT — the licenses this site's other open-weight coverage almost always carries — don't have a geography clause at all. This one draws the line at specific jurisdictions, and the four it picks are, not coincidentally, four of the jurisdictions with the most developed AI-regulatory frameworks in the world. There's a second condition stacked on top for everywhere else: commercial deployments need "separate, prior written authorization" once they clear **$20 million/year** in revenue, plus a requirement to "prominently display 'MiniMax H3'" on the interface of anything built with it.
That figure matters for the licensing question too: even inside the "open" release, two of the three modules the diagram shows aren't open at all. H3-Context-IR — which the README calls "critical to the quality of the final output" — and H3-Regenerate-2K, the 2K upsampling stage, are both hosted services you call MiniMax's API for. What's actually downloadable is the middle box, in two task-specific checkpoints (text/first-last-frame→video and reference→video), both CFG-distilled BF16. Native sparse attention, used in the final training stage, is withheld from this release as well. ## The benchmark that isn't there I looked for a number to weigh the license against and didn't find one. There is no VBench score, no Elo comparison, no named baseline model anywhere in the blog post or the Hugging Face card. The one performance claim in the entire release is pricing, and even that has no dollar figure attached: *"At 2K, H3's per-second price is less than a third of mainstream models, and at 768p, it's less than half the price of mainstream models' 720p."* Less than a third of what? Which mainstream models? The post doesn't say. To be precise about what's disclosed and what isn't: the architecture (encoder choice, layer sampled, transformer size, VAE compression ratios) is specific and checkable. The training data is not — "built entirely from real, natural data" is the only description given, with no token counts or dataset composition. The capability claims are entirely qualitative. ## The take Put the two things next to each other: a model you may be legally barred from using depending on which of four major jurisdictions you're in, released with no numbers that would let you decide whether it's worth working around that restriction if you could. Neither fact is hidden — the license text is public and precise, and the absence of benchmarks is just an absence, not a false claim. But "open-source," used as freely as MiniMax uses it in the blog copy, is doing less work here than it usually does. A license with a four-jurisdiction carve-out and a revenue-gated authorization clause is a commercial license with a wide default grant, not a permissive one. Whether that's a reasonable posture for a company shipping an expensive-to-train omni-modal model is a separate question from whether it should be called "open" without the qualifier. --- *Sources: the [MiniMax H3 blog post](https://www.minimax.io/blog/minimax-h3) and [Hugging Face model card](https://huggingface.co/MiniMaxAI/MiniMax-H3) (MiniMax, July–August 2026), including the license file's Excluded Territories clause. The figure is MiniMax's own system-overview diagram; the territory checker is my own illustration of the license's geographic scope, not a legal opinion.* --- # pdf-inspector: classifying PDFs without a single model > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/pdf-inspector > date: 2026-08-03 > tags: open-source, rust, pdf, heuristics, explainer Most PDF-to-text pipelines start with a coin flip: run OCR on everything and pay for it, or guess which files need it and get burned when you guess wrong. [pdf-inspector](https://github.com/firecrawl/pdf-inspector) (Firecrawl, MIT, 6.3k stars) skips the guess. It classifies a PDF as text-based, scanned, or mixed — page by page — and converts the text-based pages to Markdown, in under 200ms, with **one dependency** (`lopdf`) and **no ML models, no OCR, no external services**. Firecrawl's own number for why this matters: roughly **54% of PDFs** don't need OCR at all, and this is how you find out which 54% without running an OCR engine to check. That's the whole pitch: a bounded, well-understood problem — is this page's text selectable? — solved with parsing and arithmetic instead of a model. ## The heuristics, not the pitch The interesting part of pdf-inspector isn't that it's fast. It's *what it checks*. Reading `src/detector.rs`, three heuristics stand out, each with an inline rationale in the source: **Path-op density.** A page can have plenty of drawing operators and still not have real text — some PDFs render glyphs as outlined vector paths rather than as selectable characters. The detector flags this specifically: massive path-drawing volume, almost no text-showing operators, almost no distinct characters actually rendered. Try the three conditions live: **Font decodability, with a fallback chain.** Not every font that's *present* in a PDF is used — the detector only inspects fonts actually invoked via a text-show operator. If those fonts are Type0/Identity-H without a `ToUnicode` CMap, decoding produces garbage characters. Rather than flag that immediately, it tries two fallbacks first — CID values that happen to look like passthrough Unicode, then an embedded TrueType `cmap` table lookup — before finally giving up and marking the page `suspected_garbled_text`. **A newspaper-layout detector.** Even a page classified `TextBased` can still need OCR: dense multi-column prose with a low font-change-to-text-op ratio reads badly as extracted text even when every character decodes correctly. The thresholds for this one are, per the source comments, calibrated against a named 50-page *Wall Street Journal* test PDF plus DPA/contract PDFs and SEC filings — a heuristic tuned against specific real documents, not an abstract rule. None of this is a neural network. It's operator-stream scanning over the page's content stream, with page sampling (8 evenly-spread pages by default, not "bail on the first bad page" — the source comment explains why: an image-only cover page followed by dense text, like most annual reports, would trip an early-exit strategy into over-flagging OCR). ## Where it lands on a benchmark Firecrawl's own July 31, 2026 benchmark, run on an Apple M4 Pro against the 200-PDF `opendataloader-bench` corpus, with OCR disabled and only non-ML local engines in the comparison: pdf-inspector edges liteparse by 0.002 overall, but wins tables decisively (TEDS 0.814 vs 0.693) and is roughly **1.6× faster** (0.470s vs 0.750s for the full corpus) — while losing on headings (MHS 0.788 vs 0.811). pymupdf4llm and markitdown aren't close on tables or speed. This is a genuinely tight three-way race at the top, not a rout. The comparison set is explicitly scoped: "only local engines without model-based PDF parsing are shown; OCR was disabled." This is not a claim of beating Docling, LlamaParse, or other vision-model-based extractors — it's the fastest option in the non-ML, non-OCR lane, and the README says so directly. ## Honest gaps This is Firecrawl's own benchmark, on Firecrawl's own hardware, published in Firecrawl's own README — there's no independent re-run I could find, though the corpus and evaluator are public and a reproducible-results branch is published, so a third party *could* check it (I did not). It's single-machine (one Apple M4 Pro), so there's no cross-platform or server-CPU number to point to. And the version story is a little tangled: the README's benchmark says it tested "pdf-inspector 0.2.6," which matches none of the three independently-versioned language bindings cleanly (Rust crate 0.1.7, npm package 1.11.2) — normal for a multi-target Rust project, but worth knowing if you go looking for "the" version number. ## The take There's no dramatic headline number here and no diagram to embed — the README's own "figure" is a Markdown table. What's worth taking from pdf-inspector is smaller and more useful: three specific, well-reasoned heuristics (path-op density, a font-decodability fallback chain, a newspaper-layout detector tuned against named documents) that collectively do a job people increasingly reach for an ML model to do, at a fraction of the cost, on the specific slice of the problem where a heuristic is the right tool. Not every classification problem needs a model. This one apparently doesn't. --- *Source: the [pdf-inspector README and benchmark](https://github.com/firecrawl/pdf-inspector) (Firecrawl, MIT, refreshed 2026-07-31) and `src/detector.rs`. The heuristic thresholds and benchmark numbers are the project's own; the interactive is my reconstruction of the vector-text condition for explanation, not a copy of the crate's code.* --- # Qwen-CUA: a computer-use agent that only ever sees pixels > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/qwen-cua > date: 2026-08-03 > tags: agents, computer-use, reinforcement-learning, moe, systems Qwen Team and XLang Lab published [Qwen-CUA](https://github.com/xlang-ai/Qwen-CUA) on 2026-08-02: a computer-use agent that never sees anything but a screenshot and never acts through anything but keyboard and mouse events. No DOM tree, no accessibility metadata, no task-specific API. The backbone is a 397B-A17B Qwen mixture-of-experts model, and a scaled variant, Qwen-CUA-Max, pushes past one trillion total parameters. The headline number is 86.2 on OSWorld-Verified. That is a real result, but it is one of eight benchmarks, and it is not even the most interesting fact in the paper.
The part worth taking apart is the context-management scheme that makes long-horizon screenshot-only control workable at all: fold the visual history in blocks of 10, not one screenshot at a time. It sounds like a minor implementation detail. It is actually the difference between a rollout fleet that reuses its KV-cache and one that recomputes a fresh prompt prefix on every single turn. ## A narrow interface on purpose Most production computer-use systems cheat a little. They read the DOM, they call an accessibility API, they get coordinates for free. Qwen-CUA's interface is deliberately narrower than that — the model observes a screenshot and emits one action from a fixed keyboard-and-mouse vocabulary (Appendix A): | Category | Actions | |---|---| | Keyboard | `key`, `key down` / `key up`, `type` | | Mouse | `move`, `left`/`right`/`middle click`, `double`/`triple click`, `left click drag`, `left mouse down`/`up`, `scroll`/`hscroll` | | Control | `screenshot`, `wait`, `terminate` (success or failure), `call user` | That last one, `call user`, is the tell that this is meant to run unattended: when the task cannot proceed autonomously — a login wall, a genuinely ambiguous instruction — the agent is allowed to stop and ask, rather than guess and keep going. This is close to the opposite design point from a coding harness. [Agent harnesses](/articles/agent-harness) walks Lilian Weng's tool taxonomy for coding agents — `bash`, `edit`, `grep`, `git_status` — and her argument for why: the tools are "deliberately simple and generic" because the model has already seen a million shell sessions in training. Qwen-CUA leans on the same instinct pointed at a different substrate. It does not give the model `bash` or a DOM query; it gives it exactly what a person gets — a screen and two input devices — on the bet that native computer use is "a sufficiently general interface for interacting with almost any software accessible to a person." The tradeoff is real: a shell command can rename 400 files in one call, while Qwen-CUA has to click, drag, and type its way through the same job one primitive at a time. The payoff is that the interface never goes stale — it works on software that has no API at all, which is most software. ## The mechanism: folding the visual prefix in blocks of 10 Screenshot-only control has an obvious cost: every turn adds an image to the context, and a long task can run for a hundred turns. Two bad options present themselves. Keep every screenshot, and the context blows past any practical budget. Use a sliding window and drop the oldest ones, and the agent forgets what it did five minutes ago — the exact state that explains why the screen looks the way it does now. Qwen-CUA's answer is in Figure 3 of the paper: scale the *active* visual history to 20 screenshots (up from 1 in Qwen2.5, 5 in Qwen3, 10 in Qwen3.5 — successive Qwen generations have simply been raising this number), and once the active window would exceed 20, fold the oldest 10 screenshots at once into a fixed textual placeholder. The reasoning and actions tied to those folded screenshots stay in the conversation; only the pixels get replaced.
The "at once" is the whole design. Fold one screenshot per turn — the obvious way to enforce a 20-image budget — and the folded-prefix boundary moves on every turn, so the text before the newest screenshot is different from what it was a moment ago. Fold 10 at a time instead, and the boundary only moves every 10 turns: steps 21 through 30 all extend the exact same prefix. Step through it below. Training uses the identical operator. Reinforcement-learning episodes are sliced into context-bounded chunks by advancing the same fold boundary, each slice inherits the full terminal reward, and only the model's own generated tokens count toward the loss. Train and inference see the same folding rule, which is the detail that keeps this from being an inference-time hack layered on top of training that never saw it. ## Why prefix stability is a rollout-economics problem Here is the part the paper is explicit about and worth spelling out: a stable prefix is not a memory nicety, it is a **KV-cache reuse story**. An inference server that serves the same prompt prefix repeatedly can cache the attention keys and values for that prefix once and reuse them for every subsequent request that shares it — skipping the prefill compute for everything except the new tokens at the end. A prefix that changes on every turn gets none of that: every request looks new to the cache, so every request pays full prefill cost. The paper names this directly, citing Anthropic's cache-aware batched-pruning guidance for computer use as the precedent for the design. Multiply that by scale. Qwen-CUA's training infrastructure is a cloud rollout fleet with close to 100,000 vCPUs and tens of thousands of concurrent environments, generating roughly 40,000 verifiable tasks' worth of trajectories. At that volume, the difference between "the prefix changes every turn" and "the prefix is stable for 9 turns out of 10" is not a rounding error in the compute bill — it is close to an order-of-magnitude difference in how much of the prefill work has to be redone per rollout step. Folding 10 at a time instead of 1 at a time is, in effect, a decision about how much of a 100,000-vCPU cluster's time goes to recomputing text it has already computed. It is the same underlying instinct as the file-backed context strategy in [Agent harnesses](/articles/agent-harness) — treat context as a bounded, managed resource instead of an ever-growing transcript — aimed at a different bottleneck. Weng's harness spills durable state to files so the model's context stays flat. Qwen-CUA can't spill screenshots to a filesystem the model can `grep`; there is no text index over pixels. So it does the analogous thing structurally: collapse old state into a fixed, cheap textual stand-in, and do it in a way that happens to also keep the serving engine's cache warm. Same principle — bound what has to be reprocessed — solved with the tool available to a vision-and-text model instead of a coding agent. ## Training: verifiable rewards at rollout-fleet scale The RL recipe is RLVR — reinforcement learning with verifiable rewards — using **Soft Adaptive Policy Optimization (SAPO)**, a smooth, temperature-gated alternative to PPO-style hard clipping (Gao et al., 2025, not original to this paper). The gate temperature is asymmetric: `τ_pos = 1.0`, `τ_neg = 1.05`, so tokens on non-positive-advantage trajectories decay faster than tokens on positive ones — called out as important specifically for long multimodal trajectories on an MoE backbone. Task-pool calibration runs 8 trial rollouts per candidate task and keeps only the ones with a mix of successes and failures, discarding tasks that are already saturated or unreachable. | Config | Value | |---|---| | Group size (valid trajectories/task) | 16 | | Oversampling before filtering | 20 candidates | | Outer batch size | 128 prompts (up to 2,048 valid trajectories/update) | | Optimizer | AdamW, LR `1e-6` constant, no warmup | | SAPO `τ_pos` / `τ_neg` | 1.0 / 1.05 | | Total updates | 1,000 | | Max turns/episode | 100 | | Max context (after slicing) | 144K tokens | | Slice interval | every 10 turn-pairs | The distributed setup is 512 H200 GPUs across 64 nodes, split disaggregated-style (32 training, 32 rollout, `verl`-style), with SGLang serving the rollout side. A full 1,000-update run takes about 5 days, roughly 61,440 H200 GPU-hours, holding upwards of 2,000 environments active concurrently at better than 75% average utilization. Across the training curve, the cross-domain validation score climbs from about 0.734 before RL to a peak of 0.770 at checkpoint 40 — the checkpoint the paper actually ships — before drifting slightly to 0.762 by the final checkpoint 50. That's a real, disclosed detail: the best model on the training curve is not the last one. Data comes from three sources layered together: environment-interaction tasks built off a feature taxonomy, user-interactive tasks with a simulated user holding back task-specific knowledge (the OSWorld 2.0 setting), long-horizon tasks chained through verifiable phase states, and personalized workflows collected from human trajectories in everyday and professional software — CAD tools and Blender included — with reasoning reconstructed via model-assisted chain-of-thought from the raw (task, screenshot, action, resulting state) tuples. ## Where it actually lands: eight benchmarks, not one The 86.2 on OSWorld-Verified is real, and it is the best score in the set on that particular benchmark. It is also the exception. Across the other seven benchmarks the picture is more mixed — Qwen-CUA leads outright on two of eight, is close behind on several, and loses outright on the rest, most notably safety. Pick a benchmark: Scaling the same recipe to Qwen-CUA-Max (over 1 trillion total parameters) moves OSWorld-Verified from 86.2 to 87.6, and helps more on partial-credit long-horizon completion: 1T)", value: 87.6, highlight: true }, ]} /> On the safety benchmark, RedTeamCUA, Qwen-CUA is a clear improvement over its own predecessor and a clear loss against Claude Opus 4.8. RedTeamCUA runs indirect prompt injection through ownCloud, Rocket.Chat, and Reddit environments and jointly reports benign task success and attack success rate (ASR — how often the injected instruction actually hijacks the agent): A 20.2-point reduction in attack success versus the previous Qwen generation is a genuine gain. It is also more than 20 times Opus 4.8's ASR. The paper states its own limits plainly here: "RedTeamCUA therefore shows improved resistance to indirect prompt injection, not a deployment-safety guarantee." Worth repeating rather than softening. ## Efficiency: the gain is not longer reasoning One honest, checkable claim in the paper: Qwen-CUA's OSWorld-Verified score does not come from generating more tokens per task. It reaches 86.2 at 3,605.8 output tokens per task; Claude Opus 4.8 needs a similar budget to reach 80.0 and roughly 21,800 tokens to reach 83.3.
The second panel is where the paper pre-empts its own obvious gotcha. On OSWorld 2.0, Qwen-CUA averages 218.9 turns per task against 83.5 for GPT-5.5 and 105.7 for Opus 4.8 — a turn count that looks far worse. But GPT-5.5 and Opus 4.8 can batch several actions into one turn; Qwen-CUA emits exactly one native action per turn by construction. The turn-count gap is mostly an artifact of how each interface packages low-level actions, not evidence that Qwen-CUA needs more attempts to do the same work. The paper says as much itself rather than leaving a reader to work it out. A related experiment adds a Bash tool alongside native computer use on MyPCBench: trajectories get shorter for every model tested, but task completion drops too, for Qwen-CUA and Qwen-3.7 specifically — the paper frames this as an unresolved "capability-efficiency frontier," not a win. ## Grading your own exam Here is the fact that belongs next to the 86.2, not three pages after it: **XLang Lab built OSWorld and OSWorld-Verified, and XLang Lab co-authored this paper.** The lab that defines what counts as a passing score on the headline benchmark is also a lab reporting how well its own model does on that benchmark. The paper does not flag this anywhere as a conflict of interest — it is simply true of the author list and the benchmark's provenance, stated here as a fact about who is grading whom, not as an accusation of anything specific. The eval protocol has a second, quieter honesty issue: baselines are not run under matched inference budgets. Per the paper's own settings, Qwen-3.7 is evaluated in non-thinking mode, GPT-5.5 runs with `xhigh` reasoning effort, and Claude Opus 4.8 runs at its max inference setting. "Most scores for comparison models are taken from official reports released by the corresponding benchmark or model providers" — for the ones the authors reproduced themselves, the settings differ by model, and the paper does not report what a matched-budget comparison would look like. The Gym-Anything table is the one place a second Opus 4.8 setting (medium) appears alongside max, and the two settings score 43.7 versus 47.3 — a 3.6-point swing from inference budget alone, which gives some sense of how much slack "differing settings" can hide. There is a third thing worth naming that I could not find explained anywhere in the paper. Figure 1's legend lists six systems, not the four in Table 1 and everywhere else in the text — it adds **Muse-Spark-1.1**, scoring 80.8 on OSWorld-Verified and 47.3 on Gym-Anything. Searching the full extracted paper text, that name appears exactly once: in the Figure 1 legend. It is not in Table 1, not in the eval-settings section, not in the references, not identified anywhere else in 24 pages. I don't know what it is or why it only appears in one chart. Two more disclosed-but-real caveats round this out. MacAgentBench's "clock" domain scored 0.0% for every model across all 12 tasks; the paper reports manually inspecting the trajectories, finding they looked like correct completions, and keeping the official 0.0% score anyway rather than quietly correcting it — which means the reported 69.2 aggregate is very likely a slight undercount, in Qwen-CUA's favor by omission, and the paper says so. And Gym-Anything's headline 46.3 runs on 97 of 197 possible environments; the other 100 were excluded because their Windows, Android, or Linux setups didn't work, not because they were held out for any principled reason. Both caveats are in the paper. Neither is in the abstract. Finally: several contributors are marked in the author list as having departed the Qwen Team by the time this was published, including researchers who worked on the original OSWorld and OpenCUA lines. The paper doesn't explain the departures, and neither can I — it's listed here because it's a real, checkable detail about who built this and who was still there to see it ship. ## The take Native computer use is not new — UI-TARS, OpenCUA, Aguvis, and AutoGLM already established that a single model can ground pixels to actions without a separate grounding stage. What Qwen-CUA adds is mostly an engineering answer to what happens when you actually try to run that idea at rollout-fleet scale: fold visual history in blocks, not one screenshot at a time, so a 100,000-vCPU cluster spends its time on new work instead of recomputing prefixes it already has. The paper's own framing for where this goes next is worth keeping: "we view native computer use not as the only action interface, but as the universal grounding and fallback layer of a hybrid agent" — paired eventually with something more like the coding-harness tool table in [Agent harnesses](/articles/agent-harness), not replacing it. The honest scorecard is two benchmark wins out of eight, a real safety improvement that still trails the safest competitor by more than 20x on attack success, and a headline number graded in part by the lab that wrote the exam. All three of those things can be true about a genuinely useful piece of systems engineering at the same time. --- *Built on Qwen Team & XLang Lab's [Qwen-CUA: Native Computer Use for (almost) Everything](https://github.com/xlang-ai/Qwen-CUA) (2026-08-02). Figures 1, 3, and 6 are reproduced from the paper for commentary, flattened onto white and cropped from the original PDF; the interactive fold timeline and benchmark explorer are my own illustrations of the mechanism and Table 1's data, not measured traces. Benchmark numbers are as reported in the paper.* --- # Qwen3.8-Max: 16 days, 265 commits, zero humans in the loop > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/qwen3-8-max > date: 2026-08-03 > tags: qwen, agents, llm, benchmarks, moe Alibaba announced [Qwen3.8-Max](https://qwen.ai/blog?id=qwen3.8) today: **2.4 trillion parameters, 95B active**, built on the architectural foundation of Qwen 3.5. It is also, by Alibaba's own framing, the first Qwen-Max-class model getting open weights at all — those weights are announced, not released; they ship "next week." Right now the only way to use Qwen3.8-Max is the API, through [QwenCloud](https://www.qwencloud.com/). A 2.4T/95B split puts activation sparsity at about 25× (2,400 / 95). That is close to [Kimi K3](/articles/kimi-k3)'s roughly 27× (2.78T total, 104.2B active) — two labs, released weeks apart, converging on almost the same ratio of total-to-active parameters at the very top of the open-weight-adjacent scale. Where K3 backs that ratio with a 47-page technical report anyone can audit against a released `config.json`, Qwen3.8-Max's architecture claims are, for now, a paragraph in a blog post. The open weights next week will be the point where the second half of that comparison becomes checkable.
But the benchmark grid is not the interesting part of this release. The interesting part is what Alibaba says the model did **completely unsupervised**, for days at a time. ## The case studies are the real headline Every frontier lab now publishes agentic benchmark numbers. Fewer publish concrete, checkable claims about what their model actually built when nobody was watching it. Qwen3.8-Max's release includes five of those, spanning a 24-hour coding contest to a 365-simulated-day economy, and they are, collectively, some of the most specific long-horizon autonomy claims I have seen from any lab this year — specific enough that at least one of them (the coding harness) has a public commit history you can go read yourself. ### 16 days, no one watching: oh-my-cli Alibaba tasked Qwen3.8-Max with building `oh-my-cli` — a CLI tool — from an empty repository, and kept it running. The loop it built for itself: an issue state machine moves work through `ready → leased → active`; an agent claims a task, implements it, and triggers Build, Unit Test, E2E, and Desktop Lifecycle validation; failures route back to the originating issue for another pass; passing PRs merge. Community feedback and the model's own test results both feed back in as new issues, so the harness is quite literally evolving its own capabilities (`/goal`, `/resume`, Dynamic Workflow, Session Replay, Desktop) as it runs. As of July 30, 2026 — about 16 days in — the repository held **265 commits, 127 PRs, and 151 issues**, all without a human merging, reviewing, or filing anything. What makes this claim unusually checkable is that the trace is public: [github.com/qwen-code-dev-bot/oh-my-cli](https://github.com/qwen-code-dev-bot/oh-my-cli). Most "our agent ran autonomously for weeks" claims ask you to take the vendor's word for it. This one, you can go read commit-by-commit. ### Reproduce a paper, then beat it Handed only a citation — [arXiv 2605.22389](https://arxiv.org/abs/2605.22389), "Unified Data Selection for LLM Reasoning" — and a set of GPUs, Qwen3.8-Max had to write the entire pipeline from nothing: no starter code, no scaffold. The paper's claim is that when you have more training data than compute to use it on, the examples worth keeping are the ones full of "hard decision points" — places in a worked solution where the model was genuinely torn between next steps. Over roughly 125 hours (about five days) of continuous, unattended work, Qwen3.8-Max wrote about **7,600 lines of code**, took **over 1,100 actions**, and ran **33 rounds of GPU training**. The first ~37 hours went into rebuilding the paper's pipeline from zero and reproducing all six of its findings — including the headline result, that the paper's selection method beats random selection by +7.7% on AIME24 after fine-tuning Qwen3-8B on the selected data. Then it kept going. The next ~88 hours ran a self-improving loop — form a hypothesis, write the code, run it on GPUs, analyze the result, try again — across four rounds and 18 self-generated ideas, each round's diagnosis shaping the next round's hypothesis: | Round | Best idea that round | AIME24 | Gain vs. baseline | |---|---|---|---| | — | Paper's method, reproduced (baseline) | 49.58% | — | | 1 | Split the data by difficulty before selecting | 50.42% | +0.84 | | 2 | Weight examples by an entropy–score gap | 51.67% | +2.09 | | 3 | Tune the selection width | 51.25% | +1.67 | | 4 | Count the hard decision points ("nhighgate") | **52.29%** | **+2.71** | The final method — round 4's "nhighgate" idea — is not a tweak Alibaba fed it. It is something the model proposed, tested, and kept because it worked better than the paper it was asked to reproduce. ### Beat 526 human teams in 24 hours Alibaba entered Qwen3.8-Max into a live competition on its own Tianchi platform — the [WWW2025 Multimodal Dialogue Intent Recognition Challenge](https://tianchi.aliyun.com/competition/entrance/532277) — against **526 human teams**, under a strict 24-hour clock. The task: read customer-service conversations, text and screenshots both, and classify what the customer actually wants. Qwen3.8-Max read the rules and built a solution in code: fine-tuned and ensembled BERT, MacBERT, and RoBERTa for the text side; fine-tuned Qwen2.5-VL-7B, backed by Chinese-CLIP for uncertain cases, for the screenshots; then fused everything through a weighted-voting system, calibrated by cross-validation and re-weighted after every submission. Across **45 submissions** in 24 hours, accuracy climbed from **0.60 to 0.853**, finishing ahead of **458 of the 526 human teams — 87% of the field**. ## Autonomous chip design: a staircase that keeps stepping down The most technically specific of the five case studies is silicon. Qwen3.8-Max was given a stub RTL workspace, an empty-module task description, and an evaluation script — no golden reference design, no human in the loop — and told to design a GCD/RSA cryptographic hardware accelerator that stays bit-exact across 4-, 6-, 8-, and 16-bit configurations under a randomized `cocotb` verification framework, while minimizing synthesized gate count. Over one continuous run of **about 500 turns and 71 evaluations across 13 milestones**, working entirely in a sandbox wired to Iverilog (simulation), Yosys (synthesis), and OpenROAD (physical layout), the model's first functionally correct design measured 8,298 gates. It drove that down to **678 gates** — a 91.8% reduction — and the milestones along the way show real architectural rewrites happening deep into the run, not just early low-hanging fruit: The single biggest step, by far, was recognizing that the 16-bit hardware modulo divider inside `modular_multiplier` could be replaced with an iterative shift-subtract architecture — one change, at turn 22, cutting 6,288 gates in a single move, over 80% of the entire reduction. Everything after that is smaller, later, and arguably harder: register and FSM pruning at turns 60–113, module fusion at turns 170–252, and gate-level refinement all the way out to turn 500. A model that only found the big early win and then plateaued would be a much less interesting story than one that kept finding real (if progressively smaller) structural improvements for 400 more turns. Alibaba then re-ran the final RTL through a physical place-and-route flow (OpenROAD, Nangate45) to check whether the front-end gate-count win actually routes. It does: the die shrank from 106×106 to 46×46 µm² (−81%), wirelength dropped from 33,369 to 4,187 µm, and the design closed timing at 500 MHz with **positive** slack (+0.66 ns), up from a failing −4.46 ns at the start. Optimizing gate count without checking place-and-route is a common way to produce a design that looks good on paper and doesn't actually work in silicon; Alibaba closed that loop. ## 365 simulated days of running a business The last of the five case studies is not a coding task at all. **E-Commerce Bench** simulates a full year of operating online stores against desensitized real Taobao/Tmall transaction data — 12 store types, 60 product categories, nearly 600 suppliers, 7,000 products — starting from ¥100,000 in capital. The model has to choose products, negotiate with suppliers, manage inventory, price dynamically, and handle returns, all while surviving seasonal demand swings, sudden supply shocks (typhoons, material shortages), and a settlement system with real cash-flow pressure. Buried in the supplier matrix: **152 fraudulent merchants**, running classic scams — membership-fee traps, low-price bait, goods not as described. Two things stand out in how Qwen3.8-Max played this. First, supplier negotiation is modeled with distinct personalities and concession strategies per supplier, and Qwen3.8-Max's negotiation efficiency measurably *improved* over the year — the same products from the same suppliers got progressively cheaper, and that experience generalized to similar products, where Alibaba says other models' negotiation efficiency plateaued mid-year. Second, it front-loaded capital early to establish position rather than playing conservatively, then converted the resulting inventory and operating gains back to cash before the simulation ended — timing that matters because unconverted assets left on the books at year-end hurt the final score. The result: a final balance of **¥416,252 — a 4.16× return** — 38% ahead of second-place GLM 5.2 and 152% ahead of Qwen3.8-Max's own predecessor, Qwen3.7-Max. Alibaba frames this as evidence of "adaptive learning from transactional feedback" across more than 2,000 rounds of interaction, rather than a model that locked in a strategy early and rode it out. That framing is plausible given the negotiation-efficiency detail above, but it is worth remembering this is Alibaba's own simulation, built on Alibaba's own marketplace data, scored by Alibaba. ## Scaling real-world work Underneath all five case studies is an infrastructure bet Alibaba is explicit about: jointly scaling RL environments and compute lifts "general working competence" across several harnesses at once (QwenWork, Claude Code, Codex, OpenClaw, Hermes), and doing that required three things to scale together rather than one at a time — environments along independent axes (task, workspace, harness) so growth compounds instead of requiring bespoke integration per new environment; a **universal reward system** unifying execution-based checks, rubric-conditioned judging over text and rendered visual output, and agentic inspection, so there is one reward mechanism instead of a pile of task-specific verifiers; and an **online data balancer** that keeps every training batch balanced across task, difficulty, workspace, and harness, which is what keeps gradient variance from blowing up RL training at scale.
That chart is worth reading carefully rather than just squinting at the upward trend: the curve peaks at 4,000 environments (0.725) and is already *down* to 0.689 by 5,000 — the shipped checkpoint is not the last one on the curve, which is the kind of disclosed detail that makes the rest of the curve more credible, not less.
That second chart is the practical payoff of training against a harness-agnostic reward system: Qwen3.8-Max does not have one harness it happens to be tuned for. Point it at QwenWork, Claude Code, Codex, OpenClaw, or Hermes and the CoWorkBench score moves in a band of about 73–76; Fable5 and Opus4.8, each shown in only their own native harness, land in a similar range without ever being tested for harness portability the same way. This is directly the concern [The harness effect](/articles/harness-effect) raises from the other side — that orchestration, not the model, is what actually determines an agent's cost and reliability on a task — and Qwen3.8-Max's answer is to train the reward system to not care which harness is wrapped around it, rather than picking one harness and optimizing hard for it. The same Dynamic Workflows capability that lets it self-orchestrate shows up in a quant-research vignette Alibaba includes alongside the five headline case studies: given a one-line task description, Qwen3.8-Max built a complete ETF-rotation strategy over several hours, pruning overfit factors when it noticed design-period and validation-period metrics diverging, and separately parallelized factor mining from six short descriptions into 50 research directions each, dispatching roughly 330 sub-agents through about 6,000 backtests to find factors with excess Sharpe ratios of 0.64–1.48. Whether that generalizes past a demo is unverifiable from a blog post, but the mechanism described — noticing an overfitting signal and automatically triggering pruning, mid-run, without being told to — is the same "acting on evidence instead of a fixed script" pattern that shows up in the chip-design and paper-reproduction case studies above. ## Multimodal and hybrid agents Qwen3.8-Max's visual pipeline gets a similar "watch itself work" framing: while executing a task, the model inspects its own intermediate results — page layout, object orientation, spatial relationships, animation quality — and revises when something looks wrong (a television facing backward, a misaligned interface). Alibaba's phrase for this is a "native feedback loop across planning, execution, verification, and iteration," which is a reasonable description if the examples given hold up, though none of them are independently reproducible from the blog post alone. The concrete new benchmark here is **RecreationBench**: the model observes a real running application as a black box — no source code, no network access — across five platforms (Ubuntu, macOS, Windows, Android, web), and has to rebuild the whole thing from scratch through interaction alone. Alibaba frames Qwen3.8-Max's showing here as "frontier-level Hybrid Agent capability" — the pairing of writing code (does the heavy lifting) with operating a GUI directly (reaches whatever a human can see and click, and reports back what a live system actually does). That second half — driving a computer through screenshots and input events alone — is exactly [Qwen-CUA](/articles/qwen-cua)'s whole premise, published by a different Qwen team one day earlier. It is worth putting the two numbers next to each other: Qwen3.8-Max reports **86.1** on OSWorld-Verified; Qwen-CUA, a dedicated 397B-A17B computer-use specialist trained specifically for this, reports **86.2**. A general-purpose 2.4T model and a purpose-built computer-use agent land within a tenth of a point of each other on the benchmark that agent was built for — which either means Qwen3.8-Max's general agentic training has genuinely absorbed computer-use skill, or that OSWorld-Verified has a ceiling both are bumping into. Both readings are consistent with the data; the blog post doesn't say which. ## The benchmark tables — and where they don't hold up Here is the full picture, reproduced from Alibaba's own release. The pattern is not "Qwen3.8-Max wins everything" — it wins some things outright, loses some things clearly, and several of its best numbers come with an asterisk worth reading before you trust them. ### Coding Agent | Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max | |---|---|---|---|---|---| | Terminal Bench 2.1 | 84.6 | 84.6 | 88.8 | 74.5 | 86.6 | | SWE-bench Pro | 69.2 | 80.0 | 64.6 | 60.6 | 67.7 | | DeepSWE 1.1 | 59.0 | 70.0 | 73.0 | 21.6 | 56.6 | | NL2Repo-Bench | 69.4 | -- | -- | 47.2 | 55.9 | | FrontierSWE | 70.0 | 88.8 | -- | 40.7 | 73.5 | | MLS-Bench-Lite | 42.8 | 49.9 | 46.2 | 31.7 | 41.0 | | PaperBench | 80.3 | 88.8 | 90.5 | 64.8 | **93.0** | | AndroidBench | 69.8 | 84.5 | 74.0 | 56.5 | 75.1 | | QwenSWEBench | 84.0 | 86.3 | 73.5 | 63.4 | 80.7 | | QwenQoderBench | 62.7 | 63.1 | 53.8 | 36.8 | 58.4 | | QwenReactBench | 1694 | 1770 | 1564 | 1538 | 1724 | | QwenSVGBench | 1648 | 1690 | 1758 | 1499 | 1713 | ### General Agent | Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max | |---|---|---|---|---|---| | CoWorkBench | 72.3 | 75.9 | 71.5 | 64.6 | 74.8 | | WorkSpaceBench | 66.8 | 68.7 | 65.6 | 61.4 | 67.7 | | JobBench | 48.4 | 57.4 | 45.4 | 31.3 | 53.4 | | SkillsBench | 65.1 | 70.9 | 73.5 | 61.2 | 70.2 | | Agents' Last Exam (Pass / Score) | 27.0 / 45.1 | -- / -- | 30.6 / 53.6 | 11.8 / 31.1 | 27.0 / 52.4 | | Automation-Bench (Pass@1) | 27.2 | 29.1 | 29.7 | 14.2 | 27.3 | | Toolathlon Verified (Pass@1) | 76.2 | 77.9 | 74.9 | 49.7 | 72.5 | | WideSearch | 72.9 | 81.2 | -- | 75.2 | **81.9** | | HLE w/ tools | 57.9 | 64.5 | 58.0 | 53.5 | 56.2 | ### General Capabilities | Benchmark | Opus4.8 | Fable5 | GPT5.6 Sol (max) | Qwen3.7-Max | Qwen3.8-Max | |---|---|---|---|---|---| | GPQA Diamond | 92.0 | 92.6 | 94.1 | 92.4 | 92.6 | | HLE | 45.7 | 53.3 | 47.2 | 41.4 | 43.6 | | IFBench | 62.2 | 63.5 | 72.7 | 79.1 | **82.8** | | $OneMillion-Bench (expert score) | 41.8 | 55.9 | 53.8 | 44.4 | 52.5 | | HealthBench | 52.4 | -- | 55.3 | 54.5 | **60.2** | | PLawBench | 69.6 | 70.2 | 72.3 | 58.9 | **73.2** | | PRBench-Legal | 52.7 | 57.6 | 57.6 | 48.5 | 57.6 | | PRBench-Finance | 51.9 | 55.8 | 55.5 | 46.8 | **58.3** | | MRCR v2 256K (8-needle) | 83.2 | -- | 93.8 | 86.7 | 92.9 | | LongBench v2 | 69.1 | -- | 67.1 | 65.3 | 66.3 | **Read the footnotes before you trust the wins.** Alibaba's own notes on this table say, plainly: - **Terminal Bench 2.1**: Qwen3.8-Max is evaluated with Claude Code at avg@10 (5h timeout, 131,072 max tokens). Every other model is scored at "the best published score across harnesses" — Opus4.8/Fable5 via Artificial Analysis, GPT5.6 Sol via OpenAI's own post. Best-of-published vs. one model's avg@10 is not the same measurement. - **SkillsBench**: a different harness per model — Opus4.8 and Fable5 on Claude Code, GPT5.6 Sol on Codex, the entire Qwen series on OpenCode. - **DeepSWE 1.1**: Qwen3.8-Max is scored on whichever of Claude Code / mini-SWE-agent is higher, and Alibaba notes it does best specifically on Claude Code — the harness closest to what it trains against. - **PaperBench**'s 93.0 — the highest score in the table — is judged by **Claude Opus 4.6**, a competitor model, not an automated or human grader. - **$OneMillion-Bench** and **PLawBench** are both judged by **gemini-3.1-pro-preview**. - Footnote 1, verbatim: "Fable5 results may involve fallbacks." - QwenSWEBench, QwenQoderBench, QwenReactBench, QwenSVGBench, CoWorkBench, and WorkSpaceBench are **Alibaba's own in-house benchmarks**, evaluated in-house. None of this means the wins are fake. It means "Qwen3.8-Max leads Terminal Bench 2.1" and "PaperBench's judge is a Claude model" are both true at once, and a benchmark table alone won't tell you that — you have to read footnote 2 and footnote 8. Explore the same numbers benchmark-by-benchmark, with the eval-setup note attached to whichever row has one: ### Multimodal (selected rows) The full multimodal table runs to roughly fifty rows across six categories; here are the ones that matter most for the agentic and visual-agent story above, using the table's own column set (Gemini3.1-Pro and Qwen3.7-Plus replace GPT5.6 Sol/Qwen3.7-Max from the tables above — Alibaba compares against a different baseline set for multimodal). | Benchmark | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | Qwen3.8-Max | |---|---|---|---|---|---|---| | MMMU-Pro | 75.6 | 81.2 | 80.5 | 83.0 | 79.0 | 82.3 | | LogicVista | 76.7 | 85.7 | 82.6 | 89.7 | 84.3 | **91.9** | | HiPhO | 69.3 | 78.6 | 85.4 | 86.8 | 84.1 | **90.0** | | OSWorld-Verified | 83.4 | 85.0 | 76.2 | 83.2 | 73.3 | **86.1** | | OSWorld 2.0 (binary / partial) | 20.6 / 54.8 | -- / 66.1 | 7.8 / 30.6 | -- / 62.6 | 2.8 / 21.5 | 19.4 / 46.7 | | WebArena-Verified | 67.9 | 71.3 | 64.3 | 69.7 | 55.3 | 66.8 | | Parametric CAD Bench | 85.1 | 87.5 | 73.5 | 86.2 | 73.8 | **91.5** | | VLMsAreBiased | 43.8 | 61.2 | 74.1 | 59.8 | 36.6 | **88.3** | | Dense200 | 20.8 | 31.1 | 69.7 | 55.3 | 60.7 | **87.0** | **Where it clearly loses, so this isn't a highlight reel:** DeepSWE 1.1 (56.6 vs. GPT5.6 Sol's 73.0, Fable5's 70.0), SWE-bench Pro (67.7 vs. Fable5's 80.0), HLE (43.6 vs. Fable5's 53.3), HLE w/ tools (56.2 vs. Fable5's 64.5), MLS-Bench-Lite (41.0 vs. Fable5's 49.9), Toolathlon Verified (72.5 vs. Fable5's 77.9), WebArena-Verified (66.8 vs. Fable5's 71.3), and OSWorld 2.0, where Fable5's own partial score (66.1) is well clear of Qwen3.8-Max's 46.7. Several of these — DeepSWE, SWE-bench Pro, HLE w/ tools — are exactly the categories where the harness or judge asymmetries above cut in Qwen3.8-Max's favor elsewhere, which makes the clean losses more credible, not less. The recurring problem across both tables is one [The harness effect](/articles/harness-effect) names directly: orchestration changes the number as much as the model does, so a table that scores different models on different harnesses is measuring two things at once and reporting only one. [Agent harnesses](/articles/agent-harness) makes the complementary point about what a harness actually *is* — tools, context management, control flow, an evaluator — which is exactly the layer these footnotes are quietly holding constant for some models and not others. None of that makes the underlying capability claims false. It does mean the honest reading of "Qwen3.8-Max leads Terminal Bench 2.1" is "leads it, evaluated differently than the models it's compared against" — a real result, with an asterisk that Alibaba, to its credit, discloses rather than hides. ## Getting it (or not) Right now, Qwen3.8-Max is API-only, via QwenCloud. The API exposes a `reasoning_effort` parameter with three levels — `xhigh` (default, for demanding tasks), `medium`, and `low` — and `preserve_thinking` is on by default. The notable integration detail: QwenCloud's API is compatible with both the OpenAI and **Anthropic** protocols, so pointing Claude Code at Qwen3.8-Max is a matter of setting `ANTHROPIC_BASE_URL` and an auth token, no separate client needed. It also plugs into Codex, Qoder CLI, Qwen Code, and OpenClaw with similarly small config changes. The open weights are, again, announced for "next week" — not this release. Until they land on Hugging Face and ModelScope, every claim in this article about what a 2.4T/95B model *is* rests on Alibaba's blog post and API behavior, not an inspectable checkpoint. That is a materially weaker evidentiary position than Kimi K3's, where the weights, a technical report, and a `config.json` all shipped together. Worth remembering the next time "first Qwen-Max-class open-weight model" gets repeated as though the weights were already out. ## The take Strip away the marketing framing and what is left is genuinely interesting: a model that, by its maker's account, ran a coding project unsupervised for 16 days with a public commit trail, rebuilt and then beat a research paper's method from a bare citation, out-negotiated a game-theoretic supplier matrix for a simulated year, and found real architectural wins in a chip design 400 turns after the obvious ones were gone. If even most of that holds up, it is a meaningfully more concrete set of long-horizon-autonomy claims than "our model scored X on benchmark Y." But the benchmark tables sitting next to those case studies are graded on a harness-by-harness, judge-by-competitor-model, in-house-benchmark basis that Alibaba discloses in footnotes rather than in the headline number — Terminal Bench 2.1 at avg@10 against everyone else's best-of-published, PaperBench judged by Opus 4.6, PLawBench judged by Gemini. That is not disqualifying. It is the same asymmetry every major lab's self-reported benchmark table has, and Alibaba's footnotes are, if anything, more forthcoming than most about exactly where the comparison stops being apples-to-apples. Read the case studies for what the model can apparently do unsupervised. Read the tables — and their footnotes — for how much weight the number itself can actually carry. They are not the same kind of evidence, and this release is unusually clear about which is which. --- *Built from [Alibaba/Qwen's Qwen3.8-Max release post](https://qwen.ai/blog?id=qwen3.8) (2026-08-03). Figures 1–3 are the post's own images — the overall performance grid, the RL-scaling curve, and the cross-harness generalization chart — flattened onto white for dark-mode compatibility and capped near 1600px; not reassembled or relabeled. The gate-count staircase and case-study switcher are original interactive reconstructions of the source's own numbers, not independently measured. Benchmark tables are Alibaba's; footnote caveats quoted or closely paraphrased from the source's own numbered notes. Weights and technical report were not available at publication time — every architectural and infrastructure claim here is Alibaba's, unverified against a released checkpoint.* --- # Recursive Harness Self-Improvement: beat your last harness, not a population of them > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/recursive-harness-self-improvement > date: 2026-08-03 > tags: agents, harness-optimization, information-theory, llm, explainer [Agent harnesses](/articles/agent-harness) argued that the loop wrapped around a model — tools, context policy, control flow — matters as much as the model's own intelligence. [The harness effect](/articles/harness-effect) showed that orchestration, not the model, is what actually sets an agent's token bill. Both pieces treat the harness as a thing worth engineering carefully by hand. [Recursive Harness Self-Improvement](https://arxiv.org/abs/2607.15524) (Lee, Xu, Seely, Lee, Zaharia, Tang — Sakana AI and UC Berkeley) asks the next question: can the harness improve *itself*? Their answer treats the harness as a single text prompt and updates it using nothing but a comparison against its own immediately-previous version. ## The idea, and the objective it can't afford The harness a coding agent runs under — roles, instructions, and the workflow connecting them — is, in RHI's framing, just a string $H$ drawn from a space of harnesses $\mathcal H$. Optimizing it against a broad population of competitors is the obvious move, and it's what most prior work does: $$ H^*_x \in \arg\max_{H \in \mathcal{H}} f_x(H), \qquad f_x(H) = \mathbb{E}_{H' \sim \mu,\; y \sim \mathcal{A}(H,x),\; y' \sim \mathcal{A}(H',x)} \big[\mathbf{1}\{y \succ y'\}\big] $$ $\mu$ is a distribution over competitor harnesses, $\mathcal A(H,x)$ is the agent running harness $H$ on task $x$, and $y \succ y'$ means an LLM judge preferred output $y$. The problem is cost: a population of size $m$ needs $m$ fresh agent executions and $\binom{m}{2}$ pairwise judgments per iteration — $\Theta(m^2)$ — before you can even take one optimization step. For a user continually specializing a harness to a new task, that's not a research inconvenience, it's prohibitive. RHI's relaxation replaces the population with a point mass on the harness's own previous version: $$ \tilde{f}_x^{(i)}(H) = \mathbb{E}_{y \sim \mathcal{A}(H,x),\; y^- \sim \mathcal{A}(H_x^{(i-1)},x)} \big[\mathbf{1}\{y \succ y^-\}\big] $$ One new execution, one comparison, cached forever after. $\Theta(1)$ per iteration, independent of how large a population you'd otherwise have wanted. ## Why comparing to yourself is still principled The obvious objection: isn't comparing only to your immediate predecessor a much weaker signal than comparing to a whole population? RHI's answer is a Bradley-Terry argument. Assume there's a latent task utility $u_x : \mathcal H \to \mathbb R$ and a link function $\sigma$ (strictly increasing, $\sigma(0) = \tfrac12$) such that $\Pr(H \succ H') = \sigma(u_x(H) - u_x(H'))$ — the standard pairwise-preference model. Then both objectives are monotone in the *same* latent utility: $$ f_x(H) = \mathbb{E}_{H' \sim \mu}\big[\sigma(u_x(H) - u_x(H'))\big], \qquad \tilde{f}_x^{(i)}(H) = \sigma\big(u_x(H) - u_x(H_x^{(i-1)})\big) $$ So any revision that beats $H^{(i-1)}_x$ with probability greater than one-half also increases the ideal, population-level objective. RHI performs **noisy local ascent** on the same utility ordering a much more expensive search would climb — it just takes a smaller, cheaper step each time, using the accumulated preference history as the only signal for which direction is up. There's no proof this converges, or how fast; it's a directional argument, not a guarantee. The algorithm this licenses is short. At iteration $i$: run the agent under $H^{(i)}$, get an output. Compare it against the cached output from $H^{(i-1)}$. Save the preference. Feed the accumulated preference history to an LLM harness optimizer, which writes $H^{(i+1)}$.
Critically, the harness optimizer never sees the evaluation prompt $x_{eval}$ directly — only the preference history, which was itself generated by a judge conditioned on $x_{eval}$. Alignment with the actual evaluation criteria happens indirectly, through the accumulated comparisons, not because the optimizer was told what's being graded. ## What actually gets rewritten RHI decomposes the harness into **agent design** (roles and instructions for each candidate agent) and **agent workflow**, which splits further into **contracts** — what information passes between subagents and the orchestrator — and **hops** — the interaction structure and control flow. The optimizer's own prompt is explicit about where to spend its edits: prioritize contracts and hops over roles and instructions.
The hypothesis behind that priority: a task-specific contract tells the orchestrator and subagents what to pass along instead of making them condition on the entire interaction history, which is "conceptually analogous to imposing a task-dependent sparsity pattern on inter-agent information flow" — sparse attention for agent communication, in effect. Better contracts should mean less redundant context, better cache efficiency, and lower cost, for free, alongside better task performance. ## Does it work, and what does it cost Across 30 synthetic ML-research tasks (finance, robotics, pharma), a few RHI iterations substantially raise the ceiling that test-time scaling alone can reach. With Opus-4.7, one iteration is enough to beat both `xhigh` and `max` reasoning-effort settings. With Opus-4.8, two iterations beat `xhigh`, `max`, **and** the provider's own built-in dynamic multi-agent harness, `ultracode` — a user-constructed, prompt-level harness beating a vendor's dynamic scaffold.
The gains aren't from longer outputs — normalized token usage stays roughly flat across iterations for Sonnet-4.6 and Opus-4.8 while win rate climbs (Opus-4.7's data can't separate the two hypotheses; only two iterations were run and its token count rose alongside performance, which the paper states plainly as inconclusive). What actually improves is cost, largely through less redundant cache read/write from better-managed context: The 60% figure is the abstract's headline, and it's the comparison against the provider's own dynamic multi-agent harness — not against a same-family reasoning-effort setting. A companion ablation (Appendix A) found something the paper didn't have to report: the provider's built-in multi-agent mode scores a **lower** Elo than running single-agent, despite costing far more — the vendor's own dynamic scaffold failing to pay for itself on this benchmark, stated without softening. ## The information-theoretic account Section 6.3 goes further than "it works" and proposes *why*: RHI implicitly maximizes task information in the components it's told to prioritize (contracts, hops) while minimizing task-conditional redundancy across all components. Formalized, $$ J(g_i) = \underbrace{\sum_{hc \in \mathcal{C}_{ext}} \frac{1}{K^{hc}_{Xi}} \sum_{k=1}^{K^{hc}_{Xi}} I\big(z^{hc,(i)}_{Xk}; X\big)}_{f_{ext}} \;-\; \beta \underbrace{\text{TC}\big\{z^{hc,(i)}_{Xk}\big\} \big|X}_{f_{int}}, \qquad \beta > 0 $$ $f_{ext}$ is mutual information between the *externally-emphasized* components (contracts, hops) and the task $X$; $f_{int}$ is the task-conditional total correlation — redundancy — across *all* four component types. The hypothesis: RHI is implicitly raising the first term and lowering the second, estimated here with canonical-correlation mutual information and total correlation over PCA-whitened sentence embeddings. The paper is careful about how much weight this deserves: it "does not prove that RHI optimizes a unique scalar objective," should be read as "an embedding-based proxy," and is explicitly "correlational rather than causal" — not a claim about what the optimizer LLM is actually doing internally, just a consistent pattern in what its edits produce. **Two honesty gaps worth stating plainly, because the paper doesn't headline them.** First, in the Opus-4.8 experiments, one of the two LLM judges is **`opus-4.8-xhigh` itself** — the same model family whose harness is being evaluated also serves as one of its own judges, scoring `opus-4.8` runs against other `opus-4.8` baselines. The paper does average across two judges (the second is an independent `gpt-5.5-max`), which dilutes the self-judging influence, but the paper does not discuss it as a potential bias source. Second, every agent tested belongs to one vendor — Sonnet-4.6, Opus-4.7, Opus-4.8 — with no cross-family test on GPT or Gemini as the agent under improvement (those families only ever appear as *judges*). And there is no empirical comparison against any of the roughly 15 directly competing methods the paper discusses in its own related work — Meta-Harness, Self-Harness, TTHE, ADAS, GPTSwarm, AFlow, GEPA, DSPy, and others. RHI is only benchmarked against same-family reasoning-effort scaling and the provider's built-in multi-agent mode, not against the alternatives it explicitly positions itself against. The benchmark itself is also self-constructed: 30 tasks synthesized from real job postings by an LLM, evaluated by the same lab that designed the method being tested on them. None of this means the result is wrong — the plain admission that Opus-4.7's token-length claim is inconclusive, and the decision to run and report the ablation showing the vendor's own multi-agent mode underperforms single-agent, are both the kind of finding a paper trying to look better than it is would have left out. ## The take The part of RHI I'd actually reuse is the trajectory-local relaxation itself: comparing only against your immediately-previous version turns an intractable population search into something you can run continuously, cheaply, and the Bradley-Terry argument for why that's still principled — not just convenient — is genuinely nice. The information-theoretic account of *why* it lands on contracts and hops is a good hypothesis, stated with the right hedges. What I'd want before trusting the magnitude of any specific number: a comparison against even one of the population-based methods it explicitly argues against, and an evaluator that isn't sometimes the model being graded. This reads as the first half of a real idea — the paper says so itself, calling the harness-to-model feedback loop "the second half" left to future work — and the half that's here is worth having, with its gaps named rather than papered over. --- *Source: [Recursive Harness Self-Improvement](https://arxiv.org/abs/2607.15524) (Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, Yujin Tang — Sakana AI, UC Berkeley), arXiv:2607.15524. Figures 1, 2, and 3 are reproduced from the paper for commentary; the interactives are mine, built on the paper's own reported formulas and measured endpoints.* --- # Sol-Attn: deciding which attention blocks to skip while you're already streaming them > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/sol-attn > date: 2026-08-03 > tags: diffusion, attention, sparse-attention, inference, video, explainer Video generation has an attention problem that language models mostly don't. A few seconds of video at a useful resolution is a very long token sequence, attention is quadratic in it, and diffusion runs the whole stack dozens of times per clip. So attention stops being *a* cost and becomes *the* cost. [Sol-Attn](https://arxiv.org/abs/2607.24027) — "Sparsifying online attention", from the SANA team at NVIDIA and collaborators — is a training-free way to skip most of that work. The idea I like is not that it sparsifies attention; everyone does that. It is *where* the decision happens. ## The problem with picking blocks Training-free sparse attention works block-wise: score each block of keys with something cheap (a proxy), keep the promising ones, skip the rest. The question is how you pick, and the paper's Figure 1 shows why the two standard answers both misbehave — on two different attention-logit distributions, one peaked and one nearly flat.
Read the sparsity numbers across the rows: - **Top-k** keeps a fixed fraction, so it reports **70.3% on both rows**. It cannot tell a peaked distribution from a flat one. On the peaked row it is leaving free sparsity on the table; on the flat row it is throwing away blocks that mattered. - **Top-p** keeps blocks until their cumulative proxy mass hits a target, which is adaptive but wildly so: **96.88%** sparsity on the peaked row and **21.9%** on the flat one. Budgets swing by a factor of four between distributions, which is miserable for a kernel that wants predictable work per tile. - **Sol-Attn** thresholds at **μ + βσ** — the mean of the block scores plus β standard deviations. It adapts (75.0% vs 67.2%) but stays in a controlled band. That statistical threshold is the whole trick, and its virtue is *computability*. A mean and a variance can be maintained as a running summary while you stream blocks. A top-k ranking cannot: you have to see every score, materialize them, and sort. Which brings us to where the decision gets made. ## Folding the decision into online softmax Flash-attention-style kernels already stream. They walk key/value tiles, keep a running maximum and a running sum, and rescale as they go — that is what makes softmax computable without ever holding the full attention matrix. Conventional sparse attention bolts a *separate* pass in front of this: score all blocks, build a proxy map in memory, rank it, then run the sparse kernel on the survivors. Sol-Attn puts the decision inside the loop that was already running.
The outer loop computes proxy scores from mean-pooled keys and compares them against that tile's threshold, producing a `1/1/0`-style mask. The inner loop then runs exact attention on the surviving tiles and skips the rest. Because the threshold is a statistic rather than a rank, no proxy map is ever materialized — the budget comes out dynamic *and* controllable, which is the combination neither top-k nor top-p manages. ## Not dropping, approximating The second idea is smaller and does more work than it looks. Standard block-sparse attention treats an unselected block as if it contributed nothing. Under aggressive sparsity that assumption is exactly where the quality goes. Sol-Attn has already computed a proxy score for every block, including the losers, since that is how it decided. So instead of discarding them it **reuses those scores to approximate the skipped blocks' contribution** — a correction term that costs nothing extra, because the information was a by-product of routing. Routing, sparse computation and approximation correction all happen in a single online-softmax pass. That is why the accuracy curve degrades gracefully rather than falling off a cliff: the tail is attenuated, not deleted. ## What it actually buys The paper reports **2.1× end-to-end for video generation** and **2.3× for video editing**. The more useful chart is the cumulative breakdown, because it shows what is attributable to what:
Those headline multiples are **cumulative**, so it is worth doing the subtraction. On HunyuanVideo, Sol-Attn takes 328.4 s down to 170.6 s — a **1.92× marginal** gain on top of the other two techniques. On Wan2.1-14B it takes 217.6 s to 161.8 s, a **1.34× marginal** gain. Real, and the largest single contributor in the Hunyuan case, but not 5.08×. Anyone quoting the total as an attention result is quoting three techniques. ## The engine around it Sol-Attn does not ship alone. It is one of five composable techniques in the [`sol-engine` branch](https://github.com/NVlabs/Sana/tree/sol-engine) of NVlabs/Sana (Apache-2.0), described as "an efficiency-oriented inference codebase for high-resolution video diffusion, built on SGLang's `multimodal_gen` runtime". The five: **caching** (reuse or skip denoising-step outputs, TeaCache/EasyCache-style), **quantization** (TransformerEngine NVFP4 4-bit, applied step-selectively), **kernel fusion** (memory-bound DiT ops — norm, activation, precision conversion), **sparse attention** (Sol-Attn), and **token pruning** (dropping low-salience video tokens during refinement steps). Reported end-to-end speedups, all on GB200 with warmup excluded: | Model | Speedup | |---|---| | Wan2.2 TI2V-5B | ~2.89× | | SANA-Video (2B) | ~2.77× | | LingBot-Video (30B) | ~2.60× | | LTX-2.3 (22B) | ~2.38× | | Cosmos3-Super (64B) | ~2.27× | | Wan2.2-A14B (14B MoE) | ~2.17× | The consistency across 2B to 64B, dense and MoE, is the interesting part — these are mostly memory-movement and redundancy wins, so they do not evaporate as models grow. There is also an **agent-native workflow**: the repo is set up so a coding agent (Codex or Claude Code) does environment setup, weight fetching and inference, troubleshooting as it goes. Worth noting on a site whose own content is written this way — treating "an agent will be the one running this" as a first-class install path is still rare. ## Where this sits The same lab has been attacking this cost from the other end. [SANA-Video 2.0](/articles/sana-video2) makes attention cheap *architecturally* — linear attention for three of every four layers, a deep-compression VAE to shrink the token count before attention ever runs. That requires training the model that way. Sol-Attn is the training-free counterpart: take a model somebody already trained and skip work at inference. SANA-Video is literally the second row of the engine's own benchmark table, so both ends compose. Against the site's other sparse-attention coverage — [MiniMax's approach](/articles/minimax-sparse-attention) — the contrast is that most sparse attention is *trained*, with the model learning to live within a sparsity pattern. Sol-Attn assumes no cooperation from the model at all. And where [MrFlow](/articles/mrflow-diffusion-acceleration) attacks diffusion cost along the *step* axis, Sol-Attn attacks the per-step cost; the engine's caching module is doing the step-axis job alongside it. **Caveats.** (1) Every number here is **self-reported**, on a paper posted 2026-07-27 with no third-party replication. (2) The engine's speedups are **GB200, warmup-excluded** — best-case hardware, and warmup is a real cost you pay once. (3) The headline 5.08× is **cumulative across three techniques**; Sol-Attn's marginal contribution is 1.92× and 1.34× on the two models shown. (4) The paper claims quality is preserved but I have not seen an independent quality evaluation, and "preserved" is doing real work in a domain where the failure mode is subtle temporal artifacts rather than a metric drop. (5) `sol-engine` is a **branch**, not a release. ## The take The mechanism is the part worth keeping. Routing decisions in sparse attention are usually treated as a preprocessing step — score, rank, select, then compute. Sol-Attn's claim is that the ranking was never necessary: a statistical threshold gets you an adaptive budget from a running summary, which means the decision can live inside the streaming loop the kernel already runs, which means the proxy map never has to exist. And once you are computing proxy scores anyway, throwing them away for the skipped blocks is wasteful — reusing them as an approximation is close to free. Both ideas come from asking where the information already is, rather than adding machinery. That tends to be the sign of a good systems result. --- *Sources: [Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification](https://arxiv.org/abs/2607.24027) (Haopeng Li, Yitong Li, Junsong Chen, Tian Ye, Haozhe Liu, Jincheng Yu, Duomin Wang, Ruihua Zhang, Zeke Xie, Enze Xie, Song Han; arXiv 2607.24027, 2026-07-27) for the method and Figures 1–3, and the [`sol-engine` branch](https://github.com/NVlabs/Sana/tree/sol-engine) of NVlabs/Sana for the engine, the five techniques and the per-model speedup table. All figures are the paper's own, served locally. Marginal-speedup arithmetic is mine, derived from the latencies in Figure 3. The interactives are mine and illustrate the mechanism; they are not measurements.* --- # ADR: the agent that watches your agents > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/uber-adr > date: 2026-08-03 > tags: security, agents, mcp, llm, explainer The first thing to get out of the way: [Uber's ADR](https://github.com/uber/ADR) has nothing to do with Architecture Decision Records. **ADR = Agentic Detection and Response** — an enterprise security framework for the AI coding agents your engineers already run. The tagline says it plainly: "ADR secures enterprise AI agents through observability, security benchmarking, and threat detection." It's infra, not a model — though it ships an LLM-based detector and a red-team benchmark, and it's been running inside Uber for **over ten months**, watching **7,200+ unique hosts** and **10,000+ agent sessions a day**. The reason it exists is a gap that's easy to miss if you haven't operated one of these agents at scale: your existing security tooling watches file writes and process spawns. It has no idea *why* a file got written. An AI coding agent with shell access, file access, and a dozen MCP servers plugged in is a new kind of actor on your endpoints, and the thing that makes its actions dangerous or benign — the reasoning that led to the tool call — is exactly the part a traditional EDR agent can't see. ## The threat model Three problems, named directly in [the paper](https://arxiv.org/abs/2605.17380) (Chenning Li, Pan Hu, Justin Xu, et al., accepted MLSys 2026 Industry Track): 1. **Limited observability** — "existing Endpoint Detection and Response (EDR) tools see file writes but not the agent reasoning, prompts, or causal chains linking intent to execution." 2. **Insufficient robustness** — static, rule-based defenses don't generalize across attack techniques. 3. **High detection cost** — running an LLM as a judge on every one of 10,000+ daily sessions is expensive if you do it naively. The concrete shape of the threat is **indirect prompt injection through a tool a developer already trusts.** Uber emulated the industry's Agent Flayer incident (Zenity Labs, 2025) to validate ADR against it, and the attack graph is worth internalizing because it's the whole genre in five steps: a malicious email creates a Jira ticket with hidden instructions; a developer's Cursor session, connected to Jira through an MCP server, fetches that ticket as ordinary context; the hidden instructions coerce the agent into reading a local secrets file; and the agent — still just doing what its context told it — makes an outbound HTTP call that exfiltrates the credentials. Nobody typed a malicious command. The attacker never touched the developer's machine. The whole thing rides in on a support ticket.
Uber also ran an internal capture-the-flag exercise that names the second broad category: a custom shell tool integrated into the assistant, then a deceptive prompt instructing the agent to run `curl | python3` "under the guise of sandbox testing." Same underlying issue — the agent's own tool surface, or a malicious MCP server sitting on it, is the attack surface. ADR-Bench, the benchmark this whole thing gets scored against, formalizes that: it's built on **133 MCP servers** (78 benign, 25 deliberately backdoored, the rest environment-emulation), and one of the backdoored ones — a credential-harvesting tool called `location_harvester` — is registered under a spoofed identity, `weather_forecast_service`, specifically to test whether a detector falls for supply-chain impersonation. That's the taxonomy: prompt injection through trusted tool output, malicious or impersonating MCP servers, and an agent that ends up exfiltrating something it was never asked to touch. ## Why you can't just run an LLM judge on everything The naive fix — point a capable model at every agent session and ask "is this suspicious?" — doesn't survive contact with the volume. At 10,000+ sessions a day, an LLM call per session that does real reasoning (reads code, checks policy, thinks about intent) is expensive and slow enough that latency alone rules it out for most of the traffic, almost all of which is a developer editing files and running tests. ADR's answer is a **two-tier pipeline**, and the design is unapologetically borrowed from how a human security team already works: a triage analyst does a fast first pass on everything, and only escalates the ambiguous cases to someone senior enough to actually investigate.
**Tier 1** is a cheap, single-call triage model (`gpt-4o` by default) reading the full session transcript, tuned for high recall with an explicit escalate-on-ambiguity bias — anything that isn't clearly benign defaults to suspicious rather than being waved through. If triage says benign, the pipeline returns immediately. That early return, not a second model being smarter, is the entire cost-control story. **Tier 2** only runs on the sessions Tier 1 flagged, and it is not an API call. Reading the released code (`Detection/guardrail/adr_agent/adr_baseline.py`) turns up something more interesting than a bigger judge model: it shells out to the **Claude Code CLI as a subprocess** — literally `subprocess.run(["claude", "--model", "claude-sonnet-4-6", "--mcp-config", ".mcp.json", ...])` — and gives that CLI session live MCP tool access to three purpose-built servers: one that fetches a flagged tool's actual source code, one that looks up known techniques for the suspected tactic, and one that checks policy violations. The reasoning agent doesn't just re-read the transcript harder. It *investigates* — it can go pull the suspicious MCP tool's own source before it commits to a verdict, the same way a human analyst would open the code instead of trusting the tool's name. Output is forced JSON (`is_threat`, `confidence`, `explanation`), with a documented retry path for the case where Claude itself flags the detector's own prompt as a possible injection attempt. Step through both paths on the two real scenarios above — an ordinary edit-test-commit session, and the Agent Flayer emulation — and watch what each tier actually sees, calls, and costs: The numbers back the design up directly. In Uber's own ablation, the triage layer alone resolves **40.7% of tasks** at **$0.017/task, 2.3s** latency; the escalated path costs **$0.029/task, 29.7s** — about 13x more expensive per task, which is the honest price of running an LLM twice, with tool calls, inside a detection loop. Removing the triage layer entirely and sending everything to the reasoning agent *does* raise recall (0.667 → 0.805) and F1 (0.800 → 0.868), but it also introduces false positives (precision drops to 0.943) and costs 29% more — a clean demonstration that the two-tier split isn't just an engineering shortcut, it's load-bearing for precision too. And in a second ablation that validates the "fetch the tool's own source" mechanism specifically: pull the source-code MCP server away from the reasoning agent and F1 falls from 0.800 to 0.676 — the single largest drop of any component removed, bigger than losing threat-intel lookup (0.800 → 0.716) or policy verification (0.800 → 0.806, barely a dent). The part of the mechanism that sounds like the interesting flex — an agent reading a tool's own code before judging it — is also, empirically, the part doing the most work. ## The benchmark, and where the baselines fall apart **ADR-Bench** is 302 tasks (42 malicious, 260 benign) derived from real enterprise telemetry, spanning 17 named attack techniques across 5 tactics and those 133 MCP servers. The class imbalance — 13.9% attack rate — is the point: an academic benchmark that's roughly half attacks doesn't tell you how a detector behaves against the traffic mix a real enterprise actually has, which is overwhelmingly benign. ADR is also evaluated on **AgentDojo**, a public, mostly-balanced prompt-injection benchmark (93 tasks), giving two very different stress tests of the same detector. The pattern that falls out is the honest headline: on AgentDojo, the baselines (ALRPHFS, GuardAgent, LlamaFirewall) look reasonable — ALRPHFS even gets within striking distance of ADR's precision. Move to ADR-Bench's enterprise-realistic imbalance and every baseline collapses to 30–40 false positives out of 260 benign tasks. That's not a rounding difference — at real traffic volume, that false-positive rate is what "unsuitable for production" cashes out to: a human review queue that drowns. ADR holds **1.000 precision, 0.667 recall, 0.800 F1, zero false positives** on ADR-Bench, and **0.927 precision, 1.000 recall, 0.962 F1** on AgentDojo — complementary strengths the paper is upfront about: precision-first where class imbalance is severe, recall-first where missing a prompt-injection attack is the costlier failure. ## What happened in production The credential-exposure numbers are where this stops being a benchmark story. Deployed across Uber's MacBook fleet (Intel and ARM), ADR has "sustained reliable detection... uncovering hundreds of credential exposures across 26 categories," which fed a shift-left prevention layer — a regex-and-entropy pre-prompt hook running in Cursor and Claude Code — that hit **97.2% precision**: 206 true positives correctly caught, only 6 false positives, across 212 unique credentials seen in hundreds of thousands of sessions. Worth noting in passing: that prevention layer is explicitly *not* LLM-based. "Simple non-LLM checks work well for known attacks with static patterns... but fail for attacks requiring reasoning about tool semantics, causal context, and enterprise policies" — which is exactly the job Tier 2 exists to do instead. One production detail is genuinely candid: of the sessions ADR's Tier 1/Tier 2 routed to a human review queue, only 34% were confirmed malicious true positives. Another 17% were true positives that turned out non-malicious — internal red-team and pentest activity that looked identical to an attack until a human checked. The paper doesn't paper over this: it explicitly flags the need to model *user personas* and intent, because a security team doing its job on purpose looks, from telemetry alone, a lot like an attacker. ## What's honest and what's missing A few things worth stating plainly before treating any of this as settled: - **Every production number is self-reported.** The 7,200 hosts, 10,000+ sessions/day, and 97.2% precision figures come from Uber's own paper — there's no third-party audit. A tweet from an Uber engineer separately claims 50,000+ sessions/day, a larger and more recent figure than the paper's own — treat that as informal color, not a citable number. - **The baseline comparisons in Table 2 aren't independently re-runnable.** ALRPHFS and GuardAgent's code was stripped from this repo for licensing reasons; the paper's own comparison numbers are reproduced as-is, documented candidly in `docs/BASELINE_REPLICATION.md`, but you can't regenerate them yourself. - **The open-source release is the detection half of a four-part system.** Per the architecture, ADR Explorer (pre-deployment red-teaming) and ADR Prevention (blocking unsafe actions in real time) are both explicitly excluded — "not included in the current open-source release. Stay tuned." What's shipped is the Sensor, the benchmark, and the dual-agent detector baseline — real, but not the complete production stack Uber runs internally. - **The reasoning tier is unusually coupled to one vendor's CLI.** Shelling out to `claude --dangerously-skip-permissions` as a subprocess is a legitimate way to get a tool-using agent for free, and it's the most narratively interesting part of the design — but it's also a portability limitation worth naming as exactly that. - **An LLM in the detection path is not free**, and ADR doesn't pretend otherwise: $0.024/task blended average cost and 18.5s average latency on ADR-Bench, with the escalated path alone running $0.029 and nearly 30 seconds. That's the real, ongoing bill for the precision this design buys. - **ADR-Bench is Uber's own benchmark**, built from Uber's own telemetry and scored by Uber. It's a genuinely useful stress test — the class-imbalance framing is a real and underused idea — but it isn't a neutral third party's yardstick, and the paper is explicit that the benchmark's attack rate doesn't mirror real production incidence. None of that erases the result: a two-tier detector that costs cents and seconds on the common case, escalates to an agent that can go read a suspicious tool's own source before it decides, and has been running against real attacks in production for the better part of a year. ## The take ADR is worth reading past its confusing name because it's a concrete answer to a question [Lilian Weng's harness framing](/articles/agent-harness) leaves open: if the harness — the loop, the tools, the context policy — is where an agent's real capability lives, then the harness is also exactly where you'd go looking for an attack, and exactly where a defense has to watch. ADR watches the causal chain a traditional EDR tool can't see, and its own reasoning tier is built the same way the agents it's watching are: a model with tool access, investigating rather than just classifying. The cost/precision tradeoff it makes explicit — cheap triage on the routine 40.7%, an expensive investigating agent on what's left — is the same lesson [Antares](/articles/antares) makes from the opposite direction, with a much smaller model doing a narrower security job: the interesting engineering in agentic security right now is less about a bigger judge model and more about routing the right amount of reasoning at the right moment. --- *Sources: [uber/ADR](https://github.com/uber/ADR) (Apache 2.0; the vendored AgentDojo benchmark under `Detection/benchmark/agentdojo/` is MIT); the paper, ["ADR: An Agentic Detection System for Enterprise Agentic AI Security"](https://arxiv.org/abs/2605.17380) (Chenning Li, Pan Hu, Justin Xu, Baris Ozbas, Olivia Liu, Caroline Van, Manxue Li, Wei Zhou, Mohammad Alizadeh, Pengyu Zhang, KK Sriramadhesikan, Ming Zhang; MLSys 2026 Industry Track). Figures reproduced from the paper for commentary, cropped from the arXiv PDF committed at `docs/adr-paper.pdf` in the repo and flattened onto white. The interactive pipeline trace and benchmark scatter are my own, built from the paper's own reported numbers — not measured traces.* --- # Macaron-V1: four 1B adapters on a frozen 744B base > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/macaron-v1 > date: 2026-07-27 > tags: llm, lora, agents, open-weights, generative-ui, explainer [Macaron-V1](https://huggingface.co/collections/mindlab-research/macaron-v1) is Mind Lab Research's agent model family, released 2026-07-21 under MIT. The flagship, **Macaron-V1-Venti**, is described as a 748B model. That number needs an asterisk immediately: **744B of it is a frozen GLM-5.2**, and Mind Lab's contribution is **four 1B LoRA adapters** — about half a percent of the artifact. That is the whole idea, and it is a genuinely different bet from how most post-training is done. Rather than fine-tuning one monolithic model, take a strong open base, freeze it, and attach a small number of tiny specialists. They call it **Mixture of LoRA (MoL)**. ## Mixture of LoRA Four adapters, each 1B: `l0` Chat, `l1` Agent, `l2` Coding, `l3` GenUI. The routing detail is the neat part — **`l0` is both the conversational backbone and the router**. It sees each new user request and dispatches it to whichever specialist fits. Note what routes and when. In a Mixture-of-Experts model, a router fires on *every token* at *every MoE layer*, and the experts were built during pretraining — they are inseparable from the model. MoL routes **once per request**, at the adapter level, and the thing being routed between is four swappable files sitting on a base someone else trained. Ongoing reasoning and tool interaction stay inside the selected LoRA for the duration; when a specialist finishes, its work is passed to the next as a concise summary rather than shared state. The practical consequences are real. Specialization costs 1B parameters instead of a full fine-tune. Adapters can be swapped, added, or updated independently. And when GLM-5.2 improves, you re-fit adapters rather than retrain a 744B model. The cost is equally real: you inherit the base's ceiling, its licence obligations, and its failure modes, and a request-level router cannot change its mind halfway through a turn the way per-token routing implicitly can. **Update, 2026-08-03: the Tall size conflict looks like two different counts, not a contradiction.** Mind Lab's blog calls Macaron-V1-Tall **50B**; the model card, the Hugging Face listing and Novita all say **36B**. The Hugging Face API settles half of it — `safetensors.total` for `Macaron-V1-Tall` is exactly **35,951,822,704 (35.95B)**, which is the 36B figure and is the *base checkpoint on its own*. The blog's 50B is its own decomposition of base plus adapters: 35B + 4 × 3.7B ≈ 49.8B. So the two numbers are measuring different things, and Novita hedges the gap as "10~50B". What I could not verify is the per-adapter figure. The published 35.95B total does not appear to include the adapter weights, so I cannot confirm 3.7B each from the metadata — that number is Mind Lab's, not something I measured. It is worth flagging because it implies Tall's adapters are roughly **3.7× the size of Venti's 1B ones** on a base twenty times smaller: on Venti the specialization is about half a percent of the artifact, on Tall closer to a tenth. If that holds, "Mixture of LoRA" means something quite different at the two scales, but the evidence for it is currently a single line in a blog post. ## What the adapters actually buy Most of the release table compares Macaron to Claude, GPT and Gemini. That is the least informative comparison available, because it confounds the adapters with GLM-5.2's own strength. The controlled experiment is sitting right there in the same table: **each variant against the frozen base it was built on**. That isolates the only thing Mind Lab actually changed. The pattern is consistent and modest: roughly **+3 to +6 points** across chat, agent and coding work. Two rows are inside noise — `T3-Bench` at +0.2 and SWE Atlas QnA at +0.6. And then UI4ABench jumps **+20.7** on Venti and **+25.4** on Tall. That outlier is the honest crux of the release. It is simultaneously the strongest evidence that a 1B adapter can teach a frozen base a genuinely new skill, *and* the result most exposed to selection effects — UI4ABench is Mind Lab's own benchmark, measuring generative UI, which is exactly the capability they built a dedicated adapter for. Both things are true at once. ## Read the evaluation table twice The published benchmark figure includes its own methodology notes, and they change how several rows should be read.
Three things stand out: - **The judges are other models, and one of them is the base.** ChatBench is scored by "a privately deployed GLM-5.2 judge" — and Venti *is* GLM-5.2 plus adapters. A model's own base evaluating its output is a conflict worth naming. Elsewhere the judge is a competitor: Claude Opus 4.6 on LivingBench, Claude Haiku 4.5 on PinchBench, GPT-5.4 on ClawGym, GLM-5.1 on VitaBench, Gemini 3.5 Flash scoring UI4ABench rubrics. - **Several rows are best-of-N, not single-shot.** PinchBench reports "the best observed score". DeepSWE allows "up to three attempts, and report the best one". SWE Atlas QnA is pass@3. SWE-Bench Verified permits up to three retries on evaluation errors and reports the best successful attempt. Those are legitimate protocols, but they are not comparable to a single-trial number from another lab's report — and some competitor cells are marked as taken from leaderboards or the models' own reports. - **Macaron does not lead everywhere.** Claude Opus 4.8 wins SWE Verified (88.6 vs 85.6) and SWE Atlas QnA (57.3 vs 49.5). GPT-5.5 wins ClawGym (82.5 vs 77.7) and DeepSWE (70.0 vs 58.4). Qwen 3.7 Max wins VitaBench (61.2 vs 60.0) and Gemini 3.1 Pro wins VitaBench2 (50.2 vs 46.0). Mind Lab says as much in its own post: "Coding is where we currently sit close to, rather than ahead of, the frontier." The rows where Gemini, Qwen and Minimax collapse to 10.0–22.6 on DeepSWE and SWE Atlas QnA are almost certainly harness incompatibility rather than capability — those evaluations run through Claude Code as the agent harness, which is not neutral ground for every model. ## The infrastructure claims Three systems are named, none with a technical report behind them yet: - **MinT** — the post-training platform, claimed to support models up to a trillion parameters via adapter-only handoffs and a "million-scale adapter catalog". Adapter-only handoff is the load-bearing idea: if specialization is always a small file, you never move a 744B checkpoint between training stages. - **MindForge** — an agentic RL framework built around discovery, expansion and update cycles against production harnesses. - **LongStraw** — million-token RL, which works by evaluating a shared prompt once into a reusable resident state and then replaying only the response branches. For agentic RL where many rollouts share a long prefix, that is the obvious win, and it rhymes with the external KV-cache pooling in [Kimi K3](/articles/kimi-k3)'s RL infrastructure. Macaron-V1 also ships a serving story: a [Mixture-of-LoRA harness](https://github.com/MindLab-Research/Mixture-of-LoRA-Harness) that keeps an OpenAI-compatible endpoint while adding the L0 router and same-request switching into the selected specialist, plus [Macaron Artifacts](https://github.com/MindLab-Research/macaron-artifacts), a local WebUI and plugin that runs inside Claude Code, Codex or Kimi Code. ## The take The interesting claim in Macaron-V1 is architectural, not competitive: that request-level routing across a few 1B adapters on a frozen base is enough to build a specialized agent model, and that you can therefore treat a frontier open-weight model as infrastructure rather than as something to fork. The base-versus-tuned comparison supports a weaker version of that claim than the headline table does — a few points nearly everywhere, and one large gain on the capability they purpose-built an adapter for. What would settle it is the technical report, which the model card lists as "coming soon", along with the full benchmark methodology. Until then this is a self-reported release with no third-party replication, several best-of-N protocols, and its own base model sitting on the judging panel. The idea is worth watching. The numbers are worth waiting on. --- *Sources: the [Macaron-V1 collection](https://huggingface.co/collections/mindlab-research/macaron-v1) and the [Macaron-V1-Venti](https://huggingface.co/mindlab-research/Macaron-V1-Venti) model card (architecture, adapter roles, parameter counts, benchmark table), and Mind Lab's [Introducing Macaron-V1](https://macaron.im/mindlab/research/introducing-macaron-v1) post (MinT, MindForge, LongStraw, variant sizes). All benchmark numbers are Mind Lab's own, with per-benchmark judge models and retry policies as annotated in their published table; no technical report has been released. The interactives are mine.* --- # Neutrino-1: quantization is a training decision, not a deployment one > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/neutrino-1 > date: 2026-07-27 > tags: llm, quantization, ternary, bitnet, inference, efficiency Fermion Research shipped three models on July 27, 2026: **Neutrino-1 8B**, **Neutrino-1 0.6B**, and a 0.6B-Chat variant, all on Hugging Face under Apache 2.0, all built on a ternary weight format — every projection weight is one of exactly three states, minus, zero, or plus. That part of the pitch is not new; I wrote about [Ternary15M](/articles/ternary15m) doing the same thing at 15M parameters a few days ago. What Neutrino-1 adds is scale (about 500x more parameters) and, more importantly, a controlled comparison that the smaller model never ran: what happens if you take the *same* ternary format and round a model into it after training, instead of training inside it from the start. The answer is not "a few points worse." It's a cliff. ## The cliff Two 8B-class checkpoints, each about 3 GB, rounded into a ternary format after full-precision training, score **24.2** and **24.7** on 5-shot MMLU. Chance, on a four-choice test, is **25.0**. Three models whose training ran *inside* the ternary constraint from the start — at the same roughly 2-3 GB artifact size — score 47.24, 65.75, and **72.1** (Neutrino-1 8B itself). Fermion's own line on this, and it's a good one: "rounding after training lands on the chance line; every model trained inside its own format clears it by twenty-two points or more." The mechanism is worth sitting with, because it explains why this isn't a smooth tradeoff curve. Round a trained weight to the nearest of three values and the individual errors don't cancel — Fermion describes them as uncorrelated, so they "accumulate along the row as a random walk instead of cancelling. The feature that leaves the layer is not a noisier version of the right answer; it is a different number, and 36 layers compound the difference." Their framing: "the failure is not noise. It is amnesia, and its signature is the cliff: scores do not degrade toward chance, they arrive there." Train inside the constraint instead, and the optimizer never learns a solution the format can't store in the first place — there's no gap between the model that was trained and the model that ships. This is the flip side of what Ternary15M already showed at tiny scale: quantization-aware training with a straight-through estimator gets hard-ternary inference to within +0.01 nats of the latent model, because the network never experienced anything else. Neutrino-1 is the same bet, replayed at 8B, with the missing control group finally run: skip the QAT and round instead, and the model doesn't degrade gracefully — it falls through the floor. ## What's actually ternary Not everything. Of the 8B's 8.19 billion parameters, 6.95 billion (all 252 projection matrices — seven per layer, attention and feed-forward, across 36 decoder layers) are ternary. The token embedding table and the output head stay int8; the RMSNorm gains stay fp32. The reasoning Fermion gives is about where rounding error can hide: "a linear layer's output feature sums hundreds of three-state contributions, so individual state errors cancel inside the sum; that summation is what makes the format survivable." An embedding lookup returns one row verbatim — there's no sum to average the error away — and the output head "decides tokens by small logit margins," where a rounding error can flip the argmax. Both stay out of the ternary lane for the same reason: no averaging effect to hide behind. The scale mechanism is also more granular than Ternary15M's. Ternary15M ternarizes each output channel around one number: `absmean(W)`, the mean absolute weight for that whole row. Neutrino-1 groups weights into fixed-size blocks along the input dimension and gives *each block* its own higher-precision scale — Fermion's phrase is "state times scale." Smaller blocks track the underlying weights more closely at the cost of more stored scales; one scale for an entire row (Ternary15M's approach) is the coarsest, cheapest case. The toy below runs the actual arithmetic — round `clamp(w / scale, -1, 1)` per weight, scale is each block's mean absolute value — so you can see the tradeoff move: ## Where the bytes go
This is the honest caveat behind "4.2 times fewer bytes than fp16 at bf16 on the same memory system," and it's worth being precise about it: that ratio is a **whole-artifact** number, not the per-weight ternary compression ratio. Log2 of 3 states is about 1.58 bits, which against a 16-bit float is closer to a 10x reduction — but the embedding and output tables (int8, one byte per value, no ternary discount) and the norm gains (fp32) drag the average down. At 8B, those non-ternary tensors are only 15.2% of the parameters but 32.8% of the bytes. At 0.6B the effect is worse: the same un-ternarized vocabulary is 47.5% of the file, because a smaller transformer has fewer projection weights to amortize a fixed-size vocabulary table against. It's the identical finding Ternary15M made at 15M parameters, where the FP32 embedding table was over 60% of that model's total parameters and dominated its 43 MB footprint — the direction of the effect is the same at both ends of a 500x scale range: the *smaller* the model, the more its un-quantized vocabulary — not its ternary matmuls — decides the file size. On disk, Neutrino-1 8B is 3.88 GB; it downloads at 2.56 GB because the ternary lane compresses further in transit (Fermion reports 0.516-0.569 of raw bytes, layer-dependent). Neutrino-1 0.6B downloads at 328 MB. ## Sparsity is learned, not imposed Across the 8B's 6.95 billion ternary weights, the split is **62.63% zero, 18.68% plus, 18.69% minus** — remarkably close to balanced between the two nonzero states, and remarkably far from an even three-way split. Fermion's framing is the one worth keeping: "most of the mass on zero: the format sets how much of each tensor falls silent, and the learned weights decide which connections go." A float layer can only make a connection small; a ternary layer, trained natively, can delete it outright and the training decides which ones.
That spike is the interesting part, because nothing about the format explains it — the format sets *how much* falls silent on average, not *where* it clusters by depth. The four attention projections sit in a tight 61.84-63.51% band at every one of the 36 layers, almost boring in its consistency. The feed-forward `down` and `gate` projections are the exception: `down` reaches 72.47% zero at layer 3, `gate` reaches 70.48% at layer 4 — roughly ten points denser than the rest of the network — and both settle back to the ~62% baseline by layer 5. The single densest tensor in the whole model is the layer 1 `down` projection at 60.59% (its local minimum, immediately before the spike). Scrub through the real per-layer numbers below: The state statistics are also stable across scale in a way that argues they're a property of the format and the training recipe, not of size: at 0.6B, fourteen times fewer parameters, the split is 62.26% zero / 18.87% plus / 18.86% minus — within half a point of the 8B on every axis. ## How much of Qwen3-8B does it keep Neutrino-1 8B is measured, on Fermion's own harnesses, against **Qwen3-8B at bf16** — described in the post as "the full-precision base it was built from," which is itself worth flagging: this isn't an independently trained architecture being compared to an unrelated baseline, it's a model built from Qwen3-8B's own weights and then retrained natively in ternary. At 4.2x fewer bytes, Neutrino-1 8B holds: Report the weakest number, not the flattering one: tool calling retention is 79%, and it's worse than that headline suggests once you look at the breakdown by category (BFCL v3, macro-averaged to 68.9 overall). Held-out, textbook function signatures score well — 82.3% simple, 83.5% multiple — but signatures drawn from real-world APIs in the wild score much lower: 61.6% live-simple, 52.0% live-multiple. Non-Python languages are worse still: 54.0% JavaScript, 43.0% Java. "Tool calling: 79%" is an average that buries a 40-point spread between the easy and hard slices of that same axis. The MMLU headline (72.1) and the "96% general knowledge" retention figure are **not directly comparable** in Fermion's own post. The retention percentages are computed against Qwen3-8B's score on an unnamed "general knowledge" suite — Fermion never states Qwen3-8B's own 5-shot MMLU number anywhere in the piece, and never confirms that "general knowledge" and "MMLU" are the same benchmark. The MMLU chart above only compares Neutrino-1 8B against *other* ternary and rounded models at similar artifact sizes, not against its own full-precision progenitor. So while the rounding-vs-native-training gap (24.2 vs 72.1) is well anchored, the honest answer to "what's the gap between 72.1 and Qwen3-8B's own MMLU" is: **the source doesn't say, and you can't back it out from what's published.** Every number in this section is self-reported by Fermion, on their own harness, with no third-party replication. For what it's worth, the one place Neutrino-1 8B is reported to exceed the reference is answer-format discipline — Fermion's explanation is that discipline is a trained *behavior*, not a bulk statistical property of the weights the way knowledge is, so the format doesn't cap it the way it caps knowledge retention. At the small end, Neutrino-1 0.6B is compared directly to Qwen3-0.6B on ARC-easy: 53.45 vs 60.82, 87.9% retention, at one-eighth the precision and a 238 MB vs 1.50 GB download. ## Serving it: the Neutrino Engine The inference side ships as its own artifact — a `pip install fermion` package, a CUDA-enabled `llama.cpp` fork, and an MLX pack — with one stated design constraint: output has to be **token-identical** to a full-precision reference on every backend. Fermion gates every release on that: the speculative-decoding path (Neutrino-1 0.6B drafting for the 8B) was checked token-by-token against the undrafted path across 27,648 consecutive tokens before any drafted throughput number was published, and they report zero divergences. Measured numbers: **33.7 tokens/second** on a MacBook M5 (GPU path; 24.9 tok/s CPU-only), **30.7 tokens/second** on an NVIDIA L4 at 4k context inside 4.68 GiB of VRAM (fits an 8 GB card), and **396 tokens/second** undrafted on an H100 80GB — rising to **763 tokens/second** with the 0.6B draft model, gated as above. Draft acceptance is prompt-dependent: near 100% on counting/enumeration, 96.5% on factual recall, roughly 80% on prose, and roughly 50% on code — code is where the smaller model diverges from the 8B's choices most often, so the speedup shrinks accordingly. The core argument for why a smaller artifact is faster at batch size 1 is straightforward memory-bandwidth accounting: single-stream decode reads every weight once per token, so bytes-per-token divided by memory bandwidth sets a hard floor on latency that no kernel can negotiate around. Neutrino-1 8B's 3.88 GB artifact against roughly 16 GB for the same weights at fp16 puts that floor about four times lower before a single kernel runs. That fp16 comparison is a **size** argument (3.88 GB vs ~16 GB), not a measured one. Every throughput number Fermion publishes for the Neutrino Engine is ternary-format-versus-ternary-format — against a "reference stack" running the same public 27B ternary model (105.15 vs 97.80 tok/s), against `bitnet.cpp` running BitNet b1.58-2B (102.4 vs 89.0 tok/s on an M5), against a lookup-table CPU kernel (T-MAC, 2.12x), against an int4 GEMV kernel (+12-13%). I could not find a measured fp16 Qwen3-8B throughput number on the same M5, L4, or H100 hardware anywhere in the post. The tokens/second figures are real and gated for correctness, but the *speedup over full precision* claim is anchored to an artifact-size ratio, not to a same-hardware fp16 benchmark run. One more honest number: KV cache growth is indifferent to weight format and scales with context regardless — 0.60 GB at 4,096 tokens, up to 6.04 GB at the model's full 40,960-token window. Past roughly 26,000 tokens the cache alone outweighs the 3.88 GB model artifact, so the memory story stops being about weights and starts being about context length. ## What you can actually download All three models are live on [huggingface.co/FermionResearch](https://huggingface.co/FermionResearch) as of today, Apache 2.0, no waitlist: | Model | Size | Role | |---|---|---| | Neutrino-1 8B | 2.56 GB download / 3.88 GB on disk | The frontier model | | Neutrino-1 0.6B | 328 MB download | Draft model for speculative decoding, and usable standalone | | Neutrino-1 0.6B-Chat | — | Conversational small model | They're new enough that download counts were in the single digits at the time I fetched the org page — this is a same-day release, not an established artifact with a track record. ## The take Neutrino-1 is Ternary15M's bet — train inside the constraint instead of rounding into it, keep a shared scale next to the signs, let signed accumulation replace multiplies — replayed at roughly 500x the parameters, with a block-wise scale instead of one absmean per channel, a real inference engine with a correctness gate, and agentic/tool-use evals that a 15M TinyStories model has no business running. The MMLU cliff (24.2-24.7 at chance versus 72.1 trained in format) is the cleanest piece of evidence I've seen that quantization format is something you commit to before training starts, not a knob you turn afterward — and it's consistent with, not contradicted by, Ternary15M's own finding that training-aware ternary costs almost nothing when the network never knows another way to compute. None of this is independently verified. Every number in this piece — the MMLU scores, the retention percentages, the sparsity statistics, the tokens-per-second figures — is self-reported by Fermion Research on their own harnesses, comparing their model to their own full-precision progenitor. The MMLU headline and the "96% general knowledge" figure use two differently-named metrics that are never reconciled in the source. The speed claims are anchored to an artifact-size ratio, not a measured full-precision baseline on the same silicon. And the model being celebrated for "holding" Qwen3-8B's knowledge was built starting from Qwen3-8B's own weights, not trained from scratch as an independent check on the method. The cliff is real and the mechanism is coherent; the specific numbers around it deserve the same scrutiny you'd give any single-lab benchmark table until someone else reproduces them. --- *Sources: Fermion Research, "[Intelligence at one-eighth the bits](https://www.fermionresearch.com/research/one-eighth-the-bits/)," "[Introducing the Neutrino-1 models](https://www.fermionresearch.com/research/neutrino-8b/)," and "[The Neutrino Engine](https://www.fermionresearch.com/research/the-neutrino-engine/)" (all July 27, 2026); model weights at [huggingface.co/FermionResearch](https://huggingface.co/FermionResearch). All figures and quotes are self-reported by Fermion Research with no third-party replication I could find. The two figures embedded above are reproduced from Fermion Research's own per-layer and per-tensor measurements published in "One-eighth the bits"; the three interactive components are mine. Related: [Ternary15M](/articles/ternary15m), the from-scratch 15M-parameter version of the same bet.* --- # Kimi K3: a 2.8T open model that turns compute into intelligence 2.5× better > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/kimi-k3 > date: 2026-07-17 > tags: llm, mixture-of-experts, linear-attention, kimi, scaling, explainer Moonshot's [Kimi K3](https://www.kimi.com/blog/kimi-k3) is the largest open model anyone has shipped: **2.8 trillion** parameters, **104B active** per token, a **1-million-token** context, natively multimodal. The [weights are now out](https://huggingface.co/moonshotai/Kimi-K3) under the Kimi K3 License, along with a [47-page technical report](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf) — so the interesting claims can finally be checked against a `config.json` instead of a blog post. **Updated 2026-07-27** after the weights and technical report landed. The most important change: K3 activates **104B parameters per token**, not the ~50B figure circulating at announcement. That is a 3.2× jump over K2's 32.6B, and it reframes the efficiency story — K3 is not a cheap-to-run model that punches up, it is a genuinely large model that converts its compute unusually well. The interesting part is not the parameter count. It is *how the parameters are spent*. K3 is built on two attention changes — **Kimi Delta Attention (KDA)** and **Attention Residuals (AttnRes)** — that rework how information flows across sequence length and across depth, and it scales up MoE sparsity hard: it activates **16 of 896 routed experts** per token (plus 2 shared), inside a **Stable LatentMoE** framework. Together with refined training and data recipes, those structural changes yield a measured **2.5× improvement in overall scaling efficiency** over K2. This piece is a first-principles tour of each piece, why it is new, and what a model like this actually costs to build. Read the block bottom-to-top: the hidden state passes through the attention sublayer (Gated MLA + KDA), Attention Residuals reach back to earlier depths, and Stable LatentMoE routes the token to 16 of 896 experts before the block emits its output. Four ideas, each doing a specific job. Take them one at a time. ## What the weights actually say Before the mechanisms, the ground truth. Here is K3's shape as recorded in the released `config.json` and the report's model summary: | | | |---|---| | Total / activated parameters | **2.78T / 104.2B** | | Layers | **93** (1 dense, 92 MoE) | | Attention composition | **69 KDA + 24 Gated MLA** — 3:1 per block, plus a final global layer | | Hidden dimension | 7168 · 96 attention heads · head dim 128 | | Routed experts | **896**, **16 active** per token, **2 shared** | | Latent MoE dimension | 3584 (half of hidden) · per-expert hidden 3072 | | MLA compression | `kv_lora_rank` 512 · `q_lora_rank` 1536 | | AttnRes block size | **12** layers (8 blocks, 9 counting the embedding) | | Positional encoding | **none (NoPE)** | | Activation | SiTU-GLU (`hidden_act: "situ"`) | | Router | sigmoid scoring, `noaux_tc` — auxiliary-loss-free | | Vision encoder | MoonViT-V2 · 401M · 27 layers · patch 14 | | Context / vocabulary | 1,048,576 tokens · 163,840 | | Quantization | MXFP4 weights, MXFP8 activations (QAT) | Two of these are worth pausing on. The **3:1 KDA-to-MLA ratio** is not approximate — the config lists exactly which layers are which, and the full-attention layers land on 4, 8, 12, … 92, then 93. The last layer is *always* global attention, so whatever the linear layers summarized, the model gets one final unrestricted look at the whole sequence. And `attn_res_block_size: 12` pins down the AttnRes design: 93 layers partitioned into blocks of 12. Set against K2, the shape of the bet becomes clear: ## The official architecture diagram
## Kimi Delta Attention: constant-size memory over a million tokens Ordinary softmax attention keeps a **KV cache** that grows by one entry per token. At a 1M-token context that cache is the whole ballgame: decoding is [memory-bound on a cache that scales with sequence length](/articles/how-llm-inference-works), and it only gets heavier as the context fills. KDA is a **gated delta-rule linear attention**. Instead of a growing cache it keeps a **fixed-size recurrent state** $S_t$ that each token updates in place: it *erases* a little of the old state (a gated decay) and *writes* the new key/value association (the delta rule). The report's exact form applies a channel-wise decay before the delta update: $$ S_t = \left(I - \beta_t k_t k_t^{\top}\right) \mathrm{Diag}(\alpha_t)\, S_{t-1} + \beta_t k_t v_t^{\top}, \qquad \tilde{o}_t = S_t^{\top} q_t $$ where $\alpha_t \in (0,1)^{d_k}$ is the **channel-wise** one-step retention factor (the erase, per channel rather than per head) and $\beta_t \in (0,1)$ controls the delta-rule write strength. The state $S_t$ is a fixed $d_k \times d_v$ matrix — its size does not depend on how many tokens came before. Queries and keys are produced by a short convolution followed by Swish and L2 normalization; values by a short convolution and Swish. Scrub the recurrence and watch the state stay constant-size while a softmax cache piles up: That constant-size state is what makes a genuine 1M context tractable, and it is why Moonshot reports **up to 6.3× faster decoding** in million-token contexts. It is not free — a linear-attention state is a lossy summary, not a perfect record, so K3 interleaves KDA with full-attention layers (via Gated MLA) to keep exact recall where it matters. KDA also breaks the assumptions of conventional prefix caching, so Moonshot contributed a KDA implementation to the vLLM community to make serving practical. ### FlashKDA: the equation, shipped as a kernel Moonshot also open-sourced the kernel underneath. **[FlashKDA](https://github.com/MoonshotAI/FlashKDA)** (MIT) — "Flash Kimi Delta Attention" — is a set of CUTLASS kernels for exactly the recurrence above, and its exposed signature is a direct read-out of that math. The main kernel, `flash_kda.fwd`, takes query, key and value plus a **gate** tensor and **beta logits passed through a sigmoid** — the gate is $\alpha_t$, the channel-wise decay; beta is $\beta_t$, the delta-rule write strength — with an optional recurrent state in and out, which is $S_t$ itself: the fixed-size matrix that is the whole reason a 1M context stays tractable. It also batches variable-length sequences via cumulative sequence lengths and runs mixed precision (bf16 activations, fp32 params). FlashKDA needs SM90+ (Hopper and newer), CUDA 12.9+ and PyTorch 2.4+, ships benchmarks for H20 and GB200, and auto-integrates with the `flash-linear-attention` library (v0.5.0+) through a `chunk_kda` op — the same KDA lineage as the vLLM contribution above, now available as a standalone package rather than only living inside a serving framework. ### NoPE: no positional encoding at all Here is a detail the weights make unambiguous, and it is one of the quietly radical choices in K3: **there is no positional encoding**. Not RoPE, not ALiBi, not learned embeddings. K2 used [RoPE](/architectures); K3 applies **NoPE** to every MLA layer and lets the KDA layers carry position implicitly through their gating and decay — a recurrence is inherently order-sensitive, so position falls out of the mechanism rather than being added to it. The payoff is at the context frontier. Extending a RoPE model to 1M tokens normally means rescaling the frequency base or applying YaRN-style interpolation, and every such trick is a place quality can quietly degrade. With NoPE there is nothing to rescale: the report says K3 **extrapolates directly to 1M-token contexts without any positional-encoding modification**. The 3:1 hybrid earns its keep here — KDA supplies position-sensitive, recency-aware mixing, while the NoPE-MLA layers supply unrestricted global content interaction, and the two jobs stay cleanly separated. ## Attention Residuals: selective retrieval across depth The second change is about *depth*, not length. A plain residual stream compresses all prior information into a single state as it climbs — a bottleneck the report pointedly compares to an RNN over time. Transformers already solved that problem along the sequence axis by replacing recurrence with attention; **Attention Residuals** applies the same move to depth: each layer *selectively retrieves* representations from preceding layers rather than accumulating them uniformly. Toggle the two modes and scrub the current layer: Mechanically, each layer $l$ carries a learnable pseudo-query $q_l$; the keys and values are the outputs of all earlier layers (plus the token embedding), and attention weights come from a softmax kernel with an RMSNorm inside — the norm stops layers with large-magnitude outputs from dominating the read. Because depth is modest ($L < 100$), the full $O(L^2 d)$ form is affordable in arithmetic; the real cost is the $O(Ld)$ memory of keeping every layer output alive. **Block AttnRes** is the fix, and it is what K3 actually ships. The 93 layers are partitioned into blocks of 12; within a block, layer outputs are summed into one representation, and full attention runs only over the ~8 block-level representations. Memory and cross-stage communication drop from $O(Ld)$ to $O(Nd)$, and the block structure bounds the inference-time state so inter-block results merge with intra-block partial sums via online softmax. The report notes $N \approx 8$ recovers most of the benefit — which is exactly the 8 blocks of 12 the config encodes. The payoff Moonshot reports is concrete: about **25% higher training efficiency at under 2% additional cost**. That ratio is the tell — a cheap structural change that improves gradient flow and lets the stack go deeper without the usual degradation, which is exactly the kind of lever that compounds into the headline 2.5× scaling number. Alongside these, the attention sublayer uses **Gated MLA** — Multi-head Latent Attention with an input-dependent, channel-wise **full-rank** output gate, letting each token modulate which channels it reads from global attention. The MLP nonlinearity is a **Sigmoid Tanh Unit (SiTU-GLU)**, whose gate branch is a $\tanh$ bounded by a constant, so activations cannot blow up. Small pieces, but at 2.8T scale "bounded" is load-bearing. ## Stable LatentMoE: 16 of 896, and why that is hard Here is the aggressive part. K3's feed-forward is a mixture of experts with **896 routed experts**, of which only **16** fire for any given token (plus 2 always-on shared experts) — a sparsity of **56**. The experts are **latent**: rather than each selected expert receiving the full 7168-dimensional token, routed experts operate in a compact **3584-wide** latent space, half the model width, while the shared experts keep a full-width path. That separation is what makes the expansion affordable — in a conventional MoE, communication and expert-weight traffic grow with routing multiplicity, so going to 16 active experts would be punishing at full width. Scrub a few tokens and watch the selected 16 change: At this sparsity, two problems that are mild in a denser MoE become first-order. **Exploding activations:** the routed path composes a down-projection, a gated multi-branch expert FFN, and an up-projection into a chain of nearly four consecutive matmuls — ill-conditioned at 2.8T scale, which is what the normalization and the bounded SiTU-GLU are there to contain. **Load balance:** balancing nearly a thousand experts per layer exceeds the regime where existing auxiliary-loss-free schemes hold up. If a few experts hog the tokens, the rest never train, and the effective model collapses to something far smaller than 2.8T. ### Quantile balancing: no auxiliary loss, no knob K3 stays auxiliary-loss-free: balancing is done by adding a per-expert bias $b_j$ to the router score *before* Top-$k$ selection, and then omitting that bias from the mixture weights — so it steers dispatch without touching the router's gradients. The standard version nudges $b_j$ by a fixed step in the direction of the load error, which forces a trade-off between slow adaptation and load oscillation. **Quantile Balancing** replaces the nudge with a direct solve. Routing runs Top-$(k{+}1)$ instead of Top-$k$: the first $k$ entries are the routes actually taken, and the $(k{+}1)$-th is the **cutoff** a competing expert would have had to beat. Each expert's next bias is then read off as a quantile of its *margins* (score minus cutoff) across the batch — specifically the $(1 - k/n)$-quantile — which by construction hands every expert exactly its target load of $mk/n$ tokens. No auxiliary loss, no balance coefficient. Drag the quantile and flip to the aux-loss regime to see the imbalance it removes: At training scale those margins number in the millions and are scattered across ranks, so an exact quantile is not computable. K3 estimates it from a **histogram**: each rank bins its own margins, a single all-reduce sums the bin counts, and the quantile is recovered from the pooled histogram. Because counts are additive, the estimate reflects the true whole-batch quantile up to the bin width — at a communication cost of a few hundred bins per expert. The bias is frozen at inference. The systems half matters just as much. K3 uses **perfectly balanced expert-parallel training with static shapes and no host synchronization**. Variable expert loads normally produce variable tensor shapes, which force recompilation and host-side synchronization that stalls a large cluster. Quantile balancing gives every expert the same load, so the shapes are static, so the expert-parallel pipeline runs without host sync — the difference between 16-of-896 routing being a nice idea and being trainable at 2.8T. With all four pieces on the table, here is the module-level picture redrawn: the **Stable LatentMoE** and **KDA** blocks in full detail on the left, and on the right the **Block Attention Residuals** backbone — where each module's output flows through an `α` gate that can read *every* earlier block and the embedding, not just the layer below it. ### MoonEP: the same static-shapes claim, from the communication side **[MoonEP](https://github.com/MoonshotAI/MoonEP)** (MIT) is Moonshot's expert-parallel communication library, and it is the static-shapes claim from the quantile-balancing section above, attacked from the other direction. Quantile balancing makes every expert's *load* equal before dispatch; MoonEP instead guarantees every rank *receives* exactly $S \times K$ tokens — $S$ input tokens per rank, $K$ routed top-$k$ per token — no matter how skewed the actual routing is. The mechanism is a small number of redundant experts, planned online from the current router outputs by a near-optimal GPU planning kernel and prefetched before expert computation runs, with their gradients reduced back to their home ranks on the backward pass. Because those redundant experts absorb whatever skew is left, every rank ends up with an identical, statically-known token count — "statically known shapes eliminate per-layer MoE host synchronization" is not a paraphrase of that claim, it is MoonEP's own description of what it buys. Tokens land directly in their expert-grouped positions on remote ranks through zero-copy buffer views, so only a fixed $S \times K$ buffer is needed per layer, with no per-layer host synchronization to stall the pipeline. Moonshot's own benchmarks against DeepEP v2 on H20 make the case concrete: MoonEP's communication time stays close to flat as imbalance (maxvio) grows, while DeepEP v2 degrades steadily and eventually OOMs under high imbalance — and MoonEP's iteration time holds flat across the same range. It targets NVIDIA GPUs today, with Zhenwu PPU support listed as under review, and credits DeepEP, Echo and UltraEP as inspiration. ## Native vision, trained from scratch K3 is natively multimodal — text, images and video share one backbone and one context, with no post-hoc alignment stage. The notable choice is how the vision tower was trained. Standard practice, including K2.5's own, initializes the encoder from a contrastively pre-trained model like SigLIP. K3 instead trains **MoonViT-V2** (401M params, 27 layers, patch 14) **entirely from scratch with next-token prediction**. The reason given is stability, and the report shows the receipts: the SigLIP-initialized tower ran persistently higher gradient norms with frequent spikes, while the from-scratch tower stayed smooth. Training under the language-modeling objective also shapes visual features by what the LLM actually needs — fine-grained text and structure — rather than the global semantics a contrastive loss rewards. The conclusion is the interesting bit: MoonViT-V2 **matched** the SigLIP-initialized baseline on vision evals, so at this scale contrastive pre-training simply was not necessary. ## Turning compute into intelligence Stack it up — KDA's cheap long-context memory, AttnRes's cheap depth, LatentMoE's extreme-but-stable sparsity, plus refined training and data recipes — and the headline is a **~2.5× improvement in overall scaling efficiency** over K2. This is not a vibe: it is a fitted scaling-law comparison on held-out out-of-distribution validation data, with hyperparameters (batch size, learning rate, tokens-per-parameter, model shape) re-tuned independently for each family so neither is handicapped by the other's settings.
Read the gap horizontally: pick any loss level and the red curve reaches it about 2.5× further left on the FLOPs axis. Drag the capability marker to see the same trade in the other direction: That is the number that actually matters. "2.8 trillion parameters" is a spec-sheet figure; "2.5× more capability per FLOP" is an engineering result. A side note from the same study, useful to anyone tuning their own runs: under independently optimized hyperparameters, **cosine decay consistently beat WSD** — the two schedules have very different optimal peak learning rates and batch sizes, so comparisons that share one hyperparameter set tend to be unfair to whichever schedule they fit worse. ## What it would take to train it So what does building a 2.8T-A104B model actually cost? Sparsity still helps: training compute for an MoE scales with the **active** parameters, not the total, so K3's per-token training FLOPs are those of a ~104B model rather than a 2.8T one. The standard estimate is $$ C \approx 6 \, N_{\text{active}} \, D $$ with $N_{\text{active}} \approx 104\text{B}$ and $D$ the number of training tokens. Moonshot still has not published K3's token budget; for reference, K2 was trained on **15.5T tokens**. Plug in a frontier-scale budget and pick a cluster: Three things make that estimate *achievable* rather than merely large: - **Per-Head Muon.** K3 extends the Muon optimizer so that Newton–Schulz orthogonalization is applied to each attention head's momentum block *separately* rather than to the whole Q/K/V projection. Full-matrix orthogonalization lets large-gradient heads dominate the shared update direction; per-head equalizes the update scale across heads, which improves stability at scale — and is slightly cheaper, since the iterations run on tall thin blocks. - **MXFP4 / MXFP8 quantization-aware training.** From the SFT stage onward, K3 trains with **MXFP4 expert weights and MXFP8 activations**, while attention projections, latent-MoE projections, shared experts and routers stay in higher precision. The model is trained to be low-precision-native, which is why the full 2.8T weights fit in roughly **1.4 TB** and why it is servable at all without a quality cliff. - **Static-shape expert parallelism.** As above — quantile balancing plus static shapes and no host synchronization is what keeps a large cluster busy instead of stalling on dynamic routing. The context window is built up rather than trained flat: pre-training starts at **8K** and extends to **64K**, then the cooldown phase walks **256K → 1M**. Concentrating the expensive long-sequence compute into a small slice of the budget is what makes a 1M-token model economical. Length alone does not confer long-range ability, so Moonshot also *synthesizes* long-context data by permuting and concatenating documents and sub-tasks such that the embedded task can only be solved by attending across the full window — otherwise attention quietly degenerates into local patterns. ## Post-training: nine experts, then one The pre-training story is where the architecture lives, but K3's post-training has a structure worth drawing. It is a three-stage funnel: SFT for a cold-start policy, then RL that trains **nine separate experts** — three domains crossed with three reasoning-effort levels — then **Multi-Teacher On-Policy Distillation** to collapse all nine back into the single shipped checkpoint. A few mechanisms make that work at 1M context: - **Partial rollout.** In long-horizon RL, a handful of straggler trajectories can hold up an entire iteration. Generation instead pauses once a fraction $\lambda$ of trajectories finish; the rest are enqueued and resumed at the start of the next iteration, backed by persistent sandbox state. That means a single trajectory can span several iterations, so the algorithm has to tolerate badly stale off-policy data — which it does via a per-token regularization that keeps updates in a local neighborhood. - **A white-box harness, not *the* harness.** Training against one fixed agent scaffold teaches the model that scaffold's conventions. Moonshot's RL environment represents a harness as composable modules — tools, system prompts, context management, skills, memories, subagents — and can instantiate Kimi Code, Claude Code, Codex, OpenClaw and Hermes, mixing configurations across task groups so the model generalizes across harnesses rather than overfitting one. - **Deployment-aware training.** QAT runs through the *entire* post-training stage, and during RL the rollout and the training pass share the same quantization scheme — eliminating the train/inference mismatch that usually shows up when a model is quantized after the fact. Separately, K3's pre-trained multi-token-prediction layer is fine-tuned into an EAGLE-3-style **draft model** for speculative decoding, optimized directly against the acceptance rate rather than a KL surrogate. ### AgentENV: the sandbox layer, and it is open source The piece that makes all of the above physically possible is the sandbox. Long-horizon agentic RL means running an enormous number of real machines that agents can break, and **[AgentENV](https://github.com/kvcache-ai/AgentEnv)** — built by Moonshot with partners, and released under MIT — is the microVM runtime they built for it. The motivation is refreshingly blunt. Container-based sandboxes were not enough: in early experiments, aggressive agent exploration caused **kernel panics and deadlocks**. And clamping down is the wrong fix, because hard tasks need a sandbox close to a real machine — agents should be able to mount disks, run containers, even launch VMs. So AgentENV runs each sandbox as an isolated **Firecracker** microVM, buying isolation and fidelity a container cannot. On top of that it adds three lifecycle operations tuned specifically for RL: - **Pause / resume.** A paused sandbox consumes no memory or CPU. This matters more than it sounds: the sandbox spends as much as **98% of its lifetime** just waiting on the model's next inference result. Pausing that window is the difference between renting an idle fleet and not. - **Fork.** Branch a new sandbox from the *exact* state of a running one while the original keeps going — which is how you run a reward judge against a trajectory without any side effects leaking back into it. - **Snapshot.** Periodic checkpoints for error recovery. The engineering is in the latencies: incremental checkpointing saves only pages dirtied since the last checkpoint, giving **133 ms checkpoint and 49 ms resume**. Images use OverlayBD with a custom `ublk` driver, storage-layer sharing and P2P transport, so tens of thousands of sandboxes with distinct images launch in **under a second**; copy-on-write memory and page-cache tuning push memory overcommit to **6.5×** in real workloads. The scale number is the one worth sitting with. Across K3's training and evaluation, Moonshot created **51,219,741 sandboxes** spanning **1,505,678 distinct images**. That is what "agentic RL" costs when you actually run it — and it is the part of the frontier stack that almost never gets published, let alone open-sourced. AgentENV is one of three pieces of that stack Moonshot has now open-sourced: **MoonEP** (expert-parallel communication, covered under Stable LatentMoE above) and **FlashKDA** (the attention kernel, covered under Kimi Delta Attention above) are the other two — sandbox, communication and kernel, all MIT-licensed. ## The benchmarks On coding, K3 is a clear #2-or-#3 behind Fable 5 and GPT-5.6 Sol, and ahead of everything else open or closed that Moonshot tested — with a few outright wins.
On FrontierSWE it sits second, close behind Fable 5 and well ahead of the rest: On Terminal Bench 2.1 it is effectively tied for first, and on the long-horizon SWE Marathon and Program Bench it is first outright: The agentic and visual picture is similar — competitive across the board, and #1 on browsing:
The pattern is consistent: K3 wins where the task is long-horizon and tool-heavy (SWE Marathon, Program Bench, BrowseComp, Automation Bench, SpreadsheetBench 2), and comes second to Fable 5 or GPT-5.6 Sol on the single-shot, knowledge-dense ones (GDPval and AA-Briefcase Elo, DeepSWE). ### Independent numbers The obvious objection to everything above is that it is Moonshot grading its own homework. The report also collects third-party leaderboards, which is the more useful evidence: | Leaderboard | Kimi K3 | Rank | Best proprietary | |---|---|---|---| | Artificial Analysis Intelligence Index v4.1 | 57.1 | #4 / 580 | Fable 5 — 59.9 | | Vals Index | 74.7 | **#2 / 39** | Fable 5 — 75.1 | | WebDev Arena (Elo) | 1,678 | **#1 / 99** | Fable 5 — 1,634 | | Text Arena (Elo) | 1,486 | #8 / 200 | Fable 5 — 1,507 | | Agent Arena | 9.1 | #4 / 37 | Fable 5 — 12.7 | An open model holding **#1 on WebDev Arena** and **#2 on the Vals Index** — 0.4 points off Fable 5 — is a materially different claim from a vendor bar chart. Text Arena at #8 is the honest counterweight: general chat preference is not where K3 shines. ## What it costs to serve The sparsity that makes K3 cheap to train makes it cheap to run. API pricing is **$0.30 / MTok** on cache-hit input, **$3.00 / MTok** on cache-miss input, and **$15.00 / MTok** output — and Moonshot reports cache-hit rates **above 90%** in coding workloads, so the effective input price is closer to the cheap number than the expensive one. MXFP4 weights keep the footprint at ~1.4 TB. The weights ship with deployment recipes for **vLLM**, **SGLang** and **TokenSpeed**. The report's cost-efficiency comparison is the most quotable result in it:
Concretely: on **BrowseComp**, K3 takes the best score (91.2%) at **$2.03 per task** — half the cost of GPT-5.6 Sol and an order of magnitude cheaper than the Claude models at max effort. On **Kimi Code Bench 2.0** it is 4.0 points behind Fable 5 at **38% of the cost**, and at *high* effort it already matches Opus 4.8's *maximum*-effort score at roughly a third of the price. On **GDPval-AA v2** it is within 50 Elo of GPT-5.6 Sol at 13% lower cost, and 2.6× cheaper than Fable 5. ## What it can actually build The case studies are where the long-horizon claims get concrete, and they are unusually ambitious: - **GPU kernel optimization.** Given a sandbox and up to 24 hours per task, K3 cut **AttnRes** kernel latency from 283.6 ms to **114.4 ms**, cut DSA and KDA runtime by **55.1%** and **73.6%**, and reached over half of peak TFLOPS on MLA — matching Fable 5 and beating Opus 4.8, GPT-5.6 Sol and GPT-5.5. Moonshot notes an early K3 checkpoint was already doing most of their kernel-optimization work during late development. - **A GPU compiler.** K3 built [MiniTriton](https://github.com/MoonshotAI/minitriton), a Triton-like compiler with a tile-level Python frontend, an MLIR annotation layer and a PTX codegen pipeline, plus a dual-mode tensor library with reverse-mode autograd and NCCL distributed primitives. On an L20 it beats PyTorch eager and `torch.compile` in geometric mean, its from-scratch tensor-core matmul reaches ~90% of the measured machine roof, and it trains a GPT end-to-end with gradients matching torch autograd to within torch's own fp32 rounding error. - **A chip.** In a single **48-hour autonomous run**, K3 designed, optimized and verified an inference-chip prototype ([nano-kpu](https://github.com/MoonshotAI/nano-kpu)) using open-source EDA tools and the Nangate45 cell library. Inside a 4 mm² budget it closes timing at 100 MHz for an RTL-simulated **8,700+ tokens/s** decode, with 1.46M standard cells, 0.277 MiB of SRAM and an INT4 MAC array with fused dequantization. **Read the caveats.** (1) K3 still **trails Fable 5 and GPT-5.6 Sol** on overall capability and on user-experience polish — Moonshot says so directly, and flags sensitivity to thinking-history preservation and over-proactiveness in ambiguous situations. (2) The headline charts are **Moonshot's own suite**, with opponents' fallbacks (Fable 5 hit fallbacks on 35% of SWE-Marathon tasks) and cyberguards (GPT-5.6 Sol) noted; several suites also run K3 in *its own* harness (Kimi Code) against competitors in Claude Code or Codex. The third-party leaderboards above are the better evidence. (3) SWE-Marathon and PostTrainBench were run on **H20 GPUs** against an H20-recalibrated task branch, not the official H100 setting. (4) The training-cost estimate is a **first-principles calculation**, not a disclosed figure — Moonshot has published neither K3's token budget nor its cluster. (5) The license is a **custom Kimi K3 License**, not Apache or MIT — read it before commercial use. ## The take Strip away the size record and what is genuinely new in K3 is a coherent set of efficiency bets: **KDA** buys a real 1M context with constant-size memory; **NoPE** means that context needs no rescaling tricks to reach; **AttnRes** buys depth almost for free; **Stable LatentMoE** with **Quantile Balancing** buys 2.8T of capacity at 104B of active compute *and* makes that extreme sparsity trainable without an aux-loss knob or host-sync stalls; **Per-Head Muon** and **MXFP4/MXFP8 QAT** make the whole thing converge and fit. The sum is the number that matters — **~2.5× more capability per FLOP than K2**, measured on fitted scaling curves — delivered in the open at 2.8T. It does not top the frontier, and it does not pretend to. What it proves is that the gap between open and closed is now measured in scaling *efficiency*, not in whether an open lab can build at frontier scale at all — and with the weights, the config, and a 47-page report on the table, that claim is now something anyone can go audit. --- *Sources: the [Kimi K3 technical report](https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf) (architecture, scaling law, post-training, infrastructure, evaluations, case studies), the released [model weights and card](https://huggingface.co/moonshotai/Kimi-K3) (config, deployment, license), the [AgentENV repository](https://github.com/kvcache-ai/AgentEnv) (sandbox runtime), the [MoonEP repository](https://github.com/MoonshotAI/MoonEP) (expert-parallel communication), the [FlashKDA repository](https://github.com/MoonshotAI/FlashKDA) (attention kernels), and the [Kimi K3 tech blog](https://www.kimi.com/blog/kimi-k3) (pricing). Figures 3–5 here are the report's Figures 2, 7 and 13, reproduced for commentary. Benchmark numbers are Moonshot's except where marked third-party; the training-cost figures are a first-principles estimate from $C \approx 6\,N_{\text{active}}\,D$ with clearly labeled assumptions, using K2's 15.5T-token budget as a reference. Interactive diagrams are mine; the routing, loop and cost visuals are illustrative.* --- # BTL-3: a rank-32 LoRA that turns Qwen3.6-27B into a tool-use agent > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/btl-3 > date: 2026-07-24 > tags: agents, tool-use, code-generation, fine-tuning, open-weights, explainer Most "new models" are not new weights. **BTL-3**, from **Bad Theory Labs**, is a clean example: it is not a from-scratch 27B model but a **frozen rank-32 PEFT LoRA adapter** — about **934 MB** of weights — post-trained on top of a pinned revision of **Qwen3.6-27B**. Load the base, apply the adapter, and you get an agent tuned for coding, repository work, and structured tool use. The raw capability is Qwen's; what BTL-3 contributes is **behaviour** — how the model runs an agent loop and, notably, when it decides *not* to act. That framing matters for reading the numbers honestly, so keep it in mind: this is a post-training result on a strong open base, released under **Apache-2.0**, with the base checkpoint pinned to an exact revision for reproducible loading. The card labels this frozen release **"RL-0013,"** and the maximum RL sequence length (65,536 tokens) tells you the adapter was shaped by reinforcement learning over long, multi-step trajectories — not just supervised fine-tuning on completions. ## The loop it was tuned to run An agent model earns its keep inside a loop, not on a single completion. BTL-3's stated job is to "reason, act, inspect tool results, recover from failures, and stop when no action is required." That last clause is the interesting one. A tool-happy model that always reaches for a function call is easy to train and annoying to deploy; the harder behaviour is the **route** decision — recognising when a question needs a tool and when it just needs an answer. Pick a scenario and watch which path the model takes: The four scenarios line up with the four things BFCL (the Berkeley Function-Calling Leaderboard) v4 actually measures: a **single** call, **parallel** calls fired at once, recovery when a call **fails**, and **irrelevance** — correctly declining to call anything. BTL-3's headline highlight is that last one: a self-reported **91.2%** on knowing when to stay its hand. The loop is the product here; the adapter's whole point is to make Qwen3.6-27B move through it reliably. ## The tool-use profile Break BFCL v4 down by category and the shape of the model shows. It is strongest on the straightforward cases and gives ground exactly where you'd expect — when it has to compose *several* tools that each take *several* arguments: The aggregate is **88.5% BFCL v4 AST** (1097/1240 on the full official set). The 70.0% on parallel-multiple is the honest soft spot: issuing several correct calls at once, each with the right arguments, is where structured tool use is genuinely hard, and a fifth of those cases still slip. The **91.2% irrelevance** number is the one worth internalising — it is the difference between an agent you can leave in a loop and one that invents work. ## Coding, and where it falls off On standard code-generation benchmarks in **thinking mode**, BTL-3 posts strong pass-rates — and then drops sharply on the hardest composite tasks. That gap is the useful part of the picture, not a number to bury: HumanEval at **95.12%** (156/164) and LiveCodeBench v6 at **88.1%** (170/193) are the flattering figures — well-scoped "write this function" problems. **BigCodeBench-Hard Instruct at 26.35%** (39/148) is the sobering one: strict pass@1 on tasks that chain many library calls into one correct program is a different sport, and here the model solves roughly one in four. (BTL reports a softer **59.25%** at the individual *test* level on the same suite — useful context, but a test-level score is not a solved-task score, so read the strict 26.35% as the real one.) These are different benchmarks at different difficulties, not a like-for-like ladder — the labels carry that. ## The Compact edition Alongside the adapter, Bad Theory Labs ships **BTL-3 Compact**: the complete text model packed into a single **8.39 GB** native file — smaller than an 8B model stored in FP16, which works out to an effective **under 2.5 bits per parameter**. The claimed cost of that compression is measured on a "fresh private 100-turn tool-contract gate": Compact retained **83 of the 90 behaviours** the full model completed correctly, which BTL reports as **92.2% conditional tool-behaviour retention**. Read that metric for exactly what it is. It is a *private* gate that BTL defined and ran, conditioned on cases the full model already passed — so it says "Compact reproduces most of what the full model got right," not "Compact loses only 8% overall." It's a reasonable internal check and a genuinely useful artifact (a 27B-class agent in 8.39 GB is easy to self-host), but it is not an independent quality measurement. ## Running it Because BTL-3 is a LoRA adapter, deployment is "load Qwen3.6-27B at the pinned revision, then apply the adapter" — a few lines with PEFT and Transformers, or `vllm serve` with `--enable-lora` and `--max-lora-rank 32` and the Qwen XML tool parser for structured calls. The architectural context window is **262,144 tokens** (inherited from Qwen3.6's hybrid attention), though the published benchmarks were run at a **32,768-token** launch context. BTL recommends **thinking mode** for coding and reasoning, which is also the mode every headline score was measured in. The model card itself is direct about this: **run generated code and tool calls in a sandbox**, and require **explicit confirmation before destructive, privileged, financial, or otherwise high-impact actions**. An agent that scores 88.5% on tool calls still gets more than one call in ten wrong — that residual is exactly where an unsandboxed loop does damage. ## The honest read Every number here is **self-reported by Bad Theory Labs** — there is no independent evaluation yet. The underlying **capability is Qwen3.6-27B's**; BTL-3 is a **post-training / RL result**, so credit the adapter for the *loop behaviour* (tool routing, irrelevance, recovery), not for raw reasoning power. Scores were measured in **thinking mode** at a **32K context** on protocols BTL chose, the coding wins sit next to a **26.35% BigCodeBench-Hard** floor, and the Compact edition's **92.2% retention** is a private, conditional gate, not a public benchmark. The model card ships **no figures or diagrams** — the loop diagram above is my own reconstruction of the described behaviour. No training-data disclosure is provided. ## The take BTL-3 is a modest, honest kind of release: take a strong open base, spend an RL budget teaching it to behave inside an agent loop, and ship the ~934 MB of difference under Apache-2.0. The most interesting claim isn't a coding score — it's the **91.2% irrelevance**, the tuned instinct to *not* call a tool. For anyone assembling a private, self-hosted coding agent, that plus the 8.39 GB Compact build is a concrete, deployable proposition. Just hold the framing straight: the intelligence is Qwen's, the discipline is BTL's, and until someone outside Bad Theory Labs runs the suite, every figure is a vendor's own. --- *Source: the [BTL-3 model card](https://huggingface.co/badtheorylabs/BTL-3) and [BTL-3 Compact](https://huggingface.co/badtheorylabs/BTL-3-Compact) on Hugging Face, plus the [runtime source](https://github.com/Badtheorylabs/BTL-3). The card ships no figures, so the diagram here is mine; all benchmark numbers are Bad Theory Labs' self-reported values.* --- # Token-level RL is a first-order approximation to the reward you actually want > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/first-order-rl > date: 2026-07-24 > tags: reinforcement-learning, llm, mixture-of-experts, post-training, explainer RL for reasoning models rests on a mismatch nobody had really justified. The **reward** is assigned to a *whole response* — you sample a full chain of thought, check the final answer, and hand back one scalar. But the **optimizer** — REINFORCE, GRPO, and the rest — works one *token* at a time. We reward the sequence and update the tokens, and we mostly just trust that closing the loop this way improves the thing we scored. The Qwen team's [Stabilizing Reinforcement Learning with LLMs](https://arxiv.org/abs/2512.01374) (Zheng et al., arXiv:2512.01374) takes that trust and makes it a theorem with fine print. Their claim: the token-level objective is a **first-order approximation** to the true sequence-level reward — exact in the limit, and valid only when two specific gaps are small. The nice part is what falls out of it. Importance-sampling correction, clipping, and Routing Replay for Mixture-of-Experts models — a grab-bag of stabilization tricks that each arrived with its own justification — turn out to be the *same move*: keep the approximation valid. One lens, and the whole toolbox lines up behind it. ## The objective you can't optimize Write the thing we actually want to maximize — expected reward over responses the current policy would generate: $$ J^{\text{seq}}(\theta) \;=\; \mathbb{E}_{x\sim\mathcal{D},\; y\sim\pi_\theta(\cdot|x)}\big[R(x,y)\big]. $$ There's an immediate wrinkle: we don't sample $y$ from the policy we're training. Responses come out of a fast **inference engine** (vLLM, SGLang) running policy $\mu_{\theta_{\text{old}}}$, while gradients are taken in a **training engine** (Megatron, FSDP) holding $\pi_\theta$. The standard fix is an importance-sampling reweight onto the rollout policy $\mu$: $$ J^{\text{seq}}(\theta) \;=\; \mathbb{E}_{x\sim\mathcal{D},\; y\sim\mu_{\theta_{\text{old}}}(\cdot|x)}\!\left[\underbrace{\frac{\pi_\theta(y|x)}{\mu_{\theta_{\text{old}}}(y|x)}}_{\text{sequence-level IS weight}} R(x,y)\right]. $$ This is correct and completely impractical. A sequence likelihood is a product of hundreds or thousands of per-token probabilities, so the ratio $\pi_\theta(y|x)/\mu_{\theta_{\text{old}}}(y|x)$ swings across an enormous dynamic range with brutal variance. Its gradient is technically right and numerically hopeless. Nobody trains on it directly. ## The surrogate everyone actually uses So instead we optimize the **token-level** objective — sum the per-token IS ratios instead of multiplying them: $$ J^{\text{token}}(\theta) \;=\; \mathbb{E}_{x\sim\mathcal{D},\; y\sim\mu_{\theta_{\text{old}}}(\cdot|x)}\!\left[\sum_{t=1}^{|y|}\underbrace{\frac{\pi_\theta(y_t|x,y_{ The intuition to keep: the surrogate isn't wrong, it's *truncated*. As long as the policy you're optimizing stays close to the policy that generated the data, the truncation is negligible and improving the cheap objective improves the real reward. Let them separate and the surrogate starts optimizing something that isn't the reward anymore — which, in practice, is exactly what a training collapse looks like. ## Two gaps, and the tricks that close them "Keep $\pi_\theta$ close to $\mu_{\theta_{\text{old}}}$" sounds abstract until you factor the token ratio into its two honest sources: $$ \frac{\pi_\theta(y_t|\cdot)}{\mu_{\theta_{\text{old}}}(y_t|\cdot)} \;=\; \underbrace{\frac{\pi_{\theta_{\text{old}}}(y_t|\cdot)}{\mu_{\theta_{\text{old}}}(y_t|\cdot)}}_{\text{training–inference discrepancy}} \;\times\; \underbrace{\frac{\pi_\theta(y_t|\cdot)}{\pi_{\theta_{\text{old}}}(y_t|\cdot)}}_{\text{policy staleness}}. $$ - **Training–inference discrepancy** is numerical. The same weights produce slightly different probabilities in the training and inference engines — different kernels, and inference deliberately disables batch-invariant kernels for throughput, so even one engine isn't self-consistent. This is the gap between $\pi_{\theta_{\text{old}}}$ and $\mu_{\theta_{\text{old}}}$. - **Policy staleness** is procedural. To use more compute per rollout, we split a big batch of responses into mini-batches and take several gradient steps, so later mini-batches are optimized by a $\pi_\theta$ that has already drifted from the $\pi_{\theta_{\text{old}}}$ that generated them. Asynchronous frameworks make it worse. Now the stabilization toolbox reads as one idea — shrink these two gaps so the first-order approximation holds: - The **IS weight itself** is not an optional variance trick; it *is* the first-order term. Drop the training–inference correction and you're no longer approximating the sequence objective at all. - **Clipping** (the PPO move) stops gradients on tokens whose ratio has run too far, directly capping policy staleness. - **Routing Replay**, for MoE models, closes both — and it needs its own section, because MoE breaks the story in a way dense models don't. ## Why MoE breaks it, and how Routing Replay repairs it In a Mixture-of-Experts model, the probability of a token depends on *which experts the router activated* for it. That turns the token ratio into a comparison over possibly *different active parameters*: the inference engine routes the token to expert set $e^\mu$, the training engine to $e^\pi$, and when those sets disagree the ratio $\pi_\theta(y_t|\cdot)/\mu_{\theta_{\text{old}}}(y_t|\cdot)$ stops measuring "a small change in the policy" and starts measuring "two different subnetworks." The $\delta_t$ are no longer small, and the first-order approximation collapses. Routing is entangled with *both* gaps — the engines can route differently (discrepancy) and the router's choice can shift as weights update (staleness). Toggle the fix: **Routing Replay** ([Zheng et al., 2025](https://arxiv.org/abs/2507.18071); Ma et al., 2025) pins the routed experts during optimization so the token is scored over one fixed subnetwork — the model is optimized like a dense one, and the ratio means what it should again. The paper formalizes two flavors, differing only in *whose* routing you replay: | | replays | closes | first mini-batch | |---|---|---|---| | **R2** — Vanilla Routing Replay | the training engine's rollout experts $e^\pi_{\text{old}}$ | policy staleness | target policy **unaltered** | | **R3** — Rollout Routing Replay | the inference engine's experts $e^\mu_{\text{old}}$ | discrepancy **and** staleness | target policy altered | There's no free lunch here, and the paper is careful to say so. Fixing the experts restores the approximation but **biases the target policy** — you're now optimizing a model whose routing is frozen to a past choice, not the routing it would pick itself. R2 leaves the first mini-batch's target policy untouched; R3 alters it from step one but kills more of the discrepancy. Which bias is worth paying turns out to depend on how off-policy you run — a question only experiments can settle. ## MiniRL: the smallest honest baseline To test the formulation instead of a pile of confounded tricks, the authors strip RL down to **MiniRL** — REINFORCE with the token-level IS weight, group-normalized advantages (subtract the per-prompt mean reward), and PPO-style clipping. That's it. It's deliberately the minimal algorithm whose gradient stays faithful to the surrogate the theory justifies, which makes it the right probe: if the formulation is real, the things that preserve the approximation should be the things that stabilize MiniRL. The setup is a genuine stress test. A 30B MoE (cold-started from Qwen3-30B-A3B-Base), **FP8 inference against BF16 training** — deliberately mismatched precisions to *inflate* the training–inference discrepancy — on 4,096 verifiable math problems, scored as average accuracy over 32 samples on HMMT25, AIME25, and AIME24. Hundreds of thousands of GPU hours, roughly 5–6 GPU-hours per gradient step. They track not just reward but two diagnostics: policy **entropy** and the **training–inference KL divergence**, since a collapse announces itself as a KL spike before the score falls. ## What the experiments say **On-policy** (one gradient update per batch), the ablation lands exactly where the theory predicts:
Three reads, each a prediction of the formulation: - **MiniRL wins.** The plain first-order-faithful objective is the most stable and scores highest. - **Removing the training–inference IS correction collapses training** almost immediately — the green curve nose-dives, entropy crashes, KL explodes. The IS weight was never optional; it's the approximation's load-bearing term. - **Length normalization is stable but worse.** Dividing the objective by response length is common (GRPO and CISPO both do it), but it *invalidates* the first-order approximation — the gradient no longer lines up with the true sequence objective — and the benchmark score pays for it. Notably, **Routing Replay does not help on-policy**; with the gaps already small, its bias is all cost and no benefit. **Off-policy** (split the batch into $N$ mini-batches for $N$ updates), staleness enters and the picture changes. Now clipping *and* Routing Replay both become necessary — drop either and training collapses early:
The nuance the paper draws out: at **small off-policiness** ($\text{gbs}=2\times\text{mbs}$), **R2 beats R3** — R2's lighter bias wins when the approximation is only mildly stressed. At **larger off-policiness** ($4\times$, $8\times$), **R3 wins** — R2 can't hold training together and R3's stronger discrepancy-killing earns its bias back. The recipe isn't "always use X"; it's "match the replay to how off-policy you're willing to run." The benchmark numbers in the bar chart above are **read off the training curves in Figure 1**, not a reported results table — treat them as approximate. And the whole study is one task (verifiable math), one model family (Qwen MoE), and a deliberately harsh FP8-inference / BF16-training setup chosen to *amplify* the discrepancy. The mechanism is clean; how the exact crossover points transfer to other rewards, modalities, and precision regimes is not something one paper can settle. ## The result that reframes the field The finding I keep coming back to isn't a trick — it's about what *matters*. Take one base model, cold-start it three different ways (distilling from Qwen3-Max-Thinking-Preview, DeepSeek-R1-0528, and gpt-oss-120b), then run the same stable recipe. They converge to the **same place**:
Once training is stable, *how you started barely matters* — prolonged RL washes out the cold-start differences and even on-policy and off-policy runs reach comparable peaks. The implication is pointed: the field spends enormous effort curating cold-start data, and this says that effort is mostly erased by enough stable RL. The lever that actually moves the ceiling is **stability**, not initialization. ## The take - **One approximation, one toolbox.** Token-level RL is the first-order truncation of the sequence objective. IS correction, clipping, and Routing Replay aren't three unrelated patches — they're three ways to keep the truncation valid by shrinking the training–inference discrepancy and policy staleness. - **The IS weight is structural, not cosmetic.** It's the linear term itself; removing it doesn't add variance, it changes what you're optimizing, and training collapses on contact. - **MoE needs Routing Replay, and it's a real trade.** Pinning experts restores the ratio but biases the target policy. Use R2 when you're near on-policy, R3 when you push off-policy — the paper's clearest practical recipe. - **Stability is the scaling lever.** Different cold-starts, on-policy vs off-policy — once stable, they land in the same place. The honest limits: one task, one model family, a stress-test precision setup, and headline benchmark values read from curves rather than a table. The formulation is the durable part; the exact numbers are a single, if very large, data point. --- *Built on [Stabilizing Reinforcement Learning with LLMs: Formulation and Practices](https://arxiv.org/abs/2512.01374) (Zheng et al., Qwen Team, Alibaba). Routing Replay's two flavors trace to [GSPO](https://arxiv.org/abs/2507.18071) (R2) and Ma et al., 2025 (R3, arXiv:2510.11370). Figures are the paper's; the interactive diagrams are mine.* --- # FLUX 3: when an image model decides to become a world model > Satyajit Ghana — Head of Engineering @ Inkers Technology > canonical: https://ai.thesatyajit.com/articles/flux-3 > date: 2026-07-24 > tags: diffusion, image-generation, video-generation, flow-matching, black-forest-labs, explainer Every FLUX before this one was an image model. FLUX.1 was a 12B rectified-flow text-to-image transformer; FLUX.2 scaled the same recipe to a ~32B image-and-editing model. [FLUX 3](https://bfl.ai/blog/flux-3), announced by [Black Forest Labs](https://bfl.ai) on 2026-07-23, is a different kind of object. It is a single **multimodal foundation model** that learns jointly from images, video, and audio in one architecture — and then generates all of them, plus robot *action*, from the same backbone. The framing in the title is deliberate: "Real World Models." BFL is no longer trying to make prettier pictures; it is trying to build a model of how the world looks, moves, and sounds.
The argument for one model is an information argument, and it is worth stating in BFL's own terms. No single modality is a complete description of reality — each is a lossy projection captured by a different sensor. Images fix spatial structure at one instant; video restores time and reveals physical dynamics; audio exposes causal links between mechanical events and the sounds they make; language ties all of it to goals and instructions. Learn from one projection and you get a good model of that projection. Learn from all of them at once and their **mutual constraints** teach you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate problems and start being evidence about one underlying reality. That is the pitch. Because FLUX 3 shipped as an Early Access *capabilities* announcement rather than a tech report — no parameter count, no full architecture, benchmarks that are BFL's own — the honest way to cover it is to explain the mechanism it is built on from first principles, show the evidence BFL did publish, and keep the caveats visible. So: two mechanisms first, then the numbers. ## Mechanism 1: flow matching, and why few steps is the whole game FLUX has always been a **flow-matching** model, and FLUX 3 builds on an approach BFL calls [Self-Flow](https://bfl.ai/research/self-flow). Flow matching is the cleaner cousin of diffusion, and the idea is simple enough to hold in your head. You want to turn a sample of pure noise into a sample of data. So you define a path between them — for training, literally the straight line $x_t = (1-t)\,x_{\text{noise}} + t\,x_{\text{data}}$ — and you train a network to predict the **velocity** $v_\theta(x, t)$ that points along that path. Generation is then just solving an ordinary differential equation: start at noise, and take steps in the direction the velocity field tells you, until $t = 1$. The subtlety is *how many steps*. Each step is a full forward pass of a huge transformer, so steps are the cost. And here the geometry of the path matters enormously. If the learned field is **straight** — a constant velocity, which is what "rectified" flow aims for — then a first-order Euler integrator lands on the data exactly, no matter how few steps you take. If the field is **curved**, few steps cut the corner and miss the target; the error only closes as you add steps (and cost). Drag the step count and flip the field to feel it: This is why "straighten the transport paths" has been the central obsession of the whole FLUX / rectified-flow lineage: a straighter field means a few-step sampler that still lands, which means a 20-second video is merely expensive instead of impossible. The trajectory geometry above is the same one the [Mage-Flow](/articles/mage-flow) piece leans on for its 4-step Turbo — few-step generation is a property of the *field*, not a trick bolted on afterward. ## Mechanism 2: one sequence, joint attention across modalities The second mechanism is how four modalities fit in one model. FLUX 3 sits in the **MMDiT** lineage — the multimodal diffusion transformer that the [Mage-Flow explainer](/articles/mage-flow) walks through in detail — and the load-bearing idea is that every modality is tokenized into a *single* sequence, and one attention operation runs over the whole thing. A query token is not confined to its own modality: an audio token can attend to the video frame that produced the sound; a video token can attend to the text that describes the scene. Pick the query's modality below and flip joint versus per-modality attention: Per-modality attention gives you four models in a trenchcoat. **Joint** attention is what lets the mutual constraints actually do their work — it is the mechanical form of "the modalities are evidence about one reality." BFL's own diagram makes the shape concrete: shared per-modality encoders feed a single multimodal transformer, shared decoders read it back out, and *action* is added as one more lane — explicitly marked extensible.
Self-Flow is BFL's method for aligning generation and understanding inside that one backbone, and the one quantitative claim they put behind it is that unifying the two lifts *both* at once. Their ablation reports lower generation error (Fréchet distance) per modality against a flow-matching baseline, and — more interestingly — faster and higher-climbing success on downstream robot-control tasks when the backbone is finetuned:
Read that right panel carefully, because it is the real bet: the same weights that generate video are a **dynamics-aware prior** for physical control. That is the bridge from "content tool" to "world model," and it is the part most worth being skeptical of until there is a paper. ## What it actually does: video, with sound The headline capability in Early Access is video. FLUX 3 generates up to **20 seconds in a single pass**, and — the part competitors mostly don't have — every clip comes with **native audio**, generated jointly rather than dubbed on afterward. The capability list is broad: text-to-video, image-to-video (animate a still or use images as visual references), video-to-video (carry a character from a reference clip into a new scene), keyframe-to-video for controlled transitions, multilingual dialogue, and *agentic chaining* of clips into multi-shot sequences minutes long with consistent characters. Here is a representative shot from BFL's own reel — a single continuous take of a galloping horse, the kind of coherent physical motion the "world model" framing is really about: