2026-10-02 · 19 min · 3d · slam · robotics · point-cloud · state-estimation · benchmarks · paper · explainer
A 1:44 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.
› transcript
Hi, I'm Calloway! Two systems rebuild the world from a moving camera, one frame at a time. Both freeze the same depth model, and stop storing a dense point map per frame. They keep a small bounded summary instead. Each video frame enters a frozen geometry model that already knows depth and shape. It turns the frame into one small token, a compact summary of the view. A tiny head compares tokens in pairs and predicts camera motion, scoring how much to trust each one. The trusted links fold into one running estimate, keeping only a fixed set of key frames. Memory stops growing. The old way pins everything to frame one. Travel far, and the numbers grow without bound, so the track drifts. The first system predicts only motion between pairs, weighted by confidence, so any frame can anchor the next. The second system tracks each frame fast against a tiny local memory. Chunks of the map are stitched by near links, longer links, and loop links when a place returns. It all becomes one pose graph, with no slow bundle solve, so moving cars no longer break it. On streaming benchmarks the first system has the lowest trajectory error, and it is the smallest model. Add LiDAR on a brutal highway drive, and the error falls from fifty six metres to under three. The geometry model is now a frozen part. The real work is holding memory flat and skipping the slow solve. Freeze the depth model, keep a bounded summary, trust each link, and stream geometry from endless video. Every source is in the full article. I'm Calloway. Bye!
3D perception is my day job, so two papers from this year landed on the same nerve. R³ (Westlake, Michigan, NVIDIA) and AMB3R-SLAM (UCL) look unrelated — one is a feed-forward reconstructor, the other a kilometre-scale SLAM system — but they make the same two bets. They take Depth Anything 3 as a frozen (or almost-frozen) geometry backbone rather than training a geometry network from scratch, and they both replace the per-frame pointmap — the thing that makes a naive feed-forward reconstructor's memory grow without bound — with a streaming estimate whose memory is flat in the number of frames. That second bet is the whole game for video. This piece explains both mechanisms from first principles, then checks the headline numbers against the papers and the released code.
What "streaming, feed-forward 3D from unposed video" means
Classical structure-from-motion and visual SLAM start from matched features and solve a geometric optimisation: bundle adjustment jointly refines camera poses and 3D points so that every point reprojects into every view that saw it. It is accurate and it is slow, and it assumes the world holds still while you look at it. The feed-forward line — DUSt3R, VGGT, CUT3R, and now DA3 — throws that out. A transformer looks at the images and regresses geometry in one pass: for each pixel a 3D point (a "pointmap"), and for each frame a camera pose. No correspondences, no iterative solve. "Unposed" means you hand it raw video with no camera poses and no calibration; the model produces the poses as output.
That works beautifully for a handful of frames and breaks for a video. Two reasons, and both matter for anything that runs on a robot:
- The global-frame assumption. Most feed-forward models express every point and pose in one global coordinate frame, anchored to the first camera. Over a long sequence the camera's translation relative to that fixed origin grows without bound, and the network has to represent ever-larger numbers it never saw at training time. Accuracy drifts.
- Memory grows with the video. If the model attends over all frames, or keeps a pointmap per frame, its state is in the frame count . A 30-second clip at 30 fps is 900 frames; a drive is tens of thousands. At some point you run out of GPU, and the run dies mid-sequence.
"Streaming" (or causal) means the model consumes frames one at a time and commits to an estimate for frame using only frames up to , the way a robot has to. "Bounded memory" means the state it carries forward does not grow with : it keeps a fixed working set. That is the property that turns a nice demo on 20 frames into something you can leave running. Both papers are, underneath, answers to the same question — what is the smallest thing you can carry between frames and still stay globally consistent?
What DA3 gives them
Depth Anything 3 (ByteDance Seed, November 2025) is the shared foundation. It is a geometry foundation model: one plain DINOv2 transformer with interleaved within-view and cross-view self-attention, and a dual-DPT head that emits a depth map and a ray map, from which 3D points fall out. Given one or many images it returns depth, points and camera geometry in a single forward pass. It ships in sizes from Small (80M parameters) to Giant, and the Large checkpoint is 385M (reported, R³'s Table 2). TencentARC's GAE builds a generative model on the same backbone; the ROS 2 wrapper I took apart earlier runs its metric-depth checkpoint on a robot. Here it is used differently again: as a frozen feature extractor whose per-frame tokens are the raw material for a pose estimate, with the expensive geometry knowledge inherited rather than relearned.
R³: regress relative poses, weight them by confidence
R³ ("3D Reconstruction via Relative Regression", Congrong Xu et al., arXiv 2605.26519v2) attacks the global-frame problem head-on. Its thesis: do not regress an absolute pose in a global frame at all. Regress relative poses between pairs of frames, and let a confidence score say which pairs to trust.

The picture is the argument. VGGT (panel a) pins the world frame to camera one and supervises only the edges out of that anchor — so the anchor's errors poison everything, and far-from-origin translations blow up. π³ (panel b) regresses absolute poses in a model-chosen frame and treats every pairwise constraint as equally reliable. R³ (panel c) keeps no absolute pose: it predicts a directed relative pose for every ordered pair, each tagged with how much to believe it. The world frame becomes a choice made at aggregation time, not a bias baked into the architecture.
The mechanism

Concretely: the DA3 backbone, made causal and kept mostly frozen ("we keep most of the 3D backbone frozen and use a relative-pose head as a lightweight front-end", reported), produces one camera token per frame. A lightweight MLP reads a pair of tokens and predicts a relative rotation, a relative translation, and — the load-bearing bit — separate confidence scalars for rotation and for translation. Rotation and translation fail in different ways (pure rotation is observable with no translation; translation scale is not), so one confidence number cannot describe both.
That confidence does two jobs. In training it weights the loss: a confident pair that is wrong is penalised hard, an uncertain pair is allowed slack, with a regulariser so the model cannot just declare everything uncertain. At inference it steers aggregation: to place a new frame, the system picks the most confident already-registered frames as references and fuses their relative predictions, top- confidence-weighted. The whole collection of directed, confidence-weighted edges is a pose graph, and the trajectory is read off it. No bundle adjustment; the confidences do the arbitration an optimiser would otherwise do.
The keyframe bank is what bounds the memory
Streaming needs a bounded context. R³ keeps an active set : the first frame plus a keyframe bank
of capacity . A new frame is admitted only if it is novel
— if its encoder token is far enough from every token already in the bank
(max cos(tok_i, tok_j) < τ). When the bank is full, the least useful keyframe
is evicted, ranked by a utility that trades off how
distinctive a keyframe is against how confidently it connects to the rest. The
bank never exceeds , so the state carried between frames is , not
. This is the mechanism behind the flat memory curve.
- naive store
- 59.8 GiB OOM
- bounded bank
- 12.4 GiB ok
- bank fill
- 62 / 64 slots
Drag frames: the naive line climbs with N and OOMs; the bank is flat. Drag K: the bank moves up or down but never tilts. Schematic — only the 48 GiB line is R³'s; measured numbers are in the tables.
The widget is that contrast made draggable, and it is a schematic — only the 48 GiB line is the paper's. Drag the frame count: the naive per-frame store climbs a straight line and, past the point where it crosses the budget, the run is dead. The bounded bank is flat; it never depends on how long the video is. Drag the bank cap and the flat line moves up or down — memory is set by the cap, never by the sequence — which is the entire point.
Does the streaming ATE hold?
Yes. On camera pose estimation in the online (streaming) setting, absolute trajectory error (reported, R³ Table 2, lower is better):
| Streaming method | #Params | Sintel ATE | TUM-dyn ATE | ScanNet ATE |
|---|---|---|---|---|
| CUT3R | 793M | 0.213 | 0.046 | 0.099 |
| StreamVGGT | 1.26B | 0.251 | 0.061 | 0.161 |
| STream3R | 1.26B | 0.213 | 0.026 | 0.052 |
| TTT3R | 793M | 0.201 | 0.028 | 0.064 |
| R³ | 372M | 0.115 | 0.018 | 0.038 |
R³ has the lowest ATE on all three while being the smallest model on the board — 372M against the 1.26B of StreamVGGT and STream3R, about 30% their size (reasoned: 372 / 1260 ≈ 0.30, which is the paper's "≈⅓ of recent 1B-class models"). On Sintel its ATE is 0.115 against STream3R's 0.213, a 46% reduction (reasoned). The offline, full-context variant is also competitive: 0.130 ATE on Sintel against VGGT's 0.172 and DA3-Large's 0.140 (reported, Table 2 top).
And the memory claim — the reason any of this matters for long video:

On 7-Scenes, run to 1000 frames under a 48 GiB budget, R³'s reconstruction accuracy is flat — Acc 0.021 / 0.022 / 0.022 at 200 / 500 / 1000 frames, Comp 0.018 / 0.017 / 0.017 (reported, Table 4). CUT3R drifts over the same span (Acc 0.087 → 0.194 → 0.240) and StreamVGGT, a 1.26B per-frame model, runs at 200 frames but exceeds the 48 GiB budget by 500 (it reports no result at 500 or 1000). That one real data point — a 1.26B model out of budget by 500 frames — is what the widget above is calibrated to. R³ was trained on six 48 GB GPUs (reported).
The frame rate is the one number that will not sit still
R³'s throughput claim is a cautionary tale about reading a single headline number. Across the project's four surfaces I found three different frame rates (measured, by reading each source):
- The arXiv Figure 1 caption and the GitHub README: "20+ FPS".
- The arXiv Figure 1 image panel (same figure, above): "30+ FPS".
- The project page: "40 FPS", with a footnote, "Measured on a single NVIDIA RTX PRO 6000".
- The launch post: "Runs at 30+ FPS".
The paper disagrees with itself inside one figure — caption 20+, panel 30+ — and only the project page's 40 names hardware. None appears in a table with a resolution and a batch size. The numbers are all plausible for a 372M model and the hardware spread (an RTX PRO 6000 is far faster than whatever the 20+ was measured on), but "20+ / 30+ / 40 FPS" is a throughput range without a fixed denominator. The claim that is pinned, in a table and a figure, is the one that matters: memory stays flat. Trust that one.
AMB3R-SLAM: a hierarchical Sim(3) backend, no bundle adjustment
AMB3R-SLAM ("Kilometer-scale SLAM with Hierarchical Backend", Hengyi Wang & Lourdes Agapito, arXiv 2609.19518v1) takes the other road. It is training-free: it does not fine-tune anything, it orchestrates frozen foundation models inside a classical-feeling SLAM frame. The design principle, in the authors' words, is "to rely on the feed-forward predictions of geometric foundation models and avoid optimization that relies on post-hoc estimation".

The front-end is where DA3 earns its keep cheaply. To hold a high frame rate it uses DA3-Small (80M parameters). For each new frame it builds a compact memory — the submap's anchor keyframe plus the two most recent frames — and estimates the new pose against that, re-anchoring when it starts a fresh submap. The heavy lifting (dense mapping for the backend's constraints) uses a giant DA3 model; the authors also swap in VGGT-Ω to show the pipeline is model-agnostic ("Ours (Ω)"). The front-end is light and fast; the backend is where global consistency is bought.
The backend is a single Sim(3) pose graph — Sim(3) because a monocular system has no metric scale, so each node carries a 7-DoF similarity transform (rotation, translation, and a scale) rather than a rigid SE(3) pose. Three kinds of edge, three scales of consistency:
- Span-2 edges between overlapping neighbouring submaps, from dense local submapping — local consistency.
- Sparse long-context edges reaching further back — mid-level consistency, the part that keeps a long corridor straight.
- Loop-closure edges — global consistency. DBoW2 proposes candidate loops between a historical submap and the current one; each candidate is verified geometrically by sampling views from both submaps, jointly reconstructing them with the foundation model, and checking their 3D voxel-occupancy overlap. Loops whose implied rotation and translation correction is implausible are filtered out, which is what stops a perceptual aliasing (two identical-looking corridors) from folding the map.
The thing it deliberately does not do is bundle adjustment. BA ties points to poses under a static-world assumption; drop it, and moving cars and pedestrians stop corrupting the solve, so "dynamic scenes work out of the box". The backend's memory is the pose graph plus a bounded window of submaps, so — like R³ — it is flat in sequence length: peak 10.3 to 14.1 GB across its runs, holding steady out to 10k frames (reported, Table 9).
The numbers, checked
Throughput and memory on one RTX 4090 (reported, Table 9):
| Input | Dataset | FPS | Peak mem |
|---|---|---|---|
| Monocular | KITTI | 17.6 | 10.3 GB |
| Monocular | VBR | 10.2 | 14.0 GB |
| LiDAR | KITTI | 47.8 | 10.5 GB |
| LiDAR | VBR | 31.3 | 14.1 GB |
So 10.2-17.6 FPS monocular, 31.3-47.8 FPS with LiDAR — LiDAR is faster because it hands the system metric geometry directly and shortens the dense-mapping work. Memory sits at 10-14 GB regardless of which of the kilometre-scale sequences is running.
On accuracy, the headline is the VBR and Oxford Spires result: "reducing the ATE of previous state-of-the-art methods on VBR and Oxford Spires by over 70%" (reported). On VBR, monocular, AMB3R-SLAM averages 7.42 m ATE against the best prior online method's 26.65 m (LingBot-Map); that is a 72% reduction (reasoned: (26.65 − 7.42) / 26.65 = 0.72), matching the claim. With LiDAR the VBR average drops to 0.36 m — sub-metre on kilometre-scale drives.

The KITTI table is the most striking single row. KITTI sequence 01 — the highway stretch that breaks most visual systems, because the scene is featureless and fast — goes from 55.91 m ATE monocular to 2.95 m with LiDAR (reported, Table 2), a 95% cut (reasoned). The monocular 55.91 m is not a sign of weakness so much as a sign of the sequence: KITTI-01 breaks visual SLAM, and the weakest online baselines are far behind — DROID-SLAM scores 344.60 m and DA3-SLAM 263.25 m, five to six times AMB3R-SLAM's error (reasoned, from the table). It is not the single-sequence best, though — LoGeR's 47.91 m edges it on 01 (reported, Table 2) — the win is the average: over KITTI, AMB3R-SLAM is 13.11 m monocular and 0.95 m with LiDAR, beating even the LiDAR-native PIN-SLAM's 1.22 m (reported).
The shared trend
Lay the two side by side and the pattern is clear.
| R³ | AMB3R-SLAM | |
|---|---|---|
| Problem | Feed-forward streaming reconstruction | Kilometre-scale SLAM |
| DA3 role | Backbone, kept mostly frozen | Backbone, fully frozen (training-free) |
| DA3 size | 372M (DA3-derived) | Small (80M) front-end + giant backend |
| Bounds memory via | Novelty-gated keyframe bank, cap K | Windowed submaps + Sim(3) pose graph |
| Global consistency | Confidence-weighted relative edges | Span-2 + long-context + loop-closure edges |
| Classical solver | None (confidences arbitrate) | None (no bundle adjustment) |
| Released | Code (Apache-2.0 core) + weights (CC-BY-NC) | Placeholder repo only |
Three things are converging. First, the geometry network is now a frozen asset. Neither system trains geometry from scratch; DA3 is the ImageNet moment for 3D, and the research has moved up a level, to what you do with a foundation model's per-frame predictions. Second, the per-frame pointmap is the enemy of long video, and both answers are the same shape: keep a bounded working set (a keyframe bank, a submap window) and a cheap global structure (a pose graph of relative constraints) instead of a growing pile of dense predictions. Third, the iterative geometric solver is being retired. R³ replaces bundle adjustment with learned per-edge confidence; AMB3R-SLAM replaces it with a pose graph over foundation-model predictions and drops the static-world assumption with it. Thirty years of SLAM wisdom said the solver was the hard, essential part. These two say: if the per-frame predictions are good enough and you weight them honestly, you can bound the memory, skip the solve, and still close a kilometre-scale loop.
What you can run
| Code | Weights | Licence | |
|---|---|---|---|
| R³ | KevinXu02/R3 at e345f11 | KevinXu02/R3 · r3.safetensors 1.49 GB | Core Apache-2.0; training/ non-commercial; weights CC-BY-NC-4.0 |
| AMB3R-SLAM | Placeholder only | None | None declared |
R³'s licensing has a trap worth naming. The inference code is Apache-2.0, but
the training pipeline under R3/training/ is marked "NON-COMMERCIAL RESEARCH
USE ONLY" (measured, from R3/training/NOTICE), because it adapts DUSt3R and
CroCo (NAVER, CC-BY-NC-SA 4.0), CUT3R and VGGT. The released weights on Hugging
Face are CC-BY-NC-4.0 and declare Depth-Anything/Depth-Anything-3 as their
base (measured, from the model card; the r3.safetensors header reports
1,490,568,492 bytes). So you can read and run R³ freely; you cannot ship a
product on its released checkpoint, and you cannot retrain it commercially on
the given pipeline. AMB3R-SLAM you cannot run at all yet.
Both are built on the same DA3 backbone I traced through a ROS 2 node earlier, and both sit in the same lineage as the classical LiDAR-inertial odometry and stereo factor-graph SLAM I have written up — the difference is that the geometry now comes pre-trained, and the memory is the thing you engineer. For companion write-ups in this batch, see SurfLO and FAR-LIO.
What would change my mind
3 claims above, and what would falsify each
R³ quotes three different frame rates across its own four surfaces, none in a table with a fixed resolution and batch size.
Read from the arXiv Figure 1 caption and panel (20+ vs 30+), the GitHub README (20+), the project page (40, on an RTX PRO 6000) and the launch post (30+). A single table pinning FPS to a resolution, batch size and GPU — consistent with one of these numbers — would resolve it.
R³'s streaming memory is flat in frame count while a 1.26B per-frame baseline OOMs under the same 48 GiB budget past 500 frames.
Read from Table 4 and Figure 1: R³ holds Acc ≈ 0.022 to 1000 frames; StreamVGGT (1.26B) exceeds the 48 GiB budget past 500. A run of R³ whose memory grows with N, or of StreamVGGT staying within budget to 1000 frames, falsifies it.
AMB3R-SLAM's code is not released; the headline numbers cannot be reproduced yet.
Measured by cloning
HengyiWang/amb3r-slam: a 12-byte README and a.gitignore, no source, no licence. A commit that adds the runnable system falsifies it (and would let the KITTI-01 55.91→2.95 and VBR 7.42→0.36 numbers be checked).
Sources: arXiv 2605.26519v2 (R³) and 2609.19518v1 (AMB3R-SLAM) at the versions
linked; the R³ repository at e345f11 and its Hugging Face model card; the
AMB3R-SLAM placeholder repository. Figures are the papers' own, used for
commentary.