2026-10-06 · 22 min · 3d · slam · gaussian-splatting · point-cloud · vision-language-models · benchmarks · paper · explainer
The first roundup grouped six papers by problem. This week's batch has one shared shape. Each takes a feed-forward geometry model (VGGT, Depth Anything 3, LoGeR, TRELLIS.2, a VLM) and bolts on an idea that predates it: a coarse-to-fine pyramid, a fixed primitive budget, a pose graph, a gravity prior, bundle adjustment, a panorama. For each, the same four questions: what problem, what idea, what evidence, and what you can run.
| Item | The old idea | Headline, and what it is measured on | Released |
|---|---|---|---|
| T3lescope | Coarse-to-fine pyramid | Beats GenRecon on all seven ScanNet++ metrics; DA3 keeps the better F@2 | Paper; code "coming soon" |
| PocketSplat | Fixed primitive budget | 22.790 dB at 226K Gaussians on DL3DV; 10.20 to 10.95 s on an iPhone | Paper only |
| CLoSeR | Loop closure, pose graph | VBR mean ATE 36.16 → 11.15 m over LoGeR | Code, no licence file |
| G3T | Gravity-aligned frames | Camera-to-gravity error under GeoCalib's; ACC −37.9% on 10 TUM RGB-D sequences | Code and weights, no licence |
| EPO | Bundle adjustment, without tracks | AUC@5 77.1 against 70.9 for VGGT's own BA on ScanNet++ | Code, IVC licence |
| OneCanvas | Panoramic reprojection | SQA3D 65.3, VSI-Bench 71.3, SPBench 72.1 | Code and weights, MIT |
T3lescope: one generator, applied at every scale
Problem. Generative mesh models such as TRELLIS.2 fill in what the photos barely constrain (glass, gloss, the back of a chair), but they work on one fixed voxel grid. A room or a city block at that grid is mush; tiling it into overlapping cells, as GenRecon does, gives detail but lets distant cells disagree.
The idea. Run the same generator several times at halving cell sizes, each level starting from the one above. Preferred Networks fine-tune TRELLIS.2 (a sparse-structure flow model, then a mesh flow model, both on a grid per cell) on cells of many physical sizes, never telling it the size. At inference, Depth Anything 3 points place the coarsest cells and pick their views; the depth is used for nothing else. Each finer level takes the parent's decoded surface at twice the resolution, encodes it, noises it to and denoises it with finer image features: SDEdit between levels. Images are encoded once, as a pyramid of tiles; each cell reads the level matching its voxel size.

The evidence (reported). On the 25 largest ScanNet++ validation scenes the finest level beats GenRecon on all seven metrics at 32 and 256 views. It does not beat Depth Anything 3 everywhere:
| ScanNet++, 32 views | Depth MAE ↓ | Normal error ↓ | Chamfer (m) ↓ | F@2 cm ↑ |
|---|---|---|---|---|
| 2DGS (per-scene optimisation) | 0.1426 | 38.376° | 0.0723 | 0.4837 |
| Depth Anything 3 | 0.0811 | 28.000° | 0.0212 | 0.7765 |
| GenRecon | 0.1175 | 20.928° | 0.0604 | 0.2846 |
| T3lescope, level 2 | 0.0565 | 16.330° | 0.0212 | 0.7155 |
DA3 keeps the higher F-score at both view counts (0.7882 against 0.7176 at 256 views), and at 256 views also the lower Chamfer (0.0193 against 0.0202) and higher completeness. What T3lescope wins is depth and normals. On Tanks and Temples with 256 views, optimisation still wins F-score clearly (GaussianWrapping 0.650 against 0.354 on the small scenes); the paper says so. The convincing table is city blocks from street photos: Chamfer 0.645 m against 4.858 m for the best baseline. The cascade is load-bearing: starting finer levels from pure noise raises Chamfer from 0.0212 to 0.1051.
Two caveats. "Without per-scene optimisation" is not "fast": a large ScanNet++ scene takes 17 minutes to level 1 and 89 to level 2 on an H200, which the authors call comparable to per-scene optimisation. And glossy and transparent surfaces are shown, not measured: no metric isolates them (reasoned). The project page lists code as coming soon.

PocketSplat: an exact Gaussian budget for a phone
Problem. Feed-forward splatting (pixelSplat, MVSplat, DepthSplat) emits one or more Gaussians per input pixel, so the asset's size is set by the input resolution and view count rather than by what the phone can store and render. Their cost volumes also do not fit in an iPhone's memory.
The idea. Predict densely, keep exactly . A frozen Depth Anything 3 backbone (the iOS export is "six DA3 Core ML stage packages") gives depth and cameras for four views at ; a trained head turns each pixel into a 32-channel latent with a reliability and a support score: 451,584 candidates. Candidates are lifted to world space and binned into cells of the median scene depth wide. Each cell gets a quota from its image detail (the entropy of a patch), divided by the number of views that saw it so repeated evidence is not counted twice, filled by water-filling and rounded so the quotas sum to exactly . Survivors absorb their cell's pooled latent, and a scale correction widens each Gaussian by about , so a cell thinned from to still covers its area.

The evidence (reported). On DL3DV, 226K Gaussians reach 22.790 dB, 1.559 dB over pixelSplat at 1,376,256. Allocation beats a random subset of the same size by 2.19 to 3.69 dB, and the scale correction is worth 3.36 dB at 113K. On the phone (model unnamed), PocketSplat builds an asset in 10.20 to 10.95 s with a 2.64 to 2.89 GB peak; native MVSplat and DepthSplat run out of memory, and a streamed MVSplat takes 42.00 s.

What the budget buys. The post said PocketSplat "decodes only what fits the device budget". The appendix says otherwise: the on-device graph "retains N output rows at all budgets", gives unselected rows an opacity logit of −20, and the PLY writer drops them. The phone does the same work at every budget, hence the flat 10.2 to 11 s. On the GPU a smaller budget is slower (601.8 ms at 113K against 322.3 ms at full), peak memory is 2.705 GB at every budget, and deciding before decoding saves 1.73% of construction time. The budget sets the size of the asset; the memory win comes from not building a cost volume (reasoned). One oddity: the phone's Light row has SSIM 0.6690 and LPIPS 0.2952, identical to the DL3DV 113K row, and Balanced's LPIPS 0.2142 matches DL3DV 226K. Different datasets rarely agree to four decimals twice; with no code, I cannot check (reasoned).
CLoSeR: loop closure for a streaming backbone
Problem. Streaming feed-forward models (LoGeR, CUT3R, TTT3R) keep memory flat over thousands of frames but still drift: nothing tells them they have come back to a place. Submap systems such as VGGT-Long add loop closure, but each VGGT submap has its own scale, so their pose graphs need , or for VGGT-SLAM, and they align submaps by registering noisy point clouds.
The idea. Use a backbone whose scale is already consistent and ask it for the loop constraint directly. CLoSeR (ETH Zurich) wraps LoGeR without retraining. For each new window it computes SALAD place-recognition descriptors and keeps up to five earlier frames with cosine similarity above 0.7 and at least four windows in the past. It then builds a loop-conditioned window, half current frames and half matched old ones, and runs it through LoGeR, which does not need its inputs to be contiguous. The two halves come back in one frame, so their relative pose is the loop constraint. All poses are then optimised on , with no scale variable:
Sequential edges link each frame to its four neighbours; is a Huber loss that blunts a wrong loop; Levenberg-Marquardt solves it. This is the factor graph GTSAM solves for LiDAR SLAM, with a learned front end in place of scan matching.

The toy shows what one loop edge does. A 96 m loop of 48 odometry steps with a heading bias you set, dead-reckoned, then solved with one extra edge from the last pose to the first, by Gauss-Newton on :
In the toy (measured), a bias of 1° per step leaves dead reckoning 12.02 m from its start, with an RMSE of 6.80 m against the truth. One loop edge brings the RMSE to 0.25 m. The worst pose after optimisation, 0.40 m off, is pose 30, far from both ends, because the correction is spread along the loop. Turn on a 25% scale drift and the same graph still closes the gap, but the RMSE stays at 2.98 m with the worst pose 5.31 m off. With no heading bias at all, the optimised loop is worse than dead reckoning (3.04 m against 2.34 m): an graph can only bend the trajectory, not shrink it. That is the bet CLoSeR makes, and the paper measures LoGeR's cross-window scale error on KITTI at 5.8%, against 18.8% for VGGT-Long.
The evidence (reported), ATE RMSE in metres after a alignment:
| Benchmark | LoGeR | CLoSeR | Best other feed-forward |
|---|---|---|---|
| VBR, 7 loops, mean 2.31 km | 36.16 | 11.15 | 31.12 (LingBot-Map) |
| KITTI 00–10, mean 2.0 km | 25.44 | 12.79 | 18.65 (LoGeR*) |
| Oxford Spires, 14 sequences | 6.88 | 5.26 | 7.55 (VGGT-SLAM 2.0) |
| DROID-W, dynamic, 7 sequences | 1.44 | 0.75 | 0.91 (LingBot-Map) |
The ablation supports the scale argument: instead of gives 13.33 m on VBR, 20.36 m. It costs 7.12 against 7.60 frames a second on an RTX 4090.
What "drift-free" means here. 11.15 m on 2.31 km is less drift, not none.
Loop closure only helps where there is a loop: on KITTI's loop-free 01 and
08 CLoSeR scores 41.80 and 26.51 against LoGeR's 41.64 and 26.46. On DROID-W
the dynamic-scene SLAM it is named after still wins, 0.23 m. Two small
slips: the teaser gives 6.51 m on KITTI 00 where the table has 6.58, and "LoGeR*
second best" holds on KITTI but not on VBR. The repository (5230776) has evaluation scripts for
all four benchmarks and no licence file.
G3T: predict the pointmap upright
Problem. VGGT predicts every pointmap in the first camera's frame, including its roll and pitch. Two such submaps differ by a full 7-DoF similarity, and any rotation error tilts the floor.
The idea. Predict in a gravity-aligned frame instead. Cornell's G3T fine-tunes all of VGGT so the point head outputs points in the first image's gravity frame (its camera frame with roll and pitch removed, up). The camera head becomes two: a local head for gravity-to-camera rotation and field of view, and a relative head for yaw (one degree of freedom) and translation. Ground truth comes from five datasets; those not natively upright were aligned with COLMAP's Manhattan-world orientation aligner, so "gravity" in training is partly an estimate. Two upright pointmaps differ by scale, translation and a rotation about : 5 degrees of freedom. G3T-Long rebuilds VGGT-Long with a Procrustes that solves only that yaw, in the -plane, and a pose graph over that 5-parameter group.

The evidence (reported). On 7Scenes from one view, the first camera's gravity error falls from 6.78° (GeoCalib) to 1.92° (G3T's local head), and structure accuracy stays at VGGT's level. G3T-Long against VGGT-Long on ten TUM RGB-D sequences:

I recomputed the bars from the paper's Table 4 (measured): they are
medians over the ten sequences, ACC 0.0515 to 0.032 and COMP 0.044 to
0.0355. The mean ACC improves more, 0.0829 to 0.0386, carried by the long
pioneer_slam runs. Rotation error is worse with G3T-Long on two of ten
(fr1/360: 19.309° against 16.320°). Two limits: the long-sequence test is
indoor TUM RGB-D against VGGT-Long only, with none of the loop-closing
systems above, and "regardless of input image orientation" has failures the
paper shows, close-ups of floors and a cabinet shot sideways. The paper is
from May. Neither the repository nor the
Hugging Face weights declare a licence, and VGGT's released 1B checkpoint is
CC BY-NC 4.0.
EPO: bundle adjustment on edges instead of tracks
Problem. VGGT's poses are fast and rough. Its own fix, a tracker plus bundle adjustment, needs point tracks, takes minutes, and its refinement variant wants a GPU with at least 40 GB.
The idea. Align edges instead of matching points. For each image, EPO (Graz) takes a Canny edge map and its distance transform, in which each pixel holds the distance to the nearest edge. Edge pixels of image are lifted with the model's depth and projected into image , and the distance transform there says how far each one landed from an edge:
with a Huber loss, summed over image pairs that reproject consistently. The loss is differentiable everywhere, so AdamW does the work: first poses (through a small MLP, plus a per-camera translation offset) and focal length, then a per-pixel affine correction of depth, stopping when the 95th-percentile pose change goes quiet. No detector, matcher or track.


The evidence (reported), AUC at 5°, with total wall-clock time:
| VGGT | + VGGT's BA | + track refinement + BA | + EPO | |
|---|---|---|---|---|
| ScanNet++ (20 scenes) | 55.6 | 70.0, 171.6 s | 70.9, 303.5 s | 77.1, 52.1 s |
| TerraSky3D (9 scenes) | 56.8 | 71.1 | 75.5 | 79.2 |
| Mip-NeRF 360 (7 scenes) | 72.2 | 85.5 | 87.8 | 90.5 |
It also lifts MapAnything (ScanNet++ 37.3 to 60.4) and (70.0 to 80.3). Four caveats. The baselines are VGGT's own tracker-based BA, not COLMAP or GLOMAP. The refinement baseline was timed on an H200 and EPO on an RTX 4090. The reprojection-error table compares EPO's edge-to-edge distance with BA's point-to-keypoint error, two different quantities. And EPO loses outright on one scene, Munich Marienplatz (68.1 against 78.6). Splats trained from EPO's poses reach 23.93 dB on Mip-NeRF 360, against 26.69 from COLMAP poses. The repository has since moved past the paper: v1.4, on VGGT-Omega with radial distortion, reports a mean AUC@5 of 85.7 across four datasets, README numbers not in the paper. The licence is IVC's own: commercial use is allowed "after information to IVC".
OneCanvas: the whole room as one panorama
Problem. A VLM asked about a room sees 32 frames as 32 separate images. Most 3D VLMs add a geometry encoder or scale up spatial QA data.
The idea. Put every patch where it is. Qwen3-VL's frozen vision encoder embeds each frame; each patch token is lifted to 3D with its depth and pose, then placed at its continuous longitude and latitude as seen from a chosen origin, the agent's pose for situated questions. Those two angles go into the model's existing rotary position axes for width and height, and the frame index into the temporal axis. Overlapping patches stay separate tokens. A 136-channel sinusoidal code of each patch's metric offset, through a small MLP, puts back the depth the angles lost. Training is two LoRA stages: synthetic spatial tasks built by placing real patches on an empty canvas (rank 256), then QA (rank 64).

The evidence (reported): SQA3D 65.3 exact match, 2.3 over Ross3D; VSI-Bench 71.3, with route planning 12.3 points clear at 60.8; SPBench 72.1 zero-shot, 4.8 over SpaceMind; "an order of magnitude less training compute".
Three corrections to the post. The VLM is not frozen; only the vision encoder is, and its own figure marks the language model as fine-tuned. The canvas reads ground-truth sensor depth and poses; with Depth Anything 3 estimates instead the scores are 64.9, 70.0 and 71.3. And the panorama alone does not carry the result. In the VSI-Bench ablation, at matched LoRA capacity, the base VLM on raw frames scores 65.5 and "panorama only" 65.3; the 3D position code takes it to 68.5 and the curriculum to 71.3. The gain is the metric embedding and the synthetic pretraining, laid out on a panorama (reasoned).
SPBench's headline is the mean of its single-image and multi-view averages (81.5 and 62.8 give 72.1). Every baseline's overall fits that rule except SpaceMind's: its 73.8 and 59.7 average 66.75, not the 67.3 listed (measured). VSI-Bench is not zero-shot: stage 2 trains on VSI-Bench-style data. The released checkpoint scores 65.53, 71.16 and 74.41 by the README's protocols, and SQA3D drops from 65.3 to 61.2 if the canvas sits at the scene centre instead of the agent. Code and weights are MIT.
Briefly
Lyra 2.0 (NVIDIA) turns one image into a walkable world: a long camera-controlled video, with per-frame geometry used only to retrieve past frames, trained on its own degraded outputs so drift gets corrected, then lifted to 3D by a fine-tuned feed-forward model. It is not new: weights came out in April, the GUI and training code in July. "100% open source" holds for the Apache-2.0 code. The weights are under NVIDIA's Internal Scientific Research and Development Model License, which forbids production use and redistribution. WorldCrafter benchmarks against it.
EditHero is a benchmark of 457 part-level edit chains (2,755 edits, 252 objects, up to 30 turns) with an exact target after every turn, assembled from a part library. Its finding: whole-object metrics reward doing nothing (a no-op beats every learned editor on whole-object F-score, LPIPS and PSNR); learned editors follow 0.13 to 0.36 of instructions on the 55-chain comparison set; code-writing LLM agents keep the rest of the mesh intact and reach 0.62 (Opus 5.5), but take 1.5 to 6 minutes an edit against 20 to 50 s. The engine and data are promised, not released.
Texture Space Material Diffusion (NVIDIA) fine-tunes Wan 2.1-1.3B to
denoise PBR maps (base colour, height, roughness, metalness) directly in a
mesh's UV space, so there are no views to disagree. Photos are projected
into texture space as partial observations; 8K comes from generating at 2K
and refining shifted crops. It needs a mesh with non-overlapping UVs and
camera poses. The source link on the project page is commented out, and
git clone of NVlabs/texdiffusion returns "Repository not found"
(measured).
image-blaster reconstructs nothing. It is MIT-licensed Claude Code skills and scripts that chain hosted models. Claude reads the photo and lists the movable objects; an image editor (Nano Banana or GPT Image 2) erases them to a clean plate; World Labs' Marble 1.1 turns the plate into a Gaussian splat world with a collision mesh; each object is re-imaged and meshed by Hunyuan 3D on FAL (50,000 faces by default); ElevenLabs effects on FAL add sound; a React viewer plays it. Every 3D asset is generated, from one view, by a paid API.
What you can run today
| Item | Code | Weights | Licence |
|---|---|---|---|
| T3lescope | "Coming soon" | None | Paper CC BY 4.0 |
| PocketSplat | None found | None | arXiv licence only |
| CLoSeR | MoyangLi00/CLoSeR at 5230776 | LoGeR's and SALAD's | No licence file |
| G3T | g3t-paper/g3t at 193ce19 | thatbrguy/g3t | None declared; fine-tune of VGGT |
| EPO | mattiadurso/EPO at 13bee33 | Uses VGGT and others | IVC, commercial use after informing IVC |
| OneCanvas | baranowskibrt/onecanvas at 5c9c587 | BaranowskiBrt/OneCanvas-Qwen3-VL-8B | MIT |
A repository with no licence file is not open source in any legal sense; ask before building on CLoSeR or G3T. For where a feed-forward model's metres come from on a robot, see Depth Anything 3 in ROS 2.
What would change my mind
2 claims above, and what would falsify each
PocketSplat's output budget does not reduce on-device compute.
Read from the paper's Appendix D (all N rows kept, unselected opacity −20) and Tables 2 and 6, which show flat phone times and a 2.705 GB peak at every budget. A deployment that skips unselected rows and runs faster at smaller budgets falsifies it.
OneCanvas's panorama alone adds nothing on VSI-Bench; the 3D position code and curriculum do.
Read from the paper's Table 4: base VLM 65.5, panorama only 65.3. If those rows were not trained at matched settings, as the caption says they were, the comparison does not hold.
Sources: the six arXiv papers at the versions linked, the EditHero paper (arXiv 2610.02298), the Texture Space Material Diffusion project page and arXiv 2609.37654, and shallow clones of the repositories at the commits named. Figures are the papers' and READMEs' own, used for commentary.