2026-10-02 · 16 min · 3d · point-cloud · flow-matching · gaussian-splatting · benchmarks · licensing · explainer
A 1:40 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.
› transcript
Hi, I'm Tobiah. Let me show you Surflo. It turns a handful of ordinary photos into one clean 3D surface. Geometry looks the same from every angle, so Surflo squeezes any number of views into one fixed latent and flows the surface points out of it. A frozen VGGT reads your unposed photos into patch tokens. A Perceiver pours them into 128 latent queries. That is the whole scene, as one fixed-size global state, whatever the view count. A flow decoder transports each query point, on its own, from noise onto the surface. Here is the formula. Each point starts as noise. It should end on the surface. The decoder only predicts the velocity that walks one to the other. The latent is always 128 tokens wide, 65,536 numbers, whether two photos go in or eighty. Independent points can disagree, so guidance makes neighbours talk near the end of the flow. First, guess where each point is heading. Render them as little Gaussians, compare to the real photos, and nudge them to match. Turn that correction into a velocity, and take the step. DL3DV surface F1. A scene is one fixed-size state; the surface is as many points as you ask for. Compress the views into one state, flow the points out, and let guidance make them agree. Every source is in the full article. I'm Tobiah. Bye!
Geometry does not care where you stood to photograph it. A chair is the same chair from the doorway or the window, so thirty photos of a room are thirty redundant encodings of one fixed 3D state. Feed-forward reconstruction models mostly ignore this. The per-view family — VGGT, Depth Anything 3 — emits one pointmap per input image, so the representation grows linearly with the number of views and the outputs overlap, duplicated and slightly misaligned across frames. The global-latent family fixes the representation size but then commits to a single low-resolution output grid. You get one or the other: a representation that scales with input, or an output that is capped.
Surflo, from Antoine Guédon (École polytechnique, and the author of Gaussian Wrapping) with collaborators at UC Berkeley, Kyoto University and Kyutai, takes the invariance literally. Compress any number of unposed RGB views into one fixed-size latent — a single global state — and decode the surface out of it by transporting points from noise with a flow-matching ODE. Because each point is decoded on its own, the output is bound to no grid and no token budget: the same latent yields a few thousand points or a million, in one forward pass. The paper is arXiv 2606.13644; the project page lists it as a NeurIPS 2026 Oral, which I cannot verify independently, so I flag it as the authors' claim.

Three things that usually come tied together
The idea is easiest to state as a decoupling. In most reconstruction stacks, three quantities are chained to each other: the number of input views , the size of the intermediate representation, and the number of output primitives. Per-view methods tie all three to . Fixed-grid latent methods tie the representation and the output together and cap both.
Surflo cuts every link. The encoder maps views, for any , into a latent with tokens of width — that is 65,536 numbers, no matter whether two views went in or eighty. The decoder reads that latent and emits however many points you ask for, because it decodes each point independently. Input count, latent size, output resolution: three dials that no longer move together.
That independence is the whole trick, and later the whole problem. It is worth being precise about what "decode a point" means here before we get to why the points occasionally disagree with each other.
The encoder: a frozen VGGT, read four ways, squeezed by a Perceiver
The encoder does not learn to see. It wraps a frozen VGGT-1B backbone —
loaded straight from facebook/VGGT-1B in the released code
(surflo/model/ffm.py), weights never updated — and reads its patch tokens from
four intermediate layers, . I read those indices
out of the shipped config (configs/model/surflo.yaml,
intermediate_layer_idx: [4, 11, 17, 23]), not just the paper, so that one is
measured. Taking four depths rather than only the last gives the compressor both
coarse layout and fine texture to draw from.
Each patch token is then tagged with where it sits in 3D. VGGT already predicts a rough pointmap, so every patch has an approximate world coordinate; Surflo encodes that coordinate with Gaussian Fourier features — frequencies over 16 log-spaced bands — and adds it to the token. This is what lets a plain set-to-set compressor reason about geometry: the tokens arrive pre-stamped with position.
The compressor itself is a Perceiver. It starts from learnable
latent queries and, four times over (L_e = 4), cross-attends them into the
thousands of position-stamped VGGT tokens and then runs four self-attention
blocks (L_s = 4) among the latents to mix what they gathered. A separate
single token absorbs camera information, and the final latent is the
concatenation. The output is always the same shape — 128 tokens — whether the
input was a sparse pair or a dense sweep. Many views simply give the
cross-attention more evidence to pool into the same 128 slots.

The decoder: flow matching, one point at a time
Here is where the interesting generative machinery lives. A query is not a pixel and not a voxel. It is a point in : a 3D coordinate plus a unit normal — an oriented point, the thing you need to mesh a surface rather than just dot it. The decoder's job is to take a noisy query and push it onto the real surface.
It does this with flow matching, the same training recipe behind modern diffusion-style generators (the flow-in-a-geometry-latent idea also drives TencentARC's GAE, and the per-point transport here is close in spirit to the MAR line of work). The mechanism is a straight line with a learned speed. At training time you sample a target point on the known surface, a source point from a noise distribution, a time , and form the linear interpolant . The network is trained to predict the constant velocity that walks the source to the target:
At inference you do not know ; you only have noise and the latent. So you sample query points from the source, and integrate the learned ODE from to with a plain Euler solver, 150 steps. Each point rides its own velocity field onto the surface. Decode 100K points by default, or 2,000, or a million — independent, batched forward passes, the latent computed once and reused for all of them.
Two design choices in the decoder matter more than they look.
There is no self-attention over query points. The 12-layer decoder cross-attends each point into the latent (for the first 6 layers) and otherwise processes it alone; all the spatial reasoning was spent in the encoder, baked into . That is exactly why can be anything — the points never have to attend to each other, so adding more costs nothing but more parallel forward passes. Conditioning on time and camera enters through DiT-style Ada-LN blocks, the same modulation pattern I walked through in the Diffusion Transformer explainer.
The source distribution is not pure noise. Starting every point from wastes the model on empty space. Surflo instead seeds the points from a mixture of Gaussians centred on perturbed VGGT pointmap samples — the rough geometry VGGT already gave us — with enough spread to cover surfaces that were occluded in the input views. The flow then only has to clean up and complete, not hallucinate from scratch. The ablation earns it: swapping the mixture for a pure Gaussian source drops DL3DV from CD 0.0073 / F1 81.21 to 0.0088 / 73.57 (Table 5, reported).
Guidance: the one place the points talk to each other
Independent decoding has a price the paper is honest about: nothing forces two independently sampled points to land on the same surface. In regions occluded from every input view, the flow is ambiguous, and you get outliers — points floating just off the true surface because, locally, the velocity field had no reason to prefer one plausible answer over another.
The fix is an inference-time guidance term that, for the last sliver of the ODE (), makes neighbouring points agree. At each of those final steps it first reads off where each point thinks it is going, . It treats that whole cloud of predicted endpoints as a set of small oriented 3D Gaussians, renders them through the cameras VGGT recovered, back into the input viewpoints, and measures how wrong the renders are against the real photos:
It runs gradient-descent steps on this loss, nudging the predicted endpoints toward a cloud that actually reprojects to the images, producing a guided target and a guided velocity for the real Euler step. The important word is global: the rendering loss depends on the entire batch of points at once, so a gradient computed through it couples neighbours. Two queries that disagree about a surface both feel a shared corrective pull. It is the only channel through which the otherwise-independent points communicate, and it opens only in the last roughly 5% of the trajectory.
That timing is worth doing the arithmetic on. With 150 Euler steps, is about the last 8 steps, each running 32 inner updates — on the order of 256 differentiable renders per reconstruction (my arithmetic on the paper's numbers, so: reasoned). Optionally a monocular-depth expert — Depth Anything 3 — adds a scale-invariant depth-order term to sharpen geometry further. Camera poses are refined in the same loop.
Checking the headline table
The marquee numbers are surface metrics: Chamfer Distance (CD, lower is better) between predicted and ground-truth clouds, and F1 at 1% of the scene diagonal (higher is better). On DL3DV, the in-distribution held-out split, from 16 unposed views (Table 1, reported):
| Method | CD ↓ | F1 ↑ |
|---|---|---|
| VGGT pointmap | 0.0100 | 74.84 |
| Gaussian Wrapping | 0.0168 | 60.67 |
| NOVA3R (latent, fixed output) | 0.0459 | 30.51 |
| Surflo — no guidance | 0.0072 | 81.92 |
| Surflo — with guidance | 0.0083 | 78.55 |
Taken at face value, Surflo-no-guidance is the clear winner: about a third lower Chamfer than VGGT pointmaps and seven F1 points better, and it extends out of distribution — Tanks & Temples 0.0053 / 88.57, Mip-NeRF 360 0.0068 / 82.00, DeepBlending 0.0116 / 70.96 (all reported). NOVA3R, the one prior latent model that also flow-matches points, is pinned to 10K points and two input views, and it shows.
But stare at the last two rows. Guidance makes the DL3DV number worse — 0.0072 → 0.0083 CD, 81.92 → 78.55 F1 — which is backwards from the whole point of adding it, and worth more than a shrug.
Why the guidance "regression" is a measurement artefact
Here is the part a reconstruction engineer should catch. DL3DV has no real surface ground truth — it ships images and COLMAP poses, nothing meshed. So the authors built its reference surfaces, by running Gaussian Wrapping on the dense views (that is the auxiliary dataset contribution: ~10.5K scenes meshed, 10⁷ oriented points each). The "ground truth" Table 1 measures against is itself a Gaussian-Wrapping product.
That changes what the table means. A no-guidance Surflo that already resembles a Gaussian-Wrapping-shaped surface will score well against a Gaussian-Wrapping reference, almost by construction. Guidance pulls the points toward the actual photographs instead — which is what you want, but it moves them away from the pseudo-GT, so the metric reads it as a loss.
The datasets with native surface ground truth tell the opposite story (Table 2, 16 views, reported):
| Dataset (native GT) | No guidance | With guidance |
|---|---|---|
| ML-Hypersim (CD / F1) | 0.0097 / 77.98 | 0.0079 / 87.97 |
| DTU (CD / F1) | 0.0242 / 39.23 | 0.0240 / 42.05 |
| SCRREAM (CD / F1) | 0.0114 / 61.20 | 0.0070 / 81.11 |
On real GT, guidance helps — by nearly 20 F1 points on SCRREAM and 10 on ML-Hypersim. The paper itself says full photometric guidance is "consistently best for visual quality." So the metric and the eye only disagree on the benchmarks whose GT was manufactured the same way the no-guidance output was. When you have a true surface to grade against, the guidance earns its 256 renders. This is a good reminder to read what the ground truth is before you read the number measured against it.
One honest caveat in the other direction: per-view pointmaps are not dominated everywhere. On ML-Hypersim's native GT, raw VGGT pointmaps hit F1 91.91 — above Surflo's 87.97. The paper reports pointmap rows "for reference only" and excludes them from ranking, because duplicated per-view points are not a single coherent surface. Fair, but worth knowing: Surflo's win is "the best coherent global surface," not "beats VGGT on every scalar."
The speed claim, and the view-count claim
The abstract says "an order of magnitude faster than optimization-based methods." The runtime section is more specific, and the two are not in tension. On a single H100, encoding a 16-view scene is one forward pass, and the latent is cached. Decoding 10⁵ points from it "takes a few seconds, dominated by the ODE solve," which the paper calls two orders of magnitude faster than per-scene optimization like 2DGS or Gaussian Wrapping (reported). Turn guidance on and you add 30 seconds to 3 minutes depending on the step count — that is what pulls the end-to-end margin back to the conservative "order of magnitude" in the abstract. The decode-only path is the two-orders number; the full guided path is the one-order number. Both are stated; neither is inflated.
View-count robustness (Table 3, Tanks & Temples, reported) holds, with one caveat the table makes plain:
| Views | Surflo CD ↓ | Surflo F1 ↑ |
|---|---|---|
| 2 | 0.1345 | 9.28 |
| 4 | 0.0135 | 75.07 |
| 8 | 0.0059 | 86.59 |
| 32 | 0.0049 | 90.76 |
Surflo is the best or near-best at every count, and it degrades gracefully as views drop. But two views is genuinely hard geometry for everyone — F1 around 9% for Surflo and roughly 5% for the baselines. "Works from 2 to 32 views" is true in the sense that it is the least-bad option at two; it is not a claim that two views reconstruct a scene. The jump from 2 to 4 views (F1 9 → 75) is where the problem becomes tractable at all.
The ablation also confirms the latent size is load-bearing, not decoration: dropping from to tokens costs DL3DV 0.0073 → 0.0089 CD and 81.21 → 72.76 F1 (Table 5, reported). 128 tokens is enough state to hold a scene; 32 is not.
What it costs, and who can use it
Two caveats sit at the front, not the footnotes.
It inherits VGGT's failure modes. The poses, the camera tokens, and the source distribution for the flow all come from VGGT. When the backbone gives bad pointmaps — very few views, extreme baselines — the noisy-pointmap source and the patch tokenisation are unreliable, and the decoder has less to recover from. Guidance helps but cannot invent a view that was never captured. Surflo is a better reader of VGGT's state than VGGT is of its own, but it is reading VGGT.
The licence is non-commercial, and it propagates. The weights
(surflo_v0.pt, 5,542,438,742 bytes ≈ 5.16 GiB — measured from the Hugging Face
file metadata) are cc-by-nc-4.0, and the meshed DL3DV dataset is cc-by-nc-4.0 and
gated. The code is under the Inria / MPII Gaussian-Splatting License, which the
repo's own THIRD_PARTY_NOTICES.md summarises as "non-commercial, research and
evaluation use only," and whose §4.2 forces any derivative work to carry the
same use limitation — so the restriction propagates to Surflo as a whole, not
only the handful of files it borrows from 3D Gaussian Splatting. For anyone
doing industrial 3D perception, that is the load-bearing line: this is a
paper to learn the mechanism from, not a component to drop into a product. The
same clause that governs Gaussian Splatting, Gaussian Wrapping, and so much of
this lineage governs Surflo.
The one idea to keep
Strip the splatting and the benchmarks and Surflo is one architectural bet: a scene is a fixed-size state, and a surface is a thing you sample out of that state at whatever density you need. Decouple the input count, the latent size, and the output resolution — hold the middle one fixed at 128 tokens — and decode each point as an independent flow from noise to surface. The cost of that independence is that the points don't naturally agree, so you spend a few hundred differentiable renders at the end of the ODE to make them. On real ground truth, that trade comes out ahead, two orders of magnitude faster than optimising a scene from scratch. The licence just means you admire it from research distance.
For the neighbours in this space: the per-view pointmap baseline it beats is the same VGGT lineage covered in the 3D reconstruction roundup and used downstream in WorldCrafter and VoxelTTO; the dense-tracking cousin that also grows a global state is TrackEverything.