~/satyajit

Surflo: one latent for a scene, a surface at any resolution

mdjsonmcp

2026-10-02 · 16 min · 3d · point-cloud · flow-matching · gaussian-splatting · benchmarks · licensing · explainer

A 1:40 narrated explainer, drawn in code. Every number and picture in it is this article's own; the sources are below.

› transcript

Hi, I'm Tobiah. Let me show you Surflo. It turns a handful of ordinary photos into one clean 3D surface. Geometry looks the same from every angle, so Surflo squeezes any number of views into one fixed latent and flows the surface points out of it. A frozen VGGT reads your unposed photos into patch tokens. A Perceiver pours them into 128 latent queries. That is the whole scene, as one fixed-size global state, whatever the view count. A flow decoder transports each query point, on its own, from noise onto the surface. Here is the formula. Each point starts as noise. It should end on the surface. The decoder only predicts the velocity that walks one to the other. The latent is always 128 tokens wide, 65,536 numbers, whether two photos go in or eighty. Independent points can disagree, so guidance makes neighbours talk near the end of the flow. First, guess where each point is heading. Render them as little Gaussians, compare to the real photos, and nudge them to match. Turn that correction into a velocity, and take the step. DL3DV surface F1. A scene is one fixed-size state; the surface is as many points as you ask for. Compress the views into one state, flow the points out, and let guidance make them agree. Every source is in the full article. I'm Tobiah. Bye!

Geometry does not care where you stood to photograph it. A chair is the same chair from the doorway or the window, so thirty photos of a room are thirty redundant encodings of one fixed 3D state. Feed-forward reconstruction models mostly ignore this. The per-view family — VGGT, Depth Anything 3 — emits one pointmap per input image, so the representation grows linearly with the number of views and the outputs overlap, duplicated and slightly misaligned across frames. The global-latent family fixes the representation size but then commits to a single low-resolution output grid. You get one or the other: a representation that scales with input, or an output that is capped.

Surflo, from Antoine Guédon (École polytechnique, and the author of Gaussian Wrapping) with collaborators at UC Berkeley, Kyoto University and Kyutai, takes the invariance literally. Compress any number of unposed RGB views into one fixed-size latent — a single global state — and decode the surface out of it by transporting points from noise with a flow-matching ODE. Because each point is decoded on its own, the output is bound to no grid and no token budget: the same latent yields a few thousand points or a million, in one forward pass. The paper is arXiv 2606.13644; the project page lists it as a NeurIPS 2026 Oral, which I cannot verify independently, so I flag it as the authors' claim.

A reconstruction comparison. Far left: a small stack of in-the-wild photographs of the Gallos bronze sculpture on a clifftop at Tintagel, with a 'Surflo Global State' logo showing arrows from the photos converging into a short stack of latent tokens, then scattering into a sparse cloud of points. Centre: 'Surflo points', a clean, dense, soft-coloured point cloud of the cloaked figure standing on ground, and below it 'Surflo mesh', a smooth watertight mesh of the same figure. Right: 'VGGT points', a noisy rainbow cloud where each input view contributes its own duplicated, misaligned colour layer, and below it 'Gaussian Wrapping', a rougher per-scene mesh with holes.
From the same 16 unposed input views (left), Surflo decodes a clean global point cloud and mesh (centre), where per-view VGGT pointmaps (top right) are duplicated and misaligned across views and per-scene Gaussian Wrapping (bottom right) struggles from few inputs. (Surflo, Figure 1.)

Three things that usually come tied together

The idea is easiest to state as a decoupling. In most reconstruction stacks, three quantities are chained to each other: the number of input views NN, the size of the intermediate representation, and the number of output primitives. Per-view methods tie all three to NN. Fixed-grid latent methods tie the representation and the output together and cap both.

Surflo cuts every link. The encoder maps NN views, for any NN, into a latent z∈RK×D\mathbf{z} \in \mathbb{R}^{K \times D} with K=128K=128 tokens of width D=512D=512 — that is 65,536 numbers, no matter whether two views went in or eighty. The decoder reads that latent and emits however many points you ask for, because it decodes each point independently. Input count, latent size, output resolution: three dials that no longer move together.

one global state, decoded at any resolutionP = 128K points
the latent z
K = 128 × D = 512
65,536 numbers
fixed — independent of views in and points out
showing ~761 of 128K oriented points · one decoder pass per point
2K8K32K128K512K1M
The latent is 128 tokens of width 512 — 65,536 numbers — whatever you set the output to. Decoding 128K points is 128K independent, batched forward passes of the per-point velocity field. Illustrative surface; Surflo reconstructs real scenes.

That independence is the whole trick, and later the whole problem. It is worth being precise about what "decode a point" means here before we get to why the points occasionally disagree with each other.

The encoder: a frozen VGGT, read four ways, squeezed by a Perceiver

The encoder does not learn to see. It wraps a frozen VGGT-1B backbone — loaded straight from facebook/VGGT-1B in the released code (surflo/model/ffm.py), weights never updated — and reads its patch tokens from four intermediate layers, ℓ∈{4,11,17,23}\ell \in \{4, 11, 17, 23\}. I read those indices out of the shipped config (configs/model/surflo.yaml, intermediate_layer_idx: [4, 11, 17, 23]), not just the paper, so that one is measured. Taking four depths rather than only the last gives the compressor both coarse layout and fine texture to draw from.

Each patch token is then tagged with where it sits in 3D. VGGT already predicts a rough pointmap, so every patch has an approximate world coordinate; Surflo encodes that coordinate with Gaussian Fourier features — F=512F=512 frequencies over 16 log-spaced bands — and adds it to the token. This is what lets a plain set-to-set compressor reason about geometry: the tokens arrive pre-stamped with position.

The compressor itself is a Perceiver. It starts from K=128K=128 learnable latent queries and, four times over (L_e = 4), cross-attends them into the thousands of position-stamped VGGT tokens and then runs four self-attention blocks (L_s = 4) among the latents to mix what they gathered. A separate single token absorbs camera information, and the final latent is the concatenation. The output is always the same shape — 128 tokens — whether the input was a sparse pair or a dense sweep. Many views simply give the cross-attention more evidence to pool into the same 128 slots.

The Surflo architecture in three panels. Top left: N input images feed a frozen VGGT (marked with a snowflake) that outputs patch tokens for N images. Middle: the encoder E-phi takes 128 latent-token queries into Cross-Attn, then L_s repeats of Self-Attn and MLP, repeated L_e times, producing a latent z in R to the K by D. Bottom left: the decoder v-theta takes a noisy query point x_t in R-cubed cross S-squared, runs L repeats of Cross-Attn to the latent and MLP, outputs a velocity v-theta(x_t, t given z), and Euler-steps x to x plus velocity times dt, giving a clean oriented point cloud of a truck. Right panel, 'Communication via Guidance': step 1 estimate x_1 from x_t, step 2 adjust the estimate via a rendering loss whose gradient couples neighbouring points, repeated M times, step 3 read off the adjusted velocity v_g.
Encoder: a frozen VGGT turns N views into patch tokens; a Perceiver with K learnable queries distils them into a fixed-size latent z. Decoder: each noisy query point in R-cubed-cross-S-squared is transported independently by a velocity field conditioned on z. Right: inference-time guidance re-estimates the target points through a rendering loss so neighbours agree. (Surflo, Figure 2.)

The decoder: flow matching, one point at a time

Here is where the interesting generative machinery lives. A query is not a pixel and not a voxel. It is a point in R3×S2\mathbb{R}^3 \times \mathbb{S}^2: a 3D coordinate plus a unit normal — an oriented point, the thing you need to mesh a surface rather than just dot it. The decoder's job is to take a noisy query and push it onto the real surface.

It does this with flow matching, the same training recipe behind modern diffusion-style generators (the flow-in-a-geometry-latent idea also drives TencentARC's GAE, and the per-point transport here is close in spirit to the MAR line of work). The mechanism is a straight line with a learned speed. At training time you sample a target point x1\mathbf{x}_1 on the known surface, a source point x0\mathbf{x}_0 from a noise distribution, a time t∼LogitNormal(1,1.6)t \sim \mathrm{LogitNormal}(1, 1.6), and form the linear interpolant xt=(1−t)x0+t x1\mathbf{x}_t = (1-t)\mathbf{x}_0 + t\,\mathbf{x}_1. The network vθv_\theta is trained to predict the constant velocity that walks the source to the target:

LFM(θ)=Et, x0, x1, {In}∥ vθ(xt,t∣z)−(x1−x0) ∥22.\mathcal{L}_{\mathrm{FM}}(\theta) = \mathbb{E}_{t,\,\mathbf{x}_0,\,\mathbf{x}_1,\,\{\mathbf{I}_n\}} \big\lVert\, v_\theta(\mathbf{x}_t, t \mid \mathbf{z}) - (\mathbf{x}_1 - \mathbf{x}_0) \,\big\rVert_2^2 .

At inference you do not know x1\mathbf{x}_1; you only have noise and the latent. So you sample PP query points from the source, and integrate the learned ODE dxtdt=vθ(xt,t∣z)\frac{d\mathbf{x}_t}{dt} = v_\theta(\mathbf{x}_t, t \mid \mathbf{z}) from t=0t=0 to t=1t=1 with a plain Euler solver, 150 steps. Each point rides its own velocity field onto the surface. Decode 100K points by default, or 2,000, or a million — PP independent, batched forward passes, the latent computed once and reused for all of them.

Two design choices in the decoder matter more than they look.

There is no self-attention over query points. The 12-layer decoder cross-attends each point into the latent (for the first 6 layers) and otherwise processes it alone; all the spatial reasoning was spent in the encoder, baked into z\mathbf{z}. That is exactly why PP can be anything — the points never have to attend to each other, so adding more costs nothing but more parallel forward passes. Conditioning on time and camera enters through DiT-style Ada-LN blocks, the same modulation pattern I walked through in the Diffusion Transformer explainer.

The source distribution is not pure noise. Starting every point from N(0,I)\mathcal{N}(0, \mathbf{I}) wastes the model on empty space. Surflo instead seeds the points from a mixture of Gaussians centred on perturbed VGGT pointmap samples — the rough geometry VGGT already gave us — with enough spread to cover surfaces that were occluded in the input views. The flow then only has to clean up and complete, not hallucinate from scratch. The ablation earns it: swapping the mixture for a pure Gaussian source drops DL3DV from CD 0.0073 / F1 81.21 to 0.0088 / 73.57 (Table 5, reported).

Guidance: the one place the points talk to each other

Independent decoding has a price the paper is honest about: nothing forces two independently sampled points to land on the same surface. In regions occluded from every input view, the flow is ambiguous, and you get outliers — points floating just off the true surface because, locally, the velocity field had no reason to prefer one plausible answer over another.

The fix is an inference-time guidance term that, for the last sliver of the ODE (t≥0.95t \ge 0.95), makes neighbouring points agree. At each of those final steps it first reads off where each point thinks it is going, x^1=xt+(1−t) vθ\hat{\mathbf{x}}_1 = \mathbf{x}_t + (1-t)\,v_\theta. It treats that whole cloud of predicted endpoints as a set of small oriented 3D Gaussians, renders them through the cameras VGGT recovered, back into the input viewpoints, and measures how wrong the renders are against the real photos:

Lrender=1N∑n=1Nλ ∥I^n−In∥1+(1−λ) DSSIM ⁣(I^n,In).\mathcal{L}_{\mathrm{render}} = \frac{1}{N}\sum_{n=1}^{N} \lambda\,\big\lVert \hat{\mathbf{I}}_n - \mathbf{I}_n \big\rVert_1 + (1-\lambda)\,\mathrm{DSSIM}\!\left(\hat{\mathbf{I}}_n, \mathbf{I}_n\right).

It runs M=32M=32 gradient-descent steps on this loss, nudging the predicted endpoints toward a cloud that actually reprojects to the images, producing a guided target x^1 g\hat{\mathbf{x}}_1^{\,g} and a guided velocity vg=(x^1 g−xt)/(1−t)v_g = (\hat{\mathbf{x}}_1^{\,g} - \mathbf{x}_t)/(1-t) for the real Euler step. The important word is global: the rendering loss depends on the entire batch of points at once, so a gradient computed through it couples neighbours. Two queries that disagree about a surface both feel a shared corrective pull. It is the only channel through which the otherwise-independent points communicate, and it opens only in the last roughly 5% of the trajectory.

That timing is worth doing the arithmetic on. With 150 Euler steps, t≥0.95t \ge 0.95 is about the last 8 steps, each running 32 inner updates — on the order of 256 differentiable renders per reconstruction (my arithmetic on the paper's numbers, so: reasoned). Optionally a monocular-depth expert — Depth Anything 3 — adds a scale-invariant depth-order term to sharpen geometry further. Camera poses are refined in the same loop.

Checking the headline table

The marquee numbers are surface metrics: Chamfer Distance (CD, lower is better) between predicted and ground-truth clouds, and F1 at 1% of the scene diagonal (higher is better). On DL3DV, the in-distribution held-out split, from 16 unposed views (Table 1, reported):

MethodCD ↓F1 ↑
VGGT pointmap0.010074.84
Gaussian Wrapping0.016860.67
NOVA3R (latent, fixed output)0.045930.51
Surflo — no guidance0.007281.92
Surflo — with guidance0.008378.55

Taken at face value, Surflo-no-guidance is the clear winner: about a third lower Chamfer than VGGT pointmaps and seven F1 points better, and it extends out of distribution — Tanks & Temples 0.0053 / 88.57, Mip-NeRF 360 0.0068 / 82.00, DeepBlending 0.0116 / 70.96 (all reported). NOVA3R, the one prior latent model that also flow-matches points, is pinned to 10K points and two input views, and it shows.

But stare at the last two rows. Guidance makes the DL3DV number worse — 0.0072 → 0.0083 CD, 81.92 → 78.55 F1 — which is backwards from the whole point of adding it, and worth more than a shrug.

Why the guidance "regression" is a measurement artefact

Here is the part a reconstruction engineer should catch. DL3DV has no real surface ground truth — it ships images and COLMAP poses, nothing meshed. So the authors built its reference surfaces, by running Gaussian Wrapping on the dense views (that is the auxiliary dataset contribution: ~10.5K scenes meshed, 10⁷ oriented points each). The "ground truth" Table 1 measures against is itself a Gaussian-Wrapping product.

That changes what the table means. A no-guidance Surflo that already resembles a Gaussian-Wrapping-shaped surface will score well against a Gaussian-Wrapping reference, almost by construction. Guidance pulls the points toward the actual photographs instead — which is what you want, but it moves them away from the pseudo-GT, so the metric reads it as a loss.

The datasets with native surface ground truth tell the opposite story (Table 2, 16 views, reported):

Dataset (native GT)No guidanceWith guidance
ML-Hypersim (CD / F1)0.0097 / 77.980.0079 / 87.97
DTU (CD / F1)0.0242 / 39.230.0240 / 42.05
SCRREAM (CD / F1)0.0114 / 61.200.0070 / 81.11

On real GT, guidance helps — by nearly 20 F1 points on SCRREAM and 10 on ML-Hypersim. The paper itself says full photometric guidance is "consistently best for visual quality." So the metric and the eye only disagree on the benchmarks whose GT was manufactured the same way the no-guidance output was. When you have a true surface to grade against, the guidance earns its 256 renders. This is a good reminder to read what the ground truth is before you read the number measured against it.

One honest caveat in the other direction: per-view pointmaps are not dominated everywhere. On ML-Hypersim's native GT, raw VGGT pointmaps hit F1 91.91 — above Surflo's 87.97. The paper reports pointmap rows "for reference only" and excludes them from ranking, because duplicated per-view points are not a single coherent surface. Fair, but worth knowing: Surflo's win is "the best coherent global surface," not "beats VGGT on every scalar."

The speed claim, and the view-count claim

The abstract says "an order of magnitude faster than optimization-based methods." The runtime section is more specific, and the two are not in tension. On a single H100, encoding a 16-view scene is one forward pass, and the latent is cached. Decoding 10⁵ points from it "takes a few seconds, dominated by the ODE solve," which the paper calls two orders of magnitude faster than per-scene optimization like 2DGS or Gaussian Wrapping (reported). Turn guidance on and you add 30 seconds to 3 minutes depending on the step count — that is what pulls the end-to-end margin back to the conservative "order of magnitude" in the abstract. The decode-only path is the two-orders number; the full guided path is the one-order number. Both are stated; neither is inflated.

View-count robustness (Table 3, Tanks & Temples, reported) holds, with one caveat the table makes plain:

ViewsSurflo CD ↓Surflo F1 ↑
20.13459.28
40.013575.07
80.005986.59
320.004990.76

Surflo is the best or near-best at every count, and it degrades gracefully as views drop. But two views is genuinely hard geometry for everyone — F1 around 9% for Surflo and roughly 5% for the baselines. "Works from 2 to 32 views" is true in the sense that it is the least-bad option at two; it is not a claim that two views reconstruct a scene. The jump from 2 to 4 views (F1 9 → 75) is where the problem becomes tractable at all.

The ablation also confirms the latent size is load-bearing, not decoration: dropping from K=128K=128 to K=32K=32 tokens costs DL3DV 0.0073 → 0.0089 CD and 81.21 → 72.76 F1 (Table 5, reported). 128 tokens is enough state to hold a scene; 32 is not.

What it costs, and who can use it

Two caveats sit at the front, not the footnotes.

It inherits VGGT's failure modes. The poses, the camera tokens, and the source distribution for the flow all come from VGGT. When the backbone gives bad pointmaps — very few views, extreme baselines — the noisy-pointmap source and the patch tokenisation are unreliable, and the decoder has less to recover from. Guidance helps but cannot invent a view that was never captured. Surflo is a better reader of VGGT's state than VGGT is of its own, but it is reading VGGT.

The licence is non-commercial, and it propagates. The weights (surflo_v0.pt, 5,542,438,742 bytes ≈ 5.16 GiB — measured from the Hugging Face file metadata) are cc-by-nc-4.0, and the meshed DL3DV dataset is cc-by-nc-4.0 and gated. The code is under the Inria / MPII Gaussian-Splatting License, which the repo's own THIRD_PARTY_NOTICES.md summarises as "non-commercial, research and evaluation use only," and whose §4.2 forces any derivative work to carry the same use limitation — so the restriction propagates to Surflo as a whole, not only the handful of files it borrows from 3D Gaussian Splatting. For anyone doing industrial 3D perception, that is the load-bearing line: this is a paper to learn the mechanism from, not a component to drop into a product. The same clause that governs Gaussian Splatting, Gaussian Wrapping, and so much of this lineage governs Surflo.

The one idea to keep

Strip the splatting and the benchmarks and Surflo is one architectural bet: a scene is a fixed-size state, and a surface is a thing you sample out of that state at whatever density you need. Decouple the input count, the latent size, and the output resolution — hold the middle one fixed at 128 tokens — and decode each point as an independent flow from noise to surface. The cost of that independence is that the points don't naturally agree, so you spend a few hundred differentiable renders at the end of the ODE to make them. On real ground truth, that trade comes out ahead, two orders of magnitude faster than optimising a scene from scratch. The licence just means you admire it from research distance.

For the neighbours in this space: the per-view pointmap baseline it beats is the same VGGT lineage covered in the 3D reconstruction roundup and used downstream in WorldCrafter and VoxelTTO; the dense-tracking cousin that also grows a global state is TrackEverything.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Surflo: one latent for a scene, a surface at any resolution", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026surflo,
  author = {Satyajit Ghana},
  title  = {Surflo: one latent for a scene, a surface at any resolution},
  url    = {https://ai.thesatyajit.com/articles/surflo},
  year   = {2026}
}
share