# Surflo: one latent for a scene, a surface at any resolution

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/surflo
> date: 2026-10-02
> tags: 3d, point-cloud, flow-matching, gaussian-splatting, benchmarks, licensing, explainer

Geometry does not care where you stood to photograph it. A chair is the same
chair from the doorway or the window, so thirty photos of a room are thirty
redundant encodings of one fixed 3D state. Feed-forward reconstruction models
mostly ignore this. The per-view family — VGGT, Depth Anything 3 — emits one
pointmap per input image, so the representation grows linearly with the number
of views and the outputs overlap, duplicated and slightly misaligned across
frames. The global-latent family fixes the representation size but then commits
to a single low-resolution output grid. You get one or the other: a
representation that scales with input, or an output that is capped.

[Surflo](https://github.com/Anttwo/Surflo), from Antoine Guédon (École
polytechnique, and the author of Gaussian Wrapping) with collaborators at UC
Berkeley, Kyoto University and Kyutai, takes the invariance literally. Compress any number of unposed RGB views into one
fixed-size latent — a single global state — and decode the surface out of it by
transporting points from noise with a flow-matching ODE. Because each point is
decoded on its own, the output is bound to no grid and no token budget: the same
latent yields a few thousand points or a million, in one forward pass. The paper
is arXiv 2606.13644; the project page lists it as a NeurIPS 2026 Oral, which I
cannot verify independently, so I flag it as the authors' claim.

<Figure
  src="https://ai.thesatyajit.com/articles/surflo/fig2.jpg"
  alt="A reconstruction comparison. Far left: a small stack of in-the-wild photographs of the Gallos bronze sculpture on a clifftop at Tintagel, with a 'Surflo Global State' logo showing arrows from the photos converging into a short stack of latent tokens, then scattering into a sparse cloud of points. Centre: 'Surflo points', a clean, dense, soft-coloured point cloud of the cloaked figure standing on ground, and below it 'Surflo mesh', a smooth watertight mesh of the same figure. Right: 'VGGT points', a noisy rainbow cloud where each input view contributes its own duplicated, misaligned colour layer, and below it 'Gaussian Wrapping', a rougher per-scene mesh with holes."
  caption="From the same 16 unposed input views (left), Surflo decodes a clean global point cloud and mesh (centre), where per-view VGGT pointmaps (top right) are duplicated and misaligned across views and per-scene Gaussian Wrapping (bottom right) struggles from few inputs. (Surflo, Figure 1.)"
/>

## Three things that usually come tied together

The idea is easiest to state as a decoupling. In most reconstruction stacks,
three quantities are chained to each other: the number of input views $N$, the
size of the intermediate representation, and the number of output primitives.
Per-view methods tie all three to $N$. Fixed-grid latent methods tie the
representation and the output together and cap both.

Surflo cuts every link. The encoder maps $N$ views, for any $N$, into a latent
$\mathbf{z} \in \mathbb{R}^{K \times D}$ with $K=128$ tokens of width $D=512$ —
that is 65,536 numbers, no matter whether two views went in or eighty. The
decoder reads that latent and emits however many points you ask for, because it
decodes each point independently. Input count, latent size, output resolution:
three dials that no longer move together.

<LatentToPoints />

That independence is the whole trick, and later the whole problem. It is worth
being precise about what "decode a point" means here before we get to why the
points occasionally disagree with each other.

## The encoder: a frozen VGGT, read four ways, squeezed by a Perceiver

The encoder does not learn to see. It wraps a **frozen VGGT-1B backbone** —
loaded straight from `facebook/VGGT-1B` in the released code
(`surflo/model/ffm.py`), weights never updated — and reads its patch tokens from
four intermediate layers, $\ell \in \{4, 11, 17, 23\}$. I read those indices
out of the shipped config (`configs/model/surflo.yaml`,
`intermediate_layer_idx: [4, 11, 17, 23]`), not just the paper, so that one is
measured. Taking four depths rather than only the last gives the compressor both
coarse layout and fine texture to draw from.

Each patch token is then tagged with where it sits in 3D. VGGT already predicts
a rough pointmap, so every patch has an approximate world coordinate; Surflo
encodes that coordinate with Gaussian Fourier features — $F=512$ frequencies
over 16 log-spaced bands — and adds it to the token. This is what lets a plain
set-to-set compressor reason about geometry: the tokens arrive pre-stamped with
position.

The compressor itself is a **Perceiver**. It starts from $K=128$ learnable
latent queries and, four times over (`L_e = 4`), cross-attends them into the
thousands of position-stamped VGGT tokens and then runs four self-attention
blocks (`L_s = 4`) among the latents to mix what they gathered. A separate
single token absorbs camera information, and the final latent is the
concatenation. The output is always the same shape — 128 tokens — whether the
input was a sparse pair or a dense sweep. Many views simply give the
cross-attention more evidence to pool into the same 128 slots.

<Figure
  src="https://ai.thesatyajit.com/articles/surflo/fig1.png"
  alt="The Surflo architecture in three panels. Top left: N input images feed a frozen VGGT (marked with a snowflake) that outputs patch tokens for N images. Middle: the encoder E-phi takes 128 latent-token queries into Cross-Attn, then L_s repeats of Self-Attn and MLP, repeated L_e times, producing a latent z in R to the K by D. Bottom left: the decoder v-theta takes a noisy query point x_t in R-cubed cross S-squared, runs L repeats of Cross-Attn to the latent and MLP, outputs a velocity v-theta(x_t, t given z), and Euler-steps x to x plus velocity times dt, giving a clean oriented point cloud of a truck. Right panel, 'Communication via Guidance': step 1 estimate x_1 from x_t, step 2 adjust the estimate via a rendering loss whose gradient couples neighbouring points, repeated M times, step 3 read off the adjusted velocity v_g."
  caption="Encoder: a frozen VGGT turns N views into patch tokens; a Perceiver with K learnable queries distils them into a fixed-size latent z. Decoder: each noisy query point in R-cubed-cross-S-squared is transported independently by a velocity field conditioned on z. Right: inference-time guidance re-estimates the target points through a rendering loss so neighbours agree. (Surflo, Figure 2.)"
/>

## The decoder: flow matching, one point at a time

Here is where the interesting generative machinery lives. A query is not a pixel
and not a voxel. It is a point in $\mathbb{R}^3 \times \mathbb{S}^2$: a 3D
coordinate plus a unit normal — an *oriented* point, the thing you need to mesh
a surface rather than just dot it. The decoder's job is to take a noisy query
and push it onto the real surface.

It does this with **flow matching**, the same training recipe behind modern
diffusion-style generators (the flow-in-a-geometry-latent idea also drives
[TencentARC's GAE](/articles/gae-geometry-native), and the per-point
transport here is close in spirit to the MAR line of work). The mechanism is
a straight line with a learned speed. At training time you sample a target point
$\mathbf{x}_1$ on the known surface, a source point $\mathbf{x}_0$ from a noise
distribution, a time $t \sim \mathrm{LogitNormal}(1, 1.6)$, and form the linear
interpolant $\mathbf{x}_t = (1-t)\mathbf{x}_0 + t\,\mathbf{x}_1$. The network
$v_\theta$ is trained to predict the constant velocity that walks the source to
the target:

$$
\mathcal{L}_{\mathrm{FM}}(\theta) = \mathbb{E}_{t,\,\mathbf{x}_0,\,\mathbf{x}_1,\,\{\mathbf{I}_n\}}
\big\lVert\, v_\theta(\mathbf{x}_t, t \mid \mathbf{z}) - (\mathbf{x}_1 - \mathbf{x}_0) \,\big\rVert_2^2 .
$$

At inference you do not know $\mathbf{x}_1$; you only have noise and the latent.
So you sample $P$ query points from the source, and integrate the learned ODE
$\frac{d\mathbf{x}_t}{dt} = v_\theta(\mathbf{x}_t, t \mid \mathbf{z})$ from
$t=0$ to $t=1$ with a plain Euler solver, 150 steps. Each point rides its own
velocity field onto the surface. Decode 100K points by default, or 2,000, or a
million — $P$ independent, batched forward passes, the latent computed once and
reused for all of them.

Two design choices in the decoder matter more than they look.

**There is no self-attention over query points.** The 12-layer decoder
cross-attends each point into the latent (for the first 6 layers) and otherwise
processes it alone; all the spatial reasoning was spent in the encoder, baked
into $\mathbf{z}$. That is exactly why $P$ can be anything — the points never
have to attend to each other, so adding more costs nothing but more parallel
forward passes. Conditioning on time and camera enters through DiT-style Ada-LN
blocks, the same modulation pattern I walked through in
[the Diffusion Transformer explainer](/architectures/diffusion-transformer).

**The source distribution is not pure noise.** Starting every point from
$\mathcal{N}(0, \mathbf{I})$ wastes the model on empty space. Surflo instead
seeds the points from a mixture of Gaussians centred on *perturbed VGGT pointmap
samples* — the rough geometry VGGT already gave us — with enough spread to cover
surfaces that were occluded in the input views. The flow then only has to clean
up and complete, not hallucinate from scratch. The ablation earns it: swapping
the mixture for a pure Gaussian source drops DL3DV from CD 0.0073 / F1 81.21 to
0.0088 / 73.57 (Table 5, reported).

## Guidance: the one place the points talk to each other

Independent decoding has a price the paper is honest about: nothing forces two
independently sampled points to land on the *same* surface. In regions occluded
from every input view, the flow is ambiguous, and you get outliers — points
floating just off the true surface because, locally, the velocity field had no
reason to prefer one plausible answer over another.

The fix is an inference-time **guidance** term that, for the last sliver of the
ODE ($t \ge 0.95$), makes neighbouring points agree. At each of those final
steps it first reads off where each point thinks it is going,
$\hat{\mathbf{x}}_1 = \mathbf{x}_t + (1-t)\,v_\theta$. It treats that whole cloud
of predicted endpoints as a set of small oriented 3D Gaussians, renders them
through the cameras VGGT recovered, back into the input viewpoints, and measures
how wrong the renders are against the real photos:

$$
\mathcal{L}_{\mathrm{render}} = \frac{1}{N}\sum_{n=1}^{N}
\lambda\,\big\lVert \hat{\mathbf{I}}_n - \mathbf{I}_n \big\rVert_1
+ (1-\lambda)\,\mathrm{DSSIM}\!\left(\hat{\mathbf{I}}_n, \mathbf{I}_n\right).
$$

It runs $M=32$ gradient-descent steps on this loss, nudging the predicted
endpoints toward a cloud that actually reprojects to the images, producing a
guided target $\hat{\mathbf{x}}_1^{\,g}$ and a guided velocity
$v_g = (\hat{\mathbf{x}}_1^{\,g} - \mathbf{x}_t)/(1-t)$ for the real Euler step.
The important word is *global*: the rendering loss depends on the entire batch
of points at once, so a gradient computed through it couples neighbours. Two
queries that disagree about a surface both feel a shared corrective pull. It is
the only channel through which the otherwise-independent points communicate, and
it opens only in the last roughly 5% of the trajectory.

That timing is worth doing the arithmetic on. With 150 Euler steps, $t \ge 0.95$
is about the last 8 steps, each running 32 inner updates — on the order of 256
differentiable renders per reconstruction (my arithmetic on the paper's numbers,
so: reasoned). Optionally a monocular-depth expert —
[Depth Anything 3](/articles/depth-anything-3-ros2) — adds a scale-invariant
depth-order term to sharpen geometry further. Camera poses are refined in the
same loop.

<Callout type="note">
Guidance renders *oriented Gaussians*, not a mesh, and the cost is independent of
the number of input views $N$ — rendering 16 views or 2 costs the same splat
pass. The final mesh comes afterward, by Delaunay triangulation of the oriented
points, following Gaussian Wrapping.
</Callout>

## Checking the headline table

The marquee numbers are surface metrics: Chamfer Distance (CD, lower is better)
between predicted and ground-truth clouds, and F1 at 1% of the scene diagonal
(higher is better). On DL3DV, the in-distribution held-out split, from 16
unposed views (Table 1, reported):

| Method | CD ↓ | F1 ↑ |
|---|---|---|
| VGGT pointmap | 0.0100 | 74.84 |
| Gaussian Wrapping | 0.0168 | 60.67 |
| NOVA3R (latent, fixed output) | 0.0459 | 30.51 |
| **Surflo — no guidance** | **0.0072** | **81.92** |
| Surflo — with guidance | 0.0083 | 78.55 |

Taken at face value, Surflo-no-guidance is the clear winner: about a third lower
Chamfer than VGGT pointmaps and seven F1 points better, and it extends out of
distribution — Tanks & Temples 0.0053 / 88.57, Mip-NeRF 360 0.0068 / 82.00,
DeepBlending 0.0116 / 70.96 (all reported). NOVA3R, the one prior latent model
that also flow-matches points, is pinned to 10K points and two input views, and
it shows.

But stare at the last two rows. **Guidance makes the DL3DV number worse** —
0.0072 → 0.0083 CD, 81.92 → 78.55 F1 — which is backwards from the whole point
of adding it, and worth more than a shrug.

## Why the guidance "regression" is a measurement artefact

Here is the part a reconstruction engineer should catch. DL3DV has no real
surface ground truth — it ships images and COLMAP poses, nothing meshed. So the
authors *built* its reference surfaces, by running Gaussian Wrapping on the dense
views (that is the auxiliary dataset contribution: ~10.5K scenes meshed, 10⁷
oriented points each). The "ground truth" Table 1 measures against is itself a
Gaussian-Wrapping product.

That changes what the table means. A no-guidance Surflo that already resembles a
Gaussian-Wrapping-shaped surface will score well against a Gaussian-Wrapping
reference, almost by construction. Guidance pulls the points toward the *actual
photographs* instead — which is what you want, but it moves them away from the
pseudo-GT, so the metric reads it as a loss.

The datasets with *native* surface ground truth tell the opposite story
(Table 2, 16 views, reported):

| Dataset (native GT) | No guidance | With guidance |
|---|---|---|
| ML-Hypersim (CD / F1) | 0.0097 / 77.98 | 0.0079 / **87.97** |
| DTU (CD / F1) | 0.0242 / 39.23 | 0.0240 / 42.05 |
| SCRREAM (CD / F1) | 0.0114 / 61.20 | 0.0070 / **81.11** |

On real GT, guidance helps — by nearly 20 F1 points on SCRREAM and 10 on
ML-Hypersim. The paper itself says full photometric guidance is "consistently
best for visual quality." So the metric and the eye only disagree on the
benchmarks whose GT was manufactured the same way the no-guidance output was.
When you have a true surface to grade against, the guidance earns its 256
renders. This is a good reminder to read what the ground truth is before you
read the number measured against it.

One honest caveat in the other direction: per-view pointmaps are not dominated
everywhere. On ML-Hypersim's native GT, raw VGGT pointmaps hit F1 91.91 — above
Surflo's 87.97. The paper reports pointmap rows "for reference only" and excludes
them from ranking, because duplicated per-view points are not a single coherent
surface. Fair, but worth knowing: Surflo's win is "the best coherent global
surface," not "beats VGGT on every scalar."

## The speed claim, and the view-count claim

The abstract says "an order of magnitude faster than optimization-based
methods." The runtime section is more specific, and the two are not in tension.
On a single H100, encoding a 16-view scene is one forward pass, and the latent is
cached. Decoding 10⁵ points from it "takes a few seconds, dominated by the ODE
solve," which the paper calls **two orders of magnitude faster** than per-scene
optimization like 2DGS or Gaussian Wrapping (reported). Turn guidance on and you
add 30 seconds to 3 minutes depending on the step count — that is what pulls the
end-to-end margin back to the conservative "order of magnitude" in the abstract.
The decode-only path is the two-orders number; the full guided path is the
one-order number. Both are stated; neither is inflated.

View-count robustness (Table 3, Tanks & Temples, reported) holds, with one
caveat the table makes plain:

| Views | Surflo CD ↓ | Surflo F1 ↑ |
|---|---|---|
| 2 | 0.1345 | 9.28 |
| 4 | 0.0135 | 75.07 |
| 8 | 0.0059 | 86.59 |
| 32 | 0.0049 | 90.76 |

Surflo is the best or near-best at every count, and it degrades gracefully as
views drop. But two views is genuinely hard geometry for everyone — F1 around 9%
for Surflo and roughly 5% for the baselines. "Works from 2 to 32 views" is true
in the sense that it is the least-bad option at two; it is not a claim that two
views reconstruct a scene. The jump from 2 to 4 views (F1 9 → 75) is where the
problem becomes tractable at all.

The ablation also confirms the latent size is load-bearing, not decoration:
dropping from $K=128$ to $K=32$ tokens costs DL3DV 0.0073 → 0.0089 CD and
81.21 → 72.76 F1 (Table 5, reported). 128 tokens is enough state to hold a
scene; 32 is not.

## What it costs, and who can use it

Two caveats sit at the front, not the footnotes.

**It inherits VGGT's failure modes.** The poses, the camera tokens, and the
source distribution for the flow all come from VGGT. When the backbone gives bad
pointmaps — very few views, extreme baselines — the noisy-pointmap source and the
patch tokenisation are unreliable, and the decoder has less to recover from.
Guidance helps but cannot invent a view that was never captured. Surflo is a
better reader of VGGT's state than VGGT is of its own, but it is reading VGGT.

**The licence is non-commercial, and it propagates.** The weights
(`surflo_v0.pt`, 5,542,438,742 bytes ≈ 5.16 GiB — measured from the Hugging Face
file metadata) are cc-by-nc-4.0, and the meshed DL3DV dataset is cc-by-nc-4.0 and
gated. The code is under the **Inria / MPII Gaussian-Splatting License**, which the
repo's own `THIRD_PARTY_NOTICES.md` summarises as "non-commercial, research and
evaluation use only," and whose §4.2 forces any derivative work to carry the
same use limitation — so the restriction propagates to Surflo as a whole, not
only the handful of files it borrows from 3D Gaussian Splatting. For anyone
doing industrial 3D perception, that is the load-bearing line: this is a
paper to learn the mechanism from, not a component to drop into a product. The
same clause that governs Gaussian Splatting, Gaussian Wrapping, and so much of
this lineage governs Surflo.

## The one idea to keep

Strip the splatting and the benchmarks and Surflo is one architectural bet: *a
scene is a fixed-size state, and a surface is a thing you sample out of that
state at whatever density you need.* Decouple the input count, the latent size,
and the output resolution — hold the middle one fixed at 128 tokens — and
decode each point as an independent flow from noise to surface. The cost of that
independence is that the points don't naturally agree, so you spend a few hundred
differentiable renders at the end of the ODE to make them. On real ground truth,
that trade comes out ahead, two orders of magnitude faster than optimising a
scene from scratch. The licence just means you admire it from research distance.

For the neighbours in this space: the per-view pointmap baseline it beats is the
same VGGT lineage covered in the [3D reconstruction
roundup](/articles/3d-reconstruction-roundup) and used downstream in
[WorldCrafter](/articles/worldcrafter) and [VoxelTTO](/articles/voxel-tto); the
dense-tracking cousin that also grows a global state is
[TrackEverything](/articles/trackeverything).
