~/satyajit

OREN: real-time Euclidean SDF, because the octree carries the prior

mdjsonmcp

2026-10-02 · 14 min · 3d · slam · robotics · point-cloud · lidar · benchmarks · explainer

A planner does not ask a map "where is the wall?". It asks "how much room do I have here, and which way is more?". A signed distance function (SDF) answers both in constant time: at any query point it returns the signed distance to the nearest surface — positive in free space, negative inside an obstacle — and its gradient points straight away from that surface. Collision checks become a sign test. Clearance becomes a lookup. Trajectory optimisers love it because the gradient tells them which way to push.

The catch is that the SDF most online mappers actually build is not that function. It is a truncated SDF: distances are only stored in a thin band around the surface, and everywhere else the map says "unknown" or clamps to the truncation limit. That is fine for carving a mesh, where only the zero-crossing matters, and it is why TSDF fusion is fast. But it is the wrong object for a planner standing in the middle of an open room, which is exactly where you want the distance to be large and trustworthy.

OREN (Zhirui Dai, Qihao Qian, Tianxing Fan, Nikolay Atanasov, UCSD Existential Robotics Lab; arXiv 2510.18999, accepted to IROS 2026) reconstructs the untruncated thing — a continuous, differentiable, Euclidean SDF — online, from streaming point clouds, on a single GPU. It does so by splitting the problem in two: an explicit octree prior does the heavy lifting, and a small neural network cleans up what the prior cannot resolve. The code is MIT and I read it against the paper; the division of labour is the whole story, so that is where I will start.

Truncated vs Euclidean, and why the difference is load-bearing

Drag the dot below. The two modes share the same floor plan and the same analytic distance field; the only difference is what the map is allowed to store.

query (5.40, 4.25) m: map reads |d| ≥ τ → clamped at +0.30 m. true clearance 0.91 m is not stored. Drag anywhere. Red outlines are the surfaces (the zero-level set).

In Euclidean mode the field is meaningful everywhere. Stand in the centre of the room and the map still knows you are, say, 1.3 m from the nearest wall, and the arrow — the SDF gradient — points away from that surface, the direction clearance grows. In truncated mode, every cell past the band |d| < τ collapses to one grey "unknown": the map cannot distinguish a wide-open centre from a tight corner, because it threw that number away to save memory. H2-Mapping and PIN-SLAM only ever store the band near the surface. Voxblox can propagate further but pays for it. OREN keeps the whole field.

This is not a free lunch — a global Euclidean field is more to compute and more to store than a truncation band. OREN's bet is that you can get it cheaply if the explicit part of the model is accurate enough that the expensive part (the network) barely has to work.

Three families, three compromises

SDF reconstruction has historically forced a choice:

familyexamplecontinuous / differentiableEuclideanonline + large-scalememory
volumetricVoxbloxno (grid-limited)ESDF, slowlyyesgrows with resolution
Gaussian processLog-GPISyespartialpoor (cubic in points)large Gram matrix
neuraliSDF, DeepSDFyesrarelycatastrophic forgettingcompact

Volumetric methods hash voxels and run in real time, but a discrete grid is not differentiable and its accuracy is capped by voxel size. Gaussian-process SDFs are continuous and come with uncertainty, but the matrix inverse is cubic and does not scale. Neural SDFs are compact and smooth, but training one online means it forgets what it saw a hundred frames ago, and most of them only ever learned the truncated, near-surface part. OREN is a hybrid: keep the octree for speed and scale, add a network for the fine detail, and design the interface between them so neither has to do the other's job.

The explicit prior: a semi-sparse octree with gradients

An octree subdivides space adaptively — big cells in empty regions, small cells near surfaces. OREN stores, at every octant vertex, two learnable things: an SDF value dkd_k and an SDF gradient gk∈R3\mathbf{g}_k \in \mathbb{R}^3. To predict the SDF at an arbitrary query x\mathbf{x}, it finds the smallest octant containing x\mathbf{x} and interpolates from its eight corners.

The first design choice is the tree's sparsity. A fully sparse octree only creates the small child octants that actually contain surface points. That is maximally memory-efficient, but it leaves ragged boundaries: a query just outside a surface cell may have to fall back to a much coarser parent, and the interpolation jumps discontinuously across octant borders. OREN instead makes the first MM (coarse) layers semi-sparse: when a child octant is created, all its siblings are created too, even the empty ones. That costs memory but guarantees that a query near a surface always finds nearby vertices, so the field stays continuous. The finest layers stay sparse, because near the surface the prior is already accurate and the eight-fold-per-layer octant growth would be ruinous. In the shipped config this is tree_depth: 8, semi_sparse_depth: 5, resolution: 0.1 — eight levels from a 25.6 m root (0.1 × 2⁸) down to 10 cm leaves, the coarsest five semi-sparse (measured, from configs/trainer-replica.yaml).

Two 2D SDF interpolations of the same obstacle corner. The sparse octree on the left shows contour lines that kink and jump at octant boundaries; the semi-sparse octree on the right, with sibling cells filled in, produces smooth continuous contours.
Trilinear interpolation of the same corner in a sparse (left) vs semi-sparse (right) octree. Filling in the sibling octants removes the discontinuities at cell boundaries (OREN, Figure 3).

The second, more interesting choice is gradient-augmented interpolation. Ordinary trilinear interpolation blends the eight vertex values dkd_k. OREN first extrapolates each vertex using its stored gradient,

dk(x)=dk+gk⊤(x−xk),d_k(\mathbf{x}) = d_k + \mathbf{g}_k^\top (\mathbf{x} - \mathbf{x}_k),

and then blends those eight first-order estimates. The reason this matters is a clean error bound. For an octant of size LL whose SDF has Hessian spectral norm bounded by MM, the paper proves (Proposition 1):

ega≤3ML28,etl≤3 L2.e_{ga} \le \frac{3 M L^2}{8}, \qquad e_{tl} \le \frac{\sqrt{3}\,L}{2}.

Plain trilinear error is linear in the cell size; gradient-augmented error is quadratic. Since the gradient makes the first-order term of the Taylor expansion exact, only the curvature is left over. On real SDFs the curvature MM is small almost everywhere (surfaces are mostly locally flat), so the quadratic bound wins by a wide margin — and it wins most in the one place plain interpolation is worst: an octant wedged between two obstacles, where every vertex stores a small distance and the naive average sinks far below the true clearance in the middle.

octant between two obstacles, size L = 1.20 m
13077distance (cm)
truth plain, max error 53.63 cm gradient-aug, max error 12.07 cmgradient augmentation is 4.4× more accurate here; plain interpolation averages two small vertex distances and sinks below the true clearance. The gap widens as L grows (linear vs quadratic).

That widget is the mechanism in one dimension. The vertices carry small distances and the right slopes; plain interpolation (red) averages the values and underestimates the valley, while gradient augmentation (teal) climbs toward it and tracks the curve. Drag L: both errors shrink as the cell shrinks, but the plain error shrinks linearly and the augmented error quadratically, so the gap widens with cell size. This is why OREN can afford a coarse 10 cm octree and still hand the network an accurate prior.

The implicit residual: a two-layer MLP, and nothing more

The prior is good but resolution-limited; it cannot represent geometry finer than its cells. So OREN stores a second learnable thing at each vertex — an implicit feature fk∈RF\mathbf{f}_k \in \mathbb{R}^F with F=3F = 3 — interpolates it with the same weights, and feeds the interpolated feature together with the prior dga(x)d_{ga}(\mathbf{x}) into a small decoder:

d^(x)=dga(x)+δd(x),δd(x)=D(dga(x),f(x);β).\hat{d}(\mathbf{x}) = d_{ga}(\mathbf{x}) + \delta_d(\mathbf{x}), \qquad \delta_d(\mathbf{x}) = D\big(d_{ga}(\mathbf{x}), \mathbf{f}(\mathbf{x}); \beta\big).

The decoder is deliberately tiny: two 32-dimensional hidden layers with LeakyReLU (measured, residual_net.py and the config). And the output is scaled by output_sdf_scale: 0.1, so the network only ever nudges the prior by a small correction rather than predicting the distance from scratch. Because the prior is already globally accurate, the residual is a near-surface detail term — which is why the whole thing stays fast and, crucially, does not catastrophically forget: the global structure lives in the explicit octree, which grows as the scene grows, not in the fixed-size network.

OREN's five-stage pipeline: (a) key-frame selection keeping frames with small overlap, (b) sampling surface, perturbed and free-space points, (c) the gradient-augmented octree prior storing an SDF value and gradient at each vertex, (d) implicit features at vertices decoded by an MLP into a residual, (e) the prior plus residual trained with reconstruction, Eikonal and projection losses.
The full pipeline: the explicit gradient-augmented octree prior (c) plus the implicit neural residual (d) sum to the final SDF, trained with three losses (OREN, Figure 2).

Training is online. OREN keeps a set of W=8W = 8 key frames chosen to cover the observed surface with little overlap (the H2-Mapping selection rule), draws 20,480 rays per step, and samples three kinds of points along them: surface, perturbed (near-surface), and free-space. Three losses supervise the sum: an L1 reconstruction loss on surface and perturbed points, an Eikonal loss pushing ∥∇d^∥→1\lVert \nabla \hat d \rVert \to 1 (computed by numerical differentiation, not autograd, which the authors found more stable where the gradient is ill-defined), and a projection loss on free-space points that supplies the one thing the other two miss — gradient direction and scale far from any surface. The ablation confirms the projection loss is load-bearing: removing it drops Replica room1 F1 from 93.33% to 89.22% and multiplies the far-region SDF error more than tenfold (reported, Table IV).

What the numbers actually say

I checked the headline claims against Tables I–III. Hardware is an Intel i9-14900K with an NVIDIA RTX 3090; the datasets are the eight synthetic Replica scenes and the real-LiDAR Newer College sequence; all methods use ground-truth poses (PIN-SLAM with its localiser disabled). That last point matters and I will come back to it: OREN is a mapper, not a SLAM system — it is not estimating where the sensor is.

SDF accuracy. Mean absolute error of the SDF over the whole scene, against a ground-truth grid (reported, Table II):

SDF MAE, all regions [cm]Replica room0Newer College
OREN2.1556.00
HIO-SDF3.27301.56
Voxblox3.1362.92

OREN is best in both. H2-Mapping and PIN-SLAM are missing from this table on purpose: they only predict near the surface, so their "SDF valid ratio" — the fraction of the grid where they return a value at all — is small, often under a fifth of the scene on Replica (reported, Table II), and a whole-scene MAE would be meaningless. The 56 cm on Newer College looks alarming until you remember the scene: an outdoor-scale college quad evaluated on a 20 cm grid, where points far from any surface genuinely sit metres away. OREN's near-surface error there is 16.96 cm and its far-region error 59.62 cm (reported); HIO-SDF's 301.56 cm whole-scene number is a training that fell over, which the authors say happens when its underlying volumetric prior is poor. Even against the solid baseline, Voxblox, OREN is about 11% lower on Newer College and roughly 5× lower than HIO-SDF (reasoned, from the table).

Speed and memory. This is where the explicit prior pays off (reported, Table III):

Replica room0ORENH2-MappingPIN-SLAMHIO-SDFVoxblox
FPS13.9612.368.431.990.87
GPU memory [GB]1.3313.966.122.14N/A

OREN is the fastest and the lightest. On Newer College it runs at 14.28 FPS, the fastest of the five, on 1.59 GB of GPU memory, again the lowest. The GPU figure is the striking one: H2-Mapping needs 13.96 GB on Replica to OREN's 1.33, about a 10× difference, and PIN-SLAM 6.12 GB, about 4.6× (reasoned). Voxblox's ESDF is so slow that on Newer College OREN is roughly 71× its 0.20 FPS (reasoned). The authors also report 7 FPS on an NVIDIA Jetson Orin with 16 GB, which is the number that decides whether this runs on a robot rather than a workstation.

Mesh completeness. Extract a mesh from the field and OREN wins on completion and recall. On Newer College the completion ratio (fraction of ground-truth surface within δ = 20 cm) is 94.20% for OREN against 72.83% for PIN-SLAM and 61.58% for H2-Mapping (reported, Table I) — a 21-point lead over the nearest baseline (reasoned).

Six z-plane slices of the reconstructed SDF on Replica room0. Ground truth and OREN both show a full continuous field coloured everywhere with a clean zero-level contour around the furniture; H2-Mapping and PIN-SLAM show only a thin coloured band near surfaces with everything else blank, on a ±0.1 colour scale; HIO-SDF is over-smoothed with spurious blue regions; Voxblox under-estimates the distances.
z-plane SDF slices on Replica room0. OREN matches the ground truth across the whole scene; H2-Mapping and PIN-SLAM store only a truncated band (note their ±0.1 m colour scale); HIO-SDF over-smooths and Voxblox under-estimates (OREN, Figure 5).

That figure is the thesis in one picture. OREN and the ground truth are coloured everywhere; H2-Mapping and PIN-SLAM are blank except for a hairline around each object, on a colour scale ten times narrower.

Where it is honest about losing

OREN is not best at everything, and the paper says so. On the surface-fit metrics — F1 score, precision, Chamfer-L1, accuracy — H2-Mapping and PIN-SLAM edge it out on Replica, because they are specialised surface reconstructors and, as the authors note, meshes with holes can inflate precision (you are only scored on the points you did commit to). OREN is "mostly the second best" there (reported). Its CPU memory is also higher than the lightest baselines — 3.81 GB on Replica — because the semi-sparse octree allocates those extra sibling vertices; the ablation measures the cost directly: making every layer semi-sparse nearly doubles the octants (21,418 → 41,073, +92%) for no real accuracy gain (reported, Table IV), which is the empirical justification for keeping only the coarse layers dense. And the whole evaluation assumes ground-truth poses, so the hard part of a deployed system — staying localised while you map — is out of scope here.

OREN-X: one octree, three modalities

The follow-up, OREN-X ("Octree Residual Network for Real-Time Multi-Modal Mapping", arXiv 2609.29157, 24 September 2026, same lead authors plus a larger team), keeps the explicit-prior-plus-residual recipe and asks it to carry more than geometry. One shared octree now stores geometric, radiance and vision-language information at its vertices, so a robot gets a distance field, a renderable colour field, and an open-vocabulary semantic field out of a single structure — query "how far is the nearest chair" and "point me at a mug" against the same map.

The reported numbers (from the abstract; there is no OREN-X repository in the lab's GitHub org yet, so I could not check these against code): 80+ FPS for SDF alone and 30+ FPS for all four modalities together, a 33% gain in near-surface SDF accuracy over single-modality baselines, vision-language features compressed 3.7× below full per-vertex storage, and a 71% improvement in mean open-vocabulary 3D mIoU. If those hold, the interesting claim is architectural: the same insight — let an accurate explicit structure carry the load so the network stays small — generalises from distance to colour to language. Treat the OREN-X figures as the publisher's until the code lands.

The code

ExistentialRobotics/oren@56c75f2 · snapshot 2026-10-02
tracked files
265
license
MIT
branch
main
tests
none found
source
1.8 MB
commit date
2026-07-08
source by language
Python878.6 kB(61)C++718.3 kB(84)C117.2 kB(9)CUDA50.3 kB(12)HTML17.1 kB(1)JavaScript15.3 kB(3)CSS9.2 kB(3)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-02 at 56c75f2 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

The OREN release (MIT) is a working system, not a figure-reproduction dump: a trainer for Replica and Newer College, a ROS 2 mapping node for live rosbags and sensor streams, an interactive GUI that shows the SDF slice, the octree structure and the camera poses as training proceeds, and a Docker setup. The config I quoted numbers from (configs/trainer-replica.yaml) exposes every hyperparameter the paper names — the octree depth and semi-sparse depth, the feature dimension, the loss weights — so the division of labour between prior and residual is right there to tune.

The lesson I take from OREN is one that keeps recurring in 3D perception: the network is rarely the whole answer. Put the geometry in a structure that is cheap, accurate and grows with the scene — here, a semi-sparse octree with gradients — and the learned part collapses to a two-layer residual. That is what buys the Euclidean field, the 1.33 GB, and the 14 FPS at once. It is the same bet as feeding a frozen depth backbone into a bounded-memory streaming reconstructor, or moving the registration geometry onto the GPU in FAR-LIO: decide what the explicit part should carry, and the model gets small. The piece OREN leaves for someone else is localisation — and as SurfSLAM shows, once the poses are not given, the pose estimate, not the geometry, becomes the hard part. When that live map is finally published off a robot, there is still the question Depth Anything 3 in ROS 2 runs into: who supplied the metres, and under what assumption.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "OREN: real-time Euclidean SDF, because the octree carries the prior", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026orensdf,
  author = {Satyajit Ghana},
  title  = {OREN: real-time Euclidean SDF, because the octree carries the prior},
  url    = {https://ai.thesatyajit.com/articles/oren-sdf},
  year   = {2026}
}
share