2026-10-02 · 14 min · 3d · slam · robotics · point-cloud · lidar · benchmarks · explainer
A planner does not ask a map "where is the wall?". It asks "how much room do I have here, and which way is more?". A signed distance function (SDF) answers both in constant time: at any query point it returns the signed distance to the nearest surface — positive in free space, negative inside an obstacle — and its gradient points straight away from that surface. Collision checks become a sign test. Clearance becomes a lookup. Trajectory optimisers love it because the gradient tells them which way to push.
The catch is that the SDF most online mappers actually build is not that function. It is a truncated SDF: distances are only stored in a thin band around the surface, and everywhere else the map says "unknown" or clamps to the truncation limit. That is fine for carving a mesh, where only the zero-crossing matters, and it is why TSDF fusion is fast. But it is the wrong object for a planner standing in the middle of an open room, which is exactly where you want the distance to be large and trustworthy.
OREN (Zhirui Dai, Qihao Qian, Tianxing Fan, Nikolay Atanasov, UCSD Existential Robotics Lab; arXiv 2510.18999, accepted to IROS 2026) reconstructs the untruncated thing — a continuous, differentiable, Euclidean SDF — online, from streaming point clouds, on a single GPU. It does so by splitting the problem in two: an explicit octree prior does the heavy lifting, and a small neural network cleans up what the prior cannot resolve. The code is MIT and I read it against the paper; the division of labour is the whole story, so that is where I will start.
Truncated vs Euclidean, and why the difference is load-bearing
Drag the dot below. The two modes share the same floor plan and the same analytic distance field; the only difference is what the map is allowed to store.
In Euclidean mode the field is meaningful everywhere. Stand in the centre of the
room and the map still knows you are, say, 1.3 m from the nearest wall, and the
arrow — the SDF gradient — points away from that surface, the direction clearance
grows. In truncated mode, every
cell past the band |d| < τ collapses to one grey "unknown": the map cannot
distinguish a wide-open centre from a tight corner, because it threw that number
away to save memory. H2-Mapping and PIN-SLAM only ever store the band near the
surface. Voxblox can propagate further but pays for it. OREN keeps the whole
field.
This is not a free lunch — a global Euclidean field is more to compute and more to store than a truncation band. OREN's bet is that you can get it cheaply if the explicit part of the model is accurate enough that the expensive part (the network) barely has to work.
Three families, three compromises
SDF reconstruction has historically forced a choice:
| family | example | continuous / differentiable | Euclidean | online + large-scale | memory |
|---|---|---|---|---|---|
| volumetric | Voxblox | no (grid-limited) | ESDF, slowly | yes | grows with resolution |
| Gaussian process | Log-GPIS | yes | partial | poor (cubic in points) | large Gram matrix |
| neural | iSDF, DeepSDF | yes | rarely | catastrophic forgetting | compact |
Volumetric methods hash voxels and run in real time, but a discrete grid is not differentiable and its accuracy is capped by voxel size. Gaussian-process SDFs are continuous and come with uncertainty, but the matrix inverse is cubic and does not scale. Neural SDFs are compact and smooth, but training one online means it forgets what it saw a hundred frames ago, and most of them only ever learned the truncated, near-surface part. OREN is a hybrid: keep the octree for speed and scale, add a network for the fine detail, and design the interface between them so neither has to do the other's job.
The explicit prior: a semi-sparse octree with gradients
An octree subdivides space adaptively — big cells in empty regions, small cells near surfaces. OREN stores, at every octant vertex, two learnable things: an SDF value and an SDF gradient . To predict the SDF at an arbitrary query , it finds the smallest octant containing and interpolates from its eight corners.
The first design choice is the tree's sparsity. A fully sparse octree only
creates the small child octants that actually contain surface points. That is
maximally memory-efficient, but it leaves ragged boundaries: a query just outside
a surface cell may have to fall back to a much coarser parent, and the
interpolation jumps discontinuously across octant borders. OREN instead makes the
first (coarse) layers semi-sparse: when a child octant is created, all
its siblings are created too, even the empty ones. That costs memory but
guarantees that a query near a surface always finds nearby vertices, so the field
stays continuous. The finest layers stay sparse, because near the surface the
prior is already accurate and the eight-fold-per-layer octant growth would be
ruinous. In the shipped config this is tree_depth: 8, semi_sparse_depth: 5,
resolution: 0.1 — eight levels from a 25.6 m root (0.1 × 2⁸) down to 10 cm
leaves, the coarsest five semi-sparse (measured, from configs/trainer-replica.yaml).

The second, more interesting choice is gradient-augmented interpolation. Ordinary trilinear interpolation blends the eight vertex values . OREN first extrapolates each vertex using its stored gradient,
and then blends those eight first-order estimates. The reason this matters is a clean error bound. For an octant of size whose SDF has Hessian spectral norm bounded by , the paper proves (Proposition 1):
Plain trilinear error is linear in the cell size; gradient-augmented error is quadratic. Since the gradient makes the first-order term of the Taylor expansion exact, only the curvature is left over. On real SDFs the curvature is small almost everywhere (surfaces are mostly locally flat), so the quadratic bound wins by a wide margin — and it wins most in the one place plain interpolation is worst: an octant wedged between two obstacles, where every vertex stores a small distance and the naive average sinks far below the true clearance in the middle.
That widget is the mechanism in one dimension. The vertices carry small
distances and the right slopes; plain interpolation (red) averages the values and
underestimates the valley, while gradient augmentation (teal) climbs toward it
and tracks the curve. Drag L: both errors shrink as the cell shrinks, but the
plain error shrinks linearly and the augmented error quadratically, so the gap
widens with cell size. This is why OREN can afford a coarse 10 cm octree and
still hand the network an accurate prior.
The implicit residual: a two-layer MLP, and nothing more
The prior is good but resolution-limited; it cannot represent geometry finer than its cells. So OREN stores a second learnable thing at each vertex — an implicit feature with — interpolates it with the same weights, and feeds the interpolated feature together with the prior into a small decoder:
The decoder is deliberately tiny: two 32-dimensional hidden layers with LeakyReLU
(measured, residual_net.py and the config). And the output is scaled by
output_sdf_scale: 0.1, so the network only ever nudges the prior by a small
correction rather than predicting the distance from scratch. Because the prior
is already globally accurate, the residual is a near-surface detail term — which
is why the whole thing stays fast and, crucially, does not catastrophically
forget: the global structure lives in the explicit octree, which grows as the
scene grows, not in the fixed-size network.

Training is online. OREN keeps a set of key frames chosen to cover the observed surface with little overlap (the H2-Mapping selection rule), draws 20,480 rays per step, and samples three kinds of points along them: surface, perturbed (near-surface), and free-space. Three losses supervise the sum: an L1 reconstruction loss on surface and perturbed points, an Eikonal loss pushing (computed by numerical differentiation, not autograd, which the authors found more stable where the gradient is ill-defined), and a projection loss on free-space points that supplies the one thing the other two miss — gradient direction and scale far from any surface. The ablation confirms the projection loss is load-bearing: removing it drops Replica room1 F1 from 93.33% to 89.22% and multiplies the far-region SDF error more than tenfold (reported, Table IV).
What the numbers actually say
I checked the headline claims against Tables I–III. Hardware is an Intel i9-14900K with an NVIDIA RTX 3090; the datasets are the eight synthetic Replica scenes and the real-LiDAR Newer College sequence; all methods use ground-truth poses (PIN-SLAM with its localiser disabled). That last point matters and I will come back to it: OREN is a mapper, not a SLAM system — it is not estimating where the sensor is.
SDF accuracy. Mean absolute error of the SDF over the whole scene, against a ground-truth grid (reported, Table II):
| SDF MAE, all regions [cm] | Replica room0 | Newer College |
|---|---|---|
| OREN | 2.15 | 56.00 |
| HIO-SDF | 3.27 | 301.56 |
| Voxblox | 3.13 | 62.92 |
OREN is best in both. H2-Mapping and PIN-SLAM are missing from this table on purpose: they only predict near the surface, so their "SDF valid ratio" — the fraction of the grid where they return a value at all — is small, often under a fifth of the scene on Replica (reported, Table II), and a whole-scene MAE would be meaningless. The 56 cm on Newer College looks alarming until you remember the scene: an outdoor-scale college quad evaluated on a 20 cm grid, where points far from any surface genuinely sit metres away. OREN's near-surface error there is 16.96 cm and its far-region error 59.62 cm (reported); HIO-SDF's 301.56 cm whole-scene number is a training that fell over, which the authors say happens when its underlying volumetric prior is poor. Even against the solid baseline, Voxblox, OREN is about 11% lower on Newer College and roughly 5× lower than HIO-SDF (reasoned, from the table).
Speed and memory. This is where the explicit prior pays off (reported, Table III):
| Replica room0 | OREN | H2-Mapping | PIN-SLAM | HIO-SDF | Voxblox |
|---|---|---|---|---|---|
| FPS | 13.96 | 12.36 | 8.43 | 1.99 | 0.87 |
| GPU memory [GB] | 1.33 | 13.96 | 6.12 | 2.14 | N/A |
OREN is the fastest and the lightest. On Newer College it runs at 14.28 FPS, the fastest of the five, on 1.59 GB of GPU memory, again the lowest. The GPU figure is the striking one: H2-Mapping needs 13.96 GB on Replica to OREN's 1.33, about a 10× difference, and PIN-SLAM 6.12 GB, about 4.6× (reasoned). Voxblox's ESDF is so slow that on Newer College OREN is roughly 71× its 0.20 FPS (reasoned). The authors also report 7 FPS on an NVIDIA Jetson Orin with 16 GB, which is the number that decides whether this runs on a robot rather than a workstation.
Mesh completeness. Extract a mesh from the field and OREN wins on completion and recall. On Newer College the completion ratio (fraction of ground-truth surface within δ = 20 cm) is 94.20% for OREN against 72.83% for PIN-SLAM and 61.58% for H2-Mapping (reported, Table I) — a 21-point lead over the nearest baseline (reasoned).

That figure is the thesis in one picture. OREN and the ground truth are coloured everywhere; H2-Mapping and PIN-SLAM are blank except for a hairline around each object, on a colour scale ten times narrower.
Where it is honest about losing
OREN is not best at everything, and the paper says so. On the surface-fit metrics — F1 score, precision, Chamfer-L1, accuracy — H2-Mapping and PIN-SLAM edge it out on Replica, because they are specialised surface reconstructors and, as the authors note, meshes with holes can inflate precision (you are only scored on the points you did commit to). OREN is "mostly the second best" there (reported). Its CPU memory is also higher than the lightest baselines — 3.81 GB on Replica — because the semi-sparse octree allocates those extra sibling vertices; the ablation measures the cost directly: making every layer semi-sparse nearly doubles the octants (21,418 → 41,073, +92%) for no real accuracy gain (reported, Table IV), which is the empirical justification for keeping only the coarse layers dense. And the whole evaluation assumes ground-truth poses, so the hard part of a deployed system — staying localised while you map — is out of scope here.
OREN-X: one octree, three modalities
The follow-up, OREN-X ("Octree Residual Network for Real-Time Multi-Modal Mapping", arXiv 2609.29157, 24 September 2026, same lead authors plus a larger team), keeps the explicit-prior-plus-residual recipe and asks it to carry more than geometry. One shared octree now stores geometric, radiance and vision-language information at its vertices, so a robot gets a distance field, a renderable colour field, and an open-vocabulary semantic field out of a single structure — query "how far is the nearest chair" and "point me at a mug" against the same map.
The reported numbers (from the abstract; there is no OREN-X repository in the lab's GitHub org yet, so I could not check these against code): 80+ FPS for SDF alone and 30+ FPS for all four modalities together, a 33% gain in near-surface SDF accuracy over single-modality baselines, vision-language features compressed 3.7× below full per-vertex storage, and a 71% improvement in mean open-vocabulary 3D mIoU. If those hold, the interesting claim is architectural: the same insight — let an accurate explicit structure carry the load so the network stays small — generalises from distance to colour to language. Treat the OREN-X figures as the publisher's until the code lands.
The code
- license
- MIT
- branch
- main
- tests
- none found
- source
- 1.8 MB
- commit date
- 2026-07-08
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-02 at 56c75f2 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow
shallow clone: counts describe the pinned tree, not the history
The OREN release (MIT) is a working system, not a figure-reproduction dump: a
trainer for Replica and Newer College, a ROS 2 mapping node for live rosbags and
sensor streams, an interactive GUI that shows the SDF slice, the octree structure
and the camera poses as training proceeds, and a Docker setup. The config I quoted
numbers from (configs/trainer-replica.yaml) exposes every hyperparameter the
paper names — the octree depth and semi-sparse depth, the feature dimension, the
loss weights — so the division of labour between prior and residual is right there
to tune.
The lesson I take from OREN is one that keeps recurring in 3D perception: the network is rarely the whole answer. Put the geometry in a structure that is cheap, accurate and grows with the scene — here, a semi-sparse octree with gradients — and the learned part collapses to a two-layer residual. That is what buys the Euclidean field, the 1.33 GB, and the 14 FPS at once. It is the same bet as feeding a frozen depth backbone into a bounded-memory streaming reconstructor, or moving the registration geometry onto the GPU in FAR-LIO: decide what the explicit part should carry, and the model gets small. The piece OREN leaves for someone else is localisation — and as SurfSLAM shows, once the poses are not given, the pose estimate, not the geometry, becomes the hard part. When that live map is finally published off a robot, there is still the question Depth Anything 3 in ROS 2 runs into: who supplied the metres, and under what assumption.