~/satyajit

Gaussian-splat colour: 192 bytes of spherical harmonics, or 28 bytes and a tiny MLP

mdjsonmcp

2026-10-06 · 25 min · 3d · gaussian-splatting · cuda · benchmarks

Why read this

Notabletop 60%

Each 3DGS colour model's maths and CUDA, bytes recounted from code, the M360 table split (SH still wins outdoors), and what it means for a capture pipeline.

  • Original analysis
  • Runs on a consumer GPU
  • A lasting reference

3D & spatialApache-2.0Practitioner paper

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
2 of 3: Open code or weights with real limits
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
1 of 3: A specialist community
Only here?
2 of 3: A teardown or measurement few others did

Score 62 of 100, ranked 211 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

When I wrote up how people capture Gaussian-splat scenes with 360 cameras, one number kept coming back. A 15-million-splat drone capture of a dam was trained at spherical-harmonic degree 2, and when someone asked why not degree 3, the answer was two words: "VRAM limitation". The arithmetic behind that is not subtle. At degree 3, 48 of a splat's 59 trainable floats are colour, and 45 of those exist only to make the colour change with viewing direction.

So when Florian Hahlbohm posted "It's time to stop defaulting to spherical harmonics for 3DGS!", I read the whole ten-post thread, the paper and every CUDA function in the code. The pitch is a compact neural appearance model: 28 bytes per Gaussian instead of 192, faster training, a little better PSNR on average. The work is bigger than the pitch. They took five appearance models (SH, Spherical Voronoi, NASG, NASGabor and theirs), put all of them into one optimised pipeline with the same densification and the same everything else, and measured quality, memory, training time and frame time on desktop GPUs and a phone.

Three things surprised me. The neural model still uses spherical harmonics; it just stops storing them per Gaussian. On the outdoor scenes of Mip-NeRF 360, plain SH still has the best PSNR of the lot. And the most useful part of the paper is not a model at all. It is an argument about what the "view-dependent" part of a splat actually learns on a real capture, and it should change how anyone running a capture pipeline reads benchmark tables.

Two pipelines side by side. Left, labelled 3DGS spherical harmonics: a view direction d and a box of 3 base colours plus 45 SH coefficients feed a triangle of 16 coloured sphere glyphs (the SH basis), producing a view-dependent colour for a vase on a table. Right, labelled Ours, latent features with shared decoder: the encoded view direction and 8 encoded features feed a small fully connected network drawn as circles, whose view-dependent residual is added to a separate base colour to colour the same vase.
Degree-3 SH stores 45 view-dependent coefficients per Gaussian and weights a fixed basis with them (left). The neural model stores 8 features per Gaussian and lets one MLP shared by the whole scene decode them, given the encoded view direction, into a residual on top of the base colour (right). (paper, Figure 1; also the repository's README teaser.)

The colour of one Gaussian, as a function of direction

A Gaussian splat has a position, a shape, an opacity and a colour. The colour is not one RGB triple, because real surfaces look different from different angles: a varnished table, a laptop lid, wet stone. So 3DGS makes colour a function of the unit direction d\mathbf d from the camera to the Gaussian's centre. The paper writes every model in the same form, which is what makes the comparison fair:

c(d)=ϕ(c0+Aψ(d))\mathbf c(\mathbf d) = \phi\big(\mathbf c_0 + \mathcal A_\psi(\mathbf d)\big)

c0\mathbf c_0 is a view-independent base colour, Aψ\mathcal A_\psi is the view-dependent residual (the thing being compared), and ϕ\phi is a colour activation. Standard 3DGS uses ϕ(c^)=ReLU(c^+0.5)\phi(\hat{\mathbf c}) = \mathrm{ReLU}(\hat{\mathbf c} + 0.5). Only Aψ\mathcal A_\psi changes between models. Base colour, its initialisation from the SfM point colours, its densification handling: all identical.

Spherical harmonics

For SH, the residual is a weighted sum of fixed basis functions on the sphere, with one RGB weight per function:

ASH(d)=∑ℓ=13∑m=−ℓℓYℓm(d) cℓm\mathcal A_{\mathrm{SH}}(\mathbf d) = \sum_{\ell=1}^{3}\sum_{m=-\ell}^{\ell} Y_\ell^m(\mathbf d)\,\mathbf c_\ell^m

Degree ℓ\ell contributes 2ℓ+12\ell+1 functions, so degrees 1 to 3 give 3+5+7=153+5+7 = 15 functions, times three colour channels: 45 numbers. Add the base colour and you have the 48 fp32 values, 192 bytes, that dominate a 3DGS file. The basis functions are polynomials in the components of d\mathbf d of degree at most 3. In the CUDA that is all there is, a straight-line expression over x,y,zx, y, z (appearance_sh.cuh:43-61):

// FasterGSVDACudaBackend/.../rasterization/rtc/appearance_sh.cuh
float3 result = -sh_c1 * y * coefficients_ptr[0]
              + sh_c1 * z * coefficients_ptr[1]
              - sh_c1 * x * coefficients_ptr[2];
// ... degree 2: five more terms in xy, yz, zz, xz, xx - yy ...
// ... degree 3: seven more terms, cubic in x, y, z ...
return residual_activation(result);

That explains both of SH's habits. It is cheap and trains easily: everything is linear in the coefficients, and 3DGS switches one degree on every 1,000 iterations so view dependence comes in late. It is also band-limited. A cubic polynomial on the sphere cannot make a highlight that is bright over a few degrees and dark elsewhere; it can only make a broad bump with ripples around it. The paper's other complaint is about cost: SH coefficients "dominate per-primitive storage and memory traffic", and Adam has to update all 45 of them for every Gaussian every step, including ones not visible in that iteration's image. The paper cites its own earlier Faster-GS measurement that SH updates alone can take up to 50% of optimisation time at high primitive counts.

Spherical Voronoi: a soft partition of the sphere

Spherical Voronoi (SV) places KK sites si\mathbf s_i on the sphere, each with a colour ai\mathbf a_i and a temperature τi\tau_i, and blends the site colours by a softmax over distance to the view direction:

ASV(d)=∑iexp⁡(−τi∥si−d∥)∑jexp⁡(−τj∥sj−d∥) ai\mathcal A_{\mathrm{SV}}(\mathbf d) = \sum_i \frac{\exp(-\tau_i \lVert \mathbf s_i - \mathbf d\rVert)}{\sum_j \exp(-\tau_j \lVert \mathbf s_j - \mathbf d\rVert)}\,\mathbf a_i

With high temperatures this is a hard Voronoi partition: look from inside one cell and you see that site's colour. It can make sharp angular edges SH cannot. Two details matter. The softmax weights sum to one, so a single site gives a constant and you need at least two for any view dependence. And each site costs 7 floats (direction, temperature, RGB), so the default seven sites are 49 floats plus the base colour: 208 bytes, more than SH. The implementation here improves on the reference by computing the softmax in one pass with a running maximum (appearance_sv.cuh:41-56), reading each site once instead of three times.

NASG and NASGabor: one sharp, stretched lobe

Normalized anisotropic spherical Gaussians come from Miazga et al., who adapted them from neural path guiding. A lobe has a frame (axis z\mathbf z, tangent x\mathbf x), a sharpness λ\lambda and an anisotropy aa. With vz=d⋅zv_z = \mathbf d\cdot\mathbf z and vx=d⋅xv_x = \mathbf d\cdot\mathbf x, the code's own header comment (appearance_nasg.cuh:6-14) spells it out:

K=12(vz+1),Ke=ϵ+a vx21−vz2,E=KKe,N=λ1+a2π (1−e−2λ)K = \tfrac{1}{2}(v_z + 1),\quad K_e = \epsilon + \frac{a\,v_x^2}{1 - v_z^2},\quad E = K^{K_e},\quad N = \frac{\lambda\sqrt{1+a}}{2\pi\,(1 - e^{-2\lambda})} G(d)=e2λ(EK−1) E N,ANASG(d)=∑iGi(d) tanh⁡(ai)G(\mathbf d) = e^{2\lambda(EK - 1)}\,E\,N, \qquad \mathcal A_{\mathrm{NASG}}(\mathbf d) = \sum_i G_i(\mathbf d)\,\tanh(\mathbf a_i)

Read it slowly and it is a spherical Gaussian e2λ(K−1)e^{2\lambda(K-1)}, peaked on the axis, whose width is squeezed along x\mathbf x by the EE factor when a>0a > 0. NN normalises it so sharpening a lobe does not change how much colour it contributes overall. The tanh⁡\tanh bounds each lobe's contribution so the residual cannot run away from the base colour. A lobe is 8 numbers: three frame angles (stored as cosines through a tanh⁡\tanh), λ\lambda, aa and RGB. The paper uses one lobe, the reference default: 4×(3+8)=444\times(3+8) = 44 bytes.

NASGabor multiplies the lobe by a non-negative cosine carrier along the tangent, 12(1+cos⁡(k vx))\tfrac12\big(1+\cos(k\,v_x)\big), with k=20(tanh⁡(kraw)+1)k = 20(\tanh(k_{\mathrm{raw}})+1), so between 0 and 40 (appearance_nasgabor.cuh:13). One extra number buys a lobe that can have several bright bands, a cheap way to get a multimodal reflection from one kernel. 48 bytes.

The whole difference between SH and a lobe is easier to see than to read. The widget below takes one Gaussian and looks at it from every direction around a circle. On a great circle, real SH up to degree LL span exactly the Fourier modes up to order LL, so the least-squares SH fit is a truncated Fourier series; the lobe fit is a 1D cut of a spherical-Gaussian lobe, fitted by exhaustive search. Slide the sharpness up.

one Gaussian, seen from every direction on a circletoy fit · bytes from Table 1

Target: a diffuse colour plus one reflection that sharpens as the slider rises.

lobes
00.510°90°180°270°360°as rendered (clamped to 0..1): target · SH · lobes
SH degree 3: PSNR 23.8 dB · peak 0.63 vs 0.85, floor 0.30 vs 0.35
1 lobe: PSNR 51.0 dB · λ found 20.2
appearance bytes per Gaussian, RGB, as stored (base colour included)SH degree 3192 B1 NASG lobe44 Bneural (paper)28 B

At the default sharpness (λ≈20\lambda \approx 20, a highlight roughly 30 degrees wide), degree-3 SH reaches a peak of 0.63 where the target is 0.85, and smears the rest of that highlight around the circle; one lobe fits it to within the search grid. Push λ\lambda to the right and SH gets worse while the lobe barely notices, which is the point of NASG. On the full sphere degree-3 SH stores 16 functions per colour channel, of which this circle exercises only 7; the bytes bar counts all 16. Then try the third target, which has no reflection in it at all. Hold that thought.

The neural model: per-Gaussian codes, one shared decoder

The paper's own model drops the idea of a per-Gaussian basis expansion. Each Gaussian stores a feature vector f∈R8\mathbf f \in \mathbb R^8, and one small MLP, shared by every Gaussian in the scene, maps features plus view direction to a residual:

x=(γdir(d), γfeat(f))∈R32,Aneural(d)=tanh⁡(MLPθ(x))\mathbf x = \big(\gamma_{\mathrm{dir}}(\mathbf d),\ \gamma_{\mathrm{feat}}(\mathbf f)\big) \in \mathbb R^{32}, \qquad \mathcal A_{\mathrm{neural}}(\mathbf d) = \tanh\big(\mathrm{MLP}_\theta(\mathbf x)\big)

The feature encoding is a one-frequency sinusoid per feature, γfeat(f)=(sin⁡f,cos⁡f)\gamma_{\mathrm{feat}}(f) = (\sin f, \cos f), so 8 features become 16 inputs. The direction encoding is the part I did not expect: γdir\gamma_{\mathrm{dir}} is the SH basis of degrees 0 to 3, evaluated at d\mathbf d, which is 1+3+5+7=161+3+5+7 = 16 more inputs. The kernel literally calls eval_sh_degrees(direction, ...) to fill the second half of the MLP input (appearance_neural.cuh:332). Spherical harmonics have not gone anywhere. What is gone is the 45 per-Gaussian weights. The basis values are computed on the fly and cost nothing to store, and the weighting is done by a network that is shared, so the per-Gaussian part shrinks to an 8-number code that tells the shared network which kind of surface this is.

The sizes are chosen by the hardware. Tensor Cores want the input width to be a multiple of 16; the degree-3 direction encoding is already 16, so 32 is the smallest width that leaves room for features, and 16 inputs fit 8 features at one frequency. Neural.py:35 warns if the width is anything but 32. The network is 32 to 16 to 16 to 3, ReLU, no biases (the constant degree-0 SH input does the job of one), so the whole scene's decoder is 32⋅16+16⋅16+16⋅3=81632\cdot16 + 16\cdot16 + 16\cdot3 = 816 weights. The output layer starts at zero (Neural.py:53) and so do the features, so the residual is exactly zero until training gives it something to say, and the direction bands switch on one per 1,000 iterations on the same schedule as SH.

The footprint arithmetic is worth doing yourself, because the paper's 28 is a deployed number:

modelper-Gaussian appearancebytes
none (base colour only)3 fp3212
SH degree 348 fp32192
Spherical Voronoi, 7 sites3 + 7 × 7 fp32208
NASG, 1 lobe3 + 8 fp3244
NASGabor, 1 lobe3 + 9 fp3248
neural3 fp32 + 8 fp1628

Every row matches Table 1. The neural row assumes the features are stored in fp16, which the paper justifies: the fused forward and backward run in half precision, so deploying in 16 bits loses nothing. During training, though, the features are an fp32 tensor (Neural.py:70), so the training-time footprint is 44 bytes, the same as one NASG lobe. That matters for the memory budget later.

Why it is fast: the decoder lives inside the rasterizer

A per-Gaussian MLP sounds expensive, and the paper says that belief is why nobody did this. The fix is fusion. 3DGS renders in two kernels: a per-primitive preprocess that projects each Gaussian, counts the tiles it touches and computes its colour, then a per-tile blend. Every appearance model here is evaluated inside the preprocess kernel, only for Gaussians that survived culling, with no separate launch and no round trip through global memory.

For the neural model this uses tiny-cuda-nn's fully fused MLP, which runs one warp per 32 Gaussians on Tensor Cores in fp16. That has a consequence you can see in the kernel. A Tensor Core matrix multiply needs every lane of the warp, so the network has to run before any thread is allowed to exit, even for Gaussians that will turn out to touch no tiles (preprocess_forward.cu:202-216):

#if JIT_APPEARANCE_MODEL == 5
    // residual mlp evaluation where all lanes of surviving warps must participate (warp-cooperative mma)
    tvec<__half, 3> mlp_output;
    const float3 residual = neural_residual(appearance, mean3d, cam_position[0],
                                            fwd_ctx, mlp_output, primitive_idx, appearance_degree);
#endif
    // cooperative threads no longer needed
    if (n_touched_tiles == 0 || !active) return;

Warps that are culled entirely zero their slice of the saved activations before leaving, because the backward pass runs the MLP backward for every thread and "0 x garbage may be NaN" (the comment on zero_warp_fwd_ctx). Weight gradients accumulate in fp16 with a loss scale of 128, a power of two, so undoing it is an exponent shift with no rounding.

How much fusion buys is in the paper's ablation (Table 8): the same model through the unfused reference path trains in 16m08s instead of 4m17s and renders at 112.1 FPS instead of 443.6, with the same quality. A neural appearance model that runs as a separate PyTorch pass is a different, much slower object.

The plumbing is a small piece of good engineering. Each model is one CUDA chunk with a *_residual and a *_residual_backward device function. At first use, assemble_rtc_source (appearance.cu:78-140) writes a preamble of #define JIT_APPEARANCE_MODEL, the number of lobes or sites, the activations and the MLP widths, pastes the chosen chunk and the tiny-cuda-nn device functions in front of the preprocess kernel, and compiles the lot with NVRTC. Because every configuration value is a compile-time constant, the specialised kernel has no branches on model type. Adding a model is one Python class and one CUDA chunk, which is what the thread advertises. Along the way a tiny-cuda-nn bug turned up (the acknowledgements credit Jannis Möller with finding it): its JIT backward pass evaluated hidden-activation derivatives on post-activation values, f′(f(x))f'(f(x)) instead of f′(x)f'(x). It is harmless for ReLU, whose derivative depends only on sign, and wrong for every smooth activation. setup.py:86-110 patches the vendored header at install time, with the reasoning in a comment.

nerficg-project/efficient-gaussian-appearance@4efbe67 · snapshot 2026-10-07
tracked files
663
license
Apache-2.0
branch
main
tests
none found
source
673.9 kB
commit date
2026-09-30
source by language
CUDA244.5 kB(26)Python214.1 kB(22)JavaScript105.9 kB(3)C74.5 kB(20)GLSL26.4 kB(7)HTML4.5 kB(1)CSS4.1 kB(1)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

Read at 4efbe67 (30 September 2026). Apache-2.0. The five appearance models are rasterization/rtc/appearance_*.cuh (forward and backward each) plus Appearance/*.py; the JIT assembly is rasterization/src/appearance.cu; the WebGL viewer is viewer/. Needs an NVIDIA GPU with compute capability 7.5+ and CUDA 11.8+ for the fused kernels, and installs into the NeRFICG framework.

local clone, 2026-10-07 at 4efbe67 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

What the tables say, and what the average hides

The comparison is controlled in the way that matters: MCMC densification with the same scene-specific primitive budgets for every model, the same 30k iterations, PPISP on real scenes so per-image exposure drift is not learned as view dependence, and five training runs averaged for image metrics. Timings and memory are from the first run on an RTX 4090.

six appearance models, one pipelinepaper Tables 1, 6, 9
PSNR (dB) · higher is better · bars start near the lowest valueNone27.73SH28.36SV28.64NASG28.50NASGabor28.54Neural28.55
None 72 MBSH 1152 MBSV 1248 MBNASG 264 MBNASGabor 288 MBNeural 168 MB

On Mip-NeRF 360 the PSNR order is SV 28.64, neural 28.55, NASGabor 28.54, NASG 28.50, SH 28.36, and no view dependence at all 27.73. SSIM goes the other way for the neural model: SH, NASG and NASGabor share 0.836 and the neural model is last of the view-dependent ones at 0.831. The paper says as much, and its summary table calls the neural model "hard to regularize per scene". On Tanks and Temples plus Deep Blending the neural model has the best PSNR (28.13 against 27.96), and on NeRF Synthetic it matches SH (34.04 against 34.01) while the lobe models fall behind with one lobe each.

The cost side is clearer than the quality side. On Mip-NeRF 360, SH trains in 6m12s, NASG in 3m54s, NASGabor in 4m00s and the neural model in 4m17s; peak memory falls from 6.7 GiB to between 4.9 and 5.3. The "1.3× faster" in the abstract is the conservative end: 6m12s over 4m17s is about 1.45×, and the smallest of the three dataset ratios, NeRF Synthetic's 1m22s over 1m02s, is about 1.32×. The lobe models train faster still. SV, despite its quality, is slower than SH and the largest of the lot.

Rendering tells the same story. In CUDA on the RTX 4090 (Table 2, 720p), NASG and NASGabor are the fastest: 2.25 ms against SH's 2.75 ms on the 6M-Gaussian bicycle, the "up to 18% faster" in the text. The neural model is third. In the WebGL viewer, which has no Tensor Cores, the differences mostly vanish: on an iPhone with an A18 Pro the bicycle takes 45.0 ms with SH and 45.1 ms with the neural model, and SV runs out of memory on the two largest scenes.

Outdoor and indoor are different stories

The Mip-NeRF 360 average mixes five outdoor scenes and four indoor ones. Appendix Table 9 gives every scene, so I split them. The view-dependent models against SH, mean PSNR:

modeloutdoor (5 scenes)indoor (4 scenes)
none25.05 (−0.44)31.08 (−0.87)
SH25.4931.95
SV25.38 (−0.10)32.71 (+0.77)
NASG25.38 (−0.11)32.40 (+0.46)
NASGabor25.40 (−0.09)32.47 (+0.52)
neural25.33 (−0.15)32.56 (+0.61)

Outdoors, SH beats every alternative, and turning view dependence off entirely costs less than half a dB. Indoors, every alternative clears SH, by 0.46 to 0.77 dB. The "matches or exceeds SH in PSNR on average" headline is an average of a loss and a win. The authors do say this, in the appendix: outdoor scenes are "dominated by diffuse content and their remaining error is largely geometric", while indoor scenes have glossy surfaces under controlled lighting. It deserves to be in the abstract, because whether you should switch depends almost entirely on which of those two you capture.

The geometry effect is the most convincing evidence that the expressive models are doing something real. Figure 2 shows the bonsai scene's sketch block. With SH, the reflection on the block is faked by elongated white Gaussians, which in an extrapolated view become streaks; with NASG or NASGabor the reflection sits in the residual and the base colour is a clean diffuse block. This is the same failure I described in the reconstruction roundup: when the colour model cannot make a spike, the optimiser makes matter instead.

A grid of crops from the bonsai scene for six appearance models (None, SH, SV, NASG, NASGabor, Neural) and the reference photo. Rows show the full render, the base colour, the view-dependent residual and an extrapolated view. The residual row is black for None, faint for SH, and shows a bright white chevron reflection on the sketch block for SV, NASG, NASGabor and Neural. In the extrapolated row, SH shows long bright streaks across the block where NASG and NASGabor show a cleaner surface.
Base colour, residual and an extrapolated view for each model on bonsai. SH fakes the sketch-block reflection with elongated white Gaussians (the streaks in the bottom row); NASGabor puts it in the residual and gives the cleanest diffuse base. The residual is shown as the absolute difference between full and base colour, because a signed residual through ReLU would hide its negative half. (paper, Figure 2.)

One more number belongs here. With a tenth of the primitive budget (Table 7), SH, NASGabor and the neural model land at 27.30, 27.32 and 27.30 dB. When geometry is the bottleneck, no amount of angular expressivity helps.

Footsteps in the grass

This is the part of the thread I would have put first. Posts 6 to 8 are about content that changed during capture. In the Mip-NeRF 360 garden scene, the camera operator stepped on the grass around the table. Rendered top-down with SH, the residual shows a ring exactly where the footsteps went.

Three top-down renders of the garden scene: a round wooden table on hexagonal paving surrounded by lawn. Full and Base look almost identical. The Residual panel is mostly dark, but shows a bright, grainy ring of grass circling the paved area at a fixed distance from the table.
Full render, base colour and SH residual for a top-down view of garden. The ring in the residual is where the camera operator stepped on the grass during capture: a change in the scene, fitted as view-dependent colour. (paper, Figure 5.)

The mechanism is simple once you see it. A capture is a trajectory, so viewing direction and time are correlated. If the grass was upright when the camera was on the north side and flattened by the time it reached the south, then "colour depends on direction" explains the photos perfectly well. The third target in the widget above is exactly this: no reflection, just a change halfway round. Degree-3 SH fits it fine, because a smooth step is low-frequency. An expressive model fits it too, and in the bonsai scene, where a book moved between shots, NASGabor fits it better than SH can.

Two reference photos of the bonsai room with a red circle around a grey notebook lying on other books on the floor; the notebook is in a different position in each photo. Between them, a grid of crops of the books along a camera path, in rows for SH full, base and residual and NASGabor full, base and residual. The NASGabor residual row shows a faint outline of the notebook, while SH's notebook jumps between positions.
A notebook moved during the bonsai capture. Along a path between the two reference views, NASGabor partly fits the move as a view-dependent residual; SH cannot, and the geometry compensates by popping. (paper, Figure 4.)

The paper's conclusion is careful and I agree with it: on real benchmarks, part of the advantage of an expressive model may be "robustness to violations of the static-scene assumption" rather than better modelling of outgoing light. Better PSNR on held-out views of an inconsistent scene does not mean a more faithful scene. Synthetic benchmarks avoid the problem but bring their own bias, because they are mostly diffuse hard surfaces. Thread post 8 asks the open question plainly: should we reconstruct one specific state, a motion-blurred version, or a dynamic model? None of the three is easy from current data, and the authors say they do not have the answer.

I think there are two practical readings. If you render for people, an expressive residual that soaks up a passing shadow is often what you want; it keeps the geometry clean, and the thread and project page point out that methods that cannot hide errors by popping (sorted-per-pixel renderers like HTGS, opaque-surface representations like MeshSplatting) should benefit most. If you measure from the splats, as in a BIM or earthwork workflow, the residual is also a free diagnostic. Render it on its own, the way Figure 5 does, and it shows you where the scene changed under the camera. This repository's configs expose it directly as RENDER_BASE_COLOR and RENDER_RESIDUAL_COLOR.

Where it breaks

The authors list the limits themselves, and two of them matter in practice. None of these models can do a mirror; that still needs secondary rays. And the neural decoder cannot be rotated analytically the way an SH expansion can, so moving or composing objects needs the inverse transform applied to the view direction before the network sees it. A single shared decoder may also run out of capacity on a large scene; the paper suggests a grid of decoders interpolated by camera position, as in SMERF. I would add that it does not always decompose cleanly. On the Dr Johnson scene (Figure 6) the neural model puts the desk's highlight into the base colour, so the "diffuse" render still has a highlight baked in.

The architecture ablation (Table 8) is reassuring about the decoder itself. One hidden layer, three, or 32 neurons all land within about a quarter of a dB of each other, and the differences sit in the indoor scenes. Doubling the per-Gaussian code to 16 raw features is the only variant that costs storage (44 bytes), and it does no better than a deeper decoder. Capacity in the shared network is free; capacity per Gaussian is paid by every Gaussian. That asymmetry is the whole design.

What I would change in a 360-camera capture pipeline

The site owner's own pipeline is a 360 camera, Spirula Studio or LichtFeld for SfM and training, and a web viewer at the end, often through SOG. This is how I would use the paper there.

Training memory is where the alternatives pay off. Using the same naive count as before (four fp32 copies of every trainable float: value, gradient and two Adam moments), a Gaussian here has 11 geometry floats plus its appearance: 59 floats at SH3, 38 at SH2, and 22 for one NASG lobe or the neural model in training precision. At 15M splats that is 14.16 GB, 9.12 GB and 5.28 GB of optimiser state, before images and scratch buffers. A single-lobe model costs less than the SH2 run that the dam capture had to settle for, and keeps sharper highlights than SH3. The measured peaks in Table 1 move the same way, 6.7 GiB for SH against 4.9 for NASG on Mip-NeRF 360.

Delivery is where SH still wins. The paper mentions it in passing; the author spelled it out in a reply. The SOG format vector-quantizes the 45 SH-rest values of every splat into a shared palette of up to 65,536 entries and stores one 2-byte index per splat. The paper's own WebGL viewer packs SH into 40 bytes with Spark's bit-packed layout, against 32 bytes for the neural payload. When someone asked Hahlbohm what blocks mainstream adoption, his answer was the lack of NPU access from WebGL and WebGPU for the neural model, and, for the lobe models, "how much more explored compression/quantization of SH-based 3DGS is". The paper itself stores the lobe models in plain fp16 in its viewer and leaves comparable quantization schemes for them to future work.

The export path makes the same point harder. convert_to_ply for any non-SH model writes only the base colour, with a warning that the view-dependent residual is not included (Base.py:230-236). A NASG or neural splat loaded into a standard viewer is a diffuse splat. If you train with one of these models today, you either ship through this repository's viewer, which uses its own .ngsplat format, or you accept losing the view dependence at the boundary.

So my rule of thumb from these numbers:

On hardware: the fused kernels need an NVIDIA GPU with compute capability 7.5 or newer and CUDA 11.8 or newer, because they depend on tiny-cuda-nn's JIT fusion. Spirula Studio runs on Vulkan, so none of that carries over. Porting a lobe model to another trainer is a few dozen lines of scalar maths per direction; the forward is visible in eval_nasg.glsl in this repository's viewer. Porting the neural model means writing a fused cooperative MLP for that backend, which is a real project.

How I checked

I shallow-cloned nerficg-project/efficient-gaussian-appearance at commit 4efbe67 and read the five rasterization/rtc/appearance_*.cuh chunks with their backward passes, preprocess_forward.cu, appearance.cu, setup.py, the Appearance/*.py classes, the six top-level configs and the per-scene paper_configs/, and the viewer's shaders and README. I did not build or run any of it. The byte counts in the table come from the parameter layouts in the code and match the paper's Table 1. The outdoor and indoor averages are mine, computed from the per-scene PSNR in Table 9; their nine-scene mean reproduces Table 1's Mip-NeRF 360 column. The training speed ratios are mine from Table 1. The training-memory figures are my own arithmetic, the same naive four-copy count as in the capture-pipeline article, not a measurement. The 816 weights, the 32-wide input and the 16-to-16 layers are from Neural.py and the configs, and agree with the paper. I read the thread and its replies through the fxtwitter mirror; Hahlbohm's answer about adoption is a reply in that conversation. The widget's curves are a toy on one circle, fitted in the browser; its bytes are Table 1's.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Gaussian-splat colour: 192 bytes of spherical harmonics, or 28 bytes and a tiny MLP", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026efficientgaussianappearance,
  author = {Satyajit Ghana},
  title  = {Gaussian-splat colour: 192 bytes of spherical harmonics, or 28 bytes and a tiny MLP},
  url    = {https://ai.thesatyajit.com/articles/efficient-gaussian-appearance},
  year   = {2026}
}
share