2026-10-07 · 26 min · 3d · flow-matching · diffusion-transformers · lora · synthetic-data · benchmarks
Why read this
Solidtop 85%How DistScene splits a room across voxel slots and refines each object, checked against the authors' 23 result files: per-object scale and shift, no rotation.
- Original analysis
- A new technique
- Explained from first principles
3D & spatialNothing to runResearch paper
How this was scored
- Is it new?
- 2 of 3: A real new idea, method or capability
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 0 of 3: Closed, nothing to run
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 1 of 3: General advice
- Will it last?
- 1 of 3: Relevant for months
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 51 of 100, ranked 344 of 454 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
A post from @wildmindai described DistScene in two lines: one image in, an editable 3D scene out, built on TRELLIS.2 with sparse voxels, sparse DiT flow models and LoRA, plus "object-centric refinement" to restore details. The second line is what caught me. Eleven days ago, in the first 3D roundup, I wrote that holistic single-image scene methods make every object compete for one token budget, so small things come out mushy. DistScene is a holistic method that admits the problem and then spends a second pass fixing it. I wanted to see how the two passes fit together, and what "editable" turns out to mean once you open the files.
The paper is arXiv 2610.06960, from Ping Tan's group at HKUST. The project page says code, checkpoints and dataset are "come soon", so there is nothing to run. There is something better than a demo, though: a Hugging Face repo, coolbeam/DistScene-preview-assets, with 23 full-resolution output scenes as GLB files of 222 MB to 985 MB each. A GLB keeps its scene graph in a JSON chunk at the front of the file, so I read just that chunk from each with HTTP range requests. Several of the numbers below come from those headers rather than from the paper.

What TRELLIS.2 already does, and the one thing it doesn't
Everything here sits on TRELLIS.2, Microsoft's MIT-licensed 4B-parameter image-to-3D model. DistScene changes very little of it, so the details matter.
An asset in TRELLIS.2 is a set of active voxels on a regular grid, each carrying shape and material features ("O-Voxel" in their terms). A sparse VAE compresses that with 16x spatial downsampling, so a 1024-resolution asset becomes features on at most a 64 by 64 by 64 grid of latent voxels, and a 512 one on a 32 grid. Generation then runs three flow-matching transformers, each about 1.3B parameters, in order. The first is dense: it denoises a 16 by 16 by 16 latent with 8 channels, 4,096 tokens, that decodes to an occupancy grid saying which voxels exist. The second is sparse: it generates shape features only on those active voxels. The third generates material features on the same voxels, conditioned on the shape. Each sampler takes 12 Euler steps in the released pipeline.json. The DistScene paper writes this whole recipe as one line:
Here is the position of active voxel , and are the sparse VAE's encoder and decoder, and is the image-conditioned generator.
- task
- image-to-3d
- library
- trellis2
- license
- mit
- safetensors
- 9 shards
- largest file
- 2.58 GB
- files
- 22
- downloads
- 1.9M
- likes
- 1.3K
- languages
- en
The base DistScene adapts. Three 1.3B flow transformers (sparse structure, shape, material) over a 16x-downsampled sparse VAE; 512 to 1536 output resolution.
repo last modified 2025-12-27
What TRELLIS.2 does not do is place anything. Its output lives in a canonical cube: centred, normalised, one object. Hand it a whole room and you get one fused asset, with no notion of which voxels belong to which chair. The appendix of the DistScene paper says as much about using TRELLIS.2 on a segmented multi-object image: "all objects are generated as a single fused object-centric representation rather than as independently indexed scene components."
Many slots, one coordinate frame
DistScene's first stage, Scene-Frame Generation, keeps all three TRELLIS.2 stages and changes what goes into them. Instead of one noise latent, it starts from : one per object and one for the environment, the room shell or the street with its floor, walls and terrain. Each latent is flattened into a token sequence and the sequences are concatenated:
Learnable type embeddings mark which tokens are environment and which are object, and the transformer's self-attention runs over the whole concatenated sequence, so every slot sees every other slot at every denoising step. All the slots share one coordinate frame, the scene's cube, and that does most of the work. The occupancy stage decides which voxels of the shared grid each slot owns, and because positions are in the scene frame, the decoded meshes come out already arranged. Nothing estimates a pose afterwards.

The appendix figure is more precise about which pieces are frozen. Each stage is a frozen TRELLIS.2 transformer (the snowflake) with a trainable LoRA adapter (the flame), and the colours track one component through all three modules.

How many slots? Training caps a scene at 30 components including the environment, padded with empty slots so the model learns that some slots stay empty. Inference uses 20 slots by default and throws away any that decode to nothing. The authors blame the cap on "GPU memory limitation", and the arithmetic says why. If the occupancy stage keeps TRELLIS.2's 16-cubed latent, each slot costs 4,096 tokens whether it holds a sofa or nothing, and 20 slots is 81,920 tokens in one dense self-attention. The figure shows each slot as a full-size cube with one small object in a corner: every object pays for the whole room's worth of empty space at this stage.
One thing I couldn't pin down. The main text says type embeddings distinguish "object tokens and environment tokens"; the appendix says they distinguish "their corresponding slots", and the figure draws a different colour per slot. If all object slots shared one embedding and one set of positions, only their noise would tell them apart. That can still work, since the noise breaks the symmetry between slots, but which version they built isn't stated.
The budget problem, in numbers
After stage one, each object is real geometry in the right place, but at the scene's resolution. The paper says this in one sentence: "each object occupies only a small fraction of the global sparse voxel grid." I wanted the fraction, so I took it from the result files.
Every object node in the 23 GLBs has the same matrix shape. Written column-major it is [s,0,0,0, 0,0,-s,0, 0,s,0,0, tx,ty,tz,1]: a uniform scale and a translation, wrapped around a fixed swap from Z-up to Y-up, and nothing else. The object's own mesh fills a unit cube (its longest side measures 0.986 to 1.004 across all 122 objects), and the environment node has . So is simply the object's longest side as a fraction of the scene. Across the 122 objects it runs from 0.045 to 0.666, with a median of 0.205.
Now put that on TRELLIS.2's grid. At 1024 resolution the scene frame has 64 latent voxels across, so an object of size gets of them along its longest side. The median object gets about 13. The armchairs in scene C000, at 0.138, get about 9. The smallest object in the set, in scene C006 at 0.045, gets 2.9 latent voxels, or about 46 voxels of output geometry across its longest side. A plant or a lamp at that size is a few blobs. Refinement re-voxelises each object in a cube of its own, so it gets the full 64, and the gain per axis is : 4.9x for the median object, about 22x for the smallest.
The widget draws scene C000 from above with the real bounding boxes from its GLB. Pick an object, switch between "as the scene frame made it" and "after object-centric refinement", and the right panel shows how many grid cells that object had to work with. The shapes inside the boxes are rounded rectangles, not the meshes; the cell counts depend only on size, so the stand-in is fair for this question.
The files hold one more surprise. Every component, environment included, ships with between 1,845,406 and 1,999,127 triangles and two 4096-pixel PNG textures. A 0.14-wide armchair in C000 has 1,964,991 triangles; the whole room around it has 1,919,356. That flat cap looks like a decimation target in the export step rather than anything the geometry asks for, and it means 84% of the triangles in the 23 files sit on objects. It also explains the file sizes: C000 is 17.4 million triangles and 985 MB. The page's own preview copies are about a tenth of that, and the project page says so.
Refinement: cut out, re-generate, put back
Object-Centric Refinement takes each object out of the scene, re-centres it and scales it uniformly into a canonical cube, voxelises that, and runs a third sparse flow transformer on the new support:
The conditioning is what makes this more than an upsampler. sees the input image , the scene-frame latents of every component , frozen from stage one, and a learnable marker embedding added to the tokens of the object being refined. Without the marker, the model would have a crisp local cube and a whole scene of context and no way to know which of several similar chairs in the photo it is supposed to be drawing. Only the target's local latent is generated; the scene latents never change. The refined mesh then goes back through the inverse transform, the recorded scale and translation.
The appendix says the normalisation "preserves the object's orientation", and the GLBs agree: no object node has a rotation. Keeping the orientation means refinement never has to decide which way an object faces. The chair stays turned however the scene frame turned it, and only the voxel size changes.
The ablation shows the scene context matters more than the extra resolution. With scene context removed from refinement, scene-level numbers get worse than not refining at all: MIDI-test scene F-score drops from 68.86 to 66.54, and Gen3DSR-test from 73.21 to 69.82. Object-level numbers improve only slightly without context (object F-score 40.98 to 41.97). With context, the full model reaches 71.59 and 79.68 at scene level and 51.64 for objects. The paper's explanation is that without context "the refinement model has difficulty associating the target component with its corresponding object in the input image." A sharper object that looks like the wrong chair is worse than a blurry right one.

How is trained is described two ways, and they don't quite match. The main text builds coarse-to-fine pairs by degrading a good latent: downsample it, run it through the VAE, upsample it back, and learn to recover the original with the same flow-matching loss. The appendix's Algorithm 1 instead trains to recover each object's canonical shape latent, the one TRELLIS.2 produced before the object was placed in the scene, from the scene-frame latents and the canonical support. Both are cheap, because the data engine keeps the canonical latents for free. Which one, or which mix, the released weights use, I can't tell until the code ships.
What the LoRA adapts
The paper's argument for LoRA is a single sentence: the sparse-voxel generator "already possesses strong image-to-3D priors and only needs to adapt to the compositional setting." Every generation module (occupancy, scene-frame shape, refinement, texture) starts as a copy of TRELLIS.2, keeps its backbone frozen, and gets its own rank-32 adapter. Training uses AdamW at a learning rate of , a global batch of 64, and 20,000 iterations per module, all at 512 resolution. The loss is TRELLIS.2's flow-matching objective, summed over the slots:
So what does the adapter actually have to learn? In the occupancy stage, to split one image among slots: attend across slots so that two of them don't both claim the sofa, and so that the room's slot takes the walls and floor and leaves the furniture alone. In the shape stage, to generate a component that sits off-centre in a frame it doesn't fill, which TRELLIS.2 never saw. In refinement, to use a new kind of conditioning, the scene tokens plus a marker. None of these needs new knowledge of what a chair looks like; that stays in the frozen weights. The type and marker embeddings are new parameters on top of the adapters.
The paper doesn't say which layers carry adapters, so the size is a guess. If every linear layer in a TRELLIS.2 block gets one (the self-attention projections, the cross-attention projections that read the image features, and the two MLP layers), rank 32 across 30 blocks of width 1536 comes to about 37M parameters per module, under 3% of each 1.3B transformer. That is my estimate from the released configs, not a figure from the paper.
One detail the authors flag: trained only at 512, the adapted model also runs at 1024, and every headline number uses 1024. TRELLIS.2 ships separate 512 and 1024 shape transformers, and the paper doesn't say which checkpoint each adapter starts from, so I can't tell whether this is the adapter generalising or the backbone doing the work.
Where the training scenes come from
The title's "distillation" is about data, not a student model. There isn't a scene dataset with clean, separately meshed objects at the scale this needs (the paper calls 3D-FUTURE "limited in both scale and object diversity"), so the authors make one out of TRELLIS.2 itself.

A language model writes a scene description: which objects, roughly what size. A text-to-image model draws each object and an empty environment, and TRELLIS.2 turns each drawing into a complete mesh. The room comes from the same generator as the chairs, which is neat: no separate background model. Then a placement loop drops each object onto a valid support surface, rejects anything floating or sunk into the floor, filters object-object collisions with an axis-aligned bounding-box test, inspects the mesh where boxes overlap, nudges or discards, and checks objects against walls too. Accepted scenes are rendered from many viewpoints, keeping only views where every object is visible, and re-encoded in the shared frame. The canonical latents from before placement are kept as refinement targets.
The result is about 125k scenes from 70k objects and 46k empty environments, 106k indoor and 18k outdoor (those two add to 124k; I assume rounding). Neither the language model nor the image model is named.
The ablation that isolates the data is the most convincing table in the paper. The same object-only architecture, no environment slot, trained on DistScene's synthetic scenes or on MIDI's own training set, then tested on MIDI-test, which is in-domain for MIDI's data and out-of-domain for DistScene's. The synthetic data still wins on scene-level numbers there (scene F-score 67.94 against 65.74), loses on object-level ones as you'd expect, and wins by a lot on Gen3DSR-test (62.04 against 46.68). A generator taught entirely on its own outputs generalising better than one taught on designer-made layouts is a result I'd like to see reproduced.
The same engine also explains the most visible failure. An object generator rarely produces a fully enclosed empty room, so the indoor environments in training mostly have no ceiling, and DistScene learns that rooms are open-topped.

What "editable" means
The post's word, and the paper's, is "editable". The appendix uses it once: DistScene provides "directly editable component-level geometry." There are no editing experiments, so I went to the files.
Each GLB has a world node whose children are one env and one objNNN_in_scene per object, each with its own mesh, its own material and its own pair of textures. Across the 23 scenes that is 122 objects, 1 to 9 per scene, median 5, and always exactly one environment. In Blender, every one of them is a separate selectable object. Moving the armchair means changing tx and tz in its matrix. Scaling it means changing s. Deleting it leaves the floor underneath intact, because the floor belongs to the environment mesh, not to the chair. One reply under the original post guessed that "if the objects come out as separate meshes, blender kitbashing gets way faster". They do, and it would.
What you don't get is just as concrete. There is no rotation in any object's transform, so the model never states which way a chair faces; that is baked into the mesh. Nothing re-checks a moved object: the contact and collision tests ran in the data engine on training scenes, not at inference, so drag a chair into a wall and the file accepts it. And each texture was generated conditioned on the original scene. Edit the layout and you have consistent pieces in an arrangement the model never saw, the same as hand-placing assets. Try dragging an object in the widget above into another: the only thing that changes is two numbers in its node matrix.
Prior methods differ on exactly this axis. The paper says Extend3D "does not explicitly represent the scene as independently addressable components", and TRELLIS.2 on a whole scene gives one fused asset. Per-object pipelines such as SAM 3D and 3D-Fixer do give separate objects but, in the paper's comparisons, little or no environment. DistScene's contribution is having both at once: separate pieces and a room they belong to. The project page shows the obvious use, scenes dropped straight into a navigation simulator.
Against MIDI, Gen3DSR, SceneGen and the rest
The baselines fall into three camps.
Gen3DSR is the modular camp: estimate depth and segment the image, lift each object with an object-level generator, and assemble. Its own abstract calls the pipeline "highly modular with independent, self-contained modules", and errors compound across those modules. SAM 3D predicts geometry, texture and layout per object, trained on objects annotated at large scale by people and models together. 3D-Fixer, the strongest baseline here, completes each object in place, using the partial point cloud from a geometry estimator as its spatial anchor, so it avoids pose optimisation.
MIDI and SceneGen are the feed-forward camp, DistScene's closest relatives. MIDI extends an image-to-3D model to several instances denoised together with "multi-instance attention", conditioned on partial object crops plus the whole image. SceneGen takes the image and object masks, aggregates local and global features, and predicts positions with a dedicated head. Both generate objects jointly. Neither generates the room. DistScene's argument is that the room is what holds the layout together: without walls and a floor, nothing in the network says the sofa sits against the wall.
The third camp is Extend3D, used for the outdoor benchmark: training-free, it widens an object model's latent in two directions, generates overlapping patches, and refines occluded regions with SDEdit from a depth-based initialisation.
The indoor table, as published. CD is Chamfer distance (lower is better), FS is F-score in percent at a 0.1 threshold, S and O are scene- and object-level, IoU is bounding-box overlap. All baselines were run by the DistScene authors.
| Method | MIDI CD-S | MIDI FS-S | MIDI CD-O | MIDI FS-O | MIDI IoU | Gen3DSR CD-S | Gen3DSR FS-S |
|---|---|---|---|---|---|---|---|
| Gen3DSR | 0.3795 | 26.45 | 0.2609 | 28.73 | 0.0712 | 0.6701 | 11.33 |
| TRELLIS.2 | 0.1444 | 54.19 | 0.2475 | 31.07 | 0.1471 | 0.2099 | 44.44 |
| MIDI | 0.1550 | 53.46 | 0.2162 | 36.26 | 0.1639 | 0.2830 | 32.53 |
| SAM 3D | 0.1486 | 59.79 | 0.1607 | 49.72 | 0.2102 | 0.2055 | 48.73 |
| SceneGen | 0.1502 | 51.40 | 0.1904 | 40.88 | 0.1217 | 0.2040 | 44.31 |
| 3D-Fixer | 0.1295 | 65.08 | 0.1704 | 48.10 | 0.3527 | 0.1027 | 77.97 |
| DistScene | 0.0877 | 71.59 | 0.1529 | 51.64 | 0.3688 | 0.0958 | 79.68 |
The scene-level gap on MIDI-test is the headline: scene Chamfer from 0.1295 to 0.0877 against 3D-Fixer. The paper calls that 32.2%; I get 32.3%, which is rounding, not a problem. The object-level margins are much thinner. SAM 3D's object Chamfer is 0.1607 to DistScene's 0.1529, and 3D-Fixer's box IoU is 0.3527 to 0.3688, so most of the advantage is in where things are, not in how each object looks. That fits the method: the environment slot is a layout device.

Four things about the setup are worth knowing before you quote the table.
The MIDI-test inputs are not the originals. DistScene needs the whole scene in the image, background included, and the original MIDI-test images don't provide it, so the authors re-rendered every test scene from its 3D geometry, floor and walls included. They also rendered aligned object masks and depth maps from the same scenes for the baselines that need them. That cuts both ways: the baselines received perfect masks and depth, a real advantage, but on re-rendered images rather than the benchmark's originals, so these numbers can't be set beside MIDI's own published ones.
The TRELLIS.2 row doesn't agree with itself. Table 2 gives it MIDI-test scene numbers of 0.1444 and 54.19; the ablation table's "TRELLIS.2 (mask)" row, which carries the same Gen3DSR numbers (0.2099, 44.44), gives 0.1500 and 55.57 on MIDI-test, and the "no mask" row gives a third set. Probably two runs of the same baseline. The gap is small next to DistScene's margin, but it says the baseline numbers have run-to-run noise the paper doesn't report.
The resolution claim is a little stronger than its table. The text says raising resolution "consistently improves reconstruction quality", but on Gen3DSR-test the scene frame alone at 512 has the lower Chamfer distance, 0.1077 against 0.1096 at 1024. Its F-score is lower, 71.96 against 73.21, so "mostly" would be accurate.
And the human study is not head-to-head with the feed-forward rivals. It has 43 participants on 20 text-to-image-generated inputs, comparing DistScene with TRELLIS.2, 3D-Fixer, SAM 3D and Extend3D, and DistScene is picked 65.41% of the time on average. MIDI and SceneGen are not in it.
Outdoors, on UrbanScene3D, DistScene reaches a CD-L1 of 0.0772 and F-score 0.722, against 0.0832 and 0.680 for Extend3D. The more telling comparison is with its own base: plain TRELLIS.2 already scores 0.0831 and 0.701, beating Extend3D's F-score, so DistScene's outdoor gain over the model it adapts is about 7% in Chamfer. Extend3D, EvoScene and the UrbanScene3D benchmark appear in that table with no entry in the paper's reference list.
The ablation figure is a better picture of the mechanism than the comparison, because it changes one thing per column.

What it costs
Table 5 is the paper's honest corner. Scene-frame generation takes 5.7 seconds at 512 and 36.4 seconds at 1024; refinement takes 3.3 seconds per object at 512 and 38.5 seconds per object at 1024. Peak memory runs from 17.9 to 21.8 GB, which would fit a 24 GB card if the weights are ever released. The GPU isn't named.
Refinement runs once per object, in a loop in the paper's Algorithm 2, so the totals scale with the scene. For the median preview scene, 5 objects, the full 1024 pipeline is 36.4 plus 5 times 38.5, about 229 seconds. The 512 pipeline with refinement is about 22 seconds and gives up surprisingly little: Gen3DSR-test scene F-score 78.35 against 79.68, MIDI-test 68.35 against 71.59.
If I were using this, I'd run the 512 pipeline to iterate on a scene and the full one once for the final asset. The 1024 refinement buys sharper leaves and grilles, as the resolution figure shows, at roughly ten times the wall clock.
What I make of it
The environment slot is the idea worth keeping. MIDI and SceneGen generate a set of objects and leave the arrangement to the network; DistScene gives the arrangement something to lean on, and the ablation shows it (object F-score on MIDI-test from 29.98 to 40.98 just by adding the room). Refinement is the idea worth stealing. Anyone generating several things in one shared grid hits the same budget wall, and "re-generate each piece in its own cube, conditioned on the shared tokens and a marker" is a clean, reusable answer.
The rest needs code before I'd trust it fully. The benchmarks were re-rendered and re-run by the authors, and the data engine's details (which language model, which image model, how many views per scene) are missing. The outputs have no per-object orientation and no layout constraints once edited. And "editable" in practice means what it means for any scene file whose pieces are separate nodes, which is genuinely useful and a lot less than an editor.
For now there is nothing to run, 23 result scenes to inspect, and a paper whose main claim, that generating the room alongside the furniture makes the furniture land in the right place, holds up in its own ablation.
How I checked
I read the paper in full from arXiv's HTML (v1, posted 2026-10-03), including the appendix and both algorithms, and copied every number in the tables above from it. The project page lists code, checkpoints and dataset as coming soon, and no public repository exists yet, so none of the method is checked against code.
For the result files, I listed coolbeam/DistScene-preview-assets at revision dd2b867 through the Hugging Face API and, for each of the 23 files in original_output/, fetched only the 20-byte GLB header and the JSON chunk with HTTP range requests. From those I took node names and matrices, accessor bounds and counts (vertices, triangles), material and image counts, and, with three more range reads into C000's binary chunk, the PNG headers that give the 4096 by 4096 texture size. Scales, footprints, rotations and triangle totals were computed from that JSON with a short Python script. The labels "sofa", "armchair" and "coffee table" in the widget are my reading of the input photo; the file only says obj003 and so on.
For TRELLIS.2, I shallow-cloned microsoft/TRELLIS.2 at 75fbf01 and read the pipeline (trellis2/pipelines/trellis2_image_to_3d.py), the model configs in configs/gen/, and the attention and MLP modules, plus pipeline.json from the Hugging Face model. The 16x downsampling, the 16-cubed occupancy latent with 8 channels, the 1.3B transformer width and depth and the 12-step samplers come from there. The latent-voxel counts per object assume DistScene keeps TRELLIS.2's grid, which the paper implies but doesn't state. The LoRA size is my arithmetic on those configs under an assumption about which layers are adapted. Baseline descriptions come from each paper's own abstract.