Read the paper

Pipeline

  • DINOv3 -> per-point image features
  • Perception transformer outputs per-object validity score, 3D masks, pose + scale -> canonicalized point cloud
  • Raw shape/material (Ablation shows accuracy is 0.914 compressed vs 0.919 uncompressed for a 32x smaller latent)
  • Flow-matching transformer
  • 12 step feed-forward per object, up to 16 objects done in parallel
  • Material generation conditioned on the corresponding shape latent
  • VAE decodes latent -> occupancy voxel grid -> mesh (dual contouring) -> UV unwrap + texture bake

Baseline

Boxer + SAM2 + TRELLIS.2

Results

Table 2: Quantitative results on 3D scene perception. We compare the mAP and mIoU of the perception results for detection and segmentation quality. We also measure and report the average model inference time across datasets. Table 3: Quantitative results on video-based 3D scene reconstruction. Methods are evaluated under both groundtruth and inferred perception inputs. β€œ-” indicates metrics not applicable. The average runtime per object is measured across datasets. We highlight best and second best. Table 5: Qualitative and Quantitative results on VAE design comparison. We evaluate the geometry and rendering metrics in Toys4k [50] Object Dataset and Imaginarium [78] Scene Dataset. Results show that our added HC-VAE above SC-VAE further compresses the latent space and still maintains the high-quality reconstruction.

On flow-matching transformer instead of optimization methods

Optimization method (e.g. NeRF)

  • Starts from scratch every scene
  • Renders the current guess, checks how far off it is from the real images (loss func), then nudges the parameters a bit:
  • Nothing carries over to the next scene
  • Can get stuck on a bad guesses/occlusions

Flow-matching transformer

  • Trained once on lots of data (note: these are unoccluded 3D objects -> even on occluded objects it samples from its amodally complete prior). Reused for every new object + any scene
  • Learns a velocity function to go from noise toward a real shape given some conditioning (point cloud + image features). Kinda like diffusion but it’s straight-line not schedule
  • Training: given predict . No loss, backprop, fitting per-object at inference time
  • Same fixed feed-forward cost every time, a lot faster, can run in parallel for different objects