Skip to content

Geometry as a First-Class Citizen: Lessons from Five New 3D Vision Papers

#3d-vision #geometric-representation-learning #novel-view-synthesis #robot-manipulation #world-models #depth-estimation

The most interesting work in 3D vision right now is about where geometry lives in the pipeline. Five new papers, each aimed at a different problem, converge on the same argument: geometric understanding is decided by the architecture, with supervision volume playing a supporting role.

That's a useful counterweight to the reflex that says collect more data and add more supervision. Each paper instead changes what comes out of the network, or what the network can use on its way to the output. Change the output representation and you change what can be learned at all.

The robotics angle is hard to miss. Four of the five papers name manipulation, simulation, or planning as their test bed. That makes sense. Robots are where geometry stops being a benchmark and becomes a hard constraint: the simulator needs convex collision shapes, the gripper needs metric depth, the motion model needs rotation, not just translation.

SNAP: weak decoders, strong encoders ​

Novel view synthesis (NVS) should be the ideal pretext task for geometric representation learning. To render a new view, the model has to know where surfaces are. Yet encoder-based NVS methods keep producing weak representations, and the cause is not the supervisory signal.

SNAP, a self-supervised encoder-decoder transformer, pins the problem on two architectural choices. One is the spatially expressive decoder. When the decoder can synthesize the target view from local context on its own, the scene encoder's features never have to carry 3D structure, so they don't. The other is the low-level pixel-space target, which pushes the model to reproduce texture and lighting rather than geometry. SNAP replaces both with a pose-conditioned local decoder and a latent-space reconstruction objective. The decoder has to query the encoder for what it needs, and the target removes the pixel-level shortcut.

The results support the diagnosis. SNAP is task agnostic and lands competitive with special-purpose, geometry-supervised methods across visual localization, pose estimation, point correspondence, depth estimation, and robot manipulation. Its patch features develop viewpoint invariance that approaches heavily supervised models, at lower compute and data budgets. And under camera shifts that collapse standard 2D representations, SNAP degrades more gracefully. The encoder stored usable 3D structure because the architecture forced it to.

The five papers at a glance ​

PaperTaskKey architectural moveTraining signalHeadline result
SNAPGeometric representation from novel viewsPose-conditioned local decoder; latent-space targetSelf-supervised NVSViewpoint-invariant features; competitive with geometry-supervised methods
MoSE3Dense per-pixel SE(3) motion from monocular videoJoint 3D point tracks and rigidity embeddings; differentiable SE(3) fittingSynthetic Art-Kubric dataSOTA SE(3) at pixel, part, and object level; generalizes to real video
XGenActRobot world action modelDepth, normals, actions, segmentation encoded as RGB video; one diffusion transformerCross-task future prediction52% success on five held-out RLBench tasks vs 26% for baselines
I2CDConvex collision geometry from a single imageFrozen image-to-3D model plus 38M-parameter convex-slot head227 objects; under 10 GPU-hours0.5s per image, convex by construction; 11s vs 328s on a physical robot scene
HypoDepthEvent-image monocular depthDepth Hypothesis Volume; iterative multi-scale correlation searchDSEC and MVSECSOTA on both; tiny variant runs real-time on resource-limited devices

Quick Take: The representation that leaves the network does the real work, and each paper gets there by changing the output: latent targets over pixels, convex polytopes over meshes, SE(3) over point tracks, search over regression.

MoSE3: every pixel gets a rigid transform ​

Dense 3D point tracking is the default way to model motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel. It records where a point goes. It doesn't record whether its part rotated, or which pixels move together as one body.

MoSE3 is the first feed-forward model that predicts dense SE(3) motion from monocular RGB video: full 6-DoF rigid transforms at every pixel in world space. That single output bundles rotation, translation, and grouping, which is the information a manipulation policy needs and exactly what point tracks lack.

Direct SE(3) prediction is hard for two reasons. Rotations live on a curved manifold that resists Euclidean regression, and dense SE(3) annotations are nearly impossible to acquire. MoSE3 sidesteps both by predicting two Euclidean-friendly intermediates jointly, 3D point tracks and per-pixel rigidity embeddings, then recovers SE(3) by differentiably fitting rigid transforms inside each soft cluster. The fitting step is part of the network, so training runs end-to-end.

To close the data gap, the authors built Art-Kubric, a large synthetic dataset of articulated objects with dense SE(3) and rigidity labels. Trained only on that synthetic motion data, MoSE3 reaches state-of-the-art SE(3) accuracy at pixel, part, and object level on rigid and articulated benchmarks, plus top average 3D point tracking accuracy across three datasets, and it transfers to real-world videos.

There's a pattern here that carries across this batch of papers. When the output needs structure you can't supervise directly, invent an intermediate that has structure you can, and derive the final prediction instead of regressing it. That move recurs everywhere, and it's the closest thing this week has to a shared theorem.

XGenAct: one objective for all modalities ​

World action models for robot control learn to predict how observations and actions evolve together. Pure RGB and action prediction doesn't force the model to understand space, and earlier attempts to fix that added a handful of spatial prediction tasks through specialized heads and branches. That fragments both the supervision and the architecture.

XGenAct collapses all of it into one interface. RGB observations, robot actions, metric depth, surface normals, and functional role segmentation are each encoded as RGB videos through deterministic codecs. One video diffusion transformer, one objective, no modality-specific heads. Training samples perception and action tasks across these spaces.

The codec trick is the move that matters. It turns a multi-task problem into a single-task problem. The transformer never needs to know whether the current video encodes pixels or depth. Geometry is just another video, predicted with the same machinery.

The evidence for the design is concrete. On held-out RLBench tasks, structured perception training improves average closed-loop success over RGB-only training. In a five-task external comparison, XGenAct reaches 52% success versus 26% for the strongest evaluated baselines. It also predicts future depth and segmentation more accurately than pipelines that generate RGB first and then hand frames to a frozen perception expert. Predicting geometry as the thing being generated, rather than reading it out afterwards, is what forces the world model to maintain it.

I2CD: collision geometry without the repair shop ​

Physics simulators and motion planners require convex collision geometry. Image-to-3D generative models output dense, frequently non-manifold visual meshes. Bridging the two means a reconstruct-then-decompose pipeline: repair, decimate, approximate-convex-decompose. My team has spent hours on that grind, usually on one object at a time, and every step is a place for the pipeline to break.

I2CD does the decomposition instead. It freezes the pretrained Hunyuan3D-2 image-conditioned diffusion transformer and shape decoder, and trains only a lightweight cross-attention head: 38M parameters, under ten GPU-hours. The head's learned convex-slot tokens emit the halfplane parameters of K convex polytopes. The output is compact and convex by construction, loads into physics engines with no post-processing, and takes about 0.5 seconds per image. That's fast enough to run inside a data-generation loop.

The paper's cross-simulator study is the part I'd want every robotics engineer to read. In MuJoCo, PyBullet, Genesis, and Isaac Sim, I2CD geometry is used exactly as delivered. Raw generated meshes also load everywhere, but in most cases the engine silently replaces them with a different collision shape, or demands seconds to minutes of per-object preprocessing. If you've ever watched a robot in simulation collide with an invisible box, this is why. The engine accepted your mesh and quietly substituted a primitive you never inspected.

The numbers are hard to argue with. On 227 held-out OmniObject3D and Google Scanned Objects instances, I2CD gets the highest volumetric IoU among eight reconstruct-then-decompose pipelines while running 6 to 37 times faster end-to-end. On a physical xArm7, planner-ready geometry for a 20-object cluttered scene takes 11 seconds versus 328 seconds for the strongest baseline, at comparable pick-and-place execution success (85 versus 90 of 100 trials). A 30x speedup for 5 failed trials is a trade worth making.

Event cameras record brightness changes at microsecond resolution, which makes them a natural fit for depth estimation in fast motion. The trouble is that direct full-depth regression from events is ill-posed and nonlinear. Existing methods keep optimizing contextual features and still fight the regression itself.

HypoDepth reframes the problem. It introduces a discrete Depth Hypothesis Volume (DHV) and turns depth estimation into a constrained search: weight each candidate depth hypothesis instead of regressing one value. A 3D cost volume between DHV features and contextual features drives a multi-scale correlation search, and each refinement step is a stable residual update. The cost volume is lightweight, so global-to-local refinement across resolutions stays cheap.

This is MoSE3's lesson on a different axis. When the output space is awkward, change the formulation instead of the loss. The task becomes classification over a well-posed hypothesis set, and the optimization behaves.

HypoDepth reports state-of-the-art results on DSEC and MVSEC with strong zero-shot generalization. The tiny variant keeps most of the accuracy and runs in real time on resource-limited devices, so the search formulation costs you nothing at inference.

Common pitfalls ​

What trips people up across these five papers:

  • Silently replaced collision geometry. Simulators accept raw meshes on load and swap in a primitive approximation without telling you. Verify which collision shape the engine actually instantiates, not which one you think you loaded. I2CD's cross-simulator study quantifies how common this is, and it invalidates any planning result built on unverified geometry.

  • Expressive decoders with pixel-space targets. If your NVS decoder can reconstruct the view from local context alone, the encoder never has to learn geometry. Match decoder capacity and target space to the representation you actually want, which is what SNAP's pose-conditioned local decoder and latent objective do.

  • Euclidean regression on rotation manifolds. Fitting SO(3) with an L2 loss fights the math. Predict Euclidean-friendly intermediates and recover rotations by fitting, the way MoSE3 does, instead of hoping the network learns the manifold structure implicitly.

  • Spatial supervision through separate heads. Bolting a depth head onto a backbone trains a side path, not shared geometric understanding. XGenAct's codec interface gets one model to predict every space through one objective, and that shows up in the closed-loop success gap.

  • Direct full-depth regression on event streams. If your event-camera depth is unstable, reframe it as a search over depth hypotheses. HypoDepth's cost-volume formulation converges more reliably than direct regression, at real-time cost.

One thing to remember ​

The output representation is a training signal. SNAP reconstructs in latent space instead of pixels. I2CD emits halfplanes instead of triangles. MoSE3 derives SE(3) from tracks and rigidity. These aren't convenience choices. Each one predetermines what the model can learn. Pick the representation before you pick the loss, because the loss can only reinforce what the representation allows.

What this means ​

  • If you build robot manipulation policies, adopt the codec-style interface. Encode depth, normals, and segmentation as video-like channels so a single world model predicts all of them. The 52% versus 26% success gap on held-out RLBench tasks is the evidence that shared generation beats per-task readouts.

  • If you work in sim-to-real or physics planning, skip the reconstruct-then-decompose pipeline. Convex-slot heads give you planner-ready geometry in 0.5 seconds per image with no post-processing, and the collision shape the simulator actually uses is the geometry you generated.

  • If your motion stack is built on point tracks, expect dense SE(3) prediction to take over within a year. MoSE3 shows tracks are a translation-only special case of the 6-DoF problem, and synthetic-only training already transfers to real video.