Lightweight 3D Feature Pretraining by Bayesian Inversion of 2D Foundation Models
arXiv:2606.21292v1 Announce Type: new Abstract: We present Casper3D, a lightweight probabilistic framework for converting noisy multi-view 2D foundation-model embeddings into a latent 3D semantic representation. We model view-level semantic features as noisy observations of an underlying 3D semantic state and infer this state with a set-based variational model that incorporates relative pose during multi-view reasoning. Casper3D is trained by predicting held-out semantic observations from novel