Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
arXiv:2511.14271v2 Announce Type: replace Abstract: Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still face two fundamental limitations. First, they struggle with coarse semantic alignment, often failing to capture fine-grained prompt details. Second, they lack robust 3D spatial understanding, leading to geometric inconsistencies and catastrophic failures in part assembly and spatial relationships. T