Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
This study focuses on Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text, with both modalities aligned to the text conditions. Despite progress in joint audio-video training, two critical challenges remain: (1) text conditioning is a bottleneck—shared captions (TV=TA) trigger modal interference, while a gap persists between dense training captions and concise inference user prompts, and (2) the optimal fusion mechanism for cross-modal featu
![LingBot-Vision: masked boundary modeling for self-supervised pretraining (0.296 NYUv2 linear-probe RMSE at 1.1B vs 0.309 for DINOv3-7B, trails on ImageNet); weights in 4 sizes[R]](https://preview.redd.it/ha08vg49bnbh1.png?width=140&height=78&auto=webp&s=cbd1e4aed6c0571b7f0acee245c21543fe356719)




![EMNLP: All of the papers in my review pool being detected as AI [D]](https://preview.redd.it/elmu1a6d0kbh1.png?width=140&height=140&crop=1:1,smart&auto=webp&s=c44181dd02b9668433e47a57a648d98af652fbde)
