Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography
arXiv:2606.25546v1 Announce Type: new Abstract: Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these issues, we propose a tailored VLP framework featuring three key components: (1) a CNN-ViT hybrid encoder that replaces ViT's patch embedding with a 3D CNN backbone to effic