Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
arXiv:2607.19027v1 Announce Type: new Abstract: Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision-language models and multimodal large language models, existing ZMR methods still heavily depend on query-to-video content similarity, making them vulnerable to modality and language-style gaps. These gaps lead to unreliable span propo
