Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model
arXiv:2607.18958v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defense methods predominantly target single-task scenarios (e.g., zero-shot classification) and consequently lack generalizability across various multimodal tasks. To address this limitation, we propose a dual adversarial fine

![Looking for feedback on my GPU-accelerated Snake AI project [P]](https://preview.redd.it/4k0bf6wgtneh1.gif?width=640&crop=smart&s=7309dc4cdba7df36b615ed9025f212c2b34fd4b0)





