CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification
arXiv:2607.13437v1 Announce Type: new Abstract: Recent vision models such as CLIP and SAM enable training-free segmentation and semantic encoding for fine-grained classification. A common approach is to compare the representations of segmented image regions with the text prompt embeddings of the corresponding labels. However, it remains unclear how different local regions and CLIP-based scoring strategies affect the selection of discriminative evidence, especially when ground-truth labels are un