Bridging semantics and vision: text-guided feature alignment for few-shot object detection
摘要
Few-shot object detection (FSOD) remains challenged by frequent false positives and significant missed detections due to limited supervision and poor category discrimination. To address these issues, we propose a novel framework that integrates two complementary modules: semantic-aware proposal refinement (SAPR) and latent structure-aware clustering (LSAC). Specifically, SAPR aligns textual category descriptions with visual instance features in a shared semantic space, leveraging a soft assignment mechanism to model cross-modal affinity. This enables more reliable candidate region selection by capturing fine-grained semantic preferences across categories. Complementarily, LSAC learns structured, under-complete representations in a latent embedding space, where embedded features are clustered to uncover latent semantic groupings. This process enhances category coherence and provides enriched supervisory signals for better class discrimination. Together, these components strengthen the model’s semantic understanding and intra-class compactness, effectively reducing both misclassification and omission errors. Extensive experiments on benchmark datasets—PASCAL VOC and MS COCO—demonstrate that our approach achieves substantial performance gains over the baseline and outperforms current state-of-the-art methods under multiple few-shot benchmarks.