Few-Shot Object Detection (FSOD) aims to detect novel category objects using only a few annotated samples by leveraging prior knowledge learned from a large number of base categories. However, existing methods often suffer from feature misalignment and insufficient discriminative information correspondence under extremely limited data conditions. To address these challenges, this paper proposes a High-Low Frequency Feature Alignment (HLFFA) framework, which jointly models the high- and low-frequency components of support and query images to achieve more comprehensive feature alignment. The core module, Dual-Stream Interaction Modulated Attention (DIMA), enhances local detail alignment through convolutional modulated attention and enables efficient semantic interaction via cross-attention mechanisms. In HLFFA, the high-frequency branch emphasizes edge and texture information, while the low-frequency branch captures global shape representations through average pooling. These two types of features are then adaptively fused to achieve holistic and robust matching. Extensive experiments conducted on the PASCAL VOC and MS COCO datasets demonstrate that the proposed method consistently outperforms existing FSOD approaches across various K-shot settings, significantly improving detection accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

High-Low Frequency Feature Alignment for Few-Shot Object Detection

  • Xinwei Yao,
  • Jun Liu,
  • Qiang Li,
  • Zitao Tu,
  • Hengcong Zhang

摘要

Few-Shot Object Detection (FSOD) aims to detect novel category objects using only a few annotated samples by leveraging prior knowledge learned from a large number of base categories. However, existing methods often suffer from feature misalignment and insufficient discriminative information correspondence under extremely limited data conditions. To address these challenges, this paper proposes a High-Low Frequency Feature Alignment (HLFFA) framework, which jointly models the high- and low-frequency components of support and query images to achieve more comprehensive feature alignment. The core module, Dual-Stream Interaction Modulated Attention (DIMA), enhances local detail alignment through convolutional modulated attention and enables efficient semantic interaction via cross-attention mechanisms. In HLFFA, the high-frequency branch emphasizes edge and texture information, while the low-frequency branch captures global shape representations through average pooling. These two types of features are then adaptively fused to achieve holistic and robust matching. Extensive experiments conducted on the PASCAL VOC and MS COCO datasets demonstrate that the proposed method consistently outperforms existing FSOD approaches across various K-shot settings, significantly improving detection accuracy.