<p>Few-shot image classification aims to generalize prior knowledge from abundant base classes to novel categories with limited labeled samples. Current semantic alignment methods struggle with irrelevant regions interference, while task-aware approaches suffer from the deficiency of losing crucial inter-class structural information. To address the above issues, we propose an adaptive feature recalibration transformer (AFRT) for few-shot classification. During the pre-training phase, the feature encoder learns semantic information beyond image labels and contextual relationships of local regions from masked image modeling (MIM). During the meta-finetuning phase, our method comprises a task-driven salient region refinement module (TSRR) and a bidirectional interactive feature calibration module (BIFC). TSRR establishes local semantic relationships within the support set and filters out regions that contribute more to inference, weakening the expression of irrelevant regions. BIFC facilitates bidirectional interaction between local regions of support class features and query instance features, further focusing on more subtle and discriminative shared features. Extensive experiments show that our method achieves competitive performance on four widely used few-shot image classification benchmarks. Our code is available at <a href="https://github.com/leolensg/AFRT.">https://github.com/leolensg/AFRT.</a></p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive feature recalibration transformer for enhancing few-shot image classification

  • Wei Song,
  • Yaobin Huang

摘要

Few-shot image classification aims to generalize prior knowledge from abundant base classes to novel categories with limited labeled samples. Current semantic alignment methods struggle with irrelevant regions interference, while task-aware approaches suffer from the deficiency of losing crucial inter-class structural information. To address the above issues, we propose an adaptive feature recalibration transformer (AFRT) for few-shot classification. During the pre-training phase, the feature encoder learns semantic information beyond image labels and contextual relationships of local regions from masked image modeling (MIM). During the meta-finetuning phase, our method comprises a task-driven salient region refinement module (TSRR) and a bidirectional interactive feature calibration module (BIFC). TSRR establishes local semantic relationships within the support set and filters out regions that contribute more to inference, weakening the expression of irrelevant regions. BIFC facilitates bidirectional interaction between local regions of support class features and query instance features, further focusing on more subtle and discriminative shared features. Extensive experiments show that our method achieves competitive performance on four widely used few-shot image classification benchmarks. Our code is available at https://github.com/leolensg/AFRT.