Few-Shot Object Detection (FSOD) is affected by the long-tailed distribution of data and the discrepancy in sample quantities between base classes and novel classes, leading to evident data bias. As a result, the generated feature distribution struggles to represent class features effectively. In scenarios with scarce samples, irrelevant factors in features may have a more significant impact on feature distribution and even dominate feature representation. To obtain more compact and accurate class-specific feature representations, this paper introduces the disentangled representation into few-shot object detection and proposes a semantic disentanglement representation meta-learning model, referred to as FSOD-SDR. Firstly, in the feature extraction phase, a feature information aggregation module is constructed to aggregate features from different scales of the backbone, thereby enabling a more comprehensive representation of support features containing limited information. Secondly, to address highly coupled features, background-relevant and label-relevant semantic factor distributions are simultaneously disentangled from aggregated features by a semantic disentanglement representation module. The label-relevant feature distribution can more accurately represent class features. To effectively achieve disentanglement of the goal, the Evidence Lower Bound (ELBO) loss function is extended during model optimization. Lastly, experiments on the PASCAL VOC and MS COCO datasets show that FSOD-SDR has a significant performance improvement (an average improvement of 5.7% across all metrics) over the previous state-of-the-art methods, achieving comparably good detection performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Few-Shot Object Detection via Disentangling Class-Related Factors in Feature Distribution

  • Lili Wei,
  • Xiaofen Tang,
  • Jin Dang

摘要

Few-Shot Object Detection (FSOD) is affected by the long-tailed distribution of data and the discrepancy in sample quantities between base classes and novel classes, leading to evident data bias. As a result, the generated feature distribution struggles to represent class features effectively. In scenarios with scarce samples, irrelevant factors in features may have a more significant impact on feature distribution and even dominate feature representation. To obtain more compact and accurate class-specific feature representations, this paper introduces the disentangled representation into few-shot object detection and proposes a semantic disentanglement representation meta-learning model, referred to as FSOD-SDR. Firstly, in the feature extraction phase, a feature information aggregation module is constructed to aggregate features from different scales of the backbone, thereby enabling a more comprehensive representation of support features containing limited information. Secondly, to address highly coupled features, background-relevant and label-relevant semantic factor distributions are simultaneously disentangled from aggregated features by a semantic disentanglement representation module. The label-relevant feature distribution can more accurately represent class features. To effectively achieve disentanglement of the goal, the Evidence Lower Bound (ELBO) loss function is extended during model optimization. Lastly, experiments on the PASCAL VOC and MS COCO datasets show that FSOD-SDR has a significant performance improvement (an average improvement of 5.7% across all metrics) over the previous state-of-the-art methods, achieving comparably good detection performance.