Fine-grained image retrieval aims to retrieve images from the same subclass as the query images from a database containing a specific class object. To solve the problem of large inter-class differences and small intra-class differences, we propose the Multiscale Convolutional Feature Aggregation (MsCFA) method accomplished in an unsupervised environment. This method comprises an object localization submodule and a multiscale feature fusion module. In the object localization module, we perform a coarse localization of the object based on the response levels within the feature map. Specifically, we crop the high-response regions of the feature map to ensure that the resulting cropped feature map contains more object features and less background noise. Subsequently, we employ the regional maximum activation of convolutions method to extract features from this cropped region. The cropped object feature is fused with the uncropped original feature to obtain the feature representation on a single scale. In the feature fusion module, we input images of different scales into the network to obtain feature representations at various scales and then fuse these different scale features to obtain our multiscale convolutional feature aggregation descriptor. In experiments on five classical fine-grained datasets, our method outperforms other fine-grained retrieval tasks in the same type of unsupervised environment and is comparable to those supervised tasks with end-to-end training.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multiscale Convolutional Feature Aggregation for Fine-Grained Image Retrieval

  • Caolin Yang,
  • Kaifeng Ding,
  • Chengzhuan Yang

摘要

Fine-grained image retrieval aims to retrieve images from the same subclass as the query images from a database containing a specific class object. To solve the problem of large inter-class differences and small intra-class differences, we propose the Multiscale Convolutional Feature Aggregation (MsCFA) method accomplished in an unsupervised environment. This method comprises an object localization submodule and a multiscale feature fusion module. In the object localization module, we perform a coarse localization of the object based on the response levels within the feature map. Specifically, we crop the high-response regions of the feature map to ensure that the resulting cropped feature map contains more object features and less background noise. Subsequently, we employ the regional maximum activation of convolutions method to extract features from this cropped region. The cropped object feature is fused with the uncropped original feature to obtain the feature representation on a single scale. In the feature fusion module, we input images of different scales into the network to obtain feature representations at various scales and then fuse these different scale features to obtain our multiscale convolutional feature aggregation descriptor. In experiments on five classical fine-grained datasets, our method outperforms other fine-grained retrieval tasks in the same type of unsupervised environment and is comparable to those supervised tasks with end-to-end training.