<p>Fine-grained visual classification (FGVC) is defined as the finer division of sub-categories within basic categories. The task is both valuable and challenging. Its difficulty primarily arises from its intrinsic slight inter-class variations and substantial intra-class differences. The crucial solution to FGVC lies in identifying local regions with subtle yet discriminative features and effectively representing them. Nevertheless, with the increasing prevalence of deep convolutional neural networks, researchers have primarily prioritized the use of high-level, abstract, semantic features to achieve FGVC, consequently overlooking low-level, detailed information, resulting in poor feature representation capabilities. Thus, we put forward the multi-level navigation network, denoted as MLNN, to enhance feature representation by incorporating both high-level semantics and low-level details. Specifically, MLNN is composed of (1) the feature refinement and attention enhancement module, which enables the network to learn detailed feature representations and further enhance features with attention mechanisms, and (2) the triplet-enhanced multi-level fusion module, which integrates the features of different levels, leading to a more comprehensive feature representation. Experimental outcomes reveal that our approach attains state-of-the-art performance on three widely-accepted benchmark datasets.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-level navigation network: advancing fine-grained visual classification

  • Hong Liang,
  • Xian Li,
  • Mingwen Shao,
  • Qian Zhang

摘要

Fine-grained visual classification (FGVC) is defined as the finer division of sub-categories within basic categories. The task is both valuable and challenging. Its difficulty primarily arises from its intrinsic slight inter-class variations and substantial intra-class differences. The crucial solution to FGVC lies in identifying local regions with subtle yet discriminative features and effectively representing them. Nevertheless, with the increasing prevalence of deep convolutional neural networks, researchers have primarily prioritized the use of high-level, abstract, semantic features to achieve FGVC, consequently overlooking low-level, detailed information, resulting in poor feature representation capabilities. Thus, we put forward the multi-level navigation network, denoted as MLNN, to enhance feature representation by incorporating both high-level semantics and low-level details. Specifically, MLNN is composed of (1) the feature refinement and attention enhancement module, which enables the network to learn detailed feature representations and further enhance features with attention mechanisms, and (2) the triplet-enhanced multi-level fusion module, which integrates the features of different levels, leading to a more comprehensive feature representation. Experimental outcomes reveal that our approach attains state-of-the-art performance on three widely-accepted benchmark datasets.