<p>To tackle the challenges of capturing discriminative features and effectively leveraging multi-granularity information in fine-grained image classification, this paper proposes an Interactive Modeling Network with Feature Visibility Optimization (IMN-FVO), based on a visibility-guided and interactive learning strategy. IMN-FVO enhances the utilization of multi-scale features and improves the extraction of subtle discriminative details through comprehensive feature mining. It comprises three key modules: (1) the Feature Visibility Mining Unit, which enhances salient features and suppresses redundant ones to improve feature representation and discrimination; (2) the Strip Convolution Optimization Module, which focuses on precise localization by filtering out irrelevant information; and (3) the Interactive Multi-layer Perceptron Module, which models multi-level feature interactions to enrich semantic representation and fusion. The entire framework is trained end-to-end, without requiring bounding box annotations or multi-stage processing. Experiments on CUB-200–2011, Stanford Cars, and FGVC-Aircraft show that IMN-FVO outperforms existing state-of-the-art methods, demonstrating strong effectiveness and generalization.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Feature visibility optimization-based interactive modeling for fine-grained image classification

  • Kai Yang Liao,
  • Yun Fei Tan,
  • Yuan Lin Zheng,
  • Guang Feng Lin,
  • Gang Huang,
  • Ding Wen Song

摘要

To tackle the challenges of capturing discriminative features and effectively leveraging multi-granularity information in fine-grained image classification, this paper proposes an Interactive Modeling Network with Feature Visibility Optimization (IMN-FVO), based on a visibility-guided and interactive learning strategy. IMN-FVO enhances the utilization of multi-scale features and improves the extraction of subtle discriminative details through comprehensive feature mining. It comprises three key modules: (1) the Feature Visibility Mining Unit, which enhances salient features and suppresses redundant ones to improve feature representation and discrimination; (2) the Strip Convolution Optimization Module, which focuses on precise localization by filtering out irrelevant information; and (3) the Interactive Multi-layer Perceptron Module, which models multi-level feature interactions to enrich semantic representation and fusion. The entire framework is trained end-to-end, without requiring bounding box annotations or multi-stage processing. Experiments on CUB-200–2011, Stanford Cars, and FGVC-Aircraft show that IMN-FVO outperforms existing state-of-the-art methods, demonstrating strong effectiveness and generalization.