<p>Bird species classification is a critical task in ecological research and conservation, enabling accurate monitoring of biodiversity and environmental health. This study evaluates seven advanced vision transformer models and five contemporary convolutional neural networks (<i>CNN</i>)-based models that were analysed using <i>BIRDS-525</i> dataset of around 90,000 images of 525 distinct avian classes. The models are assessed based on accuracy, F1-score, and inference time. <i>Vision Transformers</i>, particularly <i>MambaVision</i> and <i>ViT</i> enhanced through fine-tuning, demonstrated superior accuracy and efficiency, with <i>MambaVision</i> achieving the highest nominal test accuracy of 99.47%; however, statistical analysis indicates that its performance is not significantly different from other top-performing models such as <i>DenseNet201</i> and <i>Vision Transformer (ViT)</i> after Bonferroni correction. Rigorous statistical validation using pairwise t-tests and Chi-squared (<InlineEquation ID="IEq1"><EquationSource Format="TEX">\(\chi^2\)</EquationSource></InlineEquation>) tests identified a statistically indistinguishable top-performing cluster of nine architectures after Bonferroni correction, with only <i>AlexNet</i> and <i>ShuffleNetV2</i> forming a significantly inferior bottom tier. A comprehensive comparative analysis of various deep learning approaches is also presented, evaluating their performance on the benchmark <i>BIRDS-525</i> dataset against several old as well as recent state-of-the-art methodologies. The top five models are further evaluated on a custom dataset- Indian Bird Species Dataset (<i>IBSD</i>) consisting of total 361 images across five species of birds created by capturing images of locally available bird species in a real-time environment. The dataset includes noisy and cluttered images, and the models’ performance in these conditions provides preliminary insights into generalization under realistic capture conditions, with <i>ViT</i> showing the strongest overall real-world performance. The findings provide valuable insights into optimizing both <i>CNN</i>-based and transformer-based models for bird species identification, contributing to biodiversity monitoring and conservation efforts.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-grained avian wildlife species classification: a comparative study of MambaVision, vision transformers, and CNN architectures

  • Rhythm Kulsrrestha,
  • Chiranjit Pal,
  • Rahul Jain,
  • Raj Jaiswal,
  • Arvinder Kaur

摘要

Bird species classification is a critical task in ecological research and conservation, enabling accurate monitoring of biodiversity and environmental health. This study evaluates seven advanced vision transformer models and five contemporary convolutional neural networks (CNN)-based models that were analysed using BIRDS-525 dataset of around 90,000 images of 525 distinct avian classes. The models are assessed based on accuracy, F1-score, and inference time. Vision Transformers, particularly MambaVision and ViT enhanced through fine-tuning, demonstrated superior accuracy and efficiency, with MambaVision achieving the highest nominal test accuracy of 99.47%; however, statistical analysis indicates that its performance is not significantly different from other top-performing models such as DenseNet201 and Vision Transformer (ViT) after Bonferroni correction. Rigorous statistical validation using pairwise t-tests and Chi-squared (\(\chi^2\)) tests identified a statistically indistinguishable top-performing cluster of nine architectures after Bonferroni correction, with only AlexNet and ShuffleNetV2 forming a significantly inferior bottom tier. A comprehensive comparative analysis of various deep learning approaches is also presented, evaluating their performance on the benchmark BIRDS-525 dataset against several old as well as recent state-of-the-art methodologies. The top five models are further evaluated on a custom dataset- Indian Bird Species Dataset (IBSD) consisting of total 361 images across five species of birds created by capturing images of locally available bird species in a real-time environment. The dataset includes noisy and cluttered images, and the models’ performance in these conditions provides preliminary insights into generalization under realistic capture conditions, with ViT showing the strongest overall real-world performance. The findings provide valuable insights into optimizing both CNN-based and transformer-based models for bird species identification, contributing to biodiversity monitoring and conservation efforts.