Fine-grained avian wildlife species classification: a comparative study of MambaVision, vision transformers, and CNN architectures
摘要
Bird species classification is a critical task in ecological research and conservation, enabling accurate monitoring of biodiversity and environmental health. This study evaluates seven advanced vision transformer models and five contemporary convolutional neural networks (CNN)-based models that were analysed using BIRDS-525 dataset of around 90,000 images of 525 distinct avian classes. The models are assessed based on accuracy, F1-score, and inference time. Vision Transformers, particularly MambaVision and ViT enhanced through fine-tuning, demonstrated superior accuracy and efficiency, with MambaVision achieving the highest nominal test accuracy of 99.47%; however, statistical analysis indicates that its performance is not significantly different from other top-performing models such as DenseNet201 and Vision Transformer (ViT) after Bonferroni correction. Rigorous statistical validation using pairwise t-tests and Chi-squared (