Recently, with the advancement of deep learning (DL), many deep models such as ResNet50 have shown remarkable performance in natural image classification tasks. In addition, contrastive language-image pre-training (CLIP) based methods have demonstrated superior performance in many image classification tasks. In this paper, to further improve the performance of CLIP-based approaches, we introduce a novel feature fusion method by combining both convolutional neural network (CNN) features and vision transformer (ViT) features. In addition, we incorporate a feature attention module to improve the performance of the models. Finally, we use traditional machine learning (ML) classifiers to replace the original classifier head. We perform extensive experiments on four datasets (three natural and one medical), and the experimental results show that our method can provide feasible test metric (e.g., 76.76% accuracy on ISIC2019) compared to other state-of-the-art (SOTA) approaches.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HVLM: Boosting Vision Language Model Using Hybrid Deep Models and Attention Module for Image Classification

  • Bingkun Ren

摘要

Recently, with the advancement of deep learning (DL), many deep models such as ResNet50 have shown remarkable performance in natural image classification tasks. In addition, contrastive language-image pre-training (CLIP) based methods have demonstrated superior performance in many image classification tasks. In this paper, to further improve the performance of CLIP-based approaches, we introduce a novel feature fusion method by combining both convolutional neural network (CNN) features and vision transformer (ViT) features. In addition, we incorporate a feature attention module to improve the performance of the models. Finally, we use traditional machine learning (ML) classifiers to replace the original classifier head. We perform extensive experiments on four datasets (three natural and one medical), and the experimental results show that our method can provide feasible test metric (e.g., 76.76% accuracy on ISIC2019) compared to other state-of-the-art (SOTA) approaches.