Gastrointestinal cancer is the leading cause of cancer-related incidence and death. Therefore, it is important to develop a novel computer-aided diagnosis system for early detection and enhanced treatment. Traditional approaches rely on the expertise of gastroenterologists to identify diseases. However, it is a subjective process, and the interpretation can vary even between expert clinicians. Considering recent progress in classifying gastrointestinal anomalies and landmarks in endoscopic and video capsule endoscopy images, this study proposes a hybrid model incorporating the advantages of Transformers and Convolutional Neural Networks (CNNs) for enhanced classification performance. Our model utilizes DenseNet201 as a CNN branch to extract local features and integrates the Swin Transformer branch for global feature understanding. Both of their features are combined to perform the classification task. For the GastroVision dataset, our proposed model demonstrates excellent performance with Precision, Recall, F1 score, Accuracy, and Matthews Correlation Coefficient (MCC) of 0.8320, 0.8386, 0.8324, 0.8386, and 0.8191, respectively, showcasing its robustness against class imbalance dataset and surpassing other CNNs as well as Swin Transformer model. Similarly, for the Kvasir-Capsule, a large video capsule endoscopy dataset, our model surpassed all other models, thereby achieving overall Precision, Recall, F1 score, Accuracy, and MCC of 0.7007, 0.7239, 0.6900, 0.7239, and 0.3871. Moreover, we generated saliency maps to explain our model’s focus areas, showing its reliable decision-making process. The results underscore the potential of our hybrid CNN-Transformer model in aiding the early and accurate detection of gastrointestinal (GI) anomalies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Classification of Endoscopy and Video Capsule Images Using CNN-Transformer Model

  • Aliza Subedi,
  • Smriti Regmi,
  • Nisha Regmi,
  • Bhumi Bhusal,
  • Ulas Bagci,
  • Debesh Jha

摘要

Gastrointestinal cancer is the leading cause of cancer-related incidence and death. Therefore, it is important to develop a novel computer-aided diagnosis system for early detection and enhanced treatment. Traditional approaches rely on the expertise of gastroenterologists to identify diseases. However, it is a subjective process, and the interpretation can vary even between expert clinicians. Considering recent progress in classifying gastrointestinal anomalies and landmarks in endoscopic and video capsule endoscopy images, this study proposes a hybrid model incorporating the advantages of Transformers and Convolutional Neural Networks (CNNs) for enhanced classification performance. Our model utilizes DenseNet201 as a CNN branch to extract local features and integrates the Swin Transformer branch for global feature understanding. Both of their features are combined to perform the classification task. For the GastroVision dataset, our proposed model demonstrates excellent performance with Precision, Recall, F1 score, Accuracy, and Matthews Correlation Coefficient (MCC) of 0.8320, 0.8386, 0.8324, 0.8386, and 0.8191, respectively, showcasing its robustness against class imbalance dataset and surpassing other CNNs as well as Swin Transformer model. Similarly, for the Kvasir-Capsule, a large video capsule endoscopy dataset, our model surpassed all other models, thereby achieving overall Precision, Recall, F1 score, Accuracy, and MCC of 0.7007, 0.7239, 0.6900, 0.7239, and 0.3871. Moreover, we generated saliency maps to explain our model’s focus areas, showing its reliable decision-making process. The results underscore the potential of our hybrid CNN-Transformer model in aiding the early and accurate detection of gastrointestinal (GI) anomalies.