<p>Automatic electrocardiogram (ECG) classification plays a crucial role in the early prevention and assisted diagnosis of cardiovascular diseases. Existing approaches typically leverage Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to extract ECG features, yet they often fail to effectively integrate spatiotemporal information, limiting the joint modelling of morphological and rhythmic patterns. Additionally, their inherent black-box nature restricts clinical interpretability. To address these challenges, we propose a dual-branch model that integrates CNN and ViT in parallel to extract local and global ECG features. The ViT backbone is designed to learn global representations, while the CNN branch focuses on multi-scale local feature extraction. Additionally, we introduce a cross-attention-based Spatial Feature Injection Module (SFIM) to address the limitations of the ViT backbone in capturing fine-grained information. The SFIM extracts local features from the CNN branch and injects them into the ViT backbone, thereby facilitating local-global feature integration and enhancing the interpretability of the model. The proposed method achieves F1 scores of 84.33% on the CPSC dataset, 96.28% on the four-class classification task of the Chapman dataset, and 54.39% on the rhythm classification task of the PTB-XL dataset. Compared to existing approaches, our model demonstrates superior robustness and improved interpretability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A multi-scale CNN-Transformer parallel network for 12-lead ECG signal classification

  • Long Wei,
  • Yang Li

摘要

Automatic electrocardiogram (ECG) classification plays a crucial role in the early prevention and assisted diagnosis of cardiovascular diseases. Existing approaches typically leverage Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to extract ECG features, yet they often fail to effectively integrate spatiotemporal information, limiting the joint modelling of morphological and rhythmic patterns. Additionally, their inherent black-box nature restricts clinical interpretability. To address these challenges, we propose a dual-branch model that integrates CNN and ViT in parallel to extract local and global ECG features. The ViT backbone is designed to learn global representations, while the CNN branch focuses on multi-scale local feature extraction. Additionally, we introduce a cross-attention-based Spatial Feature Injection Module (SFIM) to address the limitations of the ViT backbone in capturing fine-grained information. The SFIM extracts local features from the CNN branch and injects them into the ViT backbone, thereby facilitating local-global feature integration and enhancing the interpretability of the model. The proposed method achieves F1 scores of 84.33% on the CPSC dataset, 96.28% on the four-class classification task of the Chapman dataset, and 54.39% on the rhythm classification task of the PTB-XL dataset. Compared to existing approaches, our model demonstrates superior robustness and improved interpretability.