Automatic modulation recognition using vision transformers with cross-layer feature fusion and Kolmogorov–Arnold representation
摘要
In recent years, deep learning-based automatic modulation recognition (AMR) methods have garnered significant attention due to their high recognition accuracy and adaptability across various signal types. However, achieving precise AMR in scenarios involving mixed modulation categories remains challenging. Therefore, we propose a practical network called multi-scale feature fusion vision transformer (MSFF-ViT), aimed at enhancing effective feature learning and subsequently improving recognition accuracy. MSFF-ViT is achieved through the integration of vertical shortcut connections and horizontal subspace attention mechanisms. This design enhances the network’s capability to learn effective features across different information scales. MSFF-ViT leverages advanced deep learning architecture to improve recognition performance. Additionally, to enhance parameter efficiency and capture intricate data relationships, the Kolmogorov–Arnold network structure is introduced as the classifier. The comparative experimental results on the RadioML2018.01A and RadioML2016.10A datasets demonstrate that the proposed method outperforms mainstream approaches.