<p>Speaker verification technology plays a crucial role in biometric authentication and security applications. However, the development of systems for low-resource languages such as Tibetan continues to encounter unique challenges. The complex phonology and distinctive acoustics of Tibetan, combined with the limited availability of speech data, make it difficult for conventional models to achieve strong performance. To address these challenges, this paper proposes an improved CAM–SKANF (CAM–Selective Kernel and Kolmogorov–Arnold Networks Fusion) network, built on the CAM++ framework and tailored to Tibetan speaker verification. First, we integrate Selective Kernel 2D Convolution (SK-2D) into the front-end ResBlock, forming the SKResBlend module. This module employs a multi-scale parallel convolution structure with a dynamic kernel selection mechanism, enhancing acoustic feature extraction and enabling more precise modeling of fine-grained variations and long-term dependencies. Second, we replace the original transition layer after the Dense Block with a KANs layer. Using the Kolmogorov–Arnold function decomposition, this layer optimizes high-dimensional feature representation, reduces dependence on large-scale data, and helps prevent overfitting. Finally, we introduce the Multi-scale Attention Normalization Fusion (MANF) module between adjacent KANs layers to fuse shallow local features with deep global semantics, strengthening the model’s robustness and representational capacity. Experimental results show that CAM–SKANF achieves an EER (Equal Error Rate) of 5.5476%, an improvement of about 34% over CAM++. The MinDCF (Minimum Detection Cost Function) is reduced to 0.7291, an 11.5% improvement relative to the baseline. These results demonstrate that CAM–SKANF significantly enhances discrimination ability in Tibetan speaker verification, reducing errors and improving system reliability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

CAM–SKANF: A Multi-scale and Feature-Fusion Network for Tibetan Speaker Verification

  • Zhenye Gan,
  • Le Wei

摘要

Speaker verification technology plays a crucial role in biometric authentication and security applications. However, the development of systems for low-resource languages such as Tibetan continues to encounter unique challenges. The complex phonology and distinctive acoustics of Tibetan, combined with the limited availability of speech data, make it difficult for conventional models to achieve strong performance. To address these challenges, this paper proposes an improved CAM–SKANF (CAM–Selective Kernel and Kolmogorov–Arnold Networks Fusion) network, built on the CAM++ framework and tailored to Tibetan speaker verification. First, we integrate Selective Kernel 2D Convolution (SK-2D) into the front-end ResBlock, forming the SKResBlend module. This module employs a multi-scale parallel convolution structure with a dynamic kernel selection mechanism, enhancing acoustic feature extraction and enabling more precise modeling of fine-grained variations and long-term dependencies. Second, we replace the original transition layer after the Dense Block with a KANs layer. Using the Kolmogorov–Arnold function decomposition, this layer optimizes high-dimensional feature representation, reduces dependence on large-scale data, and helps prevent overfitting. Finally, we introduce the Multi-scale Attention Normalization Fusion (MANF) module between adjacent KANs layers to fuse shallow local features with deep global semantics, strengthening the model’s robustness and representational capacity. Experimental results show that CAM–SKANF achieves an EER (Equal Error Rate) of 5.5476%, an improvement of about 34% over CAM++. The MinDCF (Minimum Detection Cost Function) is reduced to 0.7291, an 11.5% improvement relative to the baseline. These results demonstrate that CAM–SKANF significantly enhances discrimination ability in Tibetan speaker verification, reducing errors and improving system reliability.