Existing speech semantic communication systems often struggle with insufficient robustness in dynamic and low signal-to-noise ratio (SNR) environments. This paper introduces the REFINE framework, designed to ensure robust data transmission under complex channel conditions. REFINE employs a low-complexity, trainable semantic encoder to extract essential speech features and dynamically suppress noise using the diffusion sampling technique within an adaptive semantic-channel gain mechanism (ASGM). This mechanism leverages noise parameters derived from channel state information (CSI) for real-time adaptation. REFINE maintains high semantic fidelity while minimizing signal distortion across diverse SNRs and channel conditions. Simulation results indicate that REFINE achieves a significant performance gain of approximately 2.5 dB compared to the state-of-the-art DSST system in low SNR scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

REFINE: Reliable Encoding and Fine-Tuned Noise Elimination for Adaptive Semantic Speech Transmission

  • Shengliang Wu,
  • Hongzhi Pan,
  • Ting Zu,
  • Weiwei Jiang,
  • Yujun Zhu,
  • Xin He,
  • Haisheng Tan

摘要

Existing speech semantic communication systems often struggle with insufficient robustness in dynamic and low signal-to-noise ratio (SNR) environments. This paper introduces the REFINE framework, designed to ensure robust data transmission under complex channel conditions. REFINE employs a low-complexity, trainable semantic encoder to extract essential speech features and dynamically suppress noise using the diffusion sampling technique within an adaptive semantic-channel gain mechanism (ASGM). This mechanism leverages noise parameters derived from channel state information (CSI) for real-time adaptation. REFINE maintains high semantic fidelity while minimizing signal distortion across diverse SNRs and channel conditions. Simulation results indicate that REFINE achieves a significant performance gain of approximately 2.5 dB compared to the state-of-the-art DSST system in low SNR scenarios.