A Robust Watermarking Method for Speech Legitimacy Authentication Based on SSVR and TFJF
摘要
In recent years, the proliferation of deepfake audio has raised new requirements for methods of speech legitimacy authentication. For robust watermarking used for legitimacy authentication, there are two key problems: How to balance robustness and imperceptibility, how to enhance resistance to de-synchronization attacks. Existing robust watermarking algorithms are difficult to resist de-synchronization attacks while maintaining better imperceptibility and robustness against other attacks. This paper proposes a robust watermarking method based on two new features time–frequency joint feature (TFJF) and segmental singular values ratio (SSVR), which not only resists de-synchronization attacks but also exhibits better robustness against common attacks while ensuring payload capacity and imperceptibility. In this method, firstly, the SSVR features are extracted by segmenting the speech frames and applying Dual‑tree complex wavelet transform (DTCWT) and singular value decomposition (SVD). Then, Mel-frequency cepstral coefficients (MFCC) and weighted energy are extracted to construct TFJF, and SSVR are adaptively modified based on the TFJF to embed the watermarking information. Experimental results show that the SSVR-based robust watermarking method can resist de-synchronization attacks while maintaining robustness against common attacks, and the TFJF-based adaptive embedding method effectively improves the imperceptibility of the proposed watermarking while ensuring robustness. The proposed method has an average BER of 10.234% for de-synchronization attacks and 3.4871% for common attacks, while maintaining SNR of 21.1652 dB and PESQ of 3.4459 with payload capacity of 50bps.