WHformer: LDCT Image Denoising Network Based on Multi-scale Feature Fusion in the Wavelet Domain and Cascaded Attention Mechanism
摘要
Low-dose computed tomography (LDCT) has become the core technology in clinical diagnosis, but quantum noise introduced by reducing radiation dose seriously affects image quality and diagnostic accuracy. To solve this problem, we propose WHformer, an LDCT denoising network based on multi-scale feature fusion in the wavelet domain and cascaded attention mechanism. Firstly, the input image is decomposed into high-frequency and low-frequency components by using discrete wavelet transform (DWT), and the multi-scale feature components are aligned by transposition convolution operation as the input of the network. Secondly, an enhanced convolutional gating block (ECGB) and multi-scale attention block (MAB) are proposed, and their cascades are used as the core components of codecs to achieve the cooperative expression of local and global features by combining deep separable convolution and multi-scale attention mechanisms. In addition, a cross-dimensional attention module (CAM) is introduced between codecs. This module not only enhances the capability of cross-level feature fusion but also overcomes the limitation of traditional attention mechanisms that focus solely on a single dimension. The experimental results of the algorithm in this paper on the AAPM Challenge dataset show that, compared with the suboptimal model, PSNR and SSIM have increased by 1.51% and 3.19% respectively, RMSE has decreased by 8.05%, and the best effect has also been achieved on the Piglet dataset. These results fully verify that WHformer model proposed by us can not only effectively suppress the noise in LDCT images, but also retain more image details and significantly improve LDCT image quality.