<p>This paper introduces a dual hybrid neural network model combining convolutional neural networks (CNNs) and artificial neural networks (ANNs) to optimize the quantization parameter (QP) for both <InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6954_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="54" /> </InlineMediaObject> <EquationSource Format="TEX">\(64\times 64\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>64</mn> <mo>×</mo> <mn>64</mn> </mrow> </math></EquationSource> </InlineEquation> and <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6954_Article_IEq2.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="54" /> </InlineMediaObject> <EquationSource Format="TEX">\(32\times 32\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>32</mn> <mo>×</mo> <mn>32</mn> </mrow> </math></EquationSource> </InlineEquation> blocks in the versatile video coding (VVC) standard, enhancing video quality and compression efficiency. The model employs CNNs for spatial feature extraction and ANNs for structured data handling, addressing the limitations of current heuristic and just noticeable distortion (JND)-based methods. A dataset of luminance channel image blocks, encoded with various QP values, is generated and preprocessed, and the dual hybrid network structure is designed with convolutional and dense layers. The QP optimization is applied at two levels: the <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6954_Article_IEq3.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="54" /> </InlineMediaObject> <EquationSource Format="TEX">\(64\times 64\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>64</mn> <mo>×</mo> <mn>64</mn> </mrow> </math></EquationSource> </InlineEquation> model provides a global QP offset, while the <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11227_2025_6954_Article_IEq4.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="54" /> </InlineMediaObject> <EquationSource Format="TEX">\(32\times 32\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mn>32</mn> <mo>×</mo> <mn>32</mn> </mrow> </math></EquationSource> </InlineEquation> model refines the QP for further partitioned blocks. Performance evaluations using model error metrics like mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), as well as perceptual metrics like weighted PSNR (WPSNR), MS-SSIM, PSNR-HVS-M, and VMAF, demonstrate the model’s effectiveness. While our approach performs competitively with state-of-the-art algorithms, it significantly outperforms in VMAF, the most advanced and widely adopted perceptual quality metric. Furthermore, the dual-model approach yields better results at lower resolutions, whereas the single-model approach is more effective at higher resolutions. These results highlight the adaptability of the proposed models, offering improvements in both compression efficiency and perceptual quality, making them highly suitable for practical applications in modern video coding.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Perceptual QP optimization for VVC with dual hybrid neural networks

  • Javier Ruiz Atencia,
  • Otoniel Mario López Granado,
  • Manuel Pérez Malumbres,
  • Miguel Onofre Martínez-Rach

摘要

This paper introduces a dual hybrid neural network model combining convolutional neural networks (CNNs) and artificial neural networks (ANNs) to optimize the quantization parameter (QP) for both \(64\times 64\) 64 × 64 and \(32\times 32\) 32 × 32 blocks in the versatile video coding (VVC) standard, enhancing video quality and compression efficiency. The model employs CNNs for spatial feature extraction and ANNs for structured data handling, addressing the limitations of current heuristic and just noticeable distortion (JND)-based methods. A dataset of luminance channel image blocks, encoded with various QP values, is generated and preprocessed, and the dual hybrid network structure is designed with convolutional and dense layers. The QP optimization is applied at two levels: the \(64\times 64\) 64 × 64 model provides a global QP offset, while the \(32\times 32\) 32 × 32 model refines the QP for further partitioned blocks. Performance evaluations using model error metrics like mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), as well as perceptual metrics like weighted PSNR (WPSNR), MS-SSIM, PSNR-HVS-M, and VMAF, demonstrate the model’s effectiveness. While our approach performs competitively with state-of-the-art algorithms, it significantly outperforms in VMAF, the most advanced and widely adopted perceptual quality metric. Furthermore, the dual-model approach yields better results at lower resolutions, whereas the single-model approach is more effective at higher resolutions. These results highlight the adaptability of the proposed models, offering improvements in both compression efficiency and perceptual quality, making them highly suitable for practical applications in modern video coding.