User-generated content video quality assessment (UGC-VQA) is fundamentally a challenging task due to diverse distortion factors and limited annotated samples. Most existing approaches focus mainly on the underlying features but neglect to explore the high-level semantic information to understand the video quality. In order to overcome this limitation, we propose a novel semantic category-aware multi-scale network model SCAMS-Net for UGC-VQA. Structurally, we integrate multi-scale feature mapping into the intermediate layers of convolutional neural networks and Transformer networks, enabling the model to exploit a variety of perceptual features ranging from low-level color and texture details to high-level semantic content. Methodologically, we employ a novel pseudo-label generation strategy to obtain semantic category information, and explore the correlation between video quality scores and semantic categories through a supervised contrastive learning strategy. Experimental results show that our proposed SCAMS-Net outperforms many state-of-the-art methods on three publicly available VQA datasets, highlighting the importance of incorporating semantic perception into video quality assessment.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

SCAMS: Semantic Category-Aware Multi-scale Network for Video Quality Assessment

  • Longgang Ren,
  • Kaibing Zhang,
  • Dandan Fan,
  • Guang Shi

摘要

User-generated content video quality assessment (UGC-VQA) is fundamentally a challenging task due to diverse distortion factors and limited annotated samples. Most existing approaches focus mainly on the underlying features but neglect to explore the high-level semantic information to understand the video quality. In order to overcome this limitation, we propose a novel semantic category-aware multi-scale network model SCAMS-Net for UGC-VQA. Structurally, we integrate multi-scale feature mapping into the intermediate layers of convolutional neural networks and Transformer networks, enabling the model to exploit a variety of perceptual features ranging from low-level color and texture details to high-level semantic content. Methodologically, we employ a novel pseudo-label generation strategy to obtain semantic category information, and explore the correlation between video quality scores and semantic categories through a supervised contrastive learning strategy. Experimental results show that our proposed SCAMS-Net outperforms many state-of-the-art methods on three publicly available VQA datasets, highlighting the importance of incorporating semantic perception into video quality assessment.