Video Quality Assessment (VQA) is crucial for enhancing user experiences in multimedia applications, particularly in video review, and quality optimization. Current methods predominantly focus on conventional visual feature extraction techniques, yet overlook the benefits of leveraging semantic information from text in enhancing video representation. To address this gap, we propose a Text-Guided Spatio-Temporal Contextualization (TG-STC) module designed to dynamically incorporate textual semantics into the video feature extraction process. Our framework combines global semantic features from CLIP with Video Swin Transformer, enhanced by a novel multimodal feature adaptation module. This module employs gated fusion and adaptive linear transformations to align textual and visual modalities, enabling the model to focus on quality-relevant regions (e.g., sharpness, exposure). Extensive experiments on MAXWELL and KoNViD-1k demonstrates the effectiveness of our method. Our work bridges the gap between textual guidance and visual feature modeling in VQA, offering enhanced interpretability and accuracy for fine-grained quality evaluation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TG-STC: Text Guided Spatio-Temporal Contextualization for Video Quality Assessment

  • Xinjie Jin,
  • Pengfei Duan,
  • Shanshan Yang,
  • Shengwu Xiong

摘要

Video Quality Assessment (VQA) is crucial for enhancing user experiences in multimedia applications, particularly in video review, and quality optimization. Current methods predominantly focus on conventional visual feature extraction techniques, yet overlook the benefits of leveraging semantic information from text in enhancing video representation. To address this gap, we propose a Text-Guided Spatio-Temporal Contextualization (TG-STC) module designed to dynamically incorporate textual semantics into the video feature extraction process. Our framework combines global semantic features from CLIP with Video Swin Transformer, enhanced by a novel multimodal feature adaptation module. This module employs gated fusion and adaptive linear transformations to align textual and visual modalities, enabling the model to focus on quality-relevant regions (e.g., sharpness, exposure). Extensive experiments on MAXWELL and KoNViD-1k demonstrates the effectiveness of our method. Our work bridges the gap between textual guidance and visual feature modeling in VQA, offering enhanced interpretability and accuracy for fine-grained quality evaluation.