TG-STC: Text Guided Spatio-Temporal Contextualization for Video Quality Assessment
摘要
Video Quality Assessment (VQA) is crucial for enhancing user experiences in multimedia applications, particularly in video review, and quality optimization. Current methods predominantly focus on conventional visual feature extraction techniques, yet overlook the benefits of leveraging semantic information from text in enhancing video representation. To address this gap, we propose a Text-Guided Spatio-Temporal Contextualization (TG-STC) module designed to dynamically incorporate textual semantics into the video feature extraction process. Our framework combines global semantic features from CLIP with Video Swin Transformer, enhanced by a novel multimodal feature adaptation module. This module employs gated fusion and adaptive linear transformations to align textual and visual modalities, enabling the model to focus on quality-relevant regions (e.g., sharpness, exposure). Extensive experiments on MAXWELL and KoNViD-1k demonstrates the effectiveness of our method. Our work bridges the gap between textual guidance and visual feature modeling in VQA, offering enhanced interpretability and accuracy for fine-grained quality evaluation.