STAGVid2C: enhancing video-based commonsense captioning with spatio-temporal action graph
摘要
Video-based commonsense captioning is a core task in the field of video understanding that aims to generate underlying commonsense knowledge. Existing studies leverage rich visual and text information for multimodel video captioning, however, neglect the fine-grained spatial and temporal features in videos. To address these issues, this paper proposes STAGVid2C, a novel model that utilizes the Spatio-Temporal Action Graph to enhance Video-based Commonsense Captioning. The graph uses the grid features of video frames as nodes, and employs the spatial connections and temporal similarity between video frames as edges. Then, a graph neural network is employed to learn the graph representation which can capture more semantic relationships between image and action in videos. After that, STAGVid2C uses a memory network to facilitate cross-modal fusion and alignment, thereby generating more accurate commonsense captions. Experimental results on the public V2C dataset demonstrate that our proposed STAGVid2C significantly outperforms the state-of-the-art methods. Especially, in terms of CIDEr, it shows an average score improvement of 7.3% in the intention part and 8.7% in the effect part.