OptiGQA: LLM-Driven Query Optimization for Efficient Visual Grounding in Adaptive Video Question Answering
摘要
Video Question Answering (VideoQA) remains challenging due to the complexity of video data and the diverse nature of questions. Despite significant advancements in large language models (LLMs) and Vision-language models (VLMs), current VideoQA systems often fall short due to their reliance on shallow reasoning, ineffective frame selection, and a one-size-fits-all approach to answering questions. These limitations hinder their ability to perform deep reasoning, precise localization, and nuanced understanding of causal relationships. In this paper, we propose OptiGQA, a novel framework that integrates the advanced reasoning capabilities of LLMs with the efficient grounding strengths of small Video Grounding Models. OptiGQA enhances VideoQA by rewriting questions to compensate for implicit causal information missed in the query and employing a video grounding model to identify key visual cues. Furthermore, it dynamically selects appropriate processing strategies tailored to different question types, ensuring optimal performance for both global and local queries. Experiments on three standard VideoQA datasets, including NExT-QA, NExT-GQA, and IntentQA, demonstrate that the proposed method outperforms strong baselines and achieves superior localization. Notably, OptiGQA enhances computational efficiency by utilizing fewer frame captions, making it both effective and efficient, advancing the capabilities of VideoQA systems.