MDIF: A multimodal dynamic inference framework for traffic video question answering
摘要
This paper presents a Multimodal Dynamic Inference Framework (MDIF) for video question answering in traffic scenarios. MDIF is different from existing methods that rely on correlation modeling and global features. It combines causal inference with object segmentation. Causal inference reduces spurious correlations and ensures unbiased reasoning. Object segmentation extracts fine-grained information such as traffic signs, vehicle movements, and pedestrian interactions. Together, they improve causal modeling, scene understanding, and event prediction accuracy. MDIF also employs Graph Convolutional Networks (GCNs) to capture spatiotemporal dependencies in complex and dynamic traffic events. We evaluate MDIF on two large-scale datasets, SUTD-TrafficQA and MSVD-QA. The results demonstrate significant progress in prediction and counterfactual tasks, validating the robustness and generalization ability of the MDIF framework. Ablation studies further confirm the role of spatiotemporal modeling, object segmentation, and causal intervention. These components are all essential to robust, interpretable, and unbiased reasoning. Overall, MDIF provides strong advantages in causal modeling and cross-modal inference.