Global-regional-local multilevel lightweight attention modeling for event-based efficient video reconstruction
摘要
High-speed motion and complex lighting conditions pose great challenges to vision tasks based on traditional frame cameras. Event cameras emerge out to address them but encounter difficulties in applying computer vision algorithms on the event streams—a novel imaging data paradigm. Existing state-of-the-art methods overly emphasize improving the quality of reconstructed videos, while neglecting the subsequent application issues of model deployment and output videos. Due to the large number of parameters, they are too inefficient for edge embedded devices. Besides, those manual design methods for batch preprocessing data are clumsy and lack generalization ability for different scenarios. This paper focuses on event-based efficient video reconstruction for further high-level vision tasks. A novel global-regional-local multilevel lightweight attention hierarchical architecture is proposed, termed Global-Regional-Local-E2VID (GRL-E2VID). This frame-work leverages chunked incoherent radiation attention and membrane potential tensor transformation to establish global, regional, and local dependencies on asynchronous and sparse events, thereby enhancing the quality and efficiency of video reconstruction. Experimental results demonstrate that, with the reduction of 60% in parameter quantity, our approach maintains a video reconstruction quality comparable to the state-of-the-art method. Moreover, it aids in object classification, raising small object recognition confidence by 60% and strengthening model stability in complex scenarios. For object tracking, the multimodal algorithm based on GRL-E2VID’s reconstructed video other than RGB video doubles tracking efficiency. Besides, its ability of generating high-frame-rate clear video under challenging environments such as complex illumination and high-speed motion promises its value in high-level visual tasks.