UST-SU: a U-shaped video prediction network based on partial autoregression
摘要
Effectively managing temporal dependencies and capturing accurate spatial features are dual challenges in video prediction. Autoregressive models typically use shallow recursive networks with limited time-step horizons, while non-autoregressive models may overlook inherent temporal dependencies. To tackle these challenges, we introduce UST-SU, a U-shaped spatiotemporal simple unit network. UST-SU utilizes a multilayer interactive sampling architecture to extract challenging spatial features for shallow autoregressive networks. The fusion of multi-dimensional spatial features during image reconstruction preserves crucial region-specific information. Our spatiotemporal simple unit (ST-SU), focusing on encoding global spatiotemporal information, achieves a balance in subsequent memory flows. By discarding redundant information and incorporating early-stage details, ST-SU efficiently handles spatiotemporal sequence tasks, making it a suitable core unit for integration. Additionally, we propose a partial autoregressive strategy to broaden the temporal receptive field, preserving temporal dependencies discarded by non-autoregressive models. Across diverse video prediction benchmarks, UST-SU consistently outperforms previous state-of-the-art models.