<p>Modern multimedia applications increasingly require vision models that are not only accurate but also computationally efficient and capable of adapting their processing effort according to input complexity. Although Vision Transformers (ViTs) have achieved state-of-the-art performance across a wide range of visual recognition tasks, their fixed-depth architectures and high computational cost remain major obstacles to deployment in real-time and resource-constrained environments. To address these limitations, we propose the Mixture-of-Recursions Vision Transformer (MoR-ViT), a unified adaptive framework that jointly integrates recursive refinement and sparse expert routing through a lightweight depth–routing controller. Unlike existing adaptive transformer architectures that optimize recursion, routing, or token adaptation independently, MoR-ViT dynamically determines both the required computation depth and the most relevant expert pathways for each input. The proposed architecture combines (i) a recursive depth controller that adaptively regulates the number of refinement iterations and (ii) a content-aware gating mechanism that selectively routes representations through specialized recursive experts, enabling efficient and interpretable input-dependent computation. Extensive experiments on CIFAR-100, ImageNet-1&#xa0;K, ImageNet-A, ImageNet-V2, and COCO-2017 demonstrate that MoR-ViT consistently outperforms representative static, sparse, and adaptive Vision Transformer baselines, including ViT-B/16, DeiT-S, DynamicViT, Evo-ViT, and TokenLearner. The proposed model achieves 84.3% Top-1 accuracy on CIFAR-100, 83.3% on ImageNet-1&#xa0;K, 21.0% on ImageNet-A, and 53.9% mAP on COCO-2017 while reducing inference FLOPs and energy consumption by up to 28%. Furthermore, the proposed adaptive computation strategy operates with an average recursion depth of only 1.74, substantially below the maximum allowable depth, thereby providing significant computational savings without compromising predictive performance. Ablation studies, complexity analysis, and recursion-depth visualizations further confirm that adaptive recursion and expert routing contribute complementary benefits to efficiency, robustness, and interpretability. Overall, MoR-ViT advances adaptive Vision Transformers by introducing a unified depth–routing optimization framework that enables efficient, scalable, and interpretable multimedia vision, making it particularly suitable for deployment in mobile, edge-AI, and real-time visual intelligence systems.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Adaptive Mixture-of-Recursions vision transformer with joint depth routing for multimedia vision

  • Rathinasamy Muthusami,
  • Kandhasamy Saritha

摘要

Modern multimedia applications increasingly require vision models that are not only accurate but also computationally efficient and capable of adapting their processing effort according to input complexity. Although Vision Transformers (ViTs) have achieved state-of-the-art performance across a wide range of visual recognition tasks, their fixed-depth architectures and high computational cost remain major obstacles to deployment in real-time and resource-constrained environments. To address these limitations, we propose the Mixture-of-Recursions Vision Transformer (MoR-ViT), a unified adaptive framework that jointly integrates recursive refinement and sparse expert routing through a lightweight depth–routing controller. Unlike existing adaptive transformer architectures that optimize recursion, routing, or token adaptation independently, MoR-ViT dynamically determines both the required computation depth and the most relevant expert pathways for each input. The proposed architecture combines (i) a recursive depth controller that adaptively regulates the number of refinement iterations and (ii) a content-aware gating mechanism that selectively routes representations through specialized recursive experts, enabling efficient and interpretable input-dependent computation. Extensive experiments on CIFAR-100, ImageNet-1 K, ImageNet-A, ImageNet-V2, and COCO-2017 demonstrate that MoR-ViT consistently outperforms representative static, sparse, and adaptive Vision Transformer baselines, including ViT-B/16, DeiT-S, DynamicViT, Evo-ViT, and TokenLearner. The proposed model achieves 84.3% Top-1 accuracy on CIFAR-100, 83.3% on ImageNet-1 K, 21.0% on ImageNet-A, and 53.9% mAP on COCO-2017 while reducing inference FLOPs and energy consumption by up to 28%. Furthermore, the proposed adaptive computation strategy operates with an average recursion depth of only 1.74, substantially below the maximum allowable depth, thereby providing significant computational savings without compromising predictive performance. Ablation studies, complexity analysis, and recursion-depth visualizations further confirm that adaptive recursion and expert routing contribute complementary benefits to efficiency, robustness, and interpretability. Overall, MoR-ViT advances adaptive Vision Transformers by introducing a unified depth–routing optimization framework that enables efficient, scalable, and interpretable multimedia vision, making it particularly suitable for deployment in mobile, edge-AI, and real-time visual intelligence systems.