<p>Search and retrieval based on visual content instead of metadata is made possible by content-based video retrieval (CBVR), which calls for effective learning, indexing, and adaption strategies to manage massive amounts of video data. This study offers a strong framework that tackles important CBVR issues such as network retraining, resource limitations, and effective indexing. We suggest a maximum ensemble weighted softmax layer to encourage fewer classification iterations and overfitting. For effective indexing, deep neural network feature embeddings are utilized. While ongoing and prototype-based learning approaches facilitate incremental retraining and compact video representation, pre-trained models and transfer learning are employed to reduce computing overhead. Our method’s superiority is demonstrated by evaluation on two benchmark datasets, where it achieved top-1 accuracy of 0.6032 and 0.8474 on HMDB51 and UCF101 datasets respectively, utilizing L2 distance for similarity measurement. The outcomes validate the suggested method’s efficacy and scalability for actual CBVR applications in resource-constrained settings.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Stacked max ensemble weighted softmax layer for deep spatio-temporal action video retrieval

  • Alina Banerjee,
  • Ravinder M,
  • Ela Kumar

摘要

Search and retrieval based on visual content instead of metadata is made possible by content-based video retrieval (CBVR), which calls for effective learning, indexing, and adaption strategies to manage massive amounts of video data. This study offers a strong framework that tackles important CBVR issues such as network retraining, resource limitations, and effective indexing. We suggest a maximum ensemble weighted softmax layer to encourage fewer classification iterations and overfitting. For effective indexing, deep neural network feature embeddings are utilized. While ongoing and prototype-based learning approaches facilitate incremental retraining and compact video representation, pre-trained models and transfer learning are employed to reduce computing overhead. Our method’s superiority is demonstrated by evaluation on two benchmark datasets, where it achieved top-1 accuracy of 0.6032 and 0.8474 on HMDB51 and UCF101 datasets respectively, utilizing L2 distance for similarity measurement. The outcomes validate the suggested method’s efficacy and scalability for actual CBVR applications in resource-constrained settings.