CVC: Contrastive Autoencoders for Video Colorization
摘要
The process of colorizing visual media has existed since the inception of photography, as color is a crucial element that enhances the quality and appeal of such representations. The advent of deep learning techniques has led to significant advancements in video colorization, with the development of deep learning video colorization (DLVC) gaining prominence. These solutions typically entail the introduction of novel architectures or methodologies that integrate color and temporal information from videos. However, these approaches do not align with the self-supervised learning methodologies prevalent in the computer vision community, such as contrastive and autoencoder methods. In this paper, we propose a novel framework for training deep learning video colorization (DLVC) models. To the best of our knowledge, this is the first self-supervision training framework in this domain. To validate our framework, we implemented two distinct architectures, ViT and ResNet50, which were trained on the DAVIS, LDV, and UVO datasets. The results obtained from this study demonstrated superior performance on these datasets when compared with state-of-the-art models in the DLVC literature.