Video coding is essential for efficient storage, transmission, and playback of video content across digital platforms. Numerous techniques have been developed to balance file size reduction with the preservation of perceptual video quality. Traditional standards like H.264 and H.265 have dominated the field, employing methods such as motion estimation and discrete cosine transform (DCT). However, these approaches face challenges in achieving significant compression gains while maintaining high visual quality, particularly at low bit rates. The introduction of deep learning has revolutionized video coding by leveraging neural networks to exploit spatial and temporal correlations within video frames, resulting in improved compression efficiency and quality. Early developments in deep learning-based video compression began in the 2010s, aided by advancements in GPUs and large datasets. This field has rapidly evolved, with various techniques emerging, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs). These deep learning methods can be categorized based on techniques like intra-frame or inter-frame coding, architecture selections, and the degree of end-to-end compression. Overall, the integration of deep learning into video compression has led to significant advancements, enhancing video storage, transmission, and streaming without sacrificing quality. The following sections will explore specific techniques and innovations in this domain.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning-based Video Coding

  • Wei Gao

摘要

Video coding is essential for efficient storage, transmission, and playback of video content across digital platforms. Numerous techniques have been developed to balance file size reduction with the preservation of perceptual video quality. Traditional standards like H.264 and H.265 have dominated the field, employing methods such as motion estimation and discrete cosine transform (DCT). However, these approaches face challenges in achieving significant compression gains while maintaining high visual quality, particularly at low bit rates. The introduction of deep learning has revolutionized video coding by leveraging neural networks to exploit spatial and temporal correlations within video frames, resulting in improved compression efficiency and quality. Early developments in deep learning-based video compression began in the 2010s, aided by advancements in GPUs and large datasets. This field has rapidly evolved, with various techniques emerging, including convolutional neural networks (CNNs) and recurrent neural networks (RNNs). These deep learning methods can be categorized based on techniques like intra-frame or inter-frame coding, architecture selections, and the degree of end-to-end compression. Overall, the integration of deep learning into video compression has led to significant advancements, enhancing video storage, transmission, and streaming without sacrificing quality. The following sections will explore specific techniques and innovations in this domain.