With the rapid popularization of electronic devices, the large amount of data generated by camera imaging poses a huge challenge to the limited storage capacity and communication bandwidth. Achieving higher compression ratios without sacrificing visual quality remains a fundamental challenge for image compression. In this paper, we propose a novel full flow bidirectional visual threshold estimation method for camera imaging perceptual compression. Specifically, we study the features from camera imaging to visual perception to semantic understanding, and characterize them with focus identification, perceptual distribution, and semantic segmentation respectively. We also carefully design feature extraction networks suitable for each feature type. In addition, we draw inspiration from the bidirectional perceptual mechanism of the human visual system and propose a feature extraction framework that adopts top-down and bottom-up methods. We further enhance our model by regulating and fusing bidirectional perceptual features through a gated decoding structure. Extensive experimental validation on benchmark datasets confirms that our FPSNet significantly improves the accuracy of visual redundancy prediction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FPSNet: Focus-Perceptual-Semantic Full Flow Visual Redundancy Predicting for Camera Image

  • Xiongwei Xiao

摘要

With the rapid popularization of electronic devices, the large amount of data generated by camera imaging poses a huge challenge to the limited storage capacity and communication bandwidth. Achieving higher compression ratios without sacrificing visual quality remains a fundamental challenge for image compression. In this paper, we propose a novel full flow bidirectional visual threshold estimation method for camera imaging perceptual compression. Specifically, we study the features from camera imaging to visual perception to semantic understanding, and characterize them with focus identification, perceptual distribution, and semantic segmentation respectively. We also carefully design feature extraction networks suitable for each feature type. In addition, we draw inspiration from the bidirectional perceptual mechanism of the human visual system and propose a feature extraction framework that adopts top-down and bottom-up methods. We further enhance our model by regulating and fusing bidirectional perceptual features through a gated decoding structure. Extensive experimental validation on benchmark datasets confirms that our FPSNet significantly improves the accuracy of visual redundancy prediction.