Preprocessing Enhanced Image Compression for Machine Vision

被引：0

作者：

Lu, Guo ^{[1
]}

Ge, Xingtong ^{[2
,3
]}

Zhong, Tianxiong ^{[2
]}

Hu, Qiang ^{[1
]}

Geng, Jing ^{[2
]}

机构：

[1] Shanghai Jiao Tong Univ, Sch Elect Informat & Elect Engn, Shanghai 200240, Peoples R China

[2] Beijing Inst Technol, Sch Comp Sci, Beijing 100081, Peoples R China

[3] SenseTime Technol, Beijing 100080, Peoples R China

来源：

IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY | 2024年 / 34卷 / 12期

关键词：

Image coding; Task analysis; Codecs; Machine vision; Bit rate; Optimization; Neural networks; Image compression; machine vision; preprocessing; deep learning; OPTIMIZATION; FRAMEWORK;

D O I：

10.1109/TCSVT.2024.3441049

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

Recently, more and more images are compressed and sent to the back-end devices for machine analysis tasks (e.g., object detection) instead of being purely watched by humans. However, most traditional or learned image codecs are designed to minimize the distortion of the human visual system without considering the increased demand from machine vision systems. In this work, we propose a preprocessing enhanced image compression method for machine vision tasks to address this challenge. Instead of relying on the learned image codecs for end-to-end optimization, our framework is built upon the traditional non-differential codecs, which means it is standard compatible and can be easily deployed in practical applications. Specifically, we propose a neural preprocessing module before the encoder to maintain the useful semantic information for the downstream tasks and suppress the irrelevant information for bitrate saving. Furthermore, our neural preprocessing module is quantization adaptive and can be used in different compression ratios. More importantly, to jointly optimize the preprocessing module with the downstream machine vision tasks, we introduce the proxy network for the traditional non-differential codecs in the back-propagation stage. We provide extensive experiments by evaluating our compression method for several representative downstream tasks with different backbone networks. Experimental results show our method achieves a better trade-off between the coding bitrate and the performance of the downstream machine vision tasks by saving about 20% bitrate.

引用

页码：13556 / 13568

页数：13

共 73 条

[1] Akbari M, 2019, INT CONF ACOUST SPEE, P2042, DOI [10.1109/icassp.2019.8683541, 10.1109/ICASSP.2019.8683541]
[2] [Anonymous], 2020, CISCO ANN INTERNET R
[3] [Anonymous], 2024, Libjpeg-Turbo
[4] Ball Johannes, 2017, INT C LEARN REPR
[5] Balle J., 2018, INT C LEARN REPR
[6] Bellard F., 2018, BPG image format
[7] Bjontegaard G., 2001, Doc. VCEG-M33
[8] Overview of the Versatile Video Coding (VVC) Standard and its Applications
Bross, Benjamin
Wang, Ye-Kui
Ye, Yan
Liu, Shan
Chen, Jianle
Sullivan, Gary J.
Ohm, Jens-Rainer
[J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2021, 31 (10) : 3736 - 3764
[9] Deep Perceptual Preprocessing for Video Coding
Chadha, Aaron
Andreopoulos, Yiannis
[J]. 2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, : 14847 - 14856
[10] Compact Temporal Trajectory Representation for Talking Face Video Compression
Chen, Bolin
Wang, Zhao
Li, Binzhe
Wang, Shiqi
Ye, Yan
[J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2023, 33 (11) : 7009 - 7023

← 1 2 3 4 5 6 7 8 →