Perceptual Quality Assessment of Face Video Compression: A Benchmark and An Effective Method

被引：4

作者：

Li, Yixuan ^{[1
]}

Chen, Bolin ^{[1
]}

Chen, Baoliang ^{[1
]}

Wang, Meng ^{[1
]}

Wang, Shiqi ^{[1
]}

Lin, Weisi ^{[2
]}

机构：

[1] City Univ Hong Kong, Dept Comp Sci, Kowloon, Hong Kong, Peoples R China

[2] Nanyang Technol Univ, Sch Comp Engn, Singapore 639798, Singapore

来源：

IEEE TRANSACTIONS ON MULTIMEDIA | 2024年 / 26卷

关键词：

Faces; Quality assessment; Video compression; Face recognition; Image coding; Video recording; Streaming media; Face video compression; video quality assessment; subjective and objective study;

D O I：

10.1109/TMM.2024.3380260

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Recent years have witnessed an exponential increase in the demand for face video compression, and the success of artificial intelligence has expanded the boundaries beyond traditional hybrid video coding. Generative coding approaches have been identified as promising alternatives with reasonable perceptual rate-distortion trade-offs, leveraging the statistical priors of face videos. However, the great diversity of distortion types in spatial and temporal domains, ranging from the traditional hybrid coding frameworks to generative models, present grand challenges in compressed face video quality assessment (VQA) that plays a crucial role in the whole delivery chain for quality monitoring and optimization. In this paper, we introduce the large-scale Compressed Face Video Quality Assessment (CFVQA) database, which is the first attempt to systematically understand the perceptual quality and diversified compression distortions in face videos. The database contains 3,240 compressed face video clips in multiple compression levels, which are derived from 135 source videos with diversified content using six representative video codecs, including two traditional methods based on hybrid coding frameworks, two end-to-end methods, and two generative methods. The unique characteristics of CFVQA, including large-scale, fine-grained, great content diversity, and cross-compression distortion types, make the benchmarking for existing image quality assessment (IQA) and VQA feasible and practical. The results reveal the weakness of existing IQA and VQA models, which challenge real-world face video applications. In addition, a FAce VideO IntegeRity (FAVOR) index for face video compression was developed to measure the perceptual quality, considering the distinct content characteristics and temporal priors of the face videos. Experimental results exhibit its superior performance on the proposed CFVQA dataset.

引用

页码：8596 / 8608

页数：13

共 97 条

[21]

Hernandez-Ortega J, 2019, INT CONF BIOMETR

[22] FVC: A New Framework towards Deep Video Compression in Feature Space [J].

Hu, Zhihao ;

Lu, Guo ;

Xu, Dong .

2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, :1502-1511

[23]

Installations T., 1999, Networks, V910, P5

[24] Content-Aware Convolutional Neural Network for In-Loop Filtering in High Efficiency Video Coding [J].

Jia, Chuanmin ;

Wang, Shiqi ;

Zhang, Xinfeng ;

Wang, Shanshe ;

Liu, Jiaying ;

Pu, Shiliang ;

Ma, Siwei .

IEEE TRANSACTIONS ON IMAGE PROCESSING, 2019, 28 (07) :3343-3356

[25] IFQA: Interpretable Face Quality Assessment [J].

Jo, Byungho ;

Cho, Donghyeon ;

Park, In Kyu ;

Hong, Sungeun .

2023 IEEE/CVF WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV), 2023, :3433-3442

[26]

Keimel C, 2012, INT WORK QUAL MULTIM, P97, DOI 10.1109/QoMEX.2012.6263865

[27]

Kim HI, 2015, IEEE IMAGE PROC, P4027, DOI 10.1109/ICIP.2015.7351562

[28] Dynamic Receptive Field Generation for Full-Reference Image Quality Assessment [J].

Kim, Woojae ;

Nguyen, Anh-Duc ;

Lee, Sanghoon ;

Bovik, Alan Conrad .

IEEE TRANSACTIONS ON IMAGE PROCESSING, 2020, 29 :4219-4231

[29] Deep Video Quality Assessor: From Spatio-Temporal Visual Sensitivity to a Convolutional Neural Aggregation Network [J].

Kim, Woojae ;

Kim, Jongyoo ;

Ahn, Sewoong ;

Kim, Jinwoo ;

Lee, Sanghoon .

COMPUTER VISION - ECCV 2018, PT I, 2018, 11205 :224-241

[30] A Super-Resolution Flexible Video Coding Solution for Improving Live Streaming Quality [J].

Li, Qing ;

Chen, Ying ;

Zhang, Aoyang ;

Jiang, Yong ;

Zou, Longhao ;

Xu, Zhimin ;

Muntean, Gabriel-Miro .

IEEE TRANSACTIONS ON MULTIMEDIA, 2023, 25 :6341-6355

← 1 2 3 4 5 6 7 8 9 10 →