Self-Supervised Adversarial Video Summarizer With Context Latent Sequence Learning

被引：4

作者：

Xu, Yifei ^{[1
]}

Li, Xiangshun ^{[1
]}

Pan, Litong ^{[1
]}

Sang, Weiguang ^{[1
]}

Wei, Pingping ^{[2
]}

Zhu, Li ^{[1
]}

机构：

[1] Xi An Jiao Tong Univ, Sch Software, Xian 710054, Shaanxi, Peoples R China

[2] Xi An Jiao Tong Univ, State Key Lab Mfg Syst Engn, Xian 710054, Shaanxi, Peoples R China

来源：

IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY | 2023年 / 33卷 / 08期

基金：

中国博士后科学基金; 中国国家自然科学基金;

关键词：

Self-supervised GAN; video summarization; context latent sequence learning;

D O I：

10.1109/TCSVT.2023.3240464

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

Video summarization attempts to create concise and complete synopsis of a video through identifying the most informative and explanatory parts while removing redundant video frames, which facilitates retrieving, managing and browsing video efficiently. Most existing video summarization approaches heavily rely on enormous high-quality human-annotated labels or fail to produce semantically meaningful video summaries with the guidance of prior information. Without any supervised labels, we propose Self-supervised Adversarial Video Summarizer (2SAVS) that exploits context latent sequence learning to generate satisfying video summary. To implement it, our model elaborates a novel pretext task of identifying latent sequences and normal frames by training self-supervised generative adversarial network (GAN) with several well-designed losses. As the core components of 2SAVS, Clip Consistency Representation (CCR) and Hybrid Feature Refinement (HFR) are developed to ensure semantic consistency and continuity of clips. Furthermore, a novel separation loss is designed to explicitly enlarge the distance between prediction frame scores to effectively enhance the model's discriminative ability. Differently, latent sequences, additional finetune operations and generators are not required when inferring video summary. Experiments on two challenging and diverse datasets demonstrate that our approach outperforms other state-of-the-art unsupervised and weakly-supervised methods, and even produces comparable results with several excellent supervised methods.

引用

页码：4122 / 4136

页数：15

共 63 条

[1] Apostolidis E., 2019, Proceedings of the 1st International Workshop on AI for Smart TV Content Production, Access and Delivery, P17
[2] Combining Global and Local Attention with Positional Encoding for Video Summarization
Apostolidis, Evlampios
Balaouras, Georgios
Mezaris, Vasileios
Patras, Ioannis
[J]. 23RD IEEE INTERNATIONAL SYMPOSIUM ON MULTIMEDIA (ISM 2021), 2021, : 226 - 234
[3] AC-SUM-GAN: Connecting Actor-Critic and Generative Adversarial Networks for Unsupervised Video Summarization
Apostolidis, Evlampios
Adamantidou, Eleni
Metsai, Alexandros, I
Mezaris, Vasileios
Patras, Ioannis
[J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2021, 31 (08) : 3278 - 3292
[4] Performance over Random: A Robust Evaluation Protocol for Video Summarization Methods
Apostolidis, Evlampios
Adamantidou, Eleni
Metsai, Alexandros, I
Mezaris, Vasileios
Patras, Ioannis
[J]. MM '20: PROCEEDINGS OF THE 28TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, 2020, : 1056 - 1064
[5] Unsupervised Video Summarization via Attention-Driven Adversarial Learning
Apostolidis, Evlampios
Adamantidou, Eleni
Metsai, Alexandros, I
Mezaris, Vasileios
Patras, Ioannis
[J]. MULTIMEDIA MODELING (MMM 2020), PT I, 2020, 11961 : 492 - 504
[6] Arjovsky M, 2017, PR MACH LEARN RES, V70
[7] Weakly-Supervised Video Summarization Using Variational Encoder-Decoder and Web Prior
Cai, Sijia
Zuo, Wangmeng
Davis, Larry S.
Zhang, Lei
[J]. COMPUTER VISION - ECCV 2018, PT XIV, 2018, 11218 : 193 - 210
[8] Chu WS, 2015, PROC CVPR IEEE, P3584, DOI 10.1109/CVPR.2015.7298981
[9] Dissimilarity-Based Sparse Subset Selection
Elhamifar, Ehsan
Sapiro, Guillermo
Sastry, S. Shankar
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2016, 38 (11) : 2182 - 2197
[10] Extractive Video Summarizer with Memory Augmented Neural Networks
Feng, Litong
Li, Ziyin
Kuang, Zhanghui
Zhang, Wayne
[J]. PROCEEDINGS OF THE 2018 ACM MULTIMEDIA CONFERENCE (MM'18), 2018, : 976 - 983

← 1 2 3 4 5 6 7 →