Self-Supervised Adversarial Video Summarizer With Context Latent Sequence Learning

被引:4
作者
Xu, Yifei [1 ]
Li, Xiangshun [1 ]
Pan, Litong [1 ]
Sang, Weiguang [1 ]
Wei, Pingping [2 ]
Zhu, Li [1 ]
机构
[1] Xi An Jiao Tong Univ, Sch Software, Xian 710054, Shaanxi, Peoples R China
[2] Xi An Jiao Tong Univ, State Key Lab Mfg Syst Engn, Xian 710054, Shaanxi, Peoples R China
基金
中国博士后科学基金; 中国国家自然科学基金;
关键词
Self-supervised GAN; video summarization; context latent sequence learning;
D O I
10.1109/TCSVT.2023.3240464
中图分类号
TM [电工技术]; TN [电子技术、通信技术];
学科分类号
0808 ; 0809 ;
摘要
Video summarization attempts to create concise and complete synopsis of a video through identifying the most informative and explanatory parts while removing redundant video frames, which facilitates retrieving, managing and browsing video efficiently. Most existing video summarization approaches heavily rely on enormous high-quality human-annotated labels or fail to produce semantically meaningful video summaries with the guidance of prior information. Without any supervised labels, we propose Self-supervised Adversarial Video Summarizer (2SAVS) that exploits context latent sequence learning to generate satisfying video summary. To implement it, our model elaborates a novel pretext task of identifying latent sequences and normal frames by training self-supervised generative adversarial network (GAN) with several well-designed losses. As the core components of 2SAVS, Clip Consistency Representation (CCR) and Hybrid Feature Refinement (HFR) are developed to ensure semantic consistency and continuity of clips. Furthermore, a novel separation loss is designed to explicitly enlarge the distance between prediction frame scores to effectively enhance the model's discriminative ability. Differently, latent sequences, additional finetune operations and generators are not required when inferring video summary. Experiments on two challenging and diverse datasets demonstrate that our approach outperforms other state-of-the-art unsupervised and weakly-supervised methods, and even produces comparable results with several excellent supervised methods.
引用
收藏
页码:4122 / 4136
页数:15
相关论文
共 63 条
  • [1] Apostolidis E., 2019, Proceedings of the 1st International Workshop on AI for Smart TV Content Production, Access and Delivery, P17
  • [2] Combining Global and Local Attention with Positional Encoding for Video Summarization
    Apostolidis, Evlampios
    Balaouras, Georgios
    Mezaris, Vasileios
    Patras, Ioannis
    [J]. 23RD IEEE INTERNATIONAL SYMPOSIUM ON MULTIMEDIA (ISM 2021), 2021, : 226 - 234
  • [3] AC-SUM-GAN: Connecting Actor-Critic and Generative Adversarial Networks for Unsupervised Video Summarization
    Apostolidis, Evlampios
    Adamantidou, Eleni
    Metsai, Alexandros, I
    Mezaris, Vasileios
    Patras, Ioannis
    [J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2021, 31 (08) : 3278 - 3292
  • [4] Performance over Random: A Robust Evaluation Protocol for Video Summarization Methods
    Apostolidis, Evlampios
    Adamantidou, Eleni
    Metsai, Alexandros, I
    Mezaris, Vasileios
    Patras, Ioannis
    [J]. MM '20: PROCEEDINGS OF THE 28TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, 2020, : 1056 - 1064
  • [5] Unsupervised Video Summarization via Attention-Driven Adversarial Learning
    Apostolidis, Evlampios
    Adamantidou, Eleni
    Metsai, Alexandros, I
    Mezaris, Vasileios
    Patras, Ioannis
    [J]. MULTIMEDIA MODELING (MMM 2020), PT I, 2020, 11961 : 492 - 504
  • [6] Arjovsky M, 2017, PR MACH LEARN RES, V70
  • [7] Weakly-Supervised Video Summarization Using Variational Encoder-Decoder and Web Prior
    Cai, Sijia
    Zuo, Wangmeng
    Davis, Larry S.
    Zhang, Lei
    [J]. COMPUTER VISION - ECCV 2018, PT XIV, 2018, 11218 : 193 - 210
  • [8] Chu WS, 2015, PROC CVPR IEEE, P3584, DOI 10.1109/CVPR.2015.7298981
  • [9] Dissimilarity-Based Sparse Subset Selection
    Elhamifar, Ehsan
    Sapiro, Guillermo
    Sastry, S. Shankar
    [J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2016, 38 (11) : 2182 - 2197
  • [10] Extractive Video Summarizer with Memory Augmented Neural Networks
    Feng, Litong
    Li, Ziyin
    Kuang, Zhanghui
    Zhang, Wayne
    [J]. PROCEEDINGS OF THE 2018 ACM MULTIMEDIA CONFERENCE (MM'18), 2018, : 976 - 983