Distilling Privileged Knowledge for Anomalous Event Detection From Weakly Labeled Videos

被引：6

作者：

Liu, Tianshan ^{[1
]}

Lam, Kin-Man ^{[1
,2
]}

Kong, Jun ^{[3
]}

机构：

[1] Hong Kong Polytech Univ, Dept Elect & Informat Engn, Hong Kong, Peoples R China

[2] Ctr Adv Reliabil & Safety, Hong Kong, Peoples R China

[3] Jiangnan Univ, Key Lab Adv Proc Control Light Ind, Minist Educ, Wuxi 214122, Peoples R China

来源：

IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS | 2024年 / 35卷 / 09期

关键词：

Videos; Task analysis; Feature extraction; Training; Knowledge engineering; Anomaly detection; Annotations; Privileged knowledge distillation (KD); teacher-student model; video anomaly detection (VAD); weakly supervised learning;

D O I：

10.1109/TNNLS.2023.3263966

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Weakly supervised video anomaly detection (WS-VAD) aims to identify the snippets involving anomalous events in long untrimmed videos, with solely video-level binary labels. A typical paradigm among the existing WS-VAD methods is to employ multiple modalities as inputs, e.g., RGB, optical flow, and audio, as they can provide sufficient discriminative clues that are robust to the diverse, complicated real-world scenes. However, such a pipeline has high reliance on the availability of multiple modalities and is computationally expensive and storage demanding in processing long sequences, which limits its use in some applications. To address this dilemma, we propose a privileged knowledge distillation (KD) framework dedicated to the WS-VAD task, which can maintain the benefits of exploiting additional modalities, while avoiding the need for using multimodal data in the inference phase. We argue that the performance of the privileged KD framework mainly depends on two factors: 1) the effectiveness of the multimodal teacher network and 2) the completeness of the useful information transfer. To obtain a reliable teacher network, we propose a cross-modal interactive learning strategy and an anomaly normal discrimination loss, which target learning task-specific cross-modal features and encourage the separability of anomalous and normal representations, respectively. Furthermore, we design both representation-and logits-level distillation loss functions, which force the unimodal student network to distill abundant privileged knowledge from the well-trained multimodal teacher network, in a snippet-to-video fashion. Extensive experimental results on three public benchmarks demonstrate that the proposed privileged KD framework can train a lightweight yet effective detector, for localizing anomaly events under the supervision of video-level annotations.

引用

页码：12627 / 12641

页数：15

共 65 条

[31] Lopez-Paz D., 2016, International Conference on Learning Representations (ICLR)
[32] Future Frame Prediction Using Convolutional VRNN for Anomaly Detection
Lu, Yiwei
Kumar K, Mahesh
Nabavi, Seyed shahabeddin
Wang, Yang
[J]. 2019 16TH IEEE INTERNATIONAL CONFERENCE ON ADVANCED VIDEO AND SIGNAL BASED SURVEILLANCE (AVSS), 2019,
[33] Future Frame Prediction Network for Video Anomaly Detection
Luo, Weixin
Liu, Wen
Lian, Dongze
Gao, Shenghua
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (11) : 7505 - 7520
[34] Luo WX, 2017, IEEE INT CON MULTI, P439, DOI 10.1109/ICME.2017.8019325
[35] Graph Distillation for Action Detection with Privileged Modalities
Luo, Zelun
Hsieh, Jun-Ting
Jiang, Lu
Niebles, Juan Carlos
Fei-Fei, Li
[J]. COMPUTER VISION - ECCV 2018, PT XIV, 2018, 11218 : 174 - 192
[36] Madan N., 2022, ARXIV
[37] W-TALC: Weakly-Supervised Temporal Activity Localization and Classification
Paul, Sujoy
Roy, Sourya
Roy-Chowdhury, Amit K.
[J]. COMPUTER VISION - ECCV 2018, PT IV, 2018, 11208 : 588 - 607
[38] Peng Wu, 2020, Computer Vision - ECCV 2020 16th European Conference. Proceedings. Lecture Notes in Computer Science (LNCS 12375), P322, DOI 10.1007/978-3-030-58577-8_20
[39] A Survey of Single-Scene Video Anomaly Detection
Ramachandra, Bharathkumar
Jones, Michael J.
Vatsavai, Ranga Raju
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (05) : 2293 - 2312
[40] Self-Supervised Predictive Convolutional Attentive Block for Anomaly Detection
Ristea, Nicolae-Catalin
Madan, Neelu
Ionescu, Radu Tudor
Nasrollahi, Kamal
Khan, Fahad Shahbaz
Moeslund, Thomas B.
Shah, Mubarak
[J]. 2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2022, : 13566 - 13576

← 1 2 3 4 5 6 7 →