Play and Rewind: Optimizing Binary Representations of Videos by Self-Supervised Temporal Hashing

被引：63

作者：

Zhang, Hanwang ^{[1
]}

Wang, Meng ^{[2
]}

Hong, Richang ^{[2
]}

Chua, Tat-Seng ^{[1
]}

机构：

[1] Natl Univ Singapore, Singapore, Singapore

[2] Hefei Univ Technol, Hefei, Peoples R China

来源：

MM'16: PROCEEDINGS OF THE 2016 ACM MULTIMEDIA CONFERENCE | 2016年

关键词：

Temporal Hashing; Binary LSTM; Sequence Learning; Video Retrieval;

D O I：

10.1145/2964284.2964308

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

We focus on hashing videos into short binary codes for efficient Content-based Video Retrieval (CBVR), which is a fundamental technique that supports access to the evergrowing abundance of videos on the Web. Existing video hash functions are built on three isolated stages: frame pooling, relaxed learning, and binarization, which have not adequately explored the temporal order of video frames in a joint binary optimization model, resulting in severe information loss. In this paper, we propose a novel unsupervised video hashing framework called Self-Supervised Temporal Hashing (SSTH) that is able to capture the temporal nature of videos in an end-to-end learning-to-hash fashion. Specifically, the hash function of SSTH is an encoder RNN equipped with the proposed Binary LSTM (BLSTM) that generates binary codes for videos. The hash function is learned in a self-supervised fashion, where a decoder RNN is proposed to reconstruct the original video frames in both forward and reverse orders. For binary code optimization, we develop a backpropagation rule that tackles the non-differentiability of BLSTM. This rule allows efficient deep network training without suffering from the binarization loss. Through extensive CBVR experiments on two real-world consumer video datasets of Youtube and Flickr, we show that SSTH consistently outperforms state-of-theart video hashing methods, e.g., in terms of mAP@20, SSTH using only 128 bits can still outperform others using 256 bits by at least 9% to 15% on both datasets.

引用

页码：781 / 790

页数：10

共 43 条

[1]

[Anonymous], TPAMI

[2]

[Anonymous], 2016, BinaryNet: Training deep neural networks with weights and activa

[3]

[Anonymous], SIGIR

[4]

[Anonymous], TPAMI

[5]

[Anonymous], 2005, NEURAL NETWORKS

[6]

[Anonymous], 2014, NIPS

[7]

[Anonymous], 2015, ARXIV PREPRINT ARXIV

[8]

[Anonymous], 2013, CoRR abs/1308.3432

[9]

[Anonymous], 2015, CVPR

[10]

[Anonymous], 2010, CVPR

← 1 2 3 4 5 →