Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency

被引：12

作者：

Wu, Haiping ^{[1
]}

Wang, Xiaolong ^{[2
]}

机构：

[1] McGill Univ, Mila, Montreal, PQ, Canada

[2] Univ Calif San Diego, La Jolla, CA USA

来源：

2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021) | 2021年

关键词：

D O I：

10.1109/ICCV48922.2021.00999

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Recent works have advanced the performance of self-supervised representation learning by a large margin. The core among these methods is intra-image invariance learning. Two different transformations of one image instance are considered as a positive sample pair, where various tasks are designed to learn invariant representations by comparing the pair. Analogically, for video data, representations of frames from the same video are trained to be closer than frames from other videos, i.e. intra-video invariance. However, cross-video relation has barely been explored for visual representation learning. Unlike intra-video invariance, ground-truth labels of cross-video relation is usually unavailable without human labors. In this paper, we propose a novel contrastive learning method which explores the cross-video relation by using cycle-consistency for general image representation learning. This allows to collect positive sample pairs across different video instances, which we hypothesize will lead to higher-level semantics. We validate our method by transferring our image representation to multiple downstream tasks including visual object tracking, image classification, and action recognition. We show significant improvement over state-of-the-art contrastive learning methods. Project page is available at https://happywu.github.io/cycle_contrast_video

引用

页码：10129 / 10139

页数：11

共 78 条

[21] Hadsell R., 2006, IEEE C COMP VIS PATT, P1735, DOI DOI 10.1109/CVPR.2006.100
[22] Video Representation Learning by Dense Predictive Coding
Han, Tengda
Xie, Weidi
Zisserman, Andrew
[J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS (ICCVW), 2019, : 1483 - 1492
[23] Han Tengda, ARXIV PREPRINT ARXIV
[24] He K, 2020, P IEEE CVF C COMP VI, P9729
[25] Deep Residual Learning for Image Recognition
He, Kaiming
Zhang, Xiangyu
Ren, Shaoqing
Sun, Jian
[J]. 2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, : 770 - 778
[26] Henaff O. J., 2019, Data-efficient image recognition with contrastive predictive coding
[27] Hjelm R.D., 2018, ARXIV PREPRINT ARXIV
[28] GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild
Huang, Lianghua
Zhao, Xin
Huang, Kaiqi
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2021, 43 (05) : 1562 - 1577
[29] Consistent Shape Maps via Semidefinite Programming
Huang, Qi-Xing
Guibas, Leonidas
[J]. COMPUTER GRAPHICS FORUM, 2013, 32 (05) : 177 - 186
[30] Jabri A., 2020, Adv. Neural Inf. Process. Syst, V33, P19545

← 1 2 3 4 5 6 7 8 →