Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency

被引:12
作者
Wu, Haiping [1 ]
Wang, Xiaolong [2 ]
机构
[1] McGill Univ, Mila, Montreal, PQ, Canada
[2] Univ Calif San Diego, La Jolla, CA USA
来源
2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021) | 2021年
关键词
D O I
10.1109/ICCV48922.2021.00999
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Recent works have advanced the performance of self-supervised representation learning by a large margin. The core among these methods is intra-image invariance learning. Two different transformations of one image instance are considered as a positive sample pair, where various tasks are designed to learn invariant representations by comparing the pair. Analogically, for video data, representations of frames from the same video are trained to be closer than frames from other videos, i.e. intra-video invariance. However, cross-video relation has barely been explored for visual representation learning. Unlike intra-video invariance, ground-truth labels of cross-video relation is usually unavailable without human labors. In this paper, we propose a novel contrastive learning method which explores the cross-video relation by using cycle-consistency for general image representation learning. This allows to collect positive sample pairs across different video instances, which we hypothesize will lead to higher-level semantics. We validate our method by transferring our image representation to multiple downstream tasks including visual object tracking, image classification, and action recognition. We show significant improvement over state-of-the-art contrastive learning methods. Project page is available at https://happywu.github.io/cycle_contrast_video
引用
收藏
页码:10129 / 10139
页数:11
相关论文
共 78 条
  • [51] Sayed Nawid, 2019, Pattern Recognition. 40th German Conference, GCPR 2018. Proceedings: Lecture Notes in Computer Science (LNCS 11269), P228, DOI 10.1007/978-3-030-12939-2_17
  • [52] Sermanet P, 2018, IEEE INT CONF ROBOT, P1134
  • [53] Soomro Khurram, 2012, ARXIV12120402
  • [54] Sun C, 2019, 2019 CONFERENCE OF THE NORTH AMERICAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: HUMAN LANGUAGE TECHNOLOGIES (NAACL HLT 2019), VOL. 1, P380
  • [55] Unsupervised Learning of Landmarks by Descriptor Vector Exchange
    Thewlis, James
    Albanie, Samuel
    Bilen, Hakan
    Vedaldi, Andrea
    [J]. 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, : 6370 - 6380
  • [56] Tian Yonglong, 2019, ARXIV190605849
  • [57] Wang Jianyi, 2020, ECCV
  • [58] Unsupervised Deep Tracking
    Wang, Ning
    Song, Yibing
    Ma, Chao
    Zhou, Wengang
    Liu, Wei
    Li, Houqiang
    [J]. 2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 1308 - 1317
  • [59] Learning Correspondence from the Cycle-consistency of Time
    Wang, Xiaolong
    Jabri, Allan
    Efros, Alexei A.
    [J]. 2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 2561 - 2571
  • [60] Transitive Invariance for Self-supervised Visual Representation Learning
    Wang, Xiaolong
    He, Kaiming
    Gupta, Abhinav
    [J]. 2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, : 1338 - 1347