Correlated Features Synthesis and Alignment for Zero-shot Cross-modal Retrieval

被引：17

作者：

Xu, Xing ^{[1
,2
]}

Lin, Kaiyi ^{[1
,2
,3
]}

Lu, Huimin ^{[4
]}

Gao, Lianli ^{[1
,2
]}

Shen, Heng Tao ^{[1
,2
]}

机构：

[1] Univ Elect Sci & Technol China, Ctr Future Multimedia, Chengdu, Peoples R China

[2] Univ Elect Sci & Technol China, Sch Comp Sci & Engn, Chengdu, Peoples R China

[3] Peking Univ, Sch Software & Microelect, Beijing, Peoples R China

[4] Kyushu Inst Technol, Dept Mech & Control Engn, Kitakyushu, Fukuoka, Japan

来源：

PROCEEDINGS OF THE 43RD INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL (SIGIR '20) | 2020年

基金：

中国国家自然科学基金;

关键词：

Cross-modal Retrieval; Zero-shot Learning; Feature Synthesis;

D O I：

10.1145/3397271.3401149

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

The goal of cross-modal retrieval is to search for semantically similar instances in one modality by using a query from another modality. Existing approaches mainly consider the standard scenario that requires the source set for training and the target set for testing share the same scope of classes. However, they may not generalize well on zero-shot cross-modal retrieval (ZS-CMR) task, where the target set contains unseen classes that are disjoint with the seen classes in the source set. This task is more challenging due to 1) the absence of the unseen classes during training, 2) inconsistent semantics across seen and unseen classes, and 3) the heterogeneous multimodal distributions between the source and target set. To address these issues, we propose a novel Correlated Feature Synthesis and Alignment (CFSA) approach to integrate multimodal feature synthesis, common space learning and knowledge transfer for ZSCMR. Our CFSA first utilizes class-level word embeddings to guide two coupled Wassertein generative adversarial networks (WGANs) to synthesize sufficient multimodal features with semantic correlation for stable training. Then the synthetic and true multimodal features are jointly mapped to a common semantic space via an effective distribution alignment scheme, where the cross-modal correlations of different semantic features are captured and the knowledge can be transferred to the unseen classes under the cycle-consistency constraint. Experiments on four benchmark datasets for image-text retrieval and two large-scale datasets for image-sketch retrieval show the remarkable improvements achieved by our CFAS method comparing with a bundle of state-of-the-art approaches.

引用

页码：1419 / 1428

页数：10

共 50 条

[1] Generalized Zero-Shot Cross-Modal Retrieval
Dutta, Titir
Biswas, Soma
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2019, 28 (12) : 5953 - 5962
[2] Zero-shot Cross-modal Retrieval by Assembling AutoEncoder and Generative Adversarial Network
Xu, Xing
Tian, Jialin
Lin, Kaiyi
Lu, Huimin
Shao, Jie
Shen, Heng Tao
ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2021, 17 (01)
[3] Multimodal Disentanglement Variational AutoEncoders for Zero-Shot Cross-Modal Retrieval
Tian, Jialin
Wang, Kai
Xu, Xing
Cao, Zuo
Shen, Fumin
Shen, Heng Tao
PROCEEDINGS OF THE 45TH INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL (SIGIR '22), 2022, : 960 - 969
[4] CROSS-MODAL ALIGNMENT OF LOCAL AND GLOBAL FEATURES FOR ZERO-SHOT CHINESE CHARACTER RECOGNITION
Cai, Hongyi
Zhu, Anna
2024 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2024, : 2041 - 2047
[5] Cross-modal Zero-shot Hashing
Liu, Xuanwu
Li, Zhao
Wang, Jun
Yu, Guoxian
Domeniconi, Carlotta
Zhang, Xiangliang
2019 19TH IEEE INTERNATIONAL CONFERENCE ON DATA MINING (ICDM 2019), 2019, : 449 - 458
[6] CHOP: An orthogonal hashing method for zero-shot cross-modal retrieval
Yuan, Xu
Wang, Guangze
Chen, Zhikui
Zhong, Fangming
PATTERN RECOGNITION LETTERS, 2021, 145 : 247 - 253
[7] Discrete asymmetric zero-shot hashing with application to cross-modal retrieval
Shu, Zhenqiu
Yong, Kailing
Yu, Jun
Gao, Shengxiang
Mao, Cunli
Yu, Zhengtao
NEUROCOMPUTING, 2022, 511 : 366 - 379
[8] INTER-MODALITY FUSION BASED ATTENTION FOR ZERO-SHOT CROSS-MODAL RETRIEVAL
Chakraborty, Bela
Wang, Peng
Wang, Lei
2021 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2021, : 2648 - 2652
[9] Semantic-Adversarial Graph Convolutional Network for Zero-Shot Cross-Modal Retrieval
Li, Chuang
Fei, Lunke
Kang, Peipei
Liang, Jiahao
Fang, Xiaozhao
Teng, Shaohua
PRICAI 2022: TRENDS IN ARTIFICIAL INTELLIGENCE, PT II, 2022, 13630 : 459 - 472
[10] Zero-shot discrete hashing with adaptive class correlation for cross-modal retrieval
Yong, Kailing
Shu, Zhenqiu
Yu, Jun
Yu, Zhengtao
KNOWLEDGE-BASED SYSTEMS, 2024, 295

← 1 2 3 4 5 →