Consistent Prompt Tuning for Generalized Category Discovery

被引：0

作者：

Yang, Muli ^{[1
]}

Yin, Jie ^{[2
]}

Gu, Yanan ^{[3
]}

Deng, Cheng ^{[2
]}

Zhang, Hanwang ^{[4
]}

Zhu, Hongyuan ^{[1
]}

机构：

[1] ASTAR, Inst Infocomm Res I2R, Singapore, Singapore

[2] Xidian Univ, Sch Elect Engn, Xian, Peoples R China

[3] Norinco Grp Testing & Res Inst, Xian, Peoples R China

[4] Nanyang Technol Univ, Coll Comp & Data Sci, Singapore, Singapore

来源：

INTERNATIONAL JOURNAL OF COMPUTER VISION | 2025年

基金：

中国国家自然科学基金; 国家重点研发计划;

关键词：

Category discovery; Prompt learning; Multimodal learning; Transfer learning;

D O I：

10.1007/s11263-024-02343-w

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Generalized Category Discovery (GCD) aims at discovering both known and unknown classes in unlabeled data, using the knowledge learned from a limited set of labeled data. Despite today's foundation models being trained with Internet-scale multi-modal corpus, we find that they still struggle in GCD due to the ambiguity in class definitions. In this paper, we present Consistent Prompt Tuning (CPT) to disambiguate the classes for large vision-language models (e.g., CLIP). To this end, CPT learns a set of "task + class" prompts for labeled and unlabeled data of both known and unknown classes, with the "task" tokens globally shared across classes, which contain a unified class definition pattern, e.g., "the foreground is an animal named" or "the background scene is". These prompts are optimized with two efficient regularization techniques that encourage consistent global and local relationships between any two matched inputs. CPT is evaluated on various existing GCD benchmarks, as well as in new practical scenarios with fewer annotations and customized class definitions, demonstrating clear superiority and broad versatility over existing state-of-the-art methods.

引用

页码：4014 / 4041

页数：28

共 166 条

[1]

Alayrac JB, 2022, ADV NEUR IN

[2]

An W., 2023, arXiv, DOI arXiv:2312.10897

[3]

An W., 2022, arXiv

[4]

An WB, 2023, Arxiv, DOI arXiv:2312.16467

[5]

An WB, 2023, AAAI CONF ARTIF INTE, P12527

[6] Masked Siamese Networks for Label-Efficient Learning [J].

Assran, Mahmoud ;

Caron, Mathilde ;

Misra, Ishan ;

Bojanowski, Piotr ;

Bordes, Florian ;

Vincent, Pascal ;

Joulin, Armand ;

Rabbat, Mike ;

Ballas, Nicolas .

COMPUTER VISION, ECCV 2022, PT XXXI, 2022, 13691 :456-473

[7] Semi-Supervised Learning of Visual Features by Non-Parametrically Predicting View Assignments with Support Samples [J].

Assran, Mahmoud ;

Caron, Mathilde ;

Misra, Ishan ;

Bojanowski, Piotr ;

Joulin, Armand ;

Ballas, Nicolas ;

Rabbat, Michael .

2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :8423-8432

[8]

Bai JH, 2024, Arxiv, DOI arXiv:2310.01376

[9]

Banerjee A, 2024, IEEE WINT CONF APPL, P2090, DOI 10.1109/WACV57701.2024.00210

[10]

Brown TB, 2020, ADV NEUR IN, V33

← 1 2 3 4 5 6 7 8 9 10 →