Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models

被引：0

作者：

Meng, Zaiqiao ^{[1
,2
]}

Liu, Fangyu ^{[1
]}

Shareghi, Ehsan ^{[1
,3
]}

Su, Yixuan ^{[1
]}

Collins, Charlotte ^{[1
]}

Collier, Nigel ^{[1
]}

机构：

[1] Univ Cambridge, Language Technol Lab, Cambridge, England

[2] Univ Glasgow, Dept Comp Sci, Glasgow, Lanark, Scotland

[3] Monash Univ, Dept Data Sci & AI, Clayton, Vic, Australia

来源：

PROCEEDINGS OF THE 60TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), VOL 1: (LONG PAPERS) | 2022年

关键词：

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Knowledge probing is crucial for understanding the knowledge transfer mechanism behind the pre-trained language models (PLMs). Despite the growing progress of probing knowledge for PLMs in the general domain, specialised areas such as biomedical domain are vastly under-explored. To facilitate this, we release a well-curated biomedical knowledge probing benchmark, MedLAMA, constructed based on the Unified Medical Language System (UMLS) Metathesaurus. We test a wide spectrum of state-of-the-art PLMs and probing approaches on our benchmark, reaching at most 3% of acc@10. While highlighting various sources of domain-specific challenges that amount to this underwhelming performance, we illustrate that the underlying PLMs have a higher potential for probing tasks. To achieve this, we propose CONTRASTIVE-P robe, a novel self-supervised contrastive probing approach, that adjusts the underlying PLMs without using any probing data. While CONTRASTIVE-P robe pushes the acc@10 to 24%, the performance gap remains notable. Our human expert evaluation suggests that the probing performance of our CONTRASTIVE-P robe is underestimated as UMLS does not comprehensively cover all existing factual knowledge. We hope MedLAMA and CONTRASTIVE-P robe facilitate further developments of more suited probing techniques for this domain.(1)

引用

页码：4798 / 4810

页数：13

共 50 条

[1] Probing Pre-Trained Language Models for Disease Knowledge
Alghanmi, Israa
Espinosa-Anke, Luis
Schockaert, Steven
FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, ACL-IJCNLP 2021, 2021, : 3023 - 3033
[2] Probing Simile Knowledge from Pre-trained Language Models
Chen, Weijie
Chang, Yongzhu
Zhang, Rongsheng
Pu, Jiashu
Chen, Guandan
Zhang, Le
Xi, Yadong
Chen, Yijiang
Su, Chang
PROCEEDINGS OF THE 60TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), VOL 1: (LONG PAPERS), 2022, : 5875 - 5887
[3] Continual knowledge infusion into pre-trained biomedical language models
Jha, Kishlay
Zhang, Aidong
BIOINFORMATICS, 2022, 38 (02) : 494 - 502
[4] Pre-trained language models with domain knowledge for biomedical extractive summarization
Xie Q.
Bishop J.A.
Tiwari P.
Ananiadou S.
Knowledge-Based Systems, 2022, 252
[5] Probing for Hyperbole in Pre-Trained Language Models
Schneidermann, Nina Skovgaard
Hershcovich, Daniel
Pedersen, Bolette Sandford
PROCEEDINGS OF THE 61ST ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, ACL-SRW 2023, VOL 4, 2023, : 200 - 211
[6] SimKGC: Simple Contrastive Knowledge Graph Completion with Pre-trained Language Models
Wang, Liang
Zhao, Wei
Wei, Zhuoyu
Liu, Jingming
PROCEEDINGS OF THE 60TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), VOL 1: (LONG PAPERS), 2022, : 4281 - 4294
[7] Knowledge Rumination for Pre-trained Language Models
Yao, Yunzhi
Wang, Peng
Mao, Shengyu
Tan, Chuanqi
Huang, Fei
Chen, Huajun
Zhang, Ningyu
2023 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING, EMNLP 2023, 2023, : 3387 - 3404
[8] Knowledge Inheritance for Pre-trained Language Models
Qin, Yujia
Lin, Yankai
Yi, Jing
Zhang, Jiajie
Han, Xu
Zhang, Zhengyan
Su, Yusheng
Liu, Zhiyuan
Li, Peng
Sun, Maosong
Zhou, Jie
NAACL 2022: THE 2022 CONFERENCE OF THE NORTH AMERICAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: HUMAN LANGUAGE TECHNOLOGIES, 2022, : 3921 - 3937
[9] Focused Contrastive Loss for Classification With Pre-Trained Language Models
He, Jiayuan
Li, Yuan
Zhai, Zenan
Fang, Biaoyan
Thorne, Camilo
Druckenbrodt, Christian
Akhondi, Saber
Verspoor, Karin
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2024, 36 (07) : 3047 - 3061
[10] Give Me the Facts! A Survey on Factual Knowledge Probing in Pre-trained Language Models
Youssef, Paul
Koras, Osman Alperen
Li, Meijie
Schlotterer, Jorg
Seifert, Christin
FINDINGS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (EMNLP 2023), 2023, : 15588 - 15605

← 1 2 3 4 5 →