Transfer Joint Embedding for Cross-Domain Named Entity Recognition

被引:26
作者
Pan, Sinno Jialin [1 ]
Toh, Zhiqiang [1 ]
Su, Jian [1 ]
机构
[1] Inst Infocomm Res, Data Analyt Dept, Singapore 138632, Singapore
关键词
Algorithms; Experimentation; Named entity recognition; transfer learning; multiclass classification;
D O I
10.1145/2457465.2457467
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Named Entity Recognition (NER) is a fundamental task in information extraction from unstructured text. Most previous machine-learning-based NER systems are domain-specific, which implies that they may only perform well on some specific domains (e.g., Newswire) but tend to adapt poorly to other related but different domains (e.g., Weblog). Recently, transfer learning techniques have been proposed to NER. However, most transfer learning approaches to NER are developed for binary classification, while NER is a multiclass classification problem in nature. Therefore, one has to first reduce the NER task to multiple binary classification tasks and solve them independently. In this article, we propose a new transfer learning method, named Transfer Joint Embedding (TJE), for cross-domain multiclass classification, which can fully exploit the relationships between classes (labels), and reduce domain difference in data distributions for transfer learning. More specifically, we aim to embed both labels (outputs) and high-dimensional features (inputs) from different domains (e.g., a source domain and a target domain) into a unified low-dimensional latent space, where 1) each label is represented by a prototype and the intrinsic relationships between labels can be measured by Euclidean distance; 2) the distance in data distributions between the source and target domains can be reduced; 3) the source domain labeled data are closer to their corresponding label-prototypes than others. After the latent space is learned, classification on the target domain data can be done with the simple nearest neighbor rule in the latent space. Furthermore, in order to scale up TJE, we propose an efficient algorithm based on stochastic gradient descent (SGD). Finally, we apply the proposed TJE method for NER across different domains on the ACE 2005 dataset, which is a benchmark in Natural Language Processing (NLP). Experimental results demonstrate the effectiveness of TJE and show that TJE can outperform state-of-the-art transfer learning approaches to NER.
引用
收藏
页数:27
相关论文
共 50 条
[31]   Towards Chinese clinical named entity recognition by dynamic embedding using domain-specific knowledge [J].
Li, Yuan ;
Du, Guodong ;
Xiang, Yan ;
Li, Shaozi ;
Ma, Lei ;
Shao, Dangguo ;
Wang, Xiongbin ;
Chen, Haoyu .
JOURNAL OF BIOMEDICAL INFORMATICS, 2020, 106
[32]   Joint Learning of Named Entity Recognition and Relation Extraction [J].
Xu, Qiuyan ;
Li, Fang .
2011 INTERNATIONAL CONFERENCE ON COMPUTER SCIENCE AND NETWORK TECHNOLOGY (ICCSNT), VOLS 1-4, 2012, :1978-1982
[33]   Domain Adaptation with Active Learning for Named Entity Recognition [J].
Sun, Huiyu ;
Grishman, Ralph ;
Wang, Yingchao .
CLOUD COMPUTING AND SECURITY, ICCCS 2016, PT II, 2016, 10040 :611-622
[34]   Creating a Dataset for Named Entity Recognition in the Archaeology Domain [J].
Brandsen, Alex ;
Verberne, Suzan ;
Wansleeben, Milco ;
Lambers, Karsten .
PROCEEDINGS OF THE 12TH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION (LREC 2020), 2020, :4573-4577
[35]   Named Entity Recognition for Tweets [J].
Liu, Xiaohua ;
Wei, Furu ;
Zhang, Shaodian ;
Zhou, Ming .
ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY, 2013, 4 (01)
[36]   An Overview of Named Entity Recognition [J].
Sun, Peng ;
Yang, Xuezhen ;
Zhao, Xiaobing ;
Wang, Zhijuan .
2018 INTERNATIONAL CONFERENCE ON ASIAN LANGUAGE PROCESSING (IALP), 2018, :273-278
[37]   Boosted Multifeature Learning for Cross-Domain Transfer [J].
Yang, Xiaoshan ;
Zhang, Tianzhu ;
Xu, Changsheng ;
Yang, Ming-Hsuan .
ACM TRANSACTIONS ON MULTIMEDIA COMPUTING COMMUNICATIONS AND APPLICATIONS, 2015, 11 (03)
[38]   Cross-Domain Expression Recognition Based on Sparse Coding and Transfer Learning [J].
Yang, Yong ;
Zhang, Weiyi ;
Huang, Yong .
MATERIALS SCIENCE, ENERGY TECHNOLOGY, AND POWER ENGINEERING I, 2017, 1839
[39]   A Research Toward Chinese Named Entity Recognition Based on Transfer Learning [J].
Hui Kang ;
Jingwu Xiao ;
Yunpeng Zhang ;
Lei Zhang ;
Xu Zhao ;
Tie Feng .
International Journal of Computational Intelligence Systems, 16
[40]   A Research Toward Chinese Named Entity Recognition Based on Transfer Learning [J].
Kang, Hui ;
Xiao, Jingwu ;
Zhang, Yunpeng ;
Zhang, Lei ;
Zhao, Xu ;
Feng, Tie .
INTERNATIONAL JOURNAL OF COMPUTATIONAL INTELLIGENCE SYSTEMS, 2023, 16 (01)