Generalized Funnelling: Ensemble Learning and Heterogeneous Document Embeddings for Cross-Lingual Text Classification

被引:3
作者
Moreo, Alejandro [1 ]
Pedrotti, Andrea [1 ]
Sebastiani, Fabrizio [1 ]
机构
[1] CNR, Ist Sci & Tecnol Informaz, Via Giuseppe Moruzzi 1, I-56124 Pisa, Italy
基金
欧盟地平线“2020”;
关键词
Transfer learning; heterogeneous transfer learning; cross-lingual text classification; ensemble learning; word embeddings; REPRESENTATION;
D O I
10.1145/3544104
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Funnelling (FUN) is a recently proposed method for cross-lingual text classification (CLTC) based on a two-tier learning ensemble for heterogeneous transfer learning (HTL). In this ensemble method, 1st-tier classifiers, each working on a different and language-dependent feature space, return a vector of calibrated posterior probabilities (with one dimension for each class) for each document, and the final classification decision is taken by a meta-classifier that uses this vector as its input. The meta-classifier can thus exploit class-class correlations, and this (among other things) gives FUN an edge over CLTC systems in which these correlations cannot be brought to bear. In this article, we describe Generalized FUNnelling (GFUN), a generalization of FUN consisting of an HTL architecture in which 1st-tier components can be arbitrary view-generating FUNctions, i.e., language-dependent FUNctions that each produce a language-independent representation ("view") of the (monolingual) document. We describe an instance of GFUN in which the meta-classifier receives as input a vector of calibrated posterior probabilities (as in FUN) aggregated to other embedded representations that embody other types of correlations, such as word-class correlations (as encoded by Word-Class Embeddings), word-word correlations (as encoded by Multilingual Unsupervised or Supervised Embeddings), and word-context correlations (as encoded by multilingual BERT). We show that this instance of GFUN substantially improves over FUN and over state-of-the-art baselines by reporting experimental results obtained on two large, standard datasets for multilingual multilabel text classification. Our code that implements GFUN is publicly available.
引用
收藏
页数:37
相关论文
共 50 条
  • [41] Ensemble Learning for Multi-Type Classification in Heterogeneous Networks
    Serafino, Francesco
    Pio, Gianvito
    Ceci, Michelangelo
    [J]. IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2018, 30 (12) : 2326 - 2339
  • [42] Optimal heterogeneous domain adaptation for text classification in transfer learning
    Khurana, Anshu
    Verma, Om Prakash
    [J]. COMPUTERS & ELECTRICAL ENGINEERING, 2024, 116
  • [43] A joint learning approach with knowledge injection for zero-shot cross-lingual hate speech detection
    Pamungkas, Endang Wahyu
    Basile, Valerio
    Patti, Viviana
    [J]. INFORMATION PROCESSING & MANAGEMENT, 2021, 58 (04)
  • [44] Cross-Lingual Few-Shot Hate Speech and Offensive Language Detection Using Meta Learning
    Mozafari, Marzieh
    Farahbakhsh, Reza
    Crespi, Noel
    [J]. IEEE ACCESS, 2022, 10 : 14880 - 14896
  • [45] Learning a Dual-Language Vector Space for Domain-Specific Cross-Lingual Question Retrieval
    Chen, Guibin
    Chen, Chunyang
    Xing, Zhenchang
    Xu, Bowen
    [J]. 2016 31ST IEEE/ACM INTERNATIONAL CONFERENCE ON AUTOMATED SOFTWARE ENGINEERING (ASE), 2016, : 744 - 755
  • [46] Cross corpus multi-lingual speech emotion recognition using ensemble learning
    Wisha Zehra
    Abdul Rehman Javed
    Zunera Jalil
    Habib Ullah Khan
    Thippa Reddy Gadekallu
    [J]. Complex & Intelligent Systems, 2021, 7 : 1845 - 1854
  • [47] Cross corpus multi-lingual speech emotion recognition using ensemble learning
    Zehra, Wisha
    Javed, Abdul Rehman
    Jalil, Zunera
    Khan, Habib Ullah
    Gadekallu, Thippa Reddy
    [J]. COMPLEX & INTELLIGENT SYSTEMS, 2021, 7 (04) : 1845 - 1854
  • [48] Genetic Programming based Transfer Learning for Document Classification with Self-taught and Ensemble Learning
    Fu, Wenlong
    Xue, Bing
    Gao, Xiaoying
    Zhang, Mengjie
    [J]. 2019 IEEE CONGRESS ON EVOLUTIONARY COMPUTATION (CEC), 2019, : 2260 - 2267
  • [49] Emotional Text Analysis Based on Ensemble Learning of Three Different Classification Algorithms
    Bian, WenShuo
    Wang, ChunZhi
    Ye, ZhiWei
    Yan, Lingyu
    [J]. PROCEEDINGS OF THE 2019 10TH IEEE INTERNATIONAL CONFERENCE ON INTELLIGENT DATA ACQUISITION AND ADVANCED COMPUTING SYSTEMS - TECHNOLOGY AND APPLICATIONS (IDAACS), VOL. 2, 2019, : 938 - 941
  • [50] Automated classification of clinical trial eligibility criteria text based on ensemble learning and metric learning
    Kun Zeng
    Yibin Xu
    Ge Lin
    Likeng Liang
    Tianyong Hao
    [J]. BMC Medical Informatics and Decision Making, 21