FEATURE SELECTION AND CLASSIFICATION INTEGRATED METHOD FOR IDENTIFYING CITED TEXT SPANS FOR CITANCES ON IMBALANCED DATA

被引:0
|
作者
Yee, Jen-Yuan [1 ]
Tsai, Cheng-Jung [2 ]
Hsu, Tien-Yu [3 ]
Lin, Jung-Yi [4 ]
Cheng, Pei-Cheng [5 ]
机构
[1] Natl Museum Nat Sci, Visitor Serv, Dept Operat, Collect & Informat Management, Taichung 40453, Taiwan
[2] Natl Changhua Univ Educ, Grad Inst Stat & Informat Sci, Changhua 50007, Taiwan
[3] Natl Museum Nat Sci, Dept Sci Educ, Taichung 40453, Taiwan
[4] Hon Hai Precis IndCo Ltd Foxconn, IP Affairs Div, Taipei 11492, Taiwan
[5] Chien Hsin Univ Sci & Technol, Dept Informat Management, Taoyuan 32097, Taiwan
关键词
Citation analysis; cited text spans identification; feature selection; classification; class imbalance; performance evaluation; scientific paper summarization;
D O I
10.22452/mjcs.vol34no4.3
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Recent studies in scientific paper summarization have explored a new form of structured summary for a reference paper by grouping all cited and citing sentences together by facet. This involves three main tasks: (1) identifying cited text spans for citances (i.e., citing sentences), (2) classifying their discourse facets, and (3) generating a structured summary from the cited text spans and citances. This paper focuses on the first task, and approaches the task as binary classification to distinguish relevant pairs of citances and reference sentences from irrelevant pairs. We propose a new method that integrates feature selection and classification techniques to enhance classification performance. The proposed method investigates combinations of six feature selection methods (chi(2)-Statistics, Information Gain, Gain Ratio, Relief-F, Significance Attribute Evaluation, and Symmetrical Uncertainty), and five classification algorithms (k-Nearest Neighbors, Decision Tree, Support Vector Machine, Naive Bayes, and Random Forest). Additionally, to address imbalanced data during training, we apply SMOTE (Synthetic Minority Over sampling Technique) to introduce synthetic biases towards the minority. Experiments are conducted using the CLSciSumm corpora to compare the effect of feature selection applied to classification. The results reveal the benefits of feature selection in significantly boosting performance of F-1 score metric, and show that our method is competitive to the state-of-the-art methods in the CL-SciSumm evaluations.
引用
收藏
页码:355 / 373
页数:19
相关论文
共 50 条
  • [41] A New Big Data Feature Selection Approach for Text Classification
    Amazal, Houda
    Kissi, Mohamed
    SCIENTIFIC PROGRAMMING, 2021, 2021
  • [42] Penalized multiple distribution selection method for imbalanced data classification
    Shi, Ge
    Feng, Chong
    Xu, Wenfu
    Liao, Lejian
    Huang, Heyan
    KNOWLEDGE-BASED SYSTEMS, 2020, 196
  • [43] The Research Of Feature Selection Of Text Classification Based On Integrated Learning Algorithm
    Xia Huosong
    Liu Jian
    2011 TENTH INTERNATIONAL SYMPOSIUM ON DISTRIBUTED COMPUTING AND APPLICATIONS TO BUSINESS, ENGINEERING AND SCIENCE (DCABES), 2011, : 20 - 22
  • [44] Integration of feature vector selection and support vector machine for classification of imbalanced data
    Liu, Jie
    Zio, Enrico
    APPLIED SOFT COMPUTING, 2019, 75 : 702 - 711
  • [45] A hybrid method of feature selection for Chinese text sentiment classification
    Wang, Suge
    Wei, Yingjie
    Li, Deyu
    Zhang, Wu
    Li, Wei
    FOURTH INTERNATIONAL CONFERENCE ON FUZZY SYSTEMS AND KNOWLEDGE DISCOVERY, VOL 3, PROCEEDINGS, 2007, : 435 - +
  • [46] Research on Feature Selection Method in Chinese Text Automatic Classification
    Hong, Ying
    Shao, Xiwen
    PROCEEDINGS OF THE 2015 INTERNATIONAL CONFERENCE ON APPLIED SCIENCE AND ENGINEERING INNOVATION, 2015, 12 : 1759 - 1763
  • [47] Research on feature selection method in Chinese text automatic classification
    Hong, Ying
    Geng, Zengmin
    ENERGY SCIENCE AND APPLIED TECHNOLOGY, 2016, : 359 - 361
  • [48] Two-stage Feature Selection Method for Text Classification
    Li Xi
    Dai Hang
    Wang Mingwen
    MINES 2009: FIRST INTERNATIONAL CONFERENCE ON MULTIMEDIA INFORMATION NETWORKING AND SECURITY, VOL 1, PROCEEDINGS, 2009, : 234 - +
  • [49] A novel filter feature selection method for text classification: Extensive Feature Selector
    Parlak, Bekir
    Uysal, Alper Kursat
    JOURNAL OF INFORMATION SCIENCE, 2023, 49 (01) : 59 - 78
  • [50] Feature Selection Method Based on Weighted Mutual Information for Imbalanced Data
    Li, Kewen
    Yu, Mingxiao
    Liu, Lu
    Li, Timing
    Zhai, Jiannan
    INTERNATIONAL JOURNAL OF SOFTWARE ENGINEERING AND KNOWLEDGE ENGINEERING, 2018, 28 (08) : 1177 - 1194