Bankruptcy prediction using machine learning models with the text-based communicative value of annual reports
被引:23
|
作者:
Chen, Tsung-Kang
论文数: 0引用数: 0
h-index: 0
机构:
Natl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
Natl Taiwan Univ, Ctr Res Econometr Theory & Applicat, New Taipei, TaiwanNatl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
Chen, Tsung-Kang
[1
,3
]
Liao, Hsien-Hsing
论文数: 0引用数: 0
h-index: 0
机构:
Natl Taiwan Univ, Dept Finance, New Taipei, TaiwanNatl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
Liao, Hsien-Hsing
[2
]
Chen, Geng-Dao
论文数: 0引用数: 0
h-index: 0
机构:
Natl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, TaiwanNatl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
Chen, Geng-Dao
[1
]
Kang, Wei-Han
论文数: 0引用数: 0
h-index: 0
机构:
Natl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, TaiwanNatl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
Kang, Wei-Han
[1
]
Lin, Yu-Chun
论文数: 0引用数: 0
h-index: 0
机构:
Natl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, TaiwanNatl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
Lin, Yu-Chun
[1
]
机构:
[1] Natl Yang Ming Chiao Tung Univ, Dept Management Sci, Hsinchu, Taiwan
[2] Natl Taiwan Univ, Dept Finance, New Taipei, Taiwan
[3] Natl Taiwan Univ, Ctr Res Econometr Theory & Applicat, New Taipei, Taiwan
We investigate whether including the text-based communicative value of annual report increases the predictive power of four machine learning models (Logistic regression, Random Forest, XGBoost, and Support Vector Machine) for corporate bankruptcy prediction using U.S. firm observations from 1994 to 2018. We find that the overall prediction effectiveness of these four models (e.g. accuracy, F1-score, AUCs) significantly improves, especially true in the performance of XGBoost and Random Forest models. In addition, we find that annual report text-based communicative value variables significantly reduce models' Type II error and keep the Type I error at a relatively small level, especially for the short-term bankruptcy forecast. The results reveal that annual report text-based communicative value effectively mitigates the model misidentification of a non-bankrupt firm as a bankrupt firm. Our results also suggest that annual report text-based communicative value is helpful for bank's corporate loan underwriting decisions. Finally, our findings still hold when considering different testing periods and random state settings, replacing by another publicly available bankruptcy dataset, and introducing neural network models.