An Integrated Cluster Detection, Optimization, and Interpretation Approach for Financial Data

被引:161
作者
Li, Tie [1 ]
Kou, Gang [2 ]
Peng, Yi [1 ]
Yu, Philip S. [3 ]
机构
[1] Univ Elect Sci & Technol China, Sch Management & Econ, Chengdu 611731, Peoples R China
[2] Southwestern Univ Finance & Econ, Sch Business Adm, Chengdu 610074, Peoples R China
[3] Univ Illinois, Dept Comp Sci, Chicago, IL 60607 USA
基金
中国博士后科学基金; 中国国家自然科学基金;
关键词
Clustering algorithms; Data models; Correlation; Shape; Optimization; Laplace equations; Feature extraction; Clustering methods; data mining; financial management; spectral analysis; VALIDATION; ALGORITHM;
D O I
10.1109/TCYB.2021.3109066
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In many financial applications, such as fraud detection, reject inference, and credit evaluation, detecting clusters automatically is critical because it helps to understand the subpatterns of the data that can be used to infer user's behaviors and identify potential risks. Due to the complexity of human behaviors and changing social environments, the distributions of financial data are usually complex and it is challenging to find clusters and give reasonable interpretations. The goal of this study is to develop an integrated approach to detect clusters in financial data, and optimize the scope of the clusters such that the clusters can be easily interpreted. Specifically, we first proposed a new cluster quality evaluation criterion, which is free from large-scale computation and can guide base clustering algorithms such as k-Means to detect hyperellipsoidal clusters adaptively. Then, we designed a new solver for a revised support vector data description model, which efficiently refines the centroids and scopes of the detected clusters to make the clusters tighter such that the data in the clusters share greater similarities, and thus, the clusters can be easily interpreted with eigenvectors. Using ten financial datasets, the experiments showed that the proposed algorithm can efficiently find reasonable number of clusters. The proposed approach is suitable for large-scale financial datasets whose features are meaningful, and also applicable to financial mining tasks, such as data distribution interpretation and anomaly detection.
引用
收藏
页码:13848 / 13861
页数:14
相关论文
共 52 条
[1]   Combining Mixture Components for Clustering [J].
Baudry, Jean-Patrick ;
Raftery, Adrian E. ;
Celeux, Gilles ;
Lo, Kenneth ;
Gottardo, Raphael .
JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 2010, 19 (02) :332-353
[2]   Noncooperative Game Strategy in Cyber-Financial Systems With Wiener and Poisson Random Fluctuations: LMIs-Constrained MOEA Approach [J].
Chen, Bor-Sen ;
Chen, Wei-Yu ;
Young, Chun-Tao ;
Yan, Zhiguo .
IEEE TRANSACTIONS ON CYBERNETICS, 2018, 48 (12) :3323-3336
[3]   Spectral Clustering of Customer Transaction Data With a Two-Level Subspace Weighting Method [J].
Chen, Xiaojun ;
Sun, Wenya ;
Wang, Bo ;
Li, Zhihui ;
Wang, Xizhao ;
Ye, Yunming .
IEEE TRANSACTIONS ON CYBERNETICS, 2019, 49 (09) :3230-3241
[4]  
CHENG XH, 1995, J CLIMATE, V8, P2631, DOI 10.1175/1520-0442(1995)008<2631:OROSPD>2.0.CO
[5]  
2
[6]   Cluster stability and the use of noise in interpretation of clustering [J].
Davidson, GS ;
Wylie, BN ;
Boyack, KW .
IEEE SYMPOSIUM ON INFORMATION VISUALIZATION 2001, PROCEEDINGS, 2001, :23-30
[7]  
Davidson I., 2018, ADV NEURAL INFORM PR
[8]   A multilinear singular value decomposition [J].
De Lathauwer, L ;
De Moor, B ;
Vandewalle, J .
SIAM JOURNAL ON MATRIX ANALYSIS AND APPLICATIONS, 2000, 21 (04) :1253-1278
[9]   Deep Feature-Based Text Clustering and its Explanation [J].
Guan, Renchu ;
Zhang, Hao ;
Liang, Yanchun ;
Giunchiglia, Fausto ;
Huang, Lan ;
Feng, Xiaoyue .
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2022, 34 (08) :3669-3680
[10]   A Unified Scheme for Distance Metric Learning and Clustering via Rank-Reduced Regression [J].
Guo, Wenzhong ;
Shi, Yiqing ;
Wang, Shiping .
IEEE TRANSACTIONS ON SYSTEMS MAN CYBERNETICS-SYSTEMS, 2021, 51 (08) :5218-5229