An Integrated Cluster Detection, Optimization, and Interpretation Approach for Financial Data

被引:162
作者
Li, Tie [1 ]
Kou, Gang [2 ]
Peng, Yi [1 ]
Yu, Philip S. [3 ]
机构
[1] Univ Elect Sci & Technol China, Sch Management & Econ, Chengdu 611731, Peoples R China
[2] Southwestern Univ Finance & Econ, Sch Business Adm, Chengdu 610074, Peoples R China
[3] Univ Illinois, Dept Comp Sci, Chicago, IL 60607 USA
基金
中国博士后科学基金; 中国国家自然科学基金;
关键词
Clustering algorithms; Data models; Correlation; Shape; Optimization; Laplace equations; Feature extraction; Clustering methods; data mining; financial management; spectral analysis; VALIDATION; ALGORITHM;
D O I
10.1109/TCYB.2021.3109066
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In many financial applications, such as fraud detection, reject inference, and credit evaluation, detecting clusters automatically is critical because it helps to understand the subpatterns of the data that can be used to infer user's behaviors and identify potential risks. Due to the complexity of human behaviors and changing social environments, the distributions of financial data are usually complex and it is challenging to find clusters and give reasonable interpretations. The goal of this study is to develop an integrated approach to detect clusters in financial data, and optimize the scope of the clusters such that the clusters can be easily interpreted. Specifically, we first proposed a new cluster quality evaluation criterion, which is free from large-scale computation and can guide base clustering algorithms such as k-Means to detect hyperellipsoidal clusters adaptively. Then, we designed a new solver for a revised support vector data description model, which efficiently refines the centroids and scopes of the detected clusters to make the clusters tighter such that the data in the clusters share greater similarities, and thus, the clusters can be easily interpreted with eigenvectors. Using ten financial datasets, the experiments showed that the proposed algorithm can efficiently find reasonable number of clusters. The proposed approach is suitable for large-scale financial datasets whose features are meaningful, and also applicable to financial mining tasks, such as data distribution interpretation and anomaly detection.
引用
收藏
页码:13848 / 13861
页数:14
相关论文
共 52 条
[51]   Scalable Mining of Contextual Outliers Using Relevant Subspace [J].
Zhang, Jifu ;
Yu, Xiaolong ;
Xun, Yaling ;
Zhang, Sulan ;
Qin, Xiao .
IEEE TRANSACTIONS ON SYSTEMS MAN CYBERNETICS-SYSTEMS, 2020, 50 (03) :988-1002
[52]   Granular Data Description: Designing Ellipsoidal Information Granules [J].
Zhu, Xiubin ;
Pedrycz, Witold ;
Li, Zhiwu .
IEEE TRANSACTIONS ON CYBERNETICS, 2017, 47 (12) :4475-4484