An Integrated Cluster Detection, Optimization, and Interpretation Approach for Financial Data

被引:164
作者
Li, Tie [1 ]
Kou, Gang [2 ]
Peng, Yi [1 ]
Yu, Philip S. [3 ]
机构
[1] Univ Elect Sci & Technol China, Sch Management & Econ, Chengdu 611731, Peoples R China
[2] Southwestern Univ Finance & Econ, Sch Business Adm, Chengdu 610074, Peoples R China
[3] Univ Illinois, Dept Comp Sci, Chicago, IL 60607 USA
基金
中国国家自然科学基金; 中国博士后科学基金;
关键词
Clustering algorithms; Data models; Correlation; Shape; Optimization; Laplace equations; Feature extraction; Clustering methods; data mining; financial management; spectral analysis; VALIDATION; ALGORITHM;
D O I
10.1109/TCYB.2021.3109066
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
In many financial applications, such as fraud detection, reject inference, and credit evaluation, detecting clusters automatically is critical because it helps to understand the subpatterns of the data that can be used to infer user's behaviors and identify potential risks. Due to the complexity of human behaviors and changing social environments, the distributions of financial data are usually complex and it is challenging to find clusters and give reasonable interpretations. The goal of this study is to develop an integrated approach to detect clusters in financial data, and optimize the scope of the clusters such that the clusters can be easily interpreted. Specifically, we first proposed a new cluster quality evaluation criterion, which is free from large-scale computation and can guide base clustering algorithms such as k-Means to detect hyperellipsoidal clusters adaptively. Then, we designed a new solver for a revised support vector data description model, which efficiently refines the centroids and scopes of the detected clusters to make the clusters tighter such that the data in the clusters share greater similarities, and thus, the clusters can be easily interpreted with eigenvectors. Using ten financial datasets, the experiments showed that the proposed algorithm can efficiently find reasonable number of clusters. The proposed approach is suitable for large-scale financial datasets whose features are meaningful, and also applicable to financial mining tasks, such as data distribution interpretation and anomaly detection.
引用
收藏
页码:13848 / 13861
页数:14
相关论文
共 50 条
[41]   Fuzzy and Real-Coded Chemical Reaction Optimization for Intrusion Detection in Industrial Big Data Environment [J].
Ding, Weiping ;
Nayak, Janmenjoy ;
Naik, Bighnaraj ;
Pelusi, Danilo ;
Mishra, Manohar .
IEEE TRANSACTIONS ON INDUSTRIAL INFORMATICS, 2021, 17 (06) :4298-4307
[42]   INTEGRATED MODELING AND OPTIMIZATION OF MATERIAL FLOW AND FINANCIAL FLOW OF SUPPLY CHAIN NETWORK CONSIDERING FINANCIAL RATIOS [J].
Liu, Qiong ;
Rezaei, Ahmad Reza ;
Wong, Kuan Yew ;
Azami, Mohammad Mahdi .
NUMERICAL ALGEBRA CONTROL AND OPTIMIZATION, 2019, 9 (02) :113-132
[43]   A hierarchical fuzzy cluster ensemble approach and its application to big data clustering [J].
Su, Pan ;
Shang, Changjing ;
Shen, Qiang .
JOURNAL OF INTELLIGENT & FUZZY SYSTEMS, 2015, 28 (06) :2409-2421
[44]   Metaheuristics and Large Language Models Join Forces: Toward an Integrated Optimization Approach [J].
Sartori, Camilo Chacon ;
Blum, Christian ;
Bistaffa, Filippo ;
Corominas, Guillem Rodriguez .
IEEE ACCESS, 2025, 13 :2058-2079
[45]   Bridging the gap: An integrated approach to motif discovery and discord detection in time-series data [J].
Hu, Wentao .
NEUROCOMPUTING, 2025, 619
[46]   Dynamic Clustering and Anomaly Detection of Train Delays in Stream Data: An Incremental Dirichlet Process Approach [J].
Shen, Pengju ;
Song, Liying ;
Wang, Bin .
IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS, 2025,
[47]   Real Time Interpretation and Optimization of Time Series Data Stream in Big Data [J].
Jiang, Zheyuan ;
Liu, Ke .
2018 IEEE 3RD INTERNATIONAL CONFERENCE ON CLOUD COMPUTING AND BIG DATA ANALYSIS (ICCCBDA), 2018, :243-247
[48]   Applying an integrated data-driven weighting system - CoCoSo approach for financial performance evaluation of Fortune 500 companies [J].
Ersoy, Nazli .
E & M EKONOMIE A MANAGEMENT, 2023, 26 (03) :92-108
[49]   Outlier Detection in Financial Data Based on Voronoi Diagram [J].
Qu, Jilin .
2008 4TH INTERNATIONAL CONFERENCE ON WIRELESS COMMUNICATIONS, NETWORKING AND MOBILE COMPUTING, VOLS 1-31, 2008, :9693-9696
[50]   Data mining techniques for the detection of fraudulent financial statements [J].
Kirkos, Efstathios ;
Spathis, Charalambos ;
Manolopoulos, Yannis .
EXPERT SYSTEMS WITH APPLICATIONS, 2007, 32 (04) :995-1003