Data-Centric Solutions for Addressing Big Data Veracity with Class Imbalance, High Dimensionality, and Class Overlapping

被引:1
作者
Bolivar, Armando [1 ]
Garcia, Vicente [2 ]
Alejo, Roberto [3 ]
Florencia-Juarez, Rogelio [2 ]
Sanchez, J. Salvador [4 ]
机构
[1] Univ Autonoma Ciudad Juarez, Inst Ingn & Tecnol, Av Charro 450 NTE, Ciudad Juarez 32310, Chihuahua, Mexico
[2] Univ Autonoma Ciudad Juarez, Div Multidisciplinaria Ciudad Univ, Av Jose de Jesus Delgado 18100, Ciudad Juarez 32579, Chihuahua, Mexico
[3] Inst Tecnol Toluca, Tecnol Nacl Mexico, Div Postgrad Studies & Res, Av Tecnol S-N, Metepec 52149, Estado De Mexic, Mexico
[4] Univ Jaume 1, Inst New Imaging Technol, Dept Comp Languages & Syst, Av Vicent Sos Baynat S-N, Castellon De La Plana 12071, Spain
来源
APPLIED SCIENCES-BASEL | 2024年 / 14卷 / 13期
关键词
big data; class imbalance; high dimensionality; fractional norms; dissimilarity representation; MACHINE-LEARNING ALGORITHMS; SMOTE; DESIGN;
D O I
10.3390/app14135845
中图分类号
O6 [化学];
学科分类号
0703 ;
摘要
An innovative strategy for organizations to obtain value from their large datasets, allowing them to guide future strategic actions and improve their initiatives, is the use of machine learning algorithms. This has led to a growing and rapid application of various machine learning algorithms with a predominant focus on building and improving the performance of these models. However, this data-centric approach ignores the fact that data quality is crucial for building robust and accurate models. Several dataset issues, such as class imbalance, high dimensionality, and class overlapping, affect data quality, introducing bias to machine learning models. Therefore, adopting a data-centric approach is essential to constructing better datasets and producing effective models. Besides data issues, Big Data imposes new challenges, such as the scalability of algorithms. This paper proposes a scalable hybrid approach to jointly addressing class imbalance, high dimensionality, and class overlapping in Big Data domains. The proposal is based on well-known data-level solutions whose main operation is calculating the nearest neighbor using the Euclidean distance as a similarity metric. However, these strategies may lose their effectiveness on datasets with high dimensionality. Hence, the data quality is achieved by combining a data transformation approach using fractional norms and SMOTE to obtain a balanced and reduced dataset. Experiments carried out on nine two-class imbalanced and high-dimensional large datasets showed that our scalable methodology implemented in Spark outperforms the traditional approach.
引用
收藏
页数:15
相关论文
共 48 条
[31]  
NASA, AVIRIS: Airborne Visible - Infrared Imaging Spectrometer
[32]  
Ng A., 2021, Harvard Business Review, Jul
[33]   Big data and machine learning algorithms for health-care delivery [J].
Ngiam, Kee Yuan ;
Khor, Ing Wei .
LANCET ONCOLOGY, 2019, 20 (05) :E262-E273
[34]  
Onyejekwe E.R., 2024, Perspect. Health Inf. Manag, V21, P43
[35]   Bias and Unfairness in Machine Learning Models: A Systematic Review on Datasets, Tools, Fairness Metrics, and Identification and Mitigation Methods [J].
Pagano, Tiago P. ;
Loureiro, Rafael B. ;
Lisboa, Fernanda V. N. ;
Peixoto, Rodrigo M. ;
Guimaraes, Guilherme A. S. ;
Cruz, Gustavo O. R. ;
Araujo, Maira M. ;
Santos, Lucas L. ;
Cruz, Marco A. S. ;
Oliveira, Ewerton L. S. ;
Winkler, Ingrid ;
Nascimento, Erick G. S. .
BIG DATA AND COGNITIVE COMPUTING, 2023, 7 (01)
[36]   Dissimilarity representations allow for building good classifiers [J].
Pekalska, E ;
Duin, RPW .
PATTERN RECOGNITION LETTERS, 2002, 23 (08) :943-956
[37]   A Survey on Graphical Methods for Classification Predictive Performance Evaluation [J].
Prati, Ronaldo C. ;
Batista, Gustavo E. A. P. A. ;
Monard, Maria Carolina .
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2011, 23 (11) :1601-1618
[38]  
Reinsel David., 2017, Don't Focus on Big Data
[39]   Data Sampling Methods to Deal With the Big Data Multi-Class Imbalance Problem [J].
Rendon, Erendira ;
Alejo, Roberto ;
Castorena, Carlos ;
Isidro-Ortega, Frank J. ;
Granda-Gutierrez, Everardo E. .
APPLIED SCIENCES-BASEL, 2020, 10 (04)
[40]  
Sambasivan Nithya, 2021, Data Cascades in High-Stakes AI, P1, DOI DOI 10.1145/3411764.3445518