Water quality prediction using machine learning models based on grid search method

被引:103
作者
Shams, Mahmoud Y. [1 ]
Elshewey, Ahmed M. [2 ]
El-kenawy, El-Sayed M. [3 ]
Ibrahim, Abdelhameed [4 ]
Talaat, Fatma M. [1 ,5 ]
Tarek, Zahraa [6 ]
机构
[1] Kafrelsheikh Univ, Fac Artificial Intelligence, Kafrelsheikh 33516, Egypt
[2] Suez Univ, Fac Comp & Informat, Comp Sci Dept, Suez, Egypt
[3] Delta Higher Inst Engn & Technol, Dept Commun & Elect, Mansoura 35111, Egypt
[4] Mansoura Univ, Fac Engn, Comp Engn & Control Syst Dept, Mansoura 35516, Egypt
[5] New Mansoura Univ, Fac Comp Sci & Engn, Mansoura 35712, Egypt
[6] Mansoura Univ, Fac Comp & Informat, Comp Sci Dept, Mansoura 35561, Egypt
关键词
Water quality; Machine learning models; Grid search; Water quality index; Water quality classification; RIVER; IDENTIFICATION; NETWORKS; SYSTEM; INDEX;
D O I
10.1007/s11042-023-16737-4
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Water quality is very dominant for humans, animals, plants, industries, and the environment. In the last decades, the quality of water has been impacted by contamination and pollution. In this paper, the challenge is to anticipate Water Quality Index (WQI) and Water Quality Classification (WQC), such that WQI is a vital indicator for water validity. In this study, parameters optimization and tuning are utilized to improve the accuracy of several machine learning models, where the machine learning techniques are utilized for the process of predicting WQI and WQC. Grid search is a vital method used for optimizing and tuning the parameters for four classification models and also, for optimizing and tuning the parameters for four regression models. Random forest (RF) model, Extreme Gradient Boosting (Xgboost) model, Gradient Boosting (GB) model, and Adaptive Boosting (AdaBoost) model are used as classification models for predicting WQC. K-nearest neighbor (KNN) regressor model, decision tree (DT) regressor model, support vector regressor (SVR) model, and multi-layer perceptron (MLP) regressor model are used as regression models for predicting WQI. In addition, preprocessing step including, data imputation (mean imputation) and data normalization were performed to fit the data and make it convenient for any further processing. The dataset used in this study includes 7 features and 1991 instances. To examine the efficacy of the classification approaches, five assessment metrics were computed: accuracy, recall, precision, Matthews's Correlation Coefficient (MCC), and F1 score. To assess the effectiveness of the regression models, four assessment metrics were computed: Mean Absolute Error (MAE), Median Absolute Error (MedAE), Mean Square Error (MSE), and coefficient of determination (R2). In terms of classification, the testing findings showed that the GB model produced the best results, with an accuracy of 99.50% when predicting WQC values. According to the experimental results, the MLP regressor model outperformed other models in regression and achieved an R2 value of 99.8% while predicting WQI values.
引用
收藏
页码:35307 / 35334
页数:28
相关论文
共 43 条
[1]   Implementation of data intelligence models coupled with ensemble machine learning for prediction of water quality index [J].
Abba, Sani Isah ;
Pham, Quoc Bao ;
Saini, Gaurav ;
Linh, Nguyen Thi Thuy ;
Ahmed, Ali Najah ;
Mohajane, Meriame ;
Khaledian, Mohammadreza ;
Abdulkadir, Rabiu Aliyu ;
Bach, Quang-Vu .
ENVIRONMENTAL SCIENCE AND POLLUTION RESEARCH, 2020, 27 (33) :41524-41539
[2]   Modelling and Prediction of Water Quality by Using Artificial Intelligence [J].
Al-Adhaileh, Mosleh Hmoud ;
Alsaade, Fawaz Waselallah .
SUSTAINABILITY, 2021, 13 (08)
[3]   RETRACTED: Water Quality Prediction Using Artificial Intelligence Algorithms (Retracted Article) [J].
Aldhyani, Theyazn H. H. ;
Al-Yaari, Mohammed ;
Alkahtani, Hasan ;
Maashi, Mashael .
APPLIED BIONICS AND BIOMECHANICS, 2020, 2020
[4]   River water quality index prediction and uncertainty analysis: A comparative study of machine learning models [J].
Asadollah, Seyed Babak Haji Seyed ;
Sharafati, Ahmad ;
Motta, Davide ;
Yaseen, Zaher Mundher .
JOURNAL OF ENVIRONMENTAL CHEMICAL ENGINEERING, 2021, 9 (01)
[5]  
Beyer K, 1999, LECT NOTES COMPUT SC, V1540, P217
[6]  
Bhardwaj D., 2017, INT J ADV RES COMPUT, V8, P2496
[7]  
Biau G, 2012, J MACH LEARN RES, V13, P1063
[8]  
Breiman L., 1999, Machinelearning202.Pbworks, P1
[9]   Partitioning of daily evapotranspiration using a modified shuttleworth-wallace model, random Forest and support vector regression, for a cabbage farmland [J].
Chen, Han ;
Huang, Jinhui Jeanne ;
McBean, Edward .
AGRICULTURAL WATER MANAGEMENT, 2020, 228
[10]   XGBoost: A Scalable Tree Boosting System [J].
Chen, Tianqi ;
Guestrin, Carlos .
KDD'16: PROCEEDINGS OF THE 22ND ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING, 2016, :785-794