Educational data mining to predict students' academic performance: A survey study

被引:57
作者
Batool, Saba [1 ]
Rashid, Junaid [2 ]
Nisar, Muhammad Wasif [1 ]
Kim, Jungeun [3 ]
Kwon, Hyuk-Yoon [4 ]
Hussain, Amir [5 ]
机构
[1] COMSATS Univ Islamabad, Dept Comp Sci, Wah Campus, Islamabad, Pakistan
[2] Kongju Natl Univ, Dept Comp Sci & Engn, Cheonan 31080, South Korea
[3] Kongju Natl Univ, Dept Software, Dept Comp Sci & Engn, Cheonan 31080, South Korea
[4] Seoul Natl Univ Sci & Technol, Dept Ind Engn, Seoul, South Korea
[5] Edinburgh Napier Univ, Data Sci & Cyber Analyt Res Grp, Edinburgh EH11 4DY, Midlothian, Scotland
基金
新加坡国家研究基金会;
关键词
Educational data mining; Predictive analysis; Students attributes; ARTIFICIAL NEURAL-NETWORK; EARLY WARNING SYSTEMS; ARCHITECTURE STUDENTS; DROPOUT PREDICTION; GENETIC ALGORITHMS; DECISION TREE; MODEL; ONLINE; COURSES; CLASSIFICATION;
D O I
10.1007/s10639-022-11152-y
中图分类号
G40 [教育学];
学科分类号
040101 ; 120403 ;
摘要
Educational data mining is an emerging interdisciplinary research area involving both education and informatics. It has become an imperative research area due to many advantages that educational institutions can achieve. Along these lines, various data mining techniques have been used to improve learning outcomes by exploring large-scale data that come from educational settings. One of the main problems is predicting the future achievements of students before taking final exams, so we can proactively help students achieve better performance and prevent dropouts. Therefore, many efforts have been made to solve the problem of student performance prediction in the context of educational data mining. In this paper, we provide readers with a comprehensive understanding of student performance prediction and compare approximately 260 studies in the last 20 years with respect to i) major factors highly affecting student performance prediction, ii) kinds of data mining techniques including prediction and feature selection algorithms, and iii) frequently used data mining tools. The findings of the comprehensive analysis show that ANN and Random Forest are mostly used data mining algorithms, while WEKA is found as a trending tool for students' performance prediction. Students' academic records and demographic factors are the best attributes to predict performance. The study proves that irrelevant features in the dataset reduce the prediction results and increase model processing time. Therefore, almost half of the studies used feature selection techniques before building prediction models. This study attempts to provide useful and valuable information to researchers interested in advancing educational data mining. The study directs future researchers to achieve highly accurate prediction results in different scenarios using different available inputs or techniques. The study also helps institutions apply data mining techniques to predict and improve student outcomes by providing additional assistance on time.
引用
收藏
页码:905 / 971
页数:67
相关论文
共 269 条
[41]  
Banu SR., 2021, 2021 2 INT C SMART E
[42]  
Baradwaj B.K., 2012, ARXIV12013417
[43]   Towards Automatic Prediction of Student Performance in STEM Undergraduate Degree Programs [J].
Barbosa Manhaes, Laci Mary ;
Serra da Cruz, Sergio Manuel ;
Zimbrao, Geraldo .
30TH ANNUAL ACM SYMPOSIUM ON APPLIED COMPUTING, VOLS I AND II, 2015, :247-253
[44]  
Batool S., 2021, 2021 MOHAMMAD ALI JI
[45]  
Bekele R., 2005, ALGORITHMS, V22, P24
[46]   A Bayesian performance prediction model for mathematics education: A prototypical approach for effective group composition [J].
Bekele, Rahel ;
McPherson, Maggie .
BRITISH JOURNAL OF EDUCATIONAL TECHNOLOGY, 2011, 42 (03) :395-416
[47]  
Bhardwaj B.K., 2012, Data Mining: A prediction for performance improvement using classification
[48]  
Borges VRP., 2018, BRAZILIAN S COMPUTER
[49]  
Bravo L. E. C., 2020, International Journal of Advanced Science and Technology, V29, P11894
[50]  
Bresfelean VP, 2007, ITI, P51