Data Mining to Identify University Student Dropout Factors
Resumen
University dropout poses academic, social, and economic challenges that call for effective prevention strategies. The objective was to identify determining factors of student dropout through educational data mining and machine learning models. A survey was administered to 527 undergraduate students, and the data were processed with classification algorithms (Adaboost, Gradient Boosting, Extra Trees, Random Forest, Decision Tree, and XGBoost), complemented with interpretation techniques such as SHAP and sensitivity analysis. The results revealed that, in addition to prior academic performance (GPA), psychological support emerged as the most influential predictor across all models, followed by institutional and socioeconomic variables, including academic program, age, and parental job stability. Integrating psychological, institutional, and family factors into predictive systems enhances model accuracy and provides practical evidence to inform educational policies, strengthen student support programs, and design early interventions to promote retention in higher education.
