metricas

Revista de Psicodidáctica (English Edition)

Suggestions
Revista de Psicodidáctica (English Edition) Understanding students’ mathematical achievement through feature selection and...
Journal Information
Vol. 31. Issue 2.
(July - December 2026)
Cite
Cite
Share
Download PDF
More article options
Visits
157
Vol. 31. Issue 2.
(July - December 2026)
Original

Understanding students’ mathematical achievement through feature selection and machine learning: Evidence from PISA 2018

Comprensión del rendimiento en matemáticas del alumnado mediante la selección de variables y el aprendizaje automático: datos de PISA 2018
Visits
157
Ulku Hilal Yildirima,
Corresponding author
hilal_yildirim@hacettepe.edu.tr

Corresponding author.
, Yasemin Kayhan Atilganb, Derya Turfanb
a Hacettepe University, Graduate School of Science and Engineering, Department of Statistics, Ankara, Turkey
b Hacettepe University, Faculty of Science, Department of Statistics, Ankara, Turkey
This item has received
Article information
Abstract
Full Text
Bibliography
Download PDF
Statistics
Abstract

The Programme for International Student Assessment (PISA) provides a framework for examining academic achievement. Using the Turkish sample of PISA 2018, this study investigates mathematics achievement, focusing on students performing below and above the baseline proficiency level (Level 2). To address this objective, a machine learning classification approach is adopted, integrating feature selection, supervised algorithms, and model interpretation techniques. The analysis includes 6,863 students from 186 schools and 62 student- and school-level variables derived from the PISA questionnaires. Three feature selection methods–Boruta, Mutual Information, and ReliefF–are used to identify subsets of variables that serve as inputs for five classification models: Decision Tree, Bagging, Random Forest, AdaBoost, and Support Vector Machine. Model performance is evaluated using repeated stratified cross-validation across plausible values and multiple classification metrics. The results indicate that models built on substantially reduced feature sets can achieve predictive performance comparable to models using all variables. Among the evaluated configurations, the combination of Mutual Information feature selection with the top 20 variable subset and the Random Forest classifier provides a stable and practical useful balance between prediction performance and model parsimony. Model predictions are examined using conditional SHAP values, partial dependence plots, individual conditional expectation plots, and accumulated local effects plots. Results show that socio-economic background, perceived test difficulty, occupational expectations, school stratum type and behaviors hindering learning are among the variables most strongly associated with model predictions. These analyses capture predominantly non-linear associations between predictors and model predictions, and these relationships vary across students, indicating heterogeneous model responses.

Keywords:
Machine learning
Feature selection
Classification
Mathematical achievement
Educational data mining
Explainable artificial intelligence
Resumen

El Programa para la Evaluación Internacional de los Estudiantes (PISA) proporciona un marco para analizar el rendimiento académico. Utilizando la muestra de Turquía de PISA 2018, el presente estudio investiga el rendimiento en matemáticas, centrándose en el alumnado con un rendimiento inferior y superior al nivel básico de competencia (Nivel 2). Para ello, se adopta el enfoque de clasificación basado en el aprendizaje automático, que integra técnicas de selección de variables, algoritmos supervisados e interpretación de modelos. El análisis incluye a 6.863 estudiantes de 186 centros educativos y 62 variables a nivel de alumnado y a nivel de centro escolar derivadas de los cuestionarios de PISA. Se emplean tres métodos de selección de variables-—Boruta, Información Mutua y ReliefF—-para identificar subconjuntos de variables que sirven como datos de entrada para cinco modelos de clasificación: árbol de decisión, bagging, bosque aleatorio, AdaBoost y máquina de vectores de soporte. El rendimiento de los modelos es evaluado mediante la validación cruzada estratificada y repetida de valores plausibles y múltiples métricas de clasificación.

Los resultados indican que los modelos basados en conjuntos de variables sustancialmente reducidos pueden alcanzar un rendimiento predictivo similar al de los modelos que utilizan todas las variables. Entre las configuraciones evaluadas, la combinación de la selección de variables mediante Información Mutua con el subconjunto de las 20 variables más relevantes y el clasificador de bosque aleatorio ofrece un equilibrio estable y práctico entre el rendimiento de la predicción y la parsimonia del modelo. Las predicciones de los modelos son analizadas mediante valores SHAP condicionales, gráficos de dependencia parcial, gráficos de expectativa condicional individual y gráficos de efectos locales acumulados. Los resultados indican que el entorno socioeconómico, la dificultad percibida de la prueba, las expectativas laborales, el tipo de estrato escolar y los comportamientos que dificultan el aprendizaje son algunas de las variables más estrechamente relacionadas con las predicciones del modelo. Estos análisis captan principalmente asociaciones no lineales entre los indicadores y las predicciones del modelo, y estas relaciones varían según el alumnado, lo que indica que las respuestas del modelo son heterogéneas.

Palabras clave:
Aprendizaje automático
Selección de variables
Clasificación
Rendimiento en matemáticas
Minería de datos educativos
Inteligencia artificial explicable

Article

These are the options to access the full texts of the publication Revista de Psicodidáctica (English Edition)
Subscriber
Subscriber

If you already have your login data, please click here .

If you have forgotten your password you can you can recover it by clicking here and selecting the option “I have forgotten my password”
Purchase
Purchase article

Purchasing article the PDF version will be downloaded

Purchase now
Contact
Phone for subscriptions and reporting of errors
From Monday to Friday from 9 a.m. to 6 p.m. (GMT + 1) except for the months of July and August which will be from 9 a.m. to 3 p.m.
Calls from Spain
932 415 960
Calls from outside Spain
+34 932 415 960
E-mail
asdasdasd
Article options
Tools
Supplemental materials