Self-report measures of persistence in science, technology, engineering and mathematics (STEM) programs are used to generate guidelines for educational interventions and justify curriculum decisions. However, the validity of these self-report items is often studied using concurrent measures of intent to persist, instead of information about whether students do persist. The present study uses graduation with a STEM degree as an outcome measure to investigate the predictive validity of self-report STEM persistence measures for undergraduate students (N=3072). This analysis evaluates the predictive accuracy, interpretability, and fairness of a logistic regression model, boosting model, and feedforward neural net when investigating item predictive validity in empirical data and a proof-of-concept simulation. Results suggest that the three models have similar predictive accuracy, that responses to the items yield largest differences in predicted probability for the feedforward neural net, and that all three models lack algorithmic fairness for one of five accuracy metrics.
Hannah K Lewis (2026) studied this question.