PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 22, 2026Digital Health1 citationsOpen Access

Explainable artificial intelligence approaches for predicting depression by combining feature selection methods and machine learning classifiers

View Full Paper
MYMinyeol YangKLKun Chang LeeKLKwanho Lee

Key Points

  • To enhance the predictive accuracy of depression classification models through feature selection combined with explainable AI techniques.
  • Analyzed microdata from the 2021 National Mental Health Survey of Korea with 5511 adults.
  • Evaluated the impact of different feature selection methods on machine learning classifier performance.
  • Compared performance across 12 machine learning classifiers using diverse feature selection methods.
  • Identified socioeconomic, psychological, and lifestyle factors associated with depression.
  • ReliefF method achieved the highest F2-score (0.9851) with the Stacking classifier.
  • Markov Blanket method performed best in ExtraTrees and LightGBM classifiers (F2-scores = 0.9848, 0.9838).
  • Social distress, reluctance to seek help, quality of life, and physical comorbidities were identified as key predictors of depression.

Abstract

Objective Depression represents a significant global health challenge, further complicated by the multifaceted and complex nature of its diagnosis and treatment. This study explores the application of multiple feature selection (FS) methodologies combined with XAI (explainable artificial intelligence) method named SHapley Additive exPlanations (SHAP) to enhance predictive accuracy in depression classification models using large-scale national survey data. Methods Leveraging microdata from the National Mental Health Survey of Korea (2021), encompassing 5511 Korean adults, this research systematically evaluates how different FS-machine learning classifier combinations affect model performance and identifies nondiagnostic socioeconomic, psychological, and lifestyle factors associated with clinically diagnosed depression. By employing diverse FS methods (e.g., ReliefF, Markov Blanket, and Information Gain) across multiple machine learning classifiers, we systematically compare their performance across 12 classifiers. Results We demonstrate that optimal FS method selection depends on machine learning classifier architecture, with ReliefF excelling in Stacking (F2-score =0.9851) and Markov Blanket performing best in ExtraTrees and LightGBM (F2-score =0.9848, 0.9838). After excluding core diagnostic criteria variables to avoid circularity, our analysis reveals that social distress (loneliness), reluctance to seek professional help, quality of life measures, and physical health comorbidities emerge as highly influential nondiagnostic predictors. Conclusion Our findings advance the field by: (1) systematically demonstrating that FS method effectiveness varies by machine learning classifier type, (2) providing a dual-layer XAI framework combining FS with SHAP for comprehensive interpretability, and (3) identifying culturally specific risk factors in an underrepresented Asian population using high-quality face-to-face collected data. These contributions provide methodological guidance for researchers developing interpretable depression prediction models and offer clinically actionable insights for identifying at-risk individuals in Korean populations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yang et al. (2026) studied this question.

synapsesocial.com/papers/6971bdad642b1836717e2573https://doi.org/10.1177/20552076251411968
Ask AI
Helpful
Bookmark
Share
View Full Paper