PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026BioScience Trends0 citationsOpen Access

Predicting non-alcoholic fatty liver disease (NAFLD) using machine learning algorithms: Evidence from a large-scale community cohort in Taiwan

TLTzu-Chun LinYWYu-Ju WeiPLPo-Cheng Liang

Key Points

  • The study aims to identify key risk factors for non-alcoholic fatty liver disease using machine learning algorithms.
  • Analyzed data from community health examinations in southern Taiwan.
  • Evaluated five machine learning algorithms: LR, RF, KNN, AdaBoost, and XGBoost.
  • Used synthetic minority over-sampling technique to balance dataset for training and testing.
  • Assessed model performance using metrics like accuracy, precision, recall, F1 score, and AUROC.
  • XGBoost achieved the highest predictive accuracy at 83.48%.
  • Identified key predictors: LDL-C, BMI, waist circumference, FPG, and triglycerides.
  • Model AUROC was 92.85%, indicating excellent discriminatory ability.

Abstract

Closely associated with metabolic disorders, non-alcoholic fatty liver disease (NAFLD) substantially increases the risk of hepatocellular carcinoma. This study aimed to apply machine learning (ML) algorithms to a community-based cohort in southern Taiwan to identify key risk factors for NAFLD and to develop predictive models with clinical applicability. Data were derived from community health examinations, and eighteen clinical and demographic features were analyzed. Five ML algorithms were evaluated: logistic regression (LR), random forest (RF), K-nearest neighbors (KNN), adaptive boosting (AdaBoost), and extreme gradient boosting (XGBoost). Model performance was assessed using accuracy, precision, recall, F1 score, and area under the receiver operating characteristic curve (AUROC). A total of 7,510 participants were included (38.8% male; mean age 50.9 ± 15.0 years). The dataset was randomly divided into training (80%) and testing (20%) subsets, with no significant differences observed between groups in most independent variables. The Synthetic Minority Over-sampling Technique (SMOTE) was employed to balance NAFLD and non-NAFLD groups in the training dataset. Among all models, XGBoost achieved the highest performance, with an accuracy of 83.48%, precision of 84.31%, recall of 81.21%, F1 score of 82.72%, and AUROC of 92.85%. Feature importance analysis identified low-density lipoprotein cholesterol (LDL-C), body mass index (BMI), waist circumference, fasting plasma glucose (FPG), and triglycerides (TG) as the most influential predictors of NAFLD. ML algorithms, particularly XGBoost, demonstrated high accuracy in predicting NAFLD and effectively identified key clinical predictors. These findings may enhance early diagnosis and facilitate the development of targeted intervention strategies in the management of NAFLD.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Lin et al. (2026) studied this question.

synapsesocial.com/papers/69a286490a974eb0d3c011achttps://doi.org/10.5582/bst.2025.01323
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Non-Alcoholic Fatty Liver Disease Prediction with Feature Optimized XGBoost Model2024 · 1 citations
  2. 2Machine learning-based mortality prediction models for non-alcoholic fatty liver disease in the general United States population2024 · 1 citations
  3. 3Development of Cost-Effective Fatty Liver Disease Prediction Models in a Chinese Population: Statistical and Machine Learning Approaches2024 · 5 citations
  4. 4Investigation of predictive factors for fatty liver in children and adolescents using artificial intelligence2025
  5. 5Precision Non-Alcoholic Fatty Liver Disease (NAFLD) Diagnosis: Leveraging Ensemble Machine Learning and Gender Insights for Cost-Effective Detection2024 · 5 citations