PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 13, 2026Discover Analytics0 citationsOpen Access

Flexible statistical approaches for modeling nonlinear relationships in diabetes prediction using splines, Bayesian kernel regression and Bayesian regression trees

TRThimani Dananjana RanathungageHBHarsha BlumerSMSaman Muthukumarana

Key Points

  • This research aims to compare flexible statistical methods for modeling nonlinear relationships in diabetes prediction.
  • Compared restricted cubic spline regression, Bayesian kernel machine regression, and Bayesian additive regression trees.
  • Applied methods to the Pima Indians diabetes dataset with 768 observations.
  • Identified significant predictors such as glucose, insulin, age, and skin thickness.
  • Evaluated model performance using six specific metrics.
  • Restricted cubic spline regression and Bayesian kernel machine regression identified key predictors.
  • Bayesian additive regression trees achieved the best predictive accuracy with an AUC of approximately 97%.
  • Predictor-response functions were created for better clinical interpretation.

Abstract

Modeling nonlinear relationships is a fundamental challenge in statistical analysis, particularly when predictors exhibit complex and interacting effects on outcomes. This study compares three flexible methods for capturing such structures: restricted cubic spline regression (RCS), Bayesian kernel machine regression (BKMR), and Bayesian additive regression trees (BART). RCS enables explicit modeling of nonlinear associations via spline basis functions, BKMR leverages kernel functions within a Bayesian framework to capture nonlinear and non-additive effects, and BART provides a nonparametric ensemble approach that flexibly accommodates interactions and nonlinearities without prior specification. To demonstrate the utility of these methods, we apply them to the Pima Indians diabetes dataset, consisting of 768 observations of women at high risk of type II diabetes. After data preprocessing, including imputation and outlier handling, each method was fitted and evaluated using six performance measures. RCS and BKMR identified glucose, insulin, age, and skin thickness as significant predictors, while BART yielded the best predictive performance (AUC 97%). Predictor-response functions were used to enhance clinical interpretability. These findings illustrate that a methodology capable of capturing nonlinear effects can substantially improve prediction accuracy in epidemiological studies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ranathungage et al. (2026) studied this question.

synapsesocial.com/papers/69b3ac9002a1e69014cce6b5https://doi.org/10.1007/s44257-026-00057-6
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Pima Indians diabetes mellitus classification based on machine learning (ML) algorithms2022 · 325 citations
  2. 2Association Between Body Mass Index and Diabetes in Northeastern China2016 · 22 citations
  3. 3Prediction of Diabetes using Classification Algorithms2018 · 901 citations
  4. 4Enhanced anomaly detection through a Bayesian framework with a novel network merging structure learning approach2025 · 1 citations
  5. 5Bayesian kernel machine regression for estimating the health effects of multi-pollutant mixtures2014 · 1,847 citations