Background/Objectives: Prediabetes is a metabolic condition involving various phenotypes of glucose metabolism. Prediabetes increases the risk of heart disease, among other conditions. Hence, we employed machine learning tools to characterize phenotypes associated with cardiovascular disease using routine laboratory biomarkers. Methods: We processed laboratory records of over 1,000,000 de-identified individuals, resulting in a sample of 3024 individuals classified as prediabetic (fasting blood glucose 100–125 mg/dL combined with HbA1c 5.7–6.4%). Lipid profile parameters (total cholesterol TC, HDL-C, LDL-C, and triglycerides) and associated indices (atherogenic index of plasma, Log10(TG/HDL-C), triglyceride–glucose index TyG, TC/HDL-C, and LDL-C/HDL-C, among others) were analyzed using the k-means algorithm. Two groups emerged based on biomarker concentrations, a pro-atherogenic cluster (P-AC; n = 1113) and a less-atherogenic cluster (L-AC; n = 1911) for cardiovascular disease. Results: We assessed the performance of biomarkers in the P-AC and L-AC clusters using a receiver operating characteristic curve. Triglycerides (area under the curve AUC 0.977), AIP (AUC 0.978), and triglyceride–glucose index (AUC 0.974) showed sensitivity and specificity >90%. The TC/HDL-C (AUC 0.903) and LDL-C/HDL-C (AUC 0.865) indices also performed well, with sensitivity and specificity of 80%. Binomial logistic regression applied to the groups generated by k-means using the biomarkers AIP and LDL-C/HDL-C showed an AUC of 0.984 and accuracy above 93%. Conclusions: The k-means algorithm enabled the identification of a P-AC for cardiovascular disease among prediabetics using cost-effective laboratory biomarkers that are widely accessible in laboratories. Individuals classified as P-AC may benefit from differentiated treatment to minimize this factor.
Signorini et al. (Fri,) studied this question.