ABSTRACT There is increasing complexity in global food supply chains and an increase in mislabeling, adulteration, and ingredient fraud. The study assesses and compares four supervised machine learning techniques: Logistic regression (LR), Decision tree (DT), Random Forest (RFT), and XGBoost (XGB) to identify fraudulent food products from the multisource data of 12,000 real‐world items comprising UK Food Standards Agency source, Open Food Facts, and Kaggle. The dataset covers seven fraud‐relevant features such as origin mismatch, absence of label items, presence of additives, allergen declaration statement, number of ingredients, etc. After applying SMOTE to address class imbalance, the dataset was divided into 80% for training and the remaining 20% for testing. The models' performance was measured through accuracy, precision, recall, F1‐score, and AUC. Among all, XGB outperforms others, with the highest Precision (0.919), Recall (0.675), F1‐score of 0.779, and AUC of 0.95, highlighting strong potential for reliable detection of food fraud. This piece of work is novel in that it combines multi‐source product‐level data, domain‐specific feature engineering, and comparative analysis of traditional and ensemble ML classifiers and overcame the shortcomings of its predecessors, which utilized small‐scale datasets more often, single‐product datasets, and lab‐based datasets. The results indicate that XGB provides a strong, scalable system of early fraudulent food product detection, which has a practical purpose in food businesses, regulators, and surveillance as a way of improving food authenticity and consumer confidence.
Agrawal et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: