In this work we present a full-lifecycle machine learning pipeline designed to forecast credit default risk, using real-world lending datasets. The system automates data ingestion, cleaning, feature engineering, model training (using CatBoost), evaluation and deployment via a web interface. We achieve accuracy of ~92.4%, AUC-ROC ~0.97 on held out data. The goal is to give lenders a reliable system to categorize borrower risk (“Low”, “Moderate”, “High”, “Critical”) and thus support more informed credit decisioning. We describe the architecture, methods, key results, practical deployment and implications for financial institutions.
Sharma et al. (Wed,) studied this question.