PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 23, 2026Computational Economics1 citationsOpen Access

Data Imputation in Large Datasets: A Comparative Study of PCA and Machine Learning Approaches

View Full Paper
SKSabuhi KhaliliHCHelena Chuliá

Key Points

  • To compare the effectiveness of machine learning techniques and PCA in imputing missing data in large financial datasets.
  • Conducted a systematic comparison of PCA and machine learning for imputation.
  • Introduced a fully linear autoencoder with a customized loss function for observed values.
  • Examined various missingness mechanisms and dataset dimensionalities.
  • The linear autoencoder performs comparably to PCA in data imputation.
  • Random forest methods show lower accuracy in large datasets but excel in lower-dimensional scenarios.
  • MissForest technique achieved the lowest imputation error and highest predictive accuracy in financial applications.

Abstract

Abstract This paper contributes to the financial econometrics literature by providing a systematic comparison of machine learning techniques and principal component analysis (PCA) methods for data imputation in large financial datasets. Missing data is pervasive in empirical finance and economics, and imputation accuracy determines whether the entire dataset can be used in analyses such as modeling credit risk. We introduce a fully linear autoencoder with a loss function tailored to observed values, which performs on par with PCA across various missingness mechanisms while offering greater flexibility. Although random forest-based methods are less accurate in large datasets with a strict factor structure, they demonstrate superior performance in lower-dimensional settings where the primary goal is outcome prediction. This is particularly relevant in applications such as credit default prediction, where identifying risk factors from incomplete borrower data is the main objective. In two such examples, the MissForest random forest technique outperforms others by achieving lower imputation error and higher predictive accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Khalili et al. (2026) studied this question.

synapsesocial.com/papers/69730f9fc8125b09b0d1f545https://doi.org/10.1007/s10614-025-11201-x
Ask AI
Helpful
Bookmark
Share
View Full Paper