PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 3, 2026Buildings1 citationsOpen Access

Prediction of Chloride Diffusion Coefficient in Concrete by Micro-Structural Parameters Based on the MLP Method by Considering Data Missing and Small Sample in Database

View Full Paper
RFRongze FuQLQimin LuJZJiaming Zhu

Key Points

  • The MLP model achieves a testing R2 of 0.85, indicating a significant enhancement in prediction accuracy.
  • Data processed through KNN imputation improves the MLP model's performance, validating this approach for missing data.
  • Analysis involves a database of 144 macro–micro property parameters and explores methods for data completion.
  • Virtual sample generation effectively mitigates randomness from small sample sizes and enhances model training.

Abstract

Chloride diffusivity of concrete is essentially determined by its microstructural parameters. Establishing a reliable and accurate prediction model for chloride diffusion has become a research hotspot. In this study, a database containing 144 sets of macro–micro property parameters of concrete is established to train a Multilayer Perceptron (MLP) model. Taking the original collected data as a benchmark, data are randomly missing to simulate data incompleteness, and the models are trained using data filled by the Lagrange, K-Nearest Neighbor (KNN), and Miceforest methods. Moreover, the original data is expanded by the virtual sample generation (VSG) algorithm, based on a Gaussian mixture model (GMM) that fits the joint probability distribution of the original data to generate virtual samples preserving statistical (mean, standard deviation) and physical (e.g., porosity range, pore size ratio) consistency, thus mitigating the randomness caused by small sample sizes. Results indicate that the MLP model demonstrates excellent predictive performance: among schemes handling missing data, the model preprocessed by normalization with KNN imputation yields the best results with testing R2 of 0.78; the baseline model (without missing value filling, normalized) achieves testing R2 of 0.83, MAE of 0.572, and MSE of 0.424. VSG-expanded data significantly enhances the MLP model’s prediction accuracy. When expanding to 3000 groups, the testing R2 reaches 0.85, a 2.4% increase compared to 1000 groups, with further improvements as the dataset expands, confirming the feasibility of the VSG algorithm for small-sample scenarios.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fu et al. (2026) studied this question.

synapsesocial.com/papers/69a75b7ec6e9836116a22e3dhttps://doi.org/10.3390/buildings16030513
Ask AI
Helpful
Bookmark
Share
View Full Paper