Interpretable Machine Learning for Concrete Strength Prediction with Uncertainty Quantification

Authors

  • Jafar Sadeghi Independent Researcher, Boca Raton, Florida, USA

DOI:

https://doi.org/10.32996/jcsts.2026.8.9.5

Keywords:

Concrete compressive strength; machine learning; XGBoost; SHAP; conformal prediction; explainable AI

Abstract

Predicting concrete compressive strength from mix design and curing age is difficult because the relationship between ingredients, age, and strength is nonlinear. This study compares eight regression models on the University of California, Irvine (UCI)/Yeh dataset of 1,030 concrete mixtures and examines three issues that are often overlooked: statistical comparison between models, the value of physics-informed features, and uncertainty for individual predictions. The evaluated models were linear regression, ridge regression, support vector regression, an artificial neural network, random forest, gradient boosting, XGBoost, and a stacking ensemble. XGBoost gave the best held-out performance (coefficient of determination (R²) = 0.931, root mean squared error (RMSE) = 4.22 MPa, and mean absolute error (MAE) = 2.79 MPa) and the highest 10-fold cross-validated R² (0.944 ± 0.019). Paired tests across cross-validation folds showed a consistent advantage over random forest and gradient boosting (p < 0.0001 for both comparisons), although the folds are not fully independent because their training data overlap. An ablation study tested four physics-informed features: water-to-binder ratio, total binder content, aggregate-to-binder ratio, and log-transformed age. These features did not produce a significant improvement in predictive accuracy (paired t-test, p = 0.74), suggesting that gradient-boosted trees can learn similar relationships from the raw variables. However, the engineered features remained useful for interpretation. SHapley Additive exPlanations (SHAP) analysis ranked curing age and water-to-binder ratio as the two most important predictors, which agrees with established concrete behavior. Split conformal prediction was also used to estimate uncertainty for individual mixtures. A single 60/20/20 train/calibration/test split produced 91.7% empirical coverage for a nominal 95% interval. Across 100 repeated random splits, mean coverage was 95.06% ± 2.15%, showing that the method was well calibrated on average. Residual analysis showed larger uncertainty in the 40–60 MPa strength range. The results support XGBoost as an accurate and interpretable model for this benchmark while also showing that external validation on an independent dataset is an important next step.

Author Biography

  • Jafar Sadeghi, Independent Researcher, Boca Raton, Florida, USA

    Civil engineer and independent researcher with research interests in concrete materials, concrete durability, sustainable construction materials, and machine learning for concrete performance prediction.

     

Downloads

Published

2026-09-10

Issue

Section

Research Article

How to Cite

Sadeghi, J. . (2026). Interpretable Machine Learning for Concrete Strength Prediction with Uncertainty Quantification. Journal of Computer Science and Technology Studies, 8(9), 38-53. https://doi.org/10.32996/jcsts.2026.8.9.5