Deep learning architectures dominate contemporary solar radiation forecasting research, yet their effectiveness under limited data regimes remains insufficiently examined. This study investigates model capacity alignment under data scarcity using daily solar irradiance datasets from Nigeria, Ghana, and Senegal. Approximately 700 observations per country are used to compare the performance of Random Forest, XGBoost, LSTM, and CNN–LSTM models under identical experimental conditions. Results show that gradient-boosted tree ensembles achieve R² values up to 0.98, while deep recurrent architectures perform near baseline levels under the same small-sample conditions. The findings are interpreted through bias–variance trade-offs, parameter-to-sample scaling, effective sample size under autocorrelation, and structural regularisation in ensemble methods. The study demonstrates that model capacity must align with dataset scale, particularly in emerging energy infrastructures where historical observations remain limited.
Majid Rasheed (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: