Global pre-harvest crop yield forecasting is challenged by diverse data characteristics and environmental conditions across different countries, crop types, and lead times. The datasets used in crop yield predictive modeling differ in coverage, sample size, and signal noise. To address this challenge, we propose a diagnostic spatiotemporal multimodal fusion network that selects model structure based on dataset characteristics to determine whether temporal trend coupling and spatial module activation are warranted. Guided by these diagnostic outcomes, we develop a spatiotemporal multimodal fusion deep learning model that conditionally predicts detrended residuals and activates geolocation encoding only when spatial autocorrelation is detected. We evaluate the approach on a crop yield benchmark (CY-Bench), covering maize across 38 countries and wheat across 29 countries, under three lead times (early, mid, and late season) against widely used baselines. Our proposed approach achieves the lowest pooled NRMSE for both crops at all lead times, and it remains in the leading group for complementary metrics including MAPE and KGE. The diagnosis shows that different crops need different model structures. Wheat always relies on the full spatiotemporal model. However, maize mostly sticks to the temporal trend-only structure (without spatial inputs) at mid-season. Ablation studies further show that diagnostic module selection provides the primary performance lift, while fusion strategies offer secondary improvements. Variance decomposition shows that performance varies more across countries than across models. This research highlights that automated data diagnostics offers a strategic advantage in managing dataset heterogeneity, serving as a critical prerequisite to model configuration in global, large-scale crop yield forecasting.
Zhuang et al. (Fri,) studied this question.