Over the past five years, GeoAI has made great progress in automating geospatial analysis and decision-making, with applications ranging from environmental monitoring to land-use classification. Recent advances in foundation models and earth embeddings have pushed these capabilities to new levels. But for all this progress, this paper argues that GeoAI has remained limited in bridging computational performance with the platial, experiential, and contextual knowledge essential to geography. This gap is embedded in a paradox: as GeoAI models become more powerful, the methods we use to evaluate them have remained focused on quantitative benchmarks and computational performance metrics, primarily accuracy-based measures and calibration errors that quantify the gap between a model's predictions and the actual outcomes.
Yue Lin (Mon,) studied this question.