Key points are not available for this paper at this time.
Deep learning–based landslide detection models trained on global benchmark datasets have been applied across regions, assuming that learned representations are directly transferable. However, performance often degrades when models are deployed in new geographic and environmental contexts due to domain shift. While prior studies primarily address this limitation through fine-tuning, the relative roles of inference-time calibration and training-time parameter adaptation in governing transferability remain insufficiently understood. This study evaluates the sensitivity of cross-domain transferability to parameter-level interventions applied at different stages of the modeling pipeline. Using Landslide4Sense as a standardized source benchmark and the Philippines as an independent target domain, U-Net and ResU-Net architectures were assessed under a controlled, multi-stage experimental framework. The framework enables systematic isolation of inference-time spectral calibration, training-time spectral adaptation, and training-time hard-negative loss reweighting while holding data inputs, model architectures, and evaluation protocols constant. Sentinel-2 B10 attenuation was examined both as a post-hoc inference-time calibration and as a training-time adaptation to address atmospheric mismatch, followed by the introduction of hard-negative weighting to target persistent false positives arising from complex background conditions. Results show that zero-shot transferability is substantially constrained, despite identical preprocessing and standardized inputs, with high overall accuracy masking poor landslide-detection performance. Inference-time B10 attenuation yielded modest improvements, reflecting decision-level sensitivity rather than improved feature learning. In contrast, incorporating B10 attenuation during fine-tuning led to substantial and stable performance gains, indicating that domain mismatch can be internalized through representation learning. Validation-based threshold calibration provided minimal additional benefit after fine-tuning. Hard-negative weighting further influenced performance by reshaping the precision–recall trade-off, but its effects were configuration-specific and non-linear, depending on its interaction with prior spectral adaptation. Residual learning improved performance only when combined with training-time adaptation and targeted supervision, and did not enhance zero-shot transferability. Model generalizability was further assessed through cross-validation on the target-domain dataset, demonstrating that both architectures maintain stable performance across the evaluated target-domain folds once adapted, and that poor zero-shot performance reflects domain mismatch rather than intrinsic model instability. Overall, the findings show that cross-domain transferability in landslide detection is governed by interactions between training- and inference-time parameters rather than by architectural complexity alone. By disentangling calibration effects from representation learning and revealing non-linear parameter interactions, this study provides a structured framework for rigorously evaluating transferability in benchmark-driven remote sensing applications, as demonstrated in the present source–target case study.
Inting et al. (2026) studied this question.