With the growing demand for precise cross-dialect syntactic analysis, this work proposes an end-to-end framework that automatically labels English syntactic variants and quantifies their distribution across dialects.The approach integrates parser-generated silver annotations, a human-audited gold subset, and a dual-head neural model combining CRF-based sequence tagging and span classification.Domain adaptation with gradient reversal, moment matching, and supervised contrastive learning enhances robustness to dialectal shift, while probability calibration ensures accurate rate estimation.Evaluations on multi-source corpora covering American, British, Australian, and Indian English show that the proposed model improves out-of-dialect macro-F1 by 6.9 points over a strong RoBERTa baseline, reduces domain divergence in encoder space by over 55%, and recovers stable, interpretable contrasts for variants such as that-complementiser drop, particle movement, and dative alternation.
Jia et al. (Thu,) studied this question.