We present a physics-based approach to audio deepfake detection that models spectrograms as dynamical systems on lattice graphs. Our method exploits a fundamental asymmetry: real speech is generated by physical systems (vocal tract, microphone, ADC) that impose cross-frequency phase correlations, while neural vocoders process frequency bands more independently. We introduce three novel techniques: (1) bipartite lattice decomposition measuring coherence between even and odd frequency bins, (2) phase-locking value (PLV) analysis quantifying temporal consistency of phase relationships, and (3) dynamical coupling field evolution identifying persistent phase frustration. Evaluated on the In-The-Wild dataset under G.711 μ-law telephony conditions, our physics-based features achieve 5.43% equal error rate (EER) without neural networks. Combined with self-supervised embeddings, we achieve 3.33% EER—a 44% relative improvement over prior state-of-the-art. Our approach is robust to codec compression because it measures relational structure rather than absolute spectral values.
Sharpe et al. (2026) studied this question.