Computational biology relies heavily on parameterized probability distributions---from Negative Binomials for transcriptomics to categorical distributions in protein language models. Yet, analytical pipelines routinely abandon these probabilistic foundations when comparing samples, defaulting to ad hoc heuristics like Euclidean distance or Bray-Curtis dissimilarity. This practice amounts to navigating a curved statistical space with a flat ruler, severing the mathematical link between the statistical model and the geometric comparison. In this critical review, we argue that the choice of distance metric is intrinsically tied to the chosen probabilistic model. By Chentsov's theorem, parameterizing a distribution equips it with a unique Riemannian manifold governed by the Fisher information metric. However, we focus heavily on the boundaries of applicability: biological data often violates strict parametric assumptions, creating a critical tension between the theoretical rigor of information geometry and the pragmatic robustness of heuristic distances under model misspecification. We provide guidance on navigating these regimes, outlining when parametric assumptions hold and when nonparametric alternatives provide a necessary fallback. We also bring these principles to the deep learning frontier, demonstrating how neural networks navigate these identical geometric constraints. Ultimately, adhering to native geometry ensures that when an analysis fails, it fails legibly as a modeling error rather than a hidden geometric artifact.
Dillion Fox (Mon,) studied this question.