We investigate whether missing training categories create detectable topological signatures in learned representations. Comparing PCA, a deterministic autoencoder (AE), and a variational autoencoder (VAE) on MNIST and Fashion-MNIST under five ablation conditions, we observe three regimes. This is an interpretive framework, not a formal derivation. First, topological gaps are already detectable in raw PCA projections (2/5 conditions). Second, nonlinear encoding without regularization reduces this signal (AE: 1/5). Third, KL regularization amplifies it in this setting (VAE: 5/5 means above null). A ghost centroid analysis suggests a mechanism consistent with directed aspiration (r = -0. 28, p < 10^-5, in VAE only). A beta sweep across six KL weights produces a dose-response curve. Five classical geometric metrics fail to predict signal strength. A replication on Fashion-MNIST confirms that the phenomenon generalizes beyond handwritten digits (3/5 conditions above null). A random ablation control indicates that approximately 30–50% of the signal is specific to categorical removal. A reconstruction error baseline confirms that MSE and topological detection capture orthogonal aspects of the gap. Our results suggest that variational regularization renders visible and amplifies certain topological signatures of missing categories.
Régis RIGAUD (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: