Scalar loss in machine learning obscures what a neural network learns and when. While the specific circuits in a network are opaque, we can observe how a network learns across different resolutions of data. Our hope is that our resolution-based approach makes theories of learning easier to formulate and test. We illustrate this by proposing a phenomenological model, key search, with a concrete methodology that decomposes the learning process across three axes: parameter time, resolution band, and band loss. Parameter time is a reparameterization of training time by distance traveled along the optimizer's path in parameter space. A resolution band is the content present at a higher resolution level but absent at a lower one, obtained as an orthogonal component of a multiresolution analysis. Band loss is the loss restricted to that band, rather than the full resolution. Key search predicts three results, which we test on 12 CelebA 64 64 autoencoding runs (6 CNN, 6 ViT): 1. the gap between a band's current loss and its floor is well fit by a Weibull curve in parameter time over the active window (R² > 0. 99 for bands 2--64) ; 2. band activation is ordered across scales, with finer bands activating later (Spearman = 0. 79), decaying over wider parameter time windows (= 0. 96), reaching higher floors (= 0. 96), and starting with smaller removable loss (= -0. 89) ; 3. the ordinary scalar loss, obtained by summing band residuals, is also well summarized by a single Weibull curve in parameter time (median R² = 0. 996, median = -28. 83 versus exponential).
Jurij Jukić (Thu,) studied this question.