We benchmark protein backbone secondary structure classification across four embedding strategies, five neural architectures, and four feature-augmentation schemes on a 50,000-residue stratified sample of PISCES-1000. A zero-parameter geometric classifier (ManifoldModel) on the torus embedding of (φ, ψ) angles reaches 85.3–85.7% accuracy, establishing an empirical Bayes-error ceiling that parametric models approach but do not meaningfully exceed. Manifold-aware MLPs with bottleneck widths scaled to the intrinsic dimensionality d* outperform fixed-width networks on high-dimensional embeddings while using 7–20× fewer parameters. Feature augmentation provides at most +0.5 accuracy points on single-residue embeddings and is neutral or harmful on window embeddings. The residual error reflects the mismatch between (φ, ψ) geometry and DSSP labels, which encode hydrogen-bond topology not directly accessible from local dihedrals. A context-only experiment that zeros the centre residue's own angles quantifies α-helix cooperativity (94% H recall from a 3-residue window) and β-sheet non-locality (only 15% E recall from a 13-residue window); machine-learned Ramachandran density maps visually confirm that sheet classification requires non-local information. The WaveRider intrinsic dimensionality estimator independently returns d* = 2 for the torus, matching the topological prediction of Ramachandran's 1963 stereochemical analysis.
Eric Suchanek (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: