PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 7, 20260 citationsOpen Access

Architectural Observability Collapse in Transformers

View Full Paper
TCThomas Carmichael

Key Points

  • To explore how activation monitoring relates to decision-quality signals in autoregressive transformers.
  • Defined observability in the context of transformer architectures.
  • Analyzed various model configurations and their performance on observability.
  • Evaluated checkpoint dynamics and their effect on decision-quality signal preservation.
  • Observability correlates strongly with model architecture, as shown by varying performance across configurations.
  • Mistral 7B maintains observability, unlike Llama 3.1 at similar architectures, indicating architecture's role in monitoring.
  • Activation monitoring captures decision-quality signals efficiently, with a significant average probe signal absorption.

Abstract

Activation monitoring can catch confident errors in autoregressive transformers only if training preserved an internal decision-quality signal that output confidence does not expose. Monitorability is an architectural property before it is a monitor-design problem. We define observability: the linear readability of per-token decision quality from frozen mid-layer activations after controlling for max-softmax confidence and activation norm. Confidence controls absorb on average 60.3% of raw probe signal across 14 models in 6 families. Observability is not a generic property of transformers. In Pythia’s controlled suite, all three tested runs at the 24-layer, 16-head configuration collapse to ρpartial ≈ 0.10 across a 3.5× parameter gap and two Pile variants, while six other configurations occupy a separated healthy band from 0.21 to 0.38. The output-controlled residual rOC collapses at the same points. Neither nonlinear probes nor layer sweeps recover healthy-range signal. Checkpoint dynamics localize the cause: both matched-width configurations form the signal at the earliest measured checkpoint, and training erases it in the 1.4B even as it reaches lower final loss than the 1B. Architecture does not prevent the signal from appearing. It determines whether training preserves or erases it. Across independent recipes the collapse map changes but the phenomenon persists. Qwen 2.5 and Llama differ by 2.9× at matched 3B scale, with probe-seed distributions that do not overlap. Mistral 7B v0.3 preserves observability where Llama 3.1 8B collapses despite identical 32-layer, 32-head, 4096-hidden shape. Within Qwen 2.5, observability persists from 0.5B through 32B. A WikiText-trained observer transfers to downstream QA without task-specific training: at 20% flag rate, exclusive catch reaches 10.9–13.4% in seven of nine model-task cells, near the 12–15% language-modeling ceiling. Architecture selection is a monitoring decision. Code and results: https://github.com/tmcarmichael/nn-observability Croissant 1.1 metadata: https://github.com/tmcarmichael/nn-observability/blob/v4.0.0/croissant.json arXiv: https://arxiv.org/abs/2604.24801

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Thomas Carmichael (2026) studied this question.

synapsesocial.com/papers/69fbe2b3164b5133a91a212fhttps://doi.org/10.5281/zenodo.20034469
Ask AI
Helpful
Bookmark
Share
View Full Paper