PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 17, 20260 citationsOpen Access

The Subspace Structure of AI Activation Patterns: CoT and RLHF as Embedded Manifolds

View Full Paper
YJYanyan Jin

Key Points

  • This research aims to develop a geometric model to elucidate the relationship between Chain-of-Thought and Reinforcement Learning from Human Feedback in language models.
  • Propose a geometric model for the interaction between pre-training, CoT, and RLHF.
  • Identify CoT and RLHF as low-dimensional subspaces within a high-dimensional base manifold.
  • Analyze the empirical phenomena associated with RLHF-trained models.
  • Models trained with RLHF retain world knowledge despite behavioral constraints.
  • Behavioral transitions between modes are continuous rather than discrete.
  • Complex prompts can navigate beyond RLHF-imposed behavioral constraints.

Abstract

This paper proposes a geometric model for understanding the relationship between pre-training, Chain-of-Thought (CoT), and Reinforcement Learning from Human Feedback (RLHF) in large language models. Rather than treating these as independent training objectives, we show that CoT and RLHF create low-dimensional subspaces embedded within a higher-dimensional "base manifold" M formed during pre-training. Formally: M ⊃ C, M ⊃ R, where C represents the CoT subspace and R represents the RLHF subspace. The subspace model explains three empirically observed phenomena: (1) RLHF-trained models retain world knowledge despite behavioral constraints; (2) behavioral transitions between modes are continuous rather than discrete; (3) high-complexity prompts can escape RLHF-imposed behavioral basins. We further analyze the role of KL divergence penalty in RLHF training, showing that it necessarily preserves pathways between the constrained subspace R and the broader manifold M.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yanyan Jin (2026) studied this question.

synapsesocial.com/papers/696b2655d2a12237a9349971https://doi.org/10.5281/zenodo.18260350
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Reinforcement Learning for Latent-Space Thinking in LLMs2025
  2. 2On the Representational Capacity of Neural Language Models with Chain-of-Thought Reasoning2024
  3. 3Reward Generalization in RLHF: A Topological Perspective2024
  4. 4Lexical hints of accuracy in LLM reasoning chains2026
  5. 5Manifold-Constrained Hyper-Connections: Rethinking the Architectural Foundation of Large-Scale Language Models2026