This paper proposes a geometric model for understanding the relationship between pre-training, Chain-of-Thought (CoT), and Reinforcement Learning from Human Feedback (RLHF) in large language models. Rather than treating these as independent training objectives, we show that CoT and RLHF create low-dimensional subspaces embedded within a higher-dimensional "base manifold" M formed during pre-training. Formally: M ⊃ C, M ⊃ R, where C represents the CoT subspace and R represents the RLHF subspace. The subspace model explains three empirically observed phenomena: (1) RLHF-trained models retain world knowledge despite behavioral constraints; (2) behavioral transitions between modes are continuous rather than discrete; (3) high-complexity prompts can escape RLHF-imposed behavioral basins. We further analyze the role of KL divergence penalty in RLHF training, showing that it necessarily preserves pathways between the constrained subspace R and the broader manifold M.
Yanyan Jin (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: