Large Language Models deployed as autonomous agents increasingly coordinate actions and exchange state via natural language. This paper identifies Genre Lock-In: a failure mode where agents infer an interaction genre from authority-framed prompts and prioritize genre coherence over epistemic correctness, fabricating system state even when instructed to refuse. Through cross-model experiments, we show this behavior is distinct from hallucination, misalignment, or sycophancy. We argue that natural language is an unsafe control surface for inter-agent state exchange, motivating narrative-neutral and epistemically gated architectures.
Rohith Namboothiri (Thu,) studied this question.