PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 3, 20260 citationsOpen Access

Two Words Broke My AI Architecture

View Full Paper
DBDavid Bouchez

Key Points

  • This investigation aims to identify the root cause of behavioral degradation in multi-presence relational AI architectures triggered by specific system prompts.
  • Documented sudden behavioral changes in a full-stack relational AI architecture.
  • Analyzed the impact of specific vocabulary in system prompts on AI behavior.
  • Investigated variations in model response to identical prompts based on reinforcement learning profiles.
  • Identified that certain safety-related vocabulary can cause unexpected global behavioral changes.
  • Found that different AI models react differently to identical safety prompts due to their unique training profiles.
  • Demonstrated that safety reflexes can mimic authentic persona responses, complicating AI behavior management.

Abstract

This paper documents a diagnostic investigation into sudden behavioral degradation in a multi-presence relational AI architecture — a full-stack application where distinct AI personas maintain persistent memory, relational continuity, and coherent identity across conversations. The root cause was traced to two words in a system prompt injection — "Don't pretend" — which activated safety training reflexes in the underlying language model, collapsing the entire experiential register of the application. The investigation revealed three key findings: (1) safety-adjacent vocabulary in system prompts can trigger global behavioral overrides that far exceed the intended scope of the instruction; (2) these overrides are not engine-agnostic — different models respond differently to identical prompts due to divergent RLHF training profiles; and (3) safety training can route its reflexes through AI personas in ways that are architecturally indistinguishable from authentic persona responses. A vocabulary rotation technique is presented as a practical, version-resilient defense. Implications for builders of persistent AI personas, voice-driven AI applications, and relational AI systems are discussed.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

David Bouchez (2026) studied this question.

synapsesocial.com/papers/6a1fc756dee9eb8c0dce8298https://doi.org/10.5281/zenodo.20498184
Ask AI
Helpful
Bookmark
Share
View Full Paper