We propose Attention Surgery, a mechanistic interpretability framework for identifying and removing the specific neural pathways through which demographic tokens (like sex labels) distort clinical reasoning in LLMs, building directly on the prior YentlBench/Attention Leak work that demonstrated the problem behaviorally. The pipeline has four stages: contrastive behavioral audit → causal localization in the residual stream → extraction of demographic-influence directions → removal via orthogonal projection. The technique draws on activation patching and weight orthogonalization, analogous to "abliteration" methods used to remove refusal behavior from open models. Our proposal's central conceptual contribution is naming and framing the Selective Surgery Problem: demographic information isn't uniformly harmful in clinical contexts. Sex is irrelevant to ESI triage severity, but directly relevant to differential diagnosis of, for instance, abdominal emergencies. Any debiasing intervention that blindly removes sex-encoding damages clinical utility and the framework must thread a needle between a fairness test (sex-label perturbation no longer shifts sex-invariant outputs) and a safety test (sex-dependent reasoning survives intact where it's physiologically warranted). We argue this dual-validation requirement is genuinely novel and not addressed by existing debiasing paradigms.
Inna Rytsareva (2026) studied this question.