PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 19, 20260 citationsOpen Access

A Mechanistic Framework for Removing Unwarranted Demographic Influence in Clinical LLMs

View Full Paper
IRInna Rytsareva

Key Points

  • The study aims to address how demographic factors distort clinical reasoning in language models while maintaining utility.
  • Implemented Attention Surgery process in four stages: behavioral audit, causal localization, extraction, and removal.
  • Utilized techniques like activation patching and weight orthogonalization.
  • Investigated the Selective Surgery Problem related to demographic information relevance.
  • Demonstrated that not all demographic data negatively influences clinical reasoning.
  • Established the need for a dual-validation requirement to balance fairness and safety in debiasing.
  • Showed that removing demographic tokens could harm clinical utility when they are relevant.

Abstract

We propose Attention Surgery, a mechanistic interpretability framework for identifying and removing the specific neural pathways through which demographic tokens (like sex labels) distort clinical reasoning in LLMs, building directly on the prior YentlBench/Attention Leak work that demonstrated the problem behaviorally. The pipeline has four stages: contrastive behavioral audit → causal localization in the residual stream → extraction of demographic-influence directions → removal via orthogonal projection. The technique draws on activation patching and weight orthogonalization, analogous to "abliteration" methods used to remove refusal behavior from open models. Our proposal's central conceptual contribution is naming and framing the Selective Surgery Problem: demographic information isn't uniformly harmful in clinical contexts. Sex is irrelevant to ESI triage severity, but directly relevant to differential diagnosis of, for instance, abdominal emergencies. Any debiasing intervention that blindly removes sex-encoding damages clinical utility and the framework must thread a needle between a fairness test (sex-label perturbation no longer shifts sex-invariant outputs) and a safety test (sex-dependent reasoning survives intact where it's physiologically warranted). We argue this dual-validation requirement is genuinely novel and not addressed by existing debiasing paradigms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Inna Rytsareva (2026) studied this question.

synapsesocial.com/papers/69e47282010ef96374d8e7e8https://doi.org/10.5281/zenodo.19633294
Ask AI
Helpful
Bookmark
Share
View Full Paper