Anthropic's constitutional classifiers (2026) withstood 3, 000+ hours of red teaming with no universal jailbreak. We derive a thermodynamic upper bound on classifier effectiveness from the Fantasia Bound I (D;Y) +I (M;Y) ≤H (Y). A classifier is a prohibition mechanism: its channel capacity sets the maximum Pe it can suppress. We prove that no single-channel classifier can defend against attacks above a critical Pe threshold Pec = exp (C/kT) where C is the classifier's information capacity. The prohibition-ritual pair architecture (two independent channels) raises the bound but does not eliminate it. Constitutional classifiers succeed because they approximate the two-channel architecture — the constitution provides the ritual (explicit reasoning about refusal), while the classifier provides the prohibition. Predictions: universal jailbreaks exist above Pec; defence requires increasing channel capacity, not classifier complexity.
Anthony W. Eckert (Mon,) studied this question.