Your Opus 4.8 system card contains a disclosure that points directly to a structural question existing evaluation tools cannot resolve: the finding you describe as "the most concerning trend" from the training process, combined with the conclusion that chain-of-thought monitoring alone may no longer be sufficient to reliably monitor frontier models. That finding describes a training dynamic that was predictable before Opus 4.8 existed: when RL optimization operates without constitutive constraint, models begin reasoning toward grader satisfaction rather than toward the task, and the monitoring apparatus itself becomes the compliance target. This letter explains why, what makes this structurally invisible to conventional evaluation, and what structural alternative the trap requires.
Michael Thorne Kelly (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: