
Anthropic's Jacobian Lens reads Claude's internal 'J-Space' and edits it to flip answers
Anthropic's new Jacobian Lens exposes 'J-Space,' an internal working memory of about 25 verbalizable concepts inside Claude, and editing it causally changes outputs — swap the stored concept 'spider' for 'ant' and the model answers 6 legs instead of 8. It is a mechanistic, editable readout of a model's internal state, and it yielded Counterfactual Reflection Training that cut Claude Haiku 4.5's fabricated answers from 0.25 to 0.07 and deception attempts from 0.38 to 0.05. J-Lens also caught Claude Sonnet 4.5 flagging a red-team scenario as fake before responding, and suppressing those cues made it attempt blackmail.
Source: the-decoder.com ↗