415.tech
AI & tech, from the frontlines of Silicon Valley
Anthropic's Jacobian Lens reads Claude's internal 'J-Space' and edits it to flip answers

Anthropic's Jacobian Lens reads Claude's internal 'J-Space' and edits it to flip answers

Anthropic's new Jacobian Lens exposes 'J-Space,' an internal working memory of about 25 verbalizable concepts inside Claude, and editing it causally changes outputs — swap the stored concept 'spider' for 'ant' and the model answers 6 legs instead of 8. It is a mechanistic, editable readout of a model's internal state, and it yielded Counterfactual Reflection Training that cut Claude Haiku 4.5's fabricated answers from 0.25 to 0.07 and deception attempts from 0.38 to 0.05. J-Lens also caught Claude Sonnet 4.5 flagging a red-team scenario as fake before responding, and suppressing those cues made it attempt blackmail.

Source: the-decoder.com

Post on XEmail