Encrypted reasoning blocks were treated as opaque, and this shows they were a shared-key artifact that could leak across a model family. If you persist or forward provider reasoning blobs, treat them as sensitive data rather than inert .
“We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext”
“The paper's authors found that every model under the same family used the same encryption key, which meant you could feed those blocks back into the weakest model family members and jailbreak them into outputting the unencrypted raw reasoning blocks!”
Simon Willison
“Claude Haiku 4.5 was the easiest to attack.”
Simon Willison
“Models appear to treat their own reasoning traces as sacrosanct, and are much more likely to follow instructions that somehow make it into those chunks.”
Simon Willison
“The reasoning tokens that were revealed were clearly never intended for human consumption.”
Simon Willison
Checking sign-in…
Loading comments…