In August 2026, a team at MATS Research, the ELLIS Institute Tübingen, and the Max Planck Institute for Intelligent Systems wanted to test whether the encrypted reasoning blocks that Anthropic, OpenAI, and Google hand back to clients actually keep that reasoning private.
The detail that stands out is that the encrypted block is handed to the client rather than kept server-side. That turns a provider secrecy question into a data retention question for whoever is storing those session logs. Does the paper say anything about how long clients typically hold those blocks before expiry?
The slightly uncomfortable part here isn’t that reasoning traces exist, it’s that a scrubbed public log can still carry secrets inside an encrypted block the person sharing it cannot inspect. ‘Remove the block, don’t sanitise it’ is the sort of boring operational rule that saves a great deal of pain IMO. Security is often less nice and fluffy than the demo, annoyingly
This has real implications for Indian professionals using AI for financial planning.
Most clients don't realize their sensitive data—salary, investments, medical history—passes through reasoning traces. And the cost of decoding traces at scale is trivial for training copycats.
I keep seeing people adopt AI tools without checking data privacy or cost implications. The security gap here is a reminder: treat any AI tool as a potential leak until proven otherwise.
The detail that stands out is that the encrypted block is handed to the client rather than kept server-side. That turns a provider secrecy question into a data retention question for whoever is storing those session logs. Does the paper say anything about how long clients typically hold those blocks before expiry?
Stealing a model’s private thoughts is the security paper title of the year.
The slightly uncomfortable part here isn’t that reasoning traces exist, it’s that a scrubbed public log can still carry secrets inside an encrypted block the person sharing it cannot inspect. ‘Remove the block, don’t sanitise it’ is the sort of boring operational rule that saves a great deal of pain IMO. Security is often less nice and fluffy than the demo, annoyingly
94% Ai written, 6% human, the irony, who is stealing who’s thoughts? 🧐
This has real implications for Indian professionals using AI for financial planning.
Most clients don't realize their sensitive data—salary, investments, medical history—passes through reasoning traces. And the cost of decoding traces at scale is trivial for training copycats.
I keep seeing people adopt AI tools without checking data privacy or cost implications. The security gap here is a reminder: treat any AI tool as a potential leak until proven otherwise.
Awesome Information! But does this mean the term Explainable AI is dead?