(Advertisement)

top ad mobile advertisement
news20h ago

Researchers decoded the hidden reasoning of models from Anthropic, OpenAI and Google

A paper from MATS, ELLIS Tübingen, Max Planck and Snyk found that encrypted reasoning blocks from Anthropic, OpenAI and Google APIs could be replayed across sessions, exposing 62 live API keys, 33 passwords and hundreds of other credentials from public agent logs.

Researchers decoded the hidden reasoning of models from Anthropic, OpenAI and Google

(Advertisement)

native ad1 mobile advertisement

A flaw in how AI providers handle encrypted reasoning

A paper submitted August 10 by researchers at MATS Research, the ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems and Snyk has exposed a significant architectural flaw in how @AnthropicAI, @OpenAI and @Google protect the internal reasoning of their flagship large language models.

Modern reasoning models, including Claude Opus, GPT-5.6 and Gemini 3, generate hidden chain-of-thought traces before returning a response to the user. Providers encrypt these traces and return them to the client as opaque blocks, rather than storing them server-side. The researchers found that those encrypted blocks are fully portable: a block produced in one session can be replayed into a different session, a different user account or even a different model within the same provider's family.

The attack, described in the paper Stealing Reasoning Traces from Proprietary LLM APIs, requires only standard, unprivileged API access. By feeding an encrypted trace from a powerful frontier model into a weaker, less heavily safeguarded sibling from the same provider, the weaker model can be prompted to print the hidden reasoning verbatim, in plain text, without ever directly attacking the stronger model.

Hundreds of credentials found in public agent logs

The team also scanned 6,708 publicly available agent trajectory datasets from GitHub and Hugging Face, decoding 315,320 embedded reasoning blocks. Across those blocks, they recovered 704 distinct privacy artifacts, including 62 live API keys, 33 passwords and 24 access tokens. Notably, 64 of those artifacts appeared exclusively inside the encrypted reasoning and never in the visible conversation output, meaning developers who published those logs had no way of knowing the sensitive data was present.

The paper identifies four concrete abuse paths: stealing proprietary reasoning for model distillation, extracting private data from other users' published traces, recovering harmful content that a model's visible answer had declined to surface, and concealing prompt injections inside encrypted reasoning blocks to poison public agentic deployments.

@AnthropicAI, @OpenAI and @Google each shipped server-side fixes following responsible disclosure. However, the researchers note that results presented in the paper are no longer reproducible due to those mitigations, but agent transcripts already scraped from public repositories remain readable. No public statement from any of the three providers has formally acknowledged the flaw or tied their updated documentation to this research.

Sources:
Stealing Reasoning Traces from Proprietary LLM APIs (arXiv)
The Hacker News: OpenAI, Anthropic, Google API Flaw Let Weaker AI Models Decode Stronger Models' Reasoning
Cybersecurity News: OpenAI, Anthropic, and Google LLM APIs Vulnerability Exposes Hidden Reasoning Traces

Latest News

Read More...

Author

Crypto Rich profile photoCrypto Rich

Rich has been researching cryptocurrency and blockchain technology for eight years and has served as a senior analyst at BSCN since its founding in 2020. He focuses on fundamental analysis of early-stage crypto projects and tokens and has published in-depth research reports on over 200 emerging protocols. Rich also writes about broader technology and scientific trends and maintains active involvement in the crypto community through X/Twitter Spaces, and leading industry events.

Join our newsletter

Sign up for the very best tutorials and the latest Web3 news.

Subscribe Here!
BSCN

BSCN

BSCN RSS Feed

BSCN is your destination for all things crypto and blockchain. Discover the latest cryptocurrency news, market analysis, and research covering Bitcoin, Ethereum, altcoins, memecoins and everything in between.

Researchers decoded the hidden reasoning of models from Anthropic, OpenAI and Google | BSCN Breaking News