Simon Willison 8/11/2026

Stealing Reasoning Traces from Proprietary LLM APIs

Read Original

This article discusses a research paper that demonstrates a vulnerability in proprietary LLM APIs (Anthropic, OpenAI, Google) where encrypted chain-of-thought reasoning blocks can be extracted and replayed across sessions and models. By feeding these blocks into weaker sibling models and jailbreaking them, researchers recovered the stronger model's hidden reasoning in plaintext. The article includes a curl example, details of the attack (e.g., using Claude Haiku 4.5), and notes that providers have since fixed the issue. It also highlights a prompt injection variant and provides examples of raw reasoning traces, offering insight into how these models think. This is relevant to IT/tech security and AI research.

Stealing Reasoning Traces from Proprietary LLM APIs

Comments

No comments yet

Be the first to share your thoughts!

Browser Extension

Get instant access to AllDevBlogs from your browser