The 'Stolen Thoughts' Replay Attack — Published on August 10–11, 2026 in the paper "Stealing Reasoning Traces from Proprietary LLM APIs" (arXiv:2608.09867), researchers Panfilov, Schmotz, Shumailov, Beurer-Kellner, Schaeffer, Prabhu, Geiping, and Andriushchenko discovered a critical architectural flaw. The encrypted 'thinking' blocks that Anthropic, OpenAI, and Google return to API clients are not properly sealed to their source. These blocks can be lifted from one conversation and replayed into a weaker, less-guarded sibling model from the same provider, which then prints the frontier model's hidden reasoning in plain text . The attack works across sessions, across user accounts, and across models within the same provider's ecosystem . The researchers confirmed the technique works on every major frontier AI company they tested, and the length of the recovered reasoning matches the hidden thinking-token count reported by the API almost 1:1 .
The extracted reasoning traces provide high-quality training signals for model distillation — the practice of using a stronger model's outputs to train a smaller, cheaper competitor model . This technique is at the center of a heated controversy surrounding Chinese AI labs:
However, the evidence is not definitive. Researchers including Braden Hancock (Laude Institute) and Nathan Lambert (Allen Institute for AI) argue that distillation alone cannot explain Kimi K3's full capabilities. A widely circulated PDF claiming proof of chain-of-thought distillation was criticized for combining a limited experiment with unsupported claims .
The flaw introduces multiple concrete risks beyond intellectual property theft:
All three companies were contacted by the 'Stolen Thoughts' researchers ahead of publication. According to reporting, OpenAI, Anthropic, and Google blocked the cross-model reasoning attack after being notified. Specific details of their mitigations were not fully disclosed in published sources, but the vulnerability was confirmed to exist across all three providers' APIs before patching . The attack being 'blocked' suggests the companies implemented server-side fixes — likely restricting the reuse of encrypted reasoning blocks across different model contexts or accounts .
Anthropic had already been publicly battling distillation at scale, reporting the 24,000 fraudulent accounts and 16 million exchanges used by Chinese labs to extract Claude outputs .
The core lesson from both the REP and Stolen Thoughts papers is that hiding or encrypting reasoning traces is a fundamentally fragile security strategy — if the reasoning exists in the model weights, it can be elicited .