
OpenAI Blocked 16,000-Request Campaign to Steal Model Reasoning
OpenAI says a July campaign tied to Moonshot AI fired 16,000 reasoning-extraction requests from 4,000+ accounts; encryption held, but fixes now close replay paths.
- By
- Tom Whitfield
- Filed
- Channel
- Consumer Tech
- Read
- 3 min read
OpenAI says it disrupted a coordinated campaign in July that fired 16,000 requests at its API attempting to extract models' encrypted chain-of-thought reasoning — an effort the company links to people associated with China-based Moonshot AI.
The disclosure came in a company blog post. OpenAI described the activity as "consistent with adversarial distillation" — the "systematic and unauthorized use of one model's outputs or reasoning" to train or improve another model. The company stopped short of attributing every participant to a single actor.
The timeline is precise. Activity began July 1 at "initially at a low volume," followed by "high-volume spikes" on July 24 and 25. Across those spikes, more than 4,000 accounts issued 16,000 requests following a common extraction pattern. OpenAI says the campaign was "fully disrupted by July 28." The company cautions the requests attempted extraction but were "not necessarily successful" — and the post does not quantify a success rate, name which models were targeted, or say how many of the 4,000 users were Moonshot-linked.
How the attack worked
OpenAI's "protected reasoning" is the model's internal record of working through a task. The system encrypts that chain of thought and hands it to the client as an encrypted block; the client returns the block with each subsequent request, so the provider never has to store it.
One technique the operators tried took encrypted reasoning from one conversation and asked a model in a separate conversation to decrypt it. That approach mirrors a vulnerability documented in an academic paper OpenAI explicitly linked: "Stealing Reasoning Traces from Proprietary LLM APIs," dated August 10. Its authors tested OpenAI, Anthropic, and Google models, feeding a frontier model's encrypted reasoning to a corresponding weaker model that then wrote out the contents in plain text.
The researchers ran their tests in early July. After the providers acknowledged their responsible-disclosure report, the researchers were "unable to launch the same attacks." OpenAI confirmed the attack paths the researchers found — along with "related cross-model and conversation-compaction vulnerabilities" — were real.
What held and what was fixed
The encryption itself was not broken, according to OpenAI. No database was compromised, and the operators gained no direct access to stored user conversations.
Still, OpenAI shipped fixes. One patch closed a pathway that let someone holding another user's encrypted reasoning replay it and recover its contents. The company added checks to detect and hold streamed output that might expose reasoning, strengthened protections for hidden reasoning across users, workspaces, organizations, and model families, and worked with third-party providers to disrupt accounts routing activity through their services.
Moonshot's pattern
OpenAI is not the first frontier lab to flag Moonshot. Earlier this month, Anthropic's report "Detecting and countering misuse of AI: September 2026" documented a single ten-day window in which Moonshot relayed nearly 300,000 customer requests to Claude through a proxy network of 5,380 fraudulent accounts. Anthropic said Moonshot saved Claude's reasoning signatures and, in new sessions, induced Claude to convert them back into full reasoning traces — attacks Anthropic calls "cross-session replay attacks."
Moonshot has denied that its Kimi K3 model was created through distillation. In July, OpenAI President Greg Brockman said it was "too early" to determine whether Moonshot had distilled OpenAI's models.
What comes next
OpenAI's next step is ensuring partner-hosted deployments carry the same protections as its first-party tools, and the company says additional checks are needed to block tool-output attacks. It shared its findings with the Frontier Model Forum and government information-sharing channels, warning that "systems that support portable or replayable reasoning artifacts may face related risks."
The company expects distillation attempts to grow more sophisticated as frontier models improve and attackers hunt for cheaper ways to mimic them — and its defensive work, it says, is ongoing.
Source: Tom's Hardware
More from Tom Whitfield
Show full bio
Staff writer covering consumer brands and retail at Chip Dispatch.
140 articles
Related articles
openai-cancels-gpt-6-1-release-over-safety-regressions-aa6214a3
OpenAI Cancels GPT-6.1 Release Over Safety Regressions
open-source-ai-platforms-race-to-become-china-s-hugging-face-ce6b57a2
Open-Source AI Platforms Race to Become China's Hugging Face
openai-lands-record-110b-round-in-week-s-top-funding-deals-474a9a9b
OpenAI Lands Record $110B Round in Week's Top Funding Deals
thieves-hunting-nvidia-ai-chips-make-off-with-18-tons-of-sand-8058597b
Thieves Hunting Nvidia AI Chips Make Off With 18 Tons of Sand



