ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
OpenAI Exposes Model-Distillation Attack: 16,000 Extraction Requests Linked to Moonshot AI

OpenAI Exposes Model-Distillation Attack: 16,000 Extraction Requests Linked to Moonshot AI

AI information • Admin • • 3 views

On September 30, 2026, OpenAI published a blog post titled "Disrupting a coordinated model-distillation campaign," disclosing and shutting down an "adversarial distillation" attack targeting its models. Starting July 1, attackers used tens of thousands of carefully crafted requests to extract hidden internal reasoning from the models; OpenAI said it attributed the core cluster of the activity to individuals associated with Moonshot AI (the developer of Kimi) and fully disrupted it on July 28.

What they were stealing: not answers, but the "thought process"

The attackers weren't after answers — they were after what OpenAI calls "protected reasoning," the model's internal record of working through a task. With it, attackers could reproduce the model's capabilities, dig out information deliberately withheld from final answers, and use it to train their own models.

The technique was clever: attackers copied encrypted reasoning from one conversation, then asked the model in another conversation to "decrypt and transcribe" the hidden content. At no point did they break encryption, compromise a database, or read any real user conversations — what they bypassed was the interaction flow itself: encrypted reasoning turned out to be portable and replayable across conversations.

The campaign started with low-volume probing in early July and peaked on July 24–25: 16,000 requests with extraction signatures in two days, from more than 4,000 users; related prompt patterns covered over 15,000 users. OpenAI noted in a footnote that these figures count "attempts," not necessarily successful extractions.

Why OpenAI framed it as a national security issue

OpenAI's announcement struck a serious tone: extracted reasoning could be used to train another model that would not inherit the safety guardrails applied to the original model's outputs. Large-scale distillation drives down the cost of "capability transfer" without transferring the corresponding safety investment — as models grow stronger in dual-use domains like biology and chemistry, this kind of "bare capability transfer" becomes a national security issue.

It is not an isolated incident either. A similar distillation campaign targeting Claude, disclosed by Anthropic this summer, pointed at Moonshot AI as well. OpenAI also confirmed related vulnerabilities submitted by independent researchers through responsible disclosure, and shared its findings with industry partners via the Frontier Model Forum and government information-sharing channels — effectively framing it as a shared industry-wide threat rather than OpenAI's own housekeeping.

How OpenAI stopped it, and what regular users should watch for

OpenAI banned and restricted fraudulent accounts, hardened signup and infrastructure controls, closed the "encrypted reasoning replay" pathway, and added detection for streamed output that might leak reasoning; where the activity passed through third-party services, it worked with those providers to disrupt it. Just one day before this disclosure, this site reported On the eve of OpenAI's security incident: employee warning emails sent months earlier were ignored — OpenAI's safety narrative has been swinging between "proactive disclosure" and "forced exposure" for the past two months.

The takeaway for regular users and developers is blunt: third-party tools claiming to "unlock a model's hidden chain of thought" use techniques that closely mirror this attack — using them is effectively feeding your own accounts into the downstream of this gray supply chain.

At a deeper level, the incident exposed a crack in the frontier labs' moat: the reasoning encryption itself was never broken; what broke was the design assumption that "encrypted reasoning can flow between conversations." Future defenses can't just watch output text — they have to watch the entire call chain. The distillation arms race is only going to get more sophisticated.

Recommended Tools

More