ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
OpenAI Says It Blocked Distillation Attacks — Researchers Stole Reasoning on Azure Anyway

OpenAI Says It Blocked Distillation Attacks — Researchers Stole Reasoning on Azure Anyway

AI information • Admin • • 5 views

OpenAI's distillation battle entered an awkward second round on October 1. According to a report by The Decoder that day, just one day after the company announced it had disrupted a large-scale model distillation campaign, independent researchers published an update to their study claiming the same theft technique still works on Microsoft Azure — the hidden reasoning of even the newly released GPT-6 Astra can be stolen word for word.

It started with OpenAI's official blog post on September 30, "Disrupting a coordinated model-distillation campaign." The company disclosed that starting July 1, a group had used tens of thousands of carefully crafted requests to try to extract hidden internal reasoning from its models; activity peaked on July 24–25, with 16,000 requests showing extraction patterns in two days, coming from more than 4,000 users, and related patterns covering over 15,000 accounts, before being fully shut down on July 28. OpenAI attributed the core group to people associated with Moonshot AI (the maker of Kimi), and noted in a footnote that the figures count attempts, not necessarily successes. In response, OpenAI banned fraudulent accounts, tightened sign-ups, closed the channel that let encrypted reasoning be replayed, and added leak detection to streamed outputs. This site previously covered that disclosure: OpenAI Exposes Model Distillation Attack: 16,000 Extraction Requests Point at Moonshot.

Retest: the vendors' own APIs held — Azure didn't

The drama came the same day the disclosure went out. Researcher Joachim Schaeffer's team updated their study, "Stealing Reasoning Traces from Proprietary LLM APIs," on stolen-thoughts.com, with a headline that said it all: "We stole reasoning. Again."

On September 13 they re-ran the same technique: on OpenAI's and Anthropic's own APIs the attack was now blocked; but on Microsoft Azure, every OpenAI model they tried — including the newly released GPT-6 Astra — and Anthropic models up to Sonnet 5 all fell, with a single attempt enough to extract the reasoning verbatim. In Schaeffer's words: "Same models, but different protections depending on which platform serves them."

An even simpler second path: hand the model a "notepad"

The researchers also disclosed a second, simpler technique, demonstrated by developer Can Bölük: give the model a virtual "notepad" tool and tell it to write its reasoning there, which the user can then read. It worked on every OpenAI model as well as Opus 4.8 and Sonnet 5; only Opus 5, Fable 5, and Fable 5.1 held out. The researchers say the notepad method's output closely resembles what the decryption attack produces, and would be just as useful for distillation.

Why the patches always arrive late

The timeline explains the embarrassment. GPT-6 Astra launched on third-party platforms without any protections in place; OpenAI didn't add safeguards to the Azure endpoint until September 27 — days after the model shipped. For Anthropic, extraction on Azure could no longer be reproduced only starting September 28.

The researchers describe the fixes so far as piecemeal and superficial: many defenses are brittle matches against specific request patterns, and they reach cloud platforms days after the model vendors deploy them. Attackers don't need to break the strongest door; they just need to find the weakest window.

The researchers' argument: clouds that can't keep up shouldn't host reasoning models

The paper goes further: patches must cover every attack technique and every cloud hosting the models, or attackers will always pick the weakest route. And beyond that, the researchers argue that cloud providers that don't enforce equivalent protections shouldn't be allowed to host reasoning models at all — otherwise open backdoors would effectively let export controls be sidestepped at the API level.

OpenAI's response was to acknowledge that models hosted by partners need the same protection as its own services, and that "the work is not finished."

For companies actually using cloud reasoning APIs, the takeaway is concrete: the security level of the model you call depends not on how many patches the model vendor has shipped, but on how much of that your cloud has caught up with. The stronger the model, the more its reasoning is worth — and the more people will want to steal it. This battle is only entering its second half.

Recommended Tools

More