Back to AI information
Anthropic discloses Claude distillation attack: DeepSeek, Moonshot, MiniMax named

Anthropic discloses Claude distillation attack: DeepSeek, Moonshot, MiniMax named

AI information Admin 131 views

nthropic announced that it has identified three AI laboratories DeepSeek, Moonshot AI and MiniMax for launching an "industrial-scale" distilled capability extraction operation against Claude: more than 16 million interactions with Claude were generated through about 24,000 fraudulent accounts to train and improve its own models, and were suspected of violating terms of service and regional access restrictions.

The announcement pointed out that "distillation" itself is a common training method, but when competitors use fraudulent accounts, proxy services and batch high-repetitive prompts to centrally extract differentiated capabilities such as reasoning, tool use and coding in a short period of time, it will weaken security guardrails and bring security risks. Among the three operations, Moonshot had about 3.4 million interactions, MiniMax had about 13 million interactions, and DeepSeek was small but included strategies such as inducing output of "step-by-step inference trajectories".

Anthropic said it has strengthened the verification of abuse-prone channels such as detection and behavioral fingerprinting, improved education and research, and shared technical indicators with other institutions. Countermeasures at the product, API and model levels are also developed to reduce the effectiveness of outputs being used for illegal distillation. The announcement also mentioned that such attacks may affect policy discussions on model exports and computing power restrictions, but the relevant external impact still depends on the follow-up actions of all parties.

FAQs

Q: Who is the subject of this incident?

A: The announcement was made by Anthropic and involved its model Claude and the named DeepSeek, Moonshot AI, and MiniMax.

Q: How is Distillation Attack different from normal distillation?

A: Normal distillation is mostly internal compression and cost reduction in the same institution; Distillation attacks refer to the extraction of competitor capabilities through fraudulent accounts and large-scale prompts.

Q: What is the scale data disclosed?

A: Approximately 24,000 fraudulent accounts and more than 16 million interactions with Claude; Among them, Moonshot has been about 3.4 million times and MiniMax has about 13 million times.

Q: What capabilities are mainly extracted from this type of attack?

A: The announcement mentions that the focus includes reasoning ability, tool use, coding and agent tasks, etc.

Q: What risks do ordinary users need to be aware of?

A: If the model proliferation without guardrails may amplify the risk of abuse; Users should also be wary of services such as agent resale and abnormally low-price "proxy access".

Claude Distillation Attack Disclosure Anthropic named DeepSeek 24 000 account draws Claude 1 6 million conversation details Claude is industrially distilled DeepSeek distilled the Claude way Moonshot AI distillation scale MiniMax distillation interaction What is a distillation attack? The difference between model distillation and abuse Claude's reasoning ability was extracted Claude tool use is targeted Claude coding ability is extracted Agency service to resell Claude Fraudulent accounts are generated in batches Anthropic anti-distillation protection Distillation flow identification technology Behavioral fingerprint detection distillation Chain inference induces data Extract risks based on inference trajectories Strengthen the verification of educational accounts API access control upgrades Sharing distillation technical indicators Model layers fight distillation Anti-distillation measures at the product layer Claude Region Access Restrictions Distillation attack and security guardrail Lack of guardrail model risk National security risk discussion AI model capabilities are outflowing Industrial-scale prompt pattern High repetition of the Prompt feature How distillation attacks work Hydra cluster architecture Batch ban and replacement mechanism Cloud platform and third-party channels Competitor ability to replicate AI conversation data is misused Cases of AI clause violations Claude Abuse Prevention Strategy The AI industry responds collaboratively Cloud service provider collaborative protection Policy and regulatory discussions Export control and distillation Advanced chips and extraction scale Model training data compliance Enterprise API security recommendations Identify abnormal call behavior Prevent account theft Claude Security Updates

Recommended Tools

More