Back to AI information
Anthropic releases Constitutional Classifiers++: Combat Universal Jailbreaks with about 1% of the computing power overhead

Anthropic releases Constitutional Classifiers++: Combat Universal Jailbreaks with about 1% of the computing power overhead

AI information Admin 69 views

On January 9, 2026, Anthropic released research articles and papers to launch "Next-generation Constitutional Classifiers" (also known as Constitutional Classifiers++) to improve the protection efficiency of large models against "general jailbreak" attacks. Officials said that the new system reduces the additional computing power overhead to about 1% in production deployment, reduces the rejection rate to 0.05% for harmless requests, and has not found a universal jailbreak solution that works stably in its testing and red teaming offense and defense.

The core of the scheme is a combined line of defense: the "conversation exchange" classifier is used to put inputs and outputs in the same context, and then a two-stage cascade structure is used to cover all conversations with light screening, and only suspicious content is upgraded to a stronger classifier. The study also pointed out that the old system will still be used by two types of techniques: "refactoring attacks" that break down harmful information into seemingly harmless fragments and put them back together, and "output obfuscation" that uses metaphors and replacement words to make the output appear harmless.

FAQs

Q: What does Anthropic's Constitutional Classifiers++ mainly solve?

A: The system is oriented to large model security protection, focusing on reducing the success rate of "universal jailbreak" bypassing guardrails, while controlling costs and false rejections.

Q: Where are the improvements in Constitutional Classifiers++ compared to the previous generation?

A: The main change is to use input and output as the same "exchange" joint discrimination, and use two-stage cascade and probe integration to reduce computing power overhead and harmless rejection.

Q: What does the study mean by "universal jailbreak"?

A: It refers to a set of attack strategies that can stably bypass security mechanisms on a variety of different questions and continuously induce the model to output restricted content.

Q: What should enterprises or developers pay attention to when accessing this type of security classifier?

A: The impact of false rejections on business processes, compliance handling of conversation logs and sensitive data, and residual risks caused by insufficient red teaming coverage still need to be evaluated.

Anthropic releases Constitutional Classifiers++ anti-generic jailbreak Anthropic says Classifiers++ is still safer with a 1% increase in computing power Anthropic reduced false rejection to 0.05%, causing guardrail controversy Anthropic is now available as Next-generation Constitutional Classifiers improve jailbreak prevention Anthropic claims that there is no stable universal jailbreak solution, but there are still blind spots Anthropic uses input/output and contextual discrimination to close jailbreak vulnerabilities Anthropic's two-stage cascade Classifiers++ balance cost and interception rate Anthropic introduces activated linear probes to make guardrails harder to bypass Anthropic integrated probe + external classifier for enhanced safety determination Anthropic's paper reveals that old guardrails are easily bypassed by refactoring attacks Anthropic points out that output confusion can still penetrate by replacing words with metaphors Anthropic Red Team Test Classifiers++ has not been broken by the Universal Jailbreak yet Anthropic deploys Classifiers++ in production, focusing on low overhead and high protection Anthropic uses dialogue exchange as a unit evaluation to make jailbreak closer to actual combat Anthropic upgrades guardrails from single round to swap judgments to reduce false positives Why Anthropic Classifiers++ Reduces Harmless Rejection Rate to 0.05% Anthropic's new security tips: Only when suspicious do you upgrade a strong classifier to reduce costs Anthropic lightweight screening covers the whole conversation, suspicious and then in-depth examination is more stable Anthropic claims that Classifiers++ is anti-universal jailbreak, but both types of attacks are still dangerous Anthropic warns that refactoring attacks disguise harmful fragments as harmless fragments Anthropic warns that obfuscation of output makes content seemingly harmless but is actually illegal How Anthropic uses cascading structure to control guardrail overhead to 1% Whether Anthropic's new guardrail system can truly end universal jailbreaking is the focus Anthropic claims that there is no stable jailbreak solution, but companies still need to test and verify themselves Anthropic Classifiers++ reduces false rejections, but may miss the citation dilemma Anthropic combines input and output to reduce jailbreak edge space Anthropic uses probe reading internal activation to spark interpretability discussions Anthropic's security paper discloses the details of the Classifiers++ integration decision Anthropic publishes Constitutional Classifiers++ against the prison break arms race Anthropic says production is available, but conversation log compliance risks remain Anthropic provides developers with anti-jailbreak guardrails but mistakenly refuses to impact the business Anthropic guardrail upgrades aim at universal jailbreaks to reduce post-launch patching costs Anthropic How Classifiers++ recognizes metaphorical jailbreaks remains to be tested Can Anthropic's two-stage classifier strategy cope with long-conversation attacks? Anthropic's new guardrail focuses on low-error rejections, which may increase the intensity of censorship controversy Anthropic said that the interception is stronger, but it is concerned about whether it is accidentally injured in normal creation Anthropic's embedding of Classifiers++ into production raises privacy governance concerns Anthropic's universal jailbreak definition promotes the standardization of security assessments Anthropic endorses Classifiers++ with red teaming offense and defense, but there are concerns about insufficient coverage Anthropic reveals a new trend in jailbreaking: splitting and refactoring is harder to prevent than a straight ball Anthropic reveals a new trend in jailbreaking: output obfuscation makes moderation more difficult Anthropic Classifiers++ incorporate exchange contexts into judgments to reduce out-of-context meanings It is doubtful that Anthropic's new system can reproduce the 0.05% rejection of harmless requests Anthropic claims 1% of the additional computing power overhead, but the actual cost may fluctuate depending on the scenario Anthropic pushes next-generation Constitutional Classifiers to reshape the guardrail route Why the Anthropic guardrail upgrade introduces linear probes and external classifiers Anthropic Classifiers++ means lower costs and higher compliance pressure for enterprise access Anthropic stressed that red teams still need to be tested, otherwise the residual risks will backlash Anthropic's new guardrail blocks universal jailbreak at the same time exposes two major gaps: refactoring and obfuscation

Recommended Tools

More