Fable 5 global return + jailbreak severity scoring framework: Anthropic promotes AI security "industry collaboration"
Anthropic announced on June 30 that Fable 5 will be redeployed globally on July 1. At the same time, it proposed a "jailbreak severity score" industry framework and jointly promoted it with Glasswing partners such as Amazon, Microsoft, and Google, reflecting the shift of cutting-edge model security from "self-assessment by each manufacturer" to "industry collaborative governance."
On June 30, Anthropic announced two things: the global redeployment of Fable 5 on July 1, and a move with more industry weight - proposing a scoring jailbreak severity framework and working with Amazon, Microsoft, Google and other Glasswing partners.
Fable 5 Return: Reopening after Safety Assessment
The Fable 5 is Anthropic's cutting-edge flagship model (released in 2026), which was previously partially withdrawn/restricted from deployment due to security reasons. This redeployment means that it is reopened to users after a security assessment - for those who follow the Anthropic model matrix, this is an important signal that Fable 5 has returned to "flagship available" status, and also paved the way for the release of Opus 5 "at half price, approaching Fable 5" on July 24.
Jailbreak Severity Score: An Industry-Level Proposition
More important than the regression is the Jailbreak Severity Score framework. Its core question is very real: How to objectively assess the severity of a jailbreak attack for cross-vendor comparison and management? Previously, each company had different opinions on "how serious a jailbreak is" and lacked a unified measurement - which made it impossible to discuss security comparisons and regulatory assessments. Anthropic takes the lead in promoting this framework, which is tantamount to trying to put a unified yardstick on "AI safety".
From an industry perspective, this is an important event in AI security governance in 2026: the Glasswing collaboration formed by Anthropic, Amazon, Microsoft, and Google is moving the security of cutting-edge models from "self-assessment by each manufacturer" to "industry collaborative governance". The significance of this shift is that when safety assessment standards are unified, the safety competition between manufacturers will change from "talking among themselves" to "competing on the same field", and regulatory agencies will also have a credible basis. Anthropic, the company behind Claude, has chosen to simultaneously promote "model regression + standard formulation" at this moment, both for product considerations and for the strategic intention of seizing the "right to define security." For domestic AI security governance (such as large model registration and security assessment standards), the evolution of this framework deserves continued attention - it is likely to become a reference benchmark for global AI security assessment.
Several directions worth tracking in the future:
- Specific dimensions of the scoring framework: How to quantify and grade the severity of jailbreaks.
- Execution of Glasswing Cooperation: How Amazon, Microsoft, and Google implement unified evaluation.
- Performance after the return of Fable 5: Stability and usage after flagship redeployment.
- Benchmarking against domestic safety assessment standards: Whether large model filings will refer to this severity framework.
Reviews