IA · 10 June 2026 · 3 min read

Anthropic Releases Claude Fable 5: Mythos-Class Power and the Guardrail Controversy

In brief: Anthropic has released Claude Fable 5, the first public model in its advanced 'Mythos' class, which was previously restricted due to high-risk cyber and biological capabilities. To enable this release, Anthropic built highly restrictive safety filters, which have drawn sharp criticism from security researchers over frequent false positives.

by Team Mocchi's

Anthropic Releases Claude Fable 5: Mythos-Class Power and the Guardrail Controversy

The Genesis of the Mythos Class

On June 9, 2026, Anthropic announced the launch of Claude Fable 5, the first publicly available model from its highly anticipated "Mythos" family. According to reporting by The Verge, this class of models was previously withheld from broad release because its capabilities in cybersecurity and advanced biology were deemed too powerful to distribute without significant restriction.

Alongside Fable 5, the company introduced Claude Mythos 5, an unrestricted version of the same underlying model. Access to Mythos 5 remains strictly controlled, limited to a selected group of cybersecurity and government entities participating in Anthropic’s private Project Glasswing initiative. For the general public and enterprise clients, Fable 5 represents Anthropic's compromise: a highly capable system for software engineering and complex reasoning, wrapped in strict guardrails to prevent exploitation.

The Safeguard Mechanism and Fallback to Opus

To facilitate the public release of a Mythos-class model, Anthropic engineered real-time monitoring safeguards that detect prompts containing sensitive topics in biology and cybersecurity. If a user request triggers these parameters, the system blocks the Fable 5 response and seamlessly redirects the query to Claude Opus 4.8—a model released in May 2026 praised for its honesty and reliability.

Anthropic stated that during internal testing, 95 percent of user sessions were processed entirely by Fable 5 without triggering the fallback mechanism. However, access to this class of reasoning comes with premium pricing. Claude Fable 5 is priced at $10 per million input tokens and $50 per million output tokens—double the rate of Claude Opus 4.8, though significantly cheaper than the private Mythos Preview.

Cybersecurity Researchers Raise Concerns

The implementation of these strict security measures has drawn immediate criticism from security professionals. As reported by TechCrunch, researchers have flagged a high volume of false positives, with the system blocking completely benign development tasks.

Valentina Palmiotti, a prominent security researcher at IBM X-Force, publicly noted that Fable 5 rejects requests with even tangential links to cybersecurity. This includes innocuous tasks such as reading technical blog posts or reviewing application code for standard bugs. When a prompt triggers the guardrails, the chat pauses, displaying a message stating that safety measures flagged the query for cybersecurity or biology risks.

Matt Suiche, a cybersecurity veteran and technical staff member at AI startup Tolmo, confirmed these observations to TechCrunch. He explained that "if you ask it to write secure code, it assumes it is cybersecurity-related work instead of software engineering best practices, and you get downgraded." The system's guardrails appear to operate largely on simple keyword matching, triggering blocks on any terminology within the lexical field of security.

Preventive Safety vs. Practical Usability

The issues surrounding Claude Fable 5 highlight a fundamental tension in frontier AI development: balancing the mitigation of high-stakes risks—such as the creation of custom malware or biological agents—with the practical utility of the tool for legitimate professional work.

Despite the frustration among technical users, some industry experts advise patience. Suiche noted that given the early phase of the Mythos deployment, Anthropic's cautious approach is logical: "It's better to catch more people than not enough when you do such a release and to relax the guardrails over time." The challenge for Anthropic moving forward will be to transition its defensive mechanisms from rigid, keyword-based blocklists to nuanced, context-aware safety systems that support developers without compromising safety.

Further reading

All articles on the Mocchi's blog