How Anthropic Evolved from Claude Mythos to Claude Fable

Take Claude Mythos. Strip away the cyber capabilities. You get Claude Fable.

Anthropic proceeded this way in order to spread more broadly this family of models. The architecture is shared. The pricing is the same. And the advertised performance is similar.

The launch of Claude Mythos dates back to early April. It was then open, in a preliminary Preview version, to a handful of American organizations. Specifically, AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks. They could experiment with it for software vulnerability detection.

As of the latest word, roughly 200 organizations are in the loop, across 15 countries. Although Anthropic has not integrated Claude Mythos into its commercial offering, it has pushed, in beta within Claude Enterprise, the Claude Security feature, which enables using its other models to analyze code bases. It also provides, on request, access to the tools used with the preview: skills, an agent harness (mapping a codebase to report writing) and a threat-model builder (identifying potential attack targets and prioritizing work accordingly).

A Small Window for Use Without Credits on Per-Seat Plans

Users of Claude Mythos Preview can now replace it with Claude Mythos 5. A newer version that is “sometimes more capable”… and above all 2.5 times cheaper ($10 per million tokens in input and $50 in output, versus $25 and $125).

Read also: Moratoire sur l’IA : Anthropic veut geler un marché qu’elle domine

Claude Fable 5, its “cyber-free” alter ego, is available on the Anthropic API and on Enterprise plans with usage-based billing.
For Enterprise plans that are still on the per-seat model, access is included through June 22. Beyond that, depending on Anthropic’s capacity, credits may need to be purchased. The same applies to Pro, Max and Team subscriptions.

A Specific Data Retention Policy

Both models have a specific data retention policy. Anthropic will retain inputs and outputs for 30 days. It cites a single purpose: safety. More precisely, the ability to analyze malicious uses that only become detectable at scale. For example, state-sponsored espionage campaigns or jailbreak attempts based on hundreds of variations of a prompt.

This does not change anything for individual Claude plans, which are already subject to a 30-day retention period. The shift concerns organizations that have activated the ZDR (zero data retention) policy on Anthropic services or on third-party services (Amazon Bedrock, Google Cloud Agent Platform, Microsoft Foundry).

A Two-Tier Filtering System for Cyber Requests

Claude Fable 5 does not itself answer cyber-related prompts, but it can forward them to Claude Opus 4.8. Its guardrail system is thus two-tiered. It begins with an internal activation check of the model. If suspicious traffic is detected, a LLM classifier takes over. It uses a mechanism Anthropic has employed since last year: trained on synthetic data generated from a “constitution.” In other words, rules written in natural language specifying what is allowed or not. A dataset is then gradually enriched with insights from automated red-teaming.

The same type of classifier applies to the chemistry and biology domains. Anthropic has opted for safety*, filtering most requests. Until now, in these areas, Claude models have traditionally rejected only those prompts revolving around weapons.

A Bug Bounty That Produced Two Jailbreaks… “Non-Universal”

A program is in the works to enable biomedical research to use Claude Fable 5 without these guardrails. The same type of initiative already exists for cybersecurity professionals, across all of Anthropic’s models. In all cases, uses considered almost always malicious and lacking legitimate defensive use are systematically blocked. Examples: massive data exfiltration and writing ransomware.

Across all tested prompts, less than 5% triggered a switch to Claude Opus 4.8, we’re assured.

Anthropic is organizing a two-pronged bug bounty. One, private, focused on Claude Fable 5. The other, public, concerning Claude Opus 4.8, but surrounded by comparable guardrails. As of June 5, the public competition had generated about 100,000 attempts — equivalent to roughly 1,000 hours of effort. They produced two jailbreaks, but each is task-specific. No universal jailbreak that would let interacting with the model occur as if a guardrail were absent.

* A note that echoes Claude Fable 5’s performance on tasks that could give rise to malicious use. Anthropic cites a hypothesis the model proposed about a protein conferring a high level of resistance to the bacterium E. coli… and an independently conducted parallel study has corroborated it.

Dawn Liphardt

Dawn Liphardt

I'm Dawn Liphardt, the founder and lead writer of this publication. With a background in philosophy and a deep interest in the social impact of technology, I started this platform to explore how innovation shapes — and sometimes disrupts — the world we live in. My work focuses on critical, human-centered storytelling at the frontier of artificial intelligence and emerging tech.