David WalshSubscribe
Ship's dispatch · AI Models · Jun 10, 2026

Sizing up Claude Fable 5

Anthropic's first public Mythos-class model is the most capable thing you can rent — and the easiest to waste money on. Here's how I'd actually route it.

On this page
  1. One model, two products
  2. The numbers
  3. In practice
  4. The safeguard fight
  5. How I'd route it
  6. Caveats
TL;DR
  • Fable 5 is a real step up. Roughly an 11-point jump over Opus 4.8 on agentic coding (SWE-Bench Pro 80.3% vs 69.2%), and the gap widens the longer and harder the task runs.
  • Fable 5 and Mythos 5 are the same model. The only difference is safety: Fable re-routes cyber, bio/chem, and distillation queries to Opus 4.8; Mythos lifts those guardrails for vetted Project Glasswing partners.
  • My verdict: aim it at the hard, long-horizon work, not the chat box.At double Opus pricing, with an invisible anti-AI-R&D safeguard and a subscription-access cliff on June 23, route by task — don't make it your default.

On June 9, Anthropic shipped one model as two products. Claude Fable 5 (claude-fable-5) is generally available everywhere — the API, claude.ai, Claude Code, AWS Bedrock, Google Vertex, Microsoft Foundry. Claude Mythos 5is the same weights with the cyber safeguards lifted, restricted to vetted Project Glasswing partners. Anthropic's own framing: Fable is “a Mythos-class model that we've made safe for general use.”

I spent a day with it and read everything else I could find. This is the practitioner's cut — what it's good at, what it costs, and where the sharp edges are.

One model, two products

The lineup is now Haiku → Sonnet → Opus → Mythos-class. Per Anthropic's own footnote, Mythos-class models “sit above our Opus class in capability.” Fable 5 is the first one you can actually buy. The April “Mythos Preview” was the proof of concept — it found a 27-year-old OpenBSD vulnerability and a 16-year-old FFmpeg bug — and Project Glasswing partners have since logged more than 10,000 high- or critical-severity flaws.

80.3%SWE-Bench Pro
$10/$50Per Mtok · in / out
1MContext window
<5%Sessions hit fallback

Pricing is exactly double Opus 4.8 — $10 in, $50 out per million tokens — with a 1M-token context window and 128K max output. Adaptive thinking is always on and cannot be disabled; Anthropic advises defaulting to high or xhigh effort.

The numbers

Anthropic published a comparison table, corroborated by The Decoder, Digital Applied, and others. The single most important thing to understand about it is the asterisk: starred rows are the restricted Mythos 5 score, not Fable. Because Fable re-routes cyber and bio queries to Opus 4.8, a Fable deployment performs closer to Opus on exactly those rows.

SWE-Bench Pro — agentic coding

Higher is better · pass rate %

Fable 580.3
Opus 4.869.2
GPT-5.558.6
Gemini 3.1 Pro54.2
BenchmarkFable / Mythos 5Opus 4.8GPT-5.5Gemini 3.1 Pro
SWE-Bench Pro80.369.258.654.2
SWE-bench Verified95.088.682.678.8
FrontierCode Diamond (xhigh)29.313.45.7
Terminal-Bench 2.1 *88.082.783.470.7
GDPval-AA (ELO)1932189017691314
Humanity's Last Exam *64.557.952.251.4
Blueprint-Bench 2 (spatial)38.614.536.226.5
Legal Agent Benchmark13.310.42.10.0
ExploitBench (Cap%) *78.040.034.0
OSWorld-Verified (computer use)85.083.478.776.2

* Restricted Mythos 5 score. Fable 5 falls back to Opus 4.8 on cyber/bio tasks — it made 0% progresson offensive cyber in blocking mode, so the headline 78% ExploitBench number belongs to a model you can't buy.

The other story in the table is how hard it scales with effort. Per the system card, SWE-Bench Pro climbs from 75.0% at low effort to 80.4% at xhigh; FrontierCode Diamond climbs from 11.5% to 30.9%. GPT-5.5 stays flat near 5–6% on FrontierCode no matter how hard it reasons. That's the tell: Fable's gains compound on long, hard work — not on quick answers.

In practice

The hands-on reports line up with the benchmarks. Simon Willison spent about five and a half hours on launch day and called it “a beast … slow, expensive and has been quite happily churning through everything I've thrown at it.” Using Claude Code, it solved his target feature in Datasette Agent and then identified and implemented four supporting features in his underlying llm library — several days of work, shipped in a day.

I used $110.42 worth of tokens today, all as part of my $100/month subscription.Simon Willison — $99.26 of it in a single Datasette Agent session

Every's “Vibe Check” (seven testers over a week) scored Fable 91/100 on their Senior Engineer benchmark, against Opus 4.8's 63 and GPT-5.5's 62, calling it “the best coding model in the world” but “a warp drive for power users — overpowered for everyone else.” Their line that stuck with me: “It rewards a clear brief and punishes a loose one.” Dan Shipper reported tasks routinely burning 500K–1M tokens.

The scale stories are the headline. Stripe (cited by Anthropic) ran a codebase-wide migration across a 50-million-line Ruby codebase in a day — work estimated at over two months for a full team. And the “Fable” name is earned: Justin Hart's reverse-chronology “butterfly” story test produced planted payoffs revealed in reverse, the clearest “yes” he'd gotten to the question of whether a model actually understands story. Every, for its part, found creative writing “mixed” — too slow for rapid iteration.

Agentic specifics worth knowing

  • Default effort is “high” in Claude Code; “ultracode” sends xhigh plus dynamic-workflow orchestration. Thinking can't be turned off.
  • A refused request returns stop_reason: "refusal" as an HTTP 200, not an error. A new beta fallbacks parameter can auto-retry on another model — but the raw Messages API does not fall back by default. Client apps route to Opus 4.8 silently; your own integration must handle the refusal.
  • Penetration testing, CTF, and biology-adjacent codebases trip fallback frequently, often on the first request.
  • Anthropic's messaging leans hard on multi-agent orchestration: Fable as the planner that delegates sub-tasks to cheaper models. Users describe the shift from “giving it tasks to giving it objectives.”

The safeguard fight

Fable's visible classifiers cover three domains — cyber, bio/chem, distillation — that fall back to Opus 4.8 and notify the user. Anthropic ran an external bug-bounty program with more than 1,000 hours of testing during which no one found a universal jailbreak (it does disclose the UK AI Security Institute “made progress towards one”). A mandatory 30-day retention policy now applies to all Mythos-class traffic; there is no zero-data-retention option.

The controversial part, surfaced from the system card by Nathan Lambert, is a fourth, invisiblesafeguard targeting frontier-LLM development. Per the card, these safeguards “will not be visible to the user” and the model “will not fall back to a different model” — instead they “limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning.” Anthropic estimates this touches ~0.03% of traffic.

An auto-degrading model is “categorically misaligned AI.”Nathan Lambert — Interconnects

Lambert argued the move looks more like competitive entrenchment than safety. Whatever your read, it's a real consideration if your work touches training infrastructure: a model that quietly does worse, without telling you, is a different tool than one that refuses.

Heads up

Anthropic did not state an ASL designation in the launch post. METR concluded Mythos 5 is “likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks,” with acceleration “concentrated in engineering execution rather than research judgment.” Code-summary honesty improved dramatically (dishonest-summary rate: Sonnet 4.6 65.2% → Fable 5 4.6%).

The access cliff

The Hacker News thread was dominated less by capability than by the subscription change: free on Pro/Max/Team/Enterprise from June 9–22, then requiring usage credits from June 23, with restoration “when capacity allows.” Commenters called it “the pharmaceutical method — get them hooked on free samples, then raise the price.” Reddit users on Max 20x reported burning ~2% of the credit limit per minute. If you're evaluating, the clock matters.

How I'd route it

  1. Default to Opus 4.8; reserve Fable for hard, long-horizon work. The clearest wins are multi-file, multi-day migrations, deep code review, large-dataset synthesis, and agent pipelines that run for hours with minimal supervision. If a representative sample of your real tasks shows a quality or completion gain large enough to offset roughly doubled token cost, route those task types to Fable.
  2. Run your evaluation inside the free window (through June 22). Build a representative task suite now, measure quality and token spend across low/medium/high/xhigh effort, and set your routing policy before the cliff.
  3. In multi-agent pipelines, use Fable as the orchestrator, not the worker. Let it plan and delegate to Sonnet/Haiku for sub-tasks; reserve xhigh for the genuinely hard steps. Rewrite CLAUDE.md to give it judgment and fewer micro-instructions.
  4. On the raw API, handle refusals explicitly. Catch stop_reason: "refusal" and either use the fallbacks parameter or your own routing. Expect frequent fallback in security, biology, or chemistry domains.
  5. Account for 30-day retention.If you relied on zero-data-retention agreements, Mythos-class models aren't eligible — factor that into any regulated-data workflow.

Caveats

  • The asterisked benchmarks are Mythos 5, not Fable 5. Don't quote ExploitBench 78% as the model you can deploy — Fable falls back to roughly 40% there, and 0% on offensive cyber in blocking mode.
  • Several speedup claims are vendor-side. 430× kernel speedups, 69× self-training, 10× drug-design acceleration — all from Anthropic/system-card interpretation. Treat as unreplicated.
  • No confirmed ASL designation.The “ASL-3 / CB-1” characterization floating around is unverified analyst commentary.
  • Model size is inferred, not confirmed.Willison's “big model smell” is a vibe, not a parameter count. Anthropic disclosed nothing.
  • Knowledge cutoff reported as end of January 2026, per the system prompt — not formally stated on the model page.
  • Launch-day enthusiasm.“Best in the world” verdicts are day-zero impressions. Validate on your own workloads.

Sources: Anthropic launch post & system card · Simon Willison · Every “Vibe Check” · Interconnects · Digital Applied · CNBC · Reuters · Hacker News.