On June 30 Anthropic shipped Claude Sonnet 5 and made it the default everywhere the next day — claude.ai Free and Pro, and Claude Code with a 1M-token window. It closes most of the gap to Opus 4.8, wins one benchmark outright, and launches at a discount with an expiry date and an asterisk.
- 01Six points behind the flagship, one ahead. 63.2% vs 69.2% on SWE-bench Pro, 81.2% vs 83.4% on OSWorld — but 80.4% vs 74.6% over Opus 4.8 on Terminal-Bench 2.1.
- 02$2 in / $10 out per million tokens until August 31, then $3/$15 — the same sticker as Sonnet 4.6. The catch: a new tokenizer that cuts the same text into up to ~1.35× the tokens.
- 03Route by lane, not by loyalty. Sonnet 5 is the new default lane; Opus 4.8 and Fable 5 stay the escalation lane; re-price your workloads before September 1.
01 What shipped
The launch itself was almost drowned out. Sonnet 5 arrived June 30 — the same day Washington lifted the export controls on Fable 5 and Mythos 5, the story I covered yesterday. By July 1 it was the default model for every Free and Pro user on claude.ai, the default in Claude Code, and live on Bedrock, Google Cloud, and Foundry. Anthropic calls it its “most agentic Sonnet model yet,” which is marketing, and then backs it with numbers, which isn’t.
Three things matter operationally. It ships with a 1M-token context window by default — no beta header, no long-context surcharge. It becomes the model most people get without choosing anything. And it lands at an introductory price with a fuse: $2 per million input tokens and $10 out through August 31, then $3/$15 — nominally identical to Sonnet 4.6.
The numbers
Read the chart honestly and the story is “close, not equal.” Six points on SWE-bench Pro is a real gap — on hard, multi-file agentic coding you will still feel the flagship’s edge. Two points on OSWorld is a coin toss. On agentic search Sonnet 5 posts 84.7% on BrowseComp (no published Opus figure to pair it with), and on GDPval knowledge work the two models are, for any practical purpose, the same model.
02 The terminal upset
The number that made me sit up isn’t any of the near-misses. It’s Terminal-Bench 2.1: 80.4% against the flagship’s 74.6%. The middleweight doesn’t approach Opus 4.8 at driving a shell — it beats it, by nearly six points, at a third of the standard price.
That specific win matters because the terminal is where agentic work actually happens: package managers, test runners, git, deploys, the unglamorous plumbing that fills an agent’s day. A model that is cheaper andbetter at the highest-volume lane of agent traffic isn’t a budget option. It’s the correct default — which is presumably why Anthropic made it one.
“Claude Sonnet 5 is our most agentic Sonnet model yet, with substantial improvements over Sonnet 4.6 in reasoning, tool use, coding, and knowledge work.”
Anthropic · launch announcement, Jun 30, 2026The pattern rhymes with what I wrote when sizing up Fable 5: capability tiers aren’t a ladder, they’re a toolbox. Each generation, one “smaller” model quietly takes a lane from the model above it. Last time it was long-context retrieval. This time it’s the shell.
The tokenizer footnote
Sonnet 5 inherits the tokenizer Anthropic introduced with Opus 4.7: the same document can map to roughly 1.0 to 1.35 times as many tokens as it did under Sonnet 4.6, depending on content. Price-per-token is not price-per-work. Re-running a Sonnet 4.6 workload at a mid-range 1.3× inflation:
- Under intro pricing$2/$10 × 1.3 ≈ $2.60 / $13.00 effective≈ 13% cheaper than 4.6
- After Aug 31$3/$15 × 1.3 ≈ $3.90 / $19.50 effective≈ 30% dearer than 4.6
So the “same price as 4.6” framing survives only until September. After the promo, an unchanged workload costs roughly a third more in real money for — to be fair — a meaningfully better model. That may be a trade worth making. It is not the trade the sticker advertises, and if your finance team budgets in tokens, the budget line will move even though the price list didn’t.
Routing, updated
Sonnet 5 is the default lane.
Coding agents, terminal work, browsing, computer use, long-context sessions — start here. It beats the flagship where volume lives and ties it where judgment lives. Let the default be the default.
Escalate on failure, not on fear.
Keep Opus 4.8 and Fable 5 for the hardest planning, gnarly multi-file refactors, and anything where a six-point SWE-bench gap compounds. Route the retry, not the whole queue.
Re-price before September 1.
Benchmark your real workloads under the new tokenizer while the discount pays you to do it. Decide with effective cost per task — dollars per merged PR, per resolved ticket — never dollars per million tokens.
Pin versions in anything that ships.
A default that changes under you is a production incident with good benchmarks. Name claude-sonnet-5 explicitly in agents and pipelines; let chat surfaces float.
A crowded week
Fable 5 pulled offline
An export-control letter takes Anthropic’s flagship dark worldwide — the backdrop everything this week plays against.
CONTEXTSonnet 5 ships · controls lifted
Anthropic releases its most agentic Sonnet with a 1M-token default window and intro pricing. The same day, Washington lifts the Fable/Mythos controls.
LAUNCH DAYDefault everywhere
Sonnet 5 becomes the default for Free and Pro on claude.ai and in Claude Code; Fable 5 returns globally alongside it.
ROLLOUTFable 5 usage promo ends
The restored flagship’s 50%-of-weekly-limits window closes for Pro, Max, and Team plans; Fable routes to usage credits.
FUSE №1Intro pricing expires
$2/$10 becomes $3/$15 at midnight. With the tokenizer’s ~1.3× count, effective cost for an unchanged workload rises ~30% over Sonnet 4.6.
FUSE №203 My read
The obvious reading is competitive: Anthropic, days from the most anticipated IPO in tech, ships a model that undercuts its own flagship for most work and prices it to make switching irresistible for exactly two months. That reading is probably right and slightly boring.
The more useful reading is architectural. With Fable 5 restored at the frontier, Opus 4.8 holding the hard-problem tier, and Sonnet 5 now owning the default lane, the “which model?” question has stopped having one answer. It has a routing table — and the teams that treat it that way, with explicit lanes, explicit escalation, and effective-cost math, will quietly pay less for better work than the teams still picking a favorite model like a football club.
- All benchmark figures are Anthropic’s launch numbers. Independent replications were not yet published as of July 2; treat small deltas as provisional.
- The 1.0–1.35× tokenizer range is content-dependent. Code, prose, and non-English text inflate differently — measure on your corpus, not the midpoint I used above.
- Effective-cost math assumes an unchanged workload.If Sonnet 5 finishes tasks in fewer turns — the vendor’s claim — cost per task can fall even as tokens per document rise.
- This piece reflects the first 48 hours after launch; pricing and defaults are as published on July 2, 2026.
