David WalshSubscribe
Sounding · Safety index · Jul 16, 2026

Soundings in shallow water

An independent panel graded nine frontier AI labs on safety this month, and the highest mark any of them earned was a C+ — to Anthropic, the lab that markets itself as the safety-first one. The same review found Anthropic, OpenAI, Google DeepMind, and Meta had all quietly weakened or voided their own pledges to pause development at specified risk thresholds, and singled out Anthropic over a disputed link to a strike that killed civilians in Iran.

On this page
  1. The soundings
  2. Six fathom lines
  3. The tide goes out
  4. What the lead line snagged
  5. My read
  6. What to watch
  7. Caveats
TL;DR — the short versionSAFETY INDEX

Nine labs, six domains, thirty-seven indicators — and not one company cleared the mark an independent panel considers responsible practice.

  • 01The Future of Life Institute’s Summer 2026 AI Safety Index graded nine frontier labs, and the top score was a C+.Anthropic took it, ahead of OpenAI and Google DeepMind (both C), Meta (D+), Z.ai and Alibaba Cloud (both D−), and xAI, DeepSeek, and Mistral, which all failed.
  • 02Not one lab scored above a D in the “Existential Safety” domain— the category closest to the extinction-level risk every major lab’s own leadership talks about in public.
  • 03Four labs have quietly walked back their own pause pledges. Anthropic, OpenAI, Google DeepMind, and Meta have all weakened or voided earlier commitments to halt development if specific risk thresholds were crossed — reviewers call it a moving goalpost.
  • 04The same four labs reversed bans on military use between 2024 and 2026.The panel singled out Anthropic specifically for “questionable military engagements,” citing a disputed report tying Claude’s use in a Palantir targeting system to a strike that killed civilians at a school in Iran.

A leadsman casts a weighted line off the bow to find out how much water is under the keel before the chart can be trusted. A short reading means shallow water — time to change course, not congratulate the crew for still floating. That’s roughly what an independent panel of AI safety researchers did to nine frontier labs this month, and the line came back short for all of them.

01 The soundings

The Future of Life Institute has run this index since 2024; the Summer 2026 edition evaluated nine model developers against thirty-seven indicators across six domains, combining public evidence with a company survey and reviews from a seven-person panel of safety researchers and governance experts. No company reached a B, let alone an A.

LabGradeWhere it led — or lagged
AnthropicC+Leads five of the six domains; the only lab publishing both a system prompt and a full behavior specification.
OpenAICLeads the Risk Assessment domain; the only lab with a published whistleblowing policy at review time.
Google DeepMindCThird place, a fraction behind OpenAI.
MetaD+Climbed from sixth place to fourth since the last index.
Z.aiD−Not singled out in coverage of the review.
Alibaba CloudD−Not singled out in coverage of the review.
xAIFScore 0.65; dropped from fourth place to seventh.
DeepSeekFScore 0.47.
MistralFScore 0.33 — the lowest in the index.
Highest grade awarded, any labC+Anthropic's overall mark — the ceiling for the entire index
Domains where no lab beat a D1 of 6Existential Safety — the one domain every reviewed lab failed to clear
Labs that walked back pause pledges4 of 9Anthropic, OpenAI, Google DeepMind, and Meta all weakened or voided earlier commitments

02 Six fathom lines

The index doesn’t hand out one grade; it marks six of them per lab, like depth readings along a lead line — Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. Anthropic leads five of the six, drawing on relatively strong transparency and a comparatively established safety framework; it’s the only lab reviewers credit with publishing both a system prompt and a full behavior specification. OpenAI takes the sixth — Risk Assessment — on the strength of a broader evaluation suite and more diverse external red-teaming, and edges ahead in Current Harms too, partly because it was the only lab with a published whistleblowing policy at review time.

The one reading that should worry every lab’s own leadership more than the rest: Existential Safety. Every frontier lab in this index has a public position on catastrophic or extinction-level AI risk — Anthropic and OpenAI both cite it as a founding reason for existing at all. None of them scored above a D in the domain built to check whether that concern shows up in reviewable practice rather than in a mission statement.

03 The tide goes out

Two commitments the labs made publicly, in calmer years, are receding at once: a pledge to pause unilaterally if a system approached a specified danger threshold, and a ban on using their models for military applications. Both have been walked back by the same four companies.

The labs' case

Competitive necessity, not abandonment

  • Anthropic’s published red lines for government use of Claude are narrow by design: no mass domestic surveillance, no fully autonomous weapons. A human retaining the final decision, the company argues, keeps a use case on the right side of both lines.
  • Several labs frame pause-pledge changes as conditional rather than abandoned — contingent on whether competitors observe similar restraint, not a unilateral walk-back.
  • Defense-sector revenue and government contracts have become a meaningful, public part of multiple labs’ stated business strategy over the same period.
“[This] is a use case that doesn’t even violate our red lines.”Dario Amodei, CEO, Anthropic, in a Bloomberg interview, June 10
The reviewers' case

A moving goalpost undermines the whole framework

  • The panel’s own language: commitments to pause, and earlier bans on military use, have been “weakened or voided” across Anthropic, OpenAI, Google DeepMind, and Meta since 2024.
  • Stuart Russell, the UC Berkeley computer scientist on the review panel, put it more bluntly: companies that once pledged to release new systems “only with safety measures appropriate for their capability levels” are “planning to release them even if it’s demonstrably unsafe to do so.”
  • A pledge that’s contingent on a rival’s behavior isn’t a red line anymore — it’s a negotiating position, and the panel graded it that way.
“AI companies are sprinting toward a cliff. Despite acknowledging the great risks of artificial superintelligence, they continue racing to build it.”Max Tegmark, president, Future of Life Institute

04 What the lead line snagged

The single most uncomfortable line in the review isn’t a grade; it’s a citation. Reviewers flagged Anthropic specifically for “questionable military engagements,” pointing to reporting that ties Claude — integrated into Palantir’s Maven targeting-analysis system — to a February 28 strike on a school in Minab, Iran, on the first day of a U.S. military campaign there. Reported death tolls vary by outlet, from roughly 120 to as high as 168 people, most of them children.

What’s established and what isn’t established here matter equally. Established: Claude is part of the Maven pipeline the strike’s targeting reportedly drew on, and some of the data feeding that pipeline was later described as outdated. Not established, by any account this piece found: a public, primary-source document tying a specific AI-generated output to the decision to strike that particular building. Asked directly, Amodei told Bloomberg he didn’t know exactly how Anthropic’s system had been used in the operation — while maintaining that, because a human made the final call, the use “doesn’t even violate our red lines.”

05 My read

Grading on a curve is still grading on a curve, even when the subject is existential risk. A C+ read in isolation sounds like a middling but decent report card; read as the ceiling nine companies couldn’t clear, it’s closer to a fleet-wide finding that nobody’s actually seaworthy. The index’s real service isn’t ranking Anthropic above OpenAI — it’s making the industry-wide retreat on pause pledges and military bans visible as a pattern, rather than four separate press stories that each look like an isolated, defensible judgment call.

The Minab citation is the sharpest version of that pattern, because it turns an abstract governance question — who gets to redraw a red line, and when — into a specific, disputed, human cost. Anthropic’s defense is procedurally coherent: a human retained the decision, and the line the company drew never covered this case. That defense is also exactly the reviewers’ complaint. A safety commitment that a company gets to redraw around its own product’s next use case isn’t a constraint on the company; it’s a description of what the company already planned to do.

What to watch

Four tells, roughly in the order they’ll arrive
1

Does Anthropic say more than Amodei's one Bloomberg line?

A single “doesn’t know exactly how” answer to a direct question about a strike that killed children is unlikely to be the company’s last word. Watch for a fuller statement, or a Maven-specific policy change.

2

Does any lab clear a B in any domain next cycle?

The Winter index is the panel’s next scheduled checkpoint. A single domain clearing a B would be the first sign the curve is moving up rather than sideways.

3

Do pause pledges get rewritten, or just quietly not mentioned?

Watch whether any of the four labs the panel flagged publishes a revised, specific pause commitment, versus letting the old one go unmentioned in future safety documentation.

4

Does a government body cite this index directly?

Safety indices are easy to cite and easy to ignore. A congressional letter or procurement policy that names the report by name would be the clearest sign it’s shaping decisions rather than just headlines.

Caveats — read before you quote
  • Domain-level detail here is drawn from secondary reporting,not a firsthand read of the Institute’s full methodology PDF, which this piece could not retrieve directly. Overall letter grades are corroborated across multiple outlets; some finer domain-by-domain distinctions are not.
  • The Minab strike’s AI-causation is explicitly unconfirmed.No source this piece found offers a public, primary document tying a specific Claude output to the decision to target that building — the claim is that Claude was integrated into the targeting pipeline involved, not that it selected the target.
  • Reported death tolls for the strike vary by outlet (roughly 120 to 168); treat the figure as a range under active reporting, not a settled number.
  • Numeric index scores are confirmed for four labs(Anthropic 2.66; xAI 0.65; DeepSeek 0.47; Mistral 0.33); the remaining five labs’ letter grades are corroborated but their underlying numeric scores were not independently found here.
  • This is a fast-moving story; status is current as of July 16, 2026.