Artificial Intelligence & Machine Learning
,
Governance & Risk Management
,
Next-Generation Technologies & Secure Development
AI Firms Weakened Safety Pledges as Risks Grow, Says Report

Even the best-performing artificial intelligence company could only manage a C+ in an industry-wide safety review.
See Also: Edge Transformation: Top 5 SASE Predictions and Trends
The Future of Life Institute’s latest AI Safety Index found every major AI developer, including Anthropic, OpenAI and Google DeepMind, falling short of what its expert panel considers responsible practice. Three companies, xAI, DeepSeek and Mistral, failed outright.
The index, compiled by an independent panel of seven AI safety and governance experts, found that leading developers are walking back earlier safety commitments. Anthropic, OpenAI, Google DeepMind and Meta have all weakened or voided pledges to pause development unilaterally if certain risk thresholds were crossed, some now contingent on what competitors do, the report said. Reviewers called the pattern a “moving goalpost” that has “undermined safety frameworks across the board.” Five companies, including those four and xAI, have published formal safety frameworks, but the report found the frameworks often lack measurable thresholds, independent audits or clear authority to halt a release.
Sabina Nong, safety investigator at the Future of Life Institute, told ISMG the competitive pressure to release increasingly capable models has intensified. Anthropic and OpenAI now tie deployment pauses to competitors’ actions, while Google DeepMind and Meta have removed unilateral pause commitments altogether, she said. “If we leave AI safety to corporate-defined rules, we are destined to race to the bottom,” she said.
The review panel, which includes University of California-Berkeley’s Stuart Russell, University of Oxford AI Governance Initiative director Robert Trager and Renmin University’s Yi Zeng, scored each company against 37 indicators across six domains: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information sharing. The grades follow the same scale as the American grading system. Every company scored an overall grade of below C+.
Anthropic led five of the six domains, drawing on strong transparency practices and a comparatively established safety framework. It is the only company that publishes both its system prompt and a behavior specification, said the report. It’s B+ in information sharing was the single highest mark awarded anywhere in the index. OpenAI took second place overall with a C. The report credits it with leading risk assessments, though the two companies received the same C+ grade there.
Google DeepMind followed close behind OpenAI, also with a C, keeping the same three companies at the top as in the prior index. Below them, Meta was the notable mover, climbing from sixth place to fourth, improving to a D+, though the report does not explain what drove the gain. xAI moved the other way, dropping from fourth to seventh as its overall grade collapsed to an F. The report does not explain the cause, but it did criticize xAI’s risk-assessment evaluations as containing “gaping holes,” including no research and development data, against the backdrop of a “regressing” Grok 4 model.
Inadequate safety is not confined to one country or region, the report said, pointing to failing grades split evenly across the U.S., China and Europe, held respectively by xAI, DeepSeek and Mistral. The finding carried a sharper edge for Mistral – despite the European Union’s leadership on AI safety regulation, the bloc’s own top AI company scored last of the nine reviewed. Z.ai and Alibaba Cloud, both Chinese companies, scored D-.
Benchmark scores told only part of the story in the current harms domain, the report found, since several companies with otherwise solid results were pulled down by documented real-world incidents.
OpenAI faces a wrongful-death lawsuit from the family of Adam Raine, a California teenager who died by suicide. The suit alleges its safeguards were loosened beforehand. Google DeepMind settled a similar Character.ai-linked suicide lawsuit earlier this year but faces a separate suit over its own Gemini chatbot. Meta was ordered to pay $375 million in civil penalties after a New Mexico jury found it misled consumers about child safety on its platform. xAI failed the category after reviewers cited reports that its Grok tool generated child sexual abuse material and sexualized images of real people without consent, including minors (see: UK Probes X Over AI Deepfake Porn).
Existential safety, covering long-term risks from highly capable or self-improving systems, was the industry’s weakest domain. No company scored above a C-, and most fell to a D or lower. The report credited constructive attempts, including Anthropic’s constitutional classifiers, OpenAI’s calls for new governance institutions, Google DeepMind’s monitoring commitments and Meta’s provisions against loss of control, but judged them “entirely inadequate.” The reviewers also challenged reliance on chain-of-thought monitoring, the practice of inspecting a model’s step-by-step reasoning for warning signs, saying that “detection is not prevention.”
Nong said existential safety has been the industry’s weakest category since the index was first published because companies lack credible technical strategies for controlling increasingly capable AI systems. A separate finding singled out Google DeepMind, OpenAI and xAI for a gap between public messaging and actual conduct. OpenAI shifted its regulatory posture repeatedly, the reviewers said, opposing a California transparency legislation before invoking “reverse federalism” and reversing course on Illinois legislation within a month. At Google DeepMind, reviewers contrasted chief executive Demis Hassabis’s public reassurances with a company they called “largely uncooperative on policy.” xAI stayed “largely silent on most policy questions” apart from Elon Musk’s personal support for a California safety bill.
A related finding concerned the industry’s shift toward military applications. Companies including Anthropic, OpenAI, Google DeepMind and Meta had previously restricted or banned military use of their systems but reversed course between 2024 and 2026, joining xAI and Mistral in pursuing defense contracts. The panel singled out Anthropic for “questionable military engagements,” citing a reported link to the Minab school strike, which caused mass civilian deaths. DeepSeek, Z.ai and Alibaba Cloud all face separate U.S. allegations of ties to China’s military. Z.ai and Alibaba Cloud have denied the claims.
The Future of Life Institute, a nonprofit focused on reducing large-scale risks from transformative technology, has run the index since 2023 to counter competitive pressure that might otherwise reward speed over safety. Its panel concluded that the current pattern “incentivizes a collective race to the bottom, rather than to the top,” and said that safety commitments can no longer rest on self-governance.
