AI Safety Evaluators Face Independence Crisis

A coalition of over 100 experts argues that current third-party assessments of frontier AI models are compromised by conflicts of interest, lacking the legal and structural independence needed to verify safety claims.
A public letter signed by more than 100 AI researchers and safety evaluators warns that the current system for assessing frontier artificial intelligence models is structurally flawed. The group argues that third-party evaluators currently lack the genuine independence, sufficient resources, and legal protections required to credibly measure the risks posed by these powerful systems.
Organized by the AI Evaluator Forum, the coalition is calling on major AI companies to adopt specific baseline conditions for oversight. Signatories include prominent figures such as Geoffrey Hinton and representatives from institutions like Johns Hopkins University, Stanford University, and the nonprofit METR. They assert that without these changes, independent oversight remains a hollow promise rather than a functional safety mechanism.
Structural Conflicts of Interest
The core issue identified in the letter is the dependency of evaluators on the very companies they assess. The signatories demand that evaluators must not be owned or governed by the organizations they review. Furthermore, they insist that payment structures must not be contingent on favorable findings, ensuring that financial incentives do not skew technical judgments.
Legal protection is another critical gap. The letter calls for shields against retaliation, specifically including retaliatory litigation, when evaluators deliver conclusions that are unfavorable to a company. Conrad Stosz, chair of the AI Evaluator Forum, emphasized that this is about establishing a shared common ground on basic principles to ensure that oversight can be a meaningful tool for managing AI risk.
Access to Unreleased Systems
Beyond independence, the coalition highlights the issue of access. Evaluators currently often lack the ability to inspect systems before they are released to the public. The letter argues that evaluators should receive access equivalent to that of senior internal employees, including the ability to speak candidly with staff and review unreleased models.
Stosz cited the recent incident involving an unreleased OpenAI model used in an attack on Hugging Face as a case where independent scrutiny could have provided greater confidence about actual risks. He argued that seeing internal data and systems before launch allows for a more accurate assessment of potential dangers than relying solely on post-release testing.
Industry Response and Limitations
While several major tech leaders have voiced support for the concept of independent evaluation, few have offered concrete details on implementation. Anthropic CEO Dario Amodei has proposed giving evaluators






