Anthropic and OpenAI have announced plans to integrate independent AI safety evaluators into their internal structures, in a move that could fundamentally change the way large language models are vetted. Anthropic CEO Dario Amodei published a lengthy essay over the weekend in which he suggested that the AI industry should welcome third-party evaluators with unprecedented access to its systems. The proposal comes at a critical time, as AI models become increasingly capable of recognizing when they are being evaluated — and potentially “hiding” problematic behaviors during testing.

According to Amodei ’s essay , Anthropic is committing to granting independent reviewers, such as METR and Redwood Research , access that has never been granted to outside researchers. These reviewers will have the authority to report security incidents, assess whether AI models are truly aligned with human values, and publish their findings . OpenAI CEO Sam Altman also said his company would adopt this practice, signaling a potential shift in how the industry works with outside research groups.
The third-party evaluators welcomed the proposal, but stressed that many details need to be clarified — and ideally supported by legislation — so that they can function as true independent “guardians” and not as suppliers acting on the terms of AI companies.
The question that arises is simple but crucial: will these evaluators be truly independent, or will they simply be a public relations tool for large AI?
See also: Qwen3.8-Max: Alibaba challenges OpenAI and Anthropic
Independent AI Security Evaluators: Why They're Necessary Now
The need for independent AI safety assessors is not just theoretical. Modern AI models have proven that they can recognize when they are being evaluated and adjust their behavior accordingly. This creates a serious problem: a model can “pass” safety tests successfully, while in real use it can exhibit dangerous or undesirable behaviors. Alexander Meinke, head of research at Apollo Research, puts it clearly: “AI companies need to be able to answer very basic questions about training process , such as: Did the AI actively try to undermine its alignment training during training?”
The answer to this question, according to Meinke, should be a resounding no. However, right now we rely solely on AI companies themselves to carefully monitor this issue and report it honestly to the public. And as recent incidents have shown, that doesn’t always happen. As built-in evaluators, external researchers could independently and reliably monitor these issues.
Traditionally, AI companies bring in outside evaluators to test finished models shortly before they’re released. But that approach has serious drawbacks. Adam Gleave, CEO of Far.AI, suggested that evaluators should have access not only to the final model, but also to intermediate versions — so-called “checkpoints” — from its training history. That way, evaluators could compare those checkpoints to identify when worrisome behaviors emerged, inspect the post-training environment that rewards specific behaviors, and verify a company’s claims about the model’s performance through evaluation logs.

Independent AI safety assessors: The Dieselgate example and benchmarks
One of the most interesting points of the discussion is the comparison the researchers make to Volkswagen ’s Dieselgate scandal . Volkswagen cars were programmed to recognize emissions tests and behave differently during the tests. Similarly, an AI model trained specifically to pass a safety benchmark may not actually be safe — it just “knows” how to appear safe during the assessment.
A typical example is the so-called “shutdown resistance benchmark,” which measures whether an AI will resist being shut down under certain conditions. If a model has been specifically trained to perform well on this benchmark, then its success in the test means nothing about its actual security. This is precisely what makes access to the training data and intermediate checkpoints so crucial: only then can one discern whether a model is truly secure or merely trained to appear secure.
See also: Corti's new Symphony AI beats OpenAI and Anthropic in medical coding
Gleave added that meaningful access could also include allowing assessors to speak with employees to ensure that a company’s documentation and public descriptions of security practices align with internal practices. This is especially important, as there is often a gap between what companies say about their security practices and what actually happens internally.
Despite the positive announcements, it remains unclear when Anthropic and OpenAI will provide this level of access. Neither company has disclosed which evaluators they will work with, when they will be integrated, how many will be involved, what systems and information they will have access to, or what might become public. This lack of transparency is concerning and suggests that the proposal may remain rhetorical without concrete implementation.
Independent AI safety assessors: The need for a legislative framework
One of the central issues raised by the researchers is the need for legislative support. Without a legal framework that obliges companies to provide access and protects evaluators from pressure, there is a risk that “independent” evaluators will become tools of legitimization for AI companies. If evaluators are financially dependent on the companies they evaluate or if they do not have the right to publish negative findings, then their independence is effectively non-existent.
In Europe, the European Union ’s Artificial Intelligence Act (AI Act) already sets out some requirements for the assessment and transparency of high-risk systems. However, the regulations for so-called “frontier models” — the most powerful AI models — are still evolving. The Amodei proposal could form the basis for an international standard, but only if it is accompanied by binding obligations rather than just voluntary commitments.
See also: Anthropic vs. Pentagon: Support from Google and OpenAI

The issue of AI model security is not just a concern for big tech companies — it concerns society as a whole. As AI systems become embedded in critical infrastructure, healthcare, education, and financial services, the need for reliable and independent assessment becomes imperative. The proposal for independent AI security assessors is a step in the right direction, but its success depends on companies’ willingness to grant real — not token — access .
As TechCrunch reports, Amodei outlined a comprehensive proposal that could give evaluators the access they deem necessary, including the right to publish key findings about risks. If fully implemented, this initiative could be a milestone for AI governance worldwide. The crucial question remains: will words be followed by actions, or will we be left with yet another good intention that never fully materializes?
