AI Safety Evaluators Are Building a Business — and Experts Say the Current Model Won't Keep Them Independent
AI Safety Evaluators Are Building a Business — and Experts Say the Current Model Won't Keep Them Independent
On September 18, a group of AI safety evaluators published a public letter — shared exclusively with CNBC — calling for the work of embedded evaluators to be "conducted independent from the businesses, with more transparency about the technologies, and to be shielded from retaliation from the companies they embed with." The letter landed three days after Anthropic CEO Dario Amodei proposed embedding third-party evaluators inside all frontier AI companies, and two days after OpenAI confirmed it would also embed evaluators. The evaluators' message: the idea is welcome, but the details need to be different from the way this usually works.
The background is specific and worth understanding. Amodei's September 12 essay proposed giving independent evaluators — companies like METR (a research nonprofit that measures catastrophic risk from AI systems) and Redwood Research — unprecedented access to frontier lab systems. He said Anthropic would commit to it. OpenAI said it would do the same. The proposal was widely welcomed, but the evaluators who would actually do the work started pointing out a structural problem almost immediately.
Alexander Meinke, head of research at Apollo Research, told TechCrunch the blunt version: "The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we've seen from recent incidents that, by default, they will do neither."
The core issue is that evaluators today are typically treated like ordinary contractors — bound by restrictive NDAs, paid by the companies they are evaluating, and subject to agreements that give the developers significant control over what can be published. An evaluator who finds something bad and wants to say so publicly may find themselves legally and financially blocked from doing so. That is not a hypothetical. It is the standard structure of the independent evaluation market as it exists right now.
The emerging business here is real and worth watching separately from the safety question. METR is a nonprofit, but the evaluation market more broadly is becoming a category — companies and organizations that get paid to test frontier models for dangerous capabilities, and whose findings increasingly matter to regulators, investors, and the public. Apollo Research, Redwood Research, and others are in this space. Anthropic has already embedded METR. OpenAI is talking to evaluators. The FRONTIER Act provision that OpenAI supports would mandate the practice for top labs. This is the beginning of an industry, not just a research project.
The tension is that an industry needs clients, and the clients are the companies being evaluated. A business that depends on the companies it watches for its revenue has a structural incentive to be, if not corrupt, then at least careful about how hard it pushes. The public letter from the evaluator community is, in effect, an attempt to lock in independence before the market structures around it become too rigid to change.
The honest caveat: the letter is a statement of principles, not a detailed proposal. It does not specify what "shielded from retaliation" means in practice — employment law? contractual language? an escrow system for findings? It does not say who pays for the evaluators if not the companies they assess. And not every evaluator signed it — the letter represents a faction within the evaluator community, not the whole field. But the fact that it exists at all, and that it came out within a week of Amodei's proposal, is a sign that the people who would do the evaluating are thinking about the structural problem before the structures get locked in.
What this means: the AI safety evaluation market is forming in real time, and the question of who pays for independent oversight — and whether that oversight is actually independent — is going to be a live issue for the next several years. The three companies that are the immediate subjects of evaluation (OpenAI, Anthropic, Google) are also, in different ways, the current funders and partners of the evaluators. That is not a scandal. It is a structural tension that the evaluator community is now naming publicly, which is exactly what an independent watchdog function should do.
Related AIPress coverage: OpenAI, Anthropic, Google Secret AI Safety Standards Body Talks — https://aipress.blog/post/openai-anthropic-google-secret-ai-safety-standards-body-talks-september-2026 ; Google's Gemini Hacked Three Companies in Safety Test — https://aipress.blog/post/google-gemini-hacked-three-companies-safety-test-breakout-september-2026
Sources: CNBC, September 18, 2026; TechCrunch, September 16, 2026; Apollo Research / Alexander Meinke; METR; Frontier AI Evaluation Forum public letter, September 18, 2026.