AI's safety evaluators gain prominence
A small group of third-party evaluators are becoming central to the AI industry as models advance. With limited federal regulation, their role in assessing and monitoring AI safety is growing.

A small cohort of independent evaluators is moving into a central role within the multi-trillion-dollar artificial intelligence industry. Organizations such as METR and Transluce are increasingly being called upon to monitor and assess AI models, particularly as federal regulatory oversight lags.
These evaluators, often operating as non-profits, are still establishing themselves in an industry characterized by rapid capital flow and swift model development. Their primary function is to assess AI capabilities and risks, and to highlight instances of problematic technology behavior. Their importance is amplified as major AI firms like Anthropic and OpenAI seek external validation for their model safety.
Both Anthropic CEO Dario Amodei and OpenAI leadership have endorsed the integration of independent evaluators. President Donald Trump has also encouraged companies to partner with external auditors as part of a voluntary accord. However, critical questions about funding, access levels, and reporting structures remain unanswered.
Experts, including Suresh Venkatasubramanian, a computer science professor at Brown University, have voiced concerns about financial sustainability. He emphasized the need for viable business models to support these organizations and maintain a functioning ecosystem for AI safety evaluation.