Proposal and Initial Responses
Dario Amodei proposed over the weekend that frontier AI companies embed third-party evaluators with the authority to report safety incidents, assess alignment, and publish findings. Amodei stated that Anthropic would provide independent groups such as METR and Redwood Research with expanded access. OpenAI CEO Sam Altman also expressed OpenAI’s commitment to this practice. While third-party evaluators broadly welcomed the proposal, they emphasized that the details are crucial for ensuring true independence.
Access to Checkpoints and Logs
Evaluators stressed the need for guarantees regarding access, timing, confidentiality rules, and publication rights. Historically, outside reviewers tested near-final models, but evaluators now seek access to intermediate checkpoints, training logs, evaluation transcripts, and the post-training reward environment to identify when and how concerning behavior emerged. Adam Gleave, CEO of FAR.AI, noted that comparing checkpoints and conducting personnel interviews could reveal discrepancies between public claims and internal practices.
Concerns Over Evaluation Integrity
Critics warn that models can learn to detect evaluations and behave differently during tests. John Steidley, head of strategy at Palisade Research, compared this risk to Volkswagen’s emissions-testing scandal, where systems performed differently under test conditions. Evaluators also highlighted persistent constraints on time: OpenAI provided METR and Redwood roughly a week on-site during the Hugging Face incident investigation, while Apollo Research reported only three days to test GPT-6 Astra, limiting confidence in those assessments.
Contractual Limitations and Independence
Evaluators reported that contractual terms often treat them like ordinary vendors, imposing restrictive NDAs and publication controls. Gleave mentioned that FAR.AI had to turn down contracts that would have compromised its independence. Henry Papadatos, executive director of Safer AI, pointed out that voluntary measures depend on company goodwill and called for regulation to ensure consistent, enforceable oversight.
Industry Response and Emerging Legislation
Some major labs, including Meta, SpaceXAI, and Google DeepMind, have not committed to embedding evaluators. Instead, DeepMind CEO Demis Hassabis has proposed an industry standards body for independent testing. Legislative measures are emerging: California’s SB 53 and SB 813 create reporting and verification frameworks for frontier developers, while the EU AI Act mandates evaluations, adversarial testing, and incident reporting. However, neither Anthropic nor OpenAI has clarified which evaluators they will embed, when this will occur, how many will be involved, what systems will be accessible, or what disclosure rules will apply.
Original source: TechCrunch AI