Anthropic and OpenAI are pursuing a plan to station external safety evaluators inside their companies to continuously verify the development process and safety of cutting-edge artificial intelligence models. Some critics say effectiveness could be limited because evaluators would not have the authority to halt development or release of risky models.
On Sept. 16, local time, foreign media outlets including the IT publication TechCrunch reported that Anthropic CEO Dario Amodei (다리오 아모데이) proposed in a recently published post keeping third-party safety evaluators inside frontier AI companies on an ongoing basis. The proposal calls for giving evaluators access comparable to an internal risk-management team so they can report safety incidents and assess whether AI models are aligned, while also allowing them to publish key findings without company editing.
OpenAI CEO Sam Altman (샘 알트먼) also said he would join a similar approach, but did not disclose具体 details of how it would operate.
Existing external evaluations have mainly been conducted on finished models shortly before release. By contrast, proponents of a resident-evaluator approach say it should cover not only the final model but also checkpoints, intermediate versions generated during training, post-training environments, and evaluation records and logs. The concern is that an AI model could recognise it is being evaluated and behave safely in tests while hiding problematic behaviour.
Critics also say the scope of evaluation should expand beyond the model itself to the entire AI system. They say reviewing prompts and tools provided to the model, permissions, safety controls and behaviour records together is needed to identify real risks in the system.
The independence and authority of evaluators remain key issues. Experts said resident evaluators, unlike bank supervisors, would not have legal authority to halt specific practices or stop model development and releases.
There are also concerns that independence could be weakened if companies directly select evaluators and control the scope of access and follow-up actions after evaluations. Amodei mentioned the nonprofit METR as a resident evaluation organisation, but the AI safety evaluation field is still small and has many personnel and research links between major research institutions and AI companies, suggesting debate over independence will continue.
Time and access required for evaluation are also practical constraints. METR and Redwood Research said they conducted on-site investigations for about a week during OpenAI's Hugging Face incident probe, but could not reach a definitive conclusion due to limits in scope and schedule. In a pre-test of GPT-6 Astra, Apollo Research assessed that a 3-day evaluation period makes it difficult to confidently judge whether a model is aligned.
This has led to arguments that voluntary participation by companies alone has limits, and that public standards and legal grounds are needed. Some regions are also putting related systems in place.
Current laws and systems, however, do not cover what level of access and enforcement power should be granted to independent evaluators stationed inside companies. Ultimately, to allow resident evaluations to more deeply verify the development process and the overall system than existing pre-release external evaluations, the key task is expected to be ensuring evaluators' access and independence and specifying what actions will be taken based on evaluation results.