AI Labs Back AEF-1 Standard for Independent Evaluators
A new evaluator standard backed by OpenAI, Anthropic and xAI sets minimum expectations for access, conflicts, autonomy, transparency and safe harbor.
A process standard emerges
The AI Evaluator Forum has published AEF-1, a proposed standard for independent third-party evaluations of advanced AI systems. The framework is backed by major model developers including OpenAI, Anthropic and xAI, while also involving external evaluation organizations.
AEF-1 defines five operating areas: sufficient technical access and resources; minimized conflicts of interest; analytic autonomy; transparent methods and results; and protection of sensitive information. It asks evaluators to disclose the model configurations they examined, the resources available to them, relevant conflicts, publication restrictions and any departures from the standard. The document also recommends legal safe harbor for actions conducted within an agreed evaluation scope.
Why the details matter
The standard addresses a recurring weakness in frontier-model safety work: public claims often describe what was tested without clearly explaining how much access the evaluator had or whether the developer could influence the result. AEF-1 attempts to turn those questions into an auditable checklist.
The document recommends that assessments of substantially novel systems generally receive enough time for access debugging, test design, execution, analysis and iteration. It cites 20 business days as a common minimum for many novel evaluations, though the appropriate period depends on the system and risk.
The unresolved issue
AEF-1 is voluntary, and the same industry that controls model access is helping define the conditions under which access is judged adequate. That creates a real independence question even if the checklist is useful. The standard will matter only if evaluators publish compliance and exceptions, labs grant meaningful access to near-deployment systems, and buyers or regulators begin treating those disclosures as evidence rather than marketing.
For now, its significance is institutional: it gives independent evaluators a common language for describing whether a safety result was genuinely independent, reproducible and properly resourced.