The AI Evaluator Forum has introduced AEF-1, establishing a binding operational framework for independent third-party safety audits. Formally titled Minimum Operating Conditions for Independent Third Party AI Evaluations, the document obligates participating organizations to adhere to uniform procedural standards. OpenAI, Anthropic, and xAI have jointly committed to the framework, bringing the three dominant United States frontier labs under a shared umbrella. This agreement marks the first time these major competitors have converged on a formal, mutually accepted audit architecture for their most capable systems. The move directly addresses mounting scrutiny regarding how advanced models are tested prior to release.
The joint agreement follows a notable commitment made by Anthropic Chief Executive Officer Dario Amodei. Amodei pledged to grant external safety evaluators continuous, employee-like on-site access to the company's offices. Under this operational policy, third-party auditors receive physical entry badges, standardized corporate laptops, and comprehensive access to internal tooling. OpenAI matched Anthropic's pledge shortly afterward, agreeing to provide independent teams with equivalent operational permissions. This reciprocal willingness transformed isolated corporate promises into a genuine inter-organizational industry standard.
AEF-1 is structured to prevent external evaluations from devolving into superficial public relations exercises. The framework establishes strict operational criteria to eliminate conflicts of interest, including mandatory disclosure requirements for all underlying funding streams. Evaluators are explicitly barred from maintaining financial dependencies on the frontier labs whose systems they assess. By establishing transparent funding registries, the forum aims to guarantee that evaluations remain objective even when assessing high-stakes capabilities. Protecting the integrity of independent scientific inquiry forms the core tenet of the protocol.
On an operational level, AEF-1 establishes legally viable pathways for conducting deep audits across model architectures and training pipelines. Specialized evaluation organizations such as METR and Transluce now possess clear contractual frameworks for their testing protocols. Rather than interacting with black-box interfaces over external networks, auditors can inspect model checkpoints within native development environments. This level of access eliminates bureaucratic hurdles and repetitive negotiations regarding permission levels between safety researchers and developers. Consequently, evaluation teams can systematically probe for emergent vulnerabilities prior to public deployment.
The organizational structure of AEF-1 intentionally mirrors the rigorous regulatory inspection models practiced throughout the banking and financial sector. Much like bank inspectors, accredited AI evaluators gain expansive insider visibility into operational pipelines while protecting critical intellectual property through established nondisclosure mechanisms. This strategic alignment represents an important pivot away from voluntary corporate declarations toward institutionalized scrutiny. Frontier labs are effectively assembling operational compliance architecture before statutory regulators mandate prescriptive legal frameworks. The approach demonstrates a proactive effort to formalize technical risk evaluation across frontier developments.
The shared participation of xAI alongside Anthropic and OpenAI ensures comprehensive coverage across the leading private foundation model developers in the United States. Historically, external evaluation organizations operated under fragmented, ad hoc agreements characterized by uneven access to internal weights and compute environments. The adoption of AEF-1 introduces a consistent, reproducible baseline for comparative empirical testing across rival laboratories. This framework sets a high bar for forthcoming generations of frontier models and will inevitably serve as a reference point for future auditing practices. Furthermore, it offers policymakers a concrete model for enforceable governance across high-risk artificial intelligence technologies.

