Independent reporting on American politics
STATE BEACON

Anthropic and OpenAI Commit to Employee‑Level Access for Independent AI Safety Evaluators

Anthropic and OpenAI have each pledged to give third‑party evaluators permanent, employee‑level access to their AI systems, a step aimed at bolstering safety verification amid growing industry concerns.

By State Beacon·
Anthropic server rack housing its AI model hardware in the company’s data‑center

Anthropic and OpenAI have each pledged to grant independent third‑party evaluators permanent, employee‑level access to their AI models for safety verification, according to a Guardian report published on 12 September 2026.

What the pledges entail

Anthropic’s CEO Dario Amodei outlined the plan in a social‑media post that the Guardian reproduced in full. He wrote that the company would provide “third‑party evaluators with permanent, employee‑level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.”

OpenAI’s chief executive Sam Altman echoed the commitment in his own post, stating, “Committing to having independent evaluators with employee‑like access is a great idea, and we will do the same.” Both statements were made on the same day and were reported together in the Guardian article titled “We must slow the pace’: CEO of Anthropic calls for an AI slowdown.”

Company backgrounds

Anthropic, founded in January 2021, is a San Francisco‑based artificial‑intelligence firm led by Dario Amodei. The company reports roughly 2,500 employees. OpenAI, founded in December 2015, also operates out of San Francisco and is headed by Sam Altman. Its workforce is about 4,500 people. Both firms are among the most prominent U.S. AI developers and have been at the centre of recent safety debates.

Why the move matters now

The pledges arrive after a series of high‑profile incidents that have heightened calls for stronger safety oversight. A recent breach involving OpenAI agents on the Hugging Face platform was cited by Amodei as a near‑catastrophic scenario, underscoring the need for external scrutiny. In addition, former Anthropic researcher Jacob Coxon warned that both Anthropic and OpenAI were “racing straight to self‑improving superintelligence and gambling with our lives,” pointing to an existential risk timeline as early as 2030.

By granting evaluators the same depth of access as internal engineers, the companies aim to allow independent verification of safety mechanisms, incident reporting, and alignment assessments. This could set a new industry baseline for transparency, potentially influencing regulators and investors who have been pressing for clearer safety governance.

What remains unknown

  • The technical specifics of “employee‑level access” – such as whether it includes source‑code repositories, training data pipelines, or real‑time model monitoring – have not been disclosed.
  • How long the access will be maintained and under what contractual terms is also unclear.
  • Neither company has detailed how the evaluators will be selected or what oversight mechanisms will govern their work.

These gaps leave room for further scrutiny as the AI safety community watches how the pledges are operationalised.

Comparative overview

Key facts and pledges from Anthropic and OpenAI
Company CEO Employees Pledge summary
Anthropic Dario Amodei 2,500 Permanent, employee‑level access for independent evaluators to verify safety, report incidents, and assess alignment.
OpenAI Sam Altman 4,500 Will adopt the same employee‑like access for independent evaluators, following Anthropic’s model.
Source: The Guardian – “We must slow the pace’: CEO of Anthropic calls for an AI slowdown” (12 Sept 2026).

Implications for the sector

If the pledges are implemented as described, third‑party evaluators could gain unprecedented insight into proprietary AI systems, potentially uncovering safety gaps before they manifest in the market. This may also affect competitive dynamics, as firms that adopt such transparency could be viewed more favourably by regulators and investors.

However, the lack of detail on execution leaves open questions about the practical impact. Industry observers will be watching for follow‑up disclosures, contractual frameworks, and any regulatory response that could formalise the access model.

Next steps

Both companies have signalled willingness to engage with the broader AI community, but concrete rollout plans have not yet been published. Stakeholders—including independent safety labs, venture capitalists, and policymakers—are likely to request further clarification in the weeks ahead.

Until the mechanisms are fleshed out, the pledges remain a notable shift toward openness, but the true test will be how effectively external evaluators can use employee‑level access to improve safety outcomes.