The White House announced on Monday that it has met its deadline to complete a voluntary framework for evaluating advanced AI models. Yet the administration declined to say what is actually in the framework, who has reviewed it, or when technology companies will begin using it. The unusual combination of secrecy and a firm deadline has left policymakers and AI safety advocates with more questions than answers.
"The voluntary framework outlined in the June 2nd executive order was complete by the deadline," a White House official said. "Discussions with industry about next steps are underway." The official added that the lack of a classified designation does not mean the document will be released to the public. "Just because things are unclassified that doesn't mean we are going to broadcast them to everyone."
What the framework is supposed to do
The framework is meant to give the government a structure for determining whether an advanced AI model under development would be covered by the June executive order. That order created a 30-day pre-release review window for frontier models. During that window, developers would give the federal government access to model information and evaluation results before deployment. The design assumes that some AI systems could pose national security risk, particularly in areas like cyber operations, biological research, or large-scale disinformation.
Under the framework, the government would decide which models are powerful enough to require review. It would also decide what level of testing is appropriate before a model is allowed to reach the public. The benchmarks used for that assessment are not just technical tools; they are the critical gate that separates models requiring government scrutiny from those that can move ahead freely. Without those thresholds, the entire review process lacks clear boundaries.
Classified benchmarks and hidden thresholds
The White House has confirmed that the benchmarks used to assess cyber capabilities are classified. The threshold that determines which models are covered is also classified. That information has been shared only with certain developers, and only "as appropriate," according to the official. This means that no independent researcher, journalist, or even many government officials outside a small circle know exactly what kind of AI capabilities trigger a federal review.
The decision to classify the benchmarks is striking for several reasons. First, the executive order itself was presented as a transparency measure, a way to ensure that the development of frontier AI does not outpace government oversight. Second, the framework is voluntary. Companies are not legally required to participate. But if the rules are secret, a company cannot make an informed choice about whether to opt in. It must either agree to the government's private judgment or risk being seen as uncooperative.
Industry involvement and the Tuesday meeting
OpenAI, Anthropic, and Google all provided feedback on a draft of the framework. These three companies are among the most advanced AI labs in the world, and each has significant experience deploying large language models and other frontier systems. Their involvement suggests that the government wanted early input from the companies most likely to be affected by the review process.
A White House official said the administration is engaging with "many more" industry partners beyond those three. A staff-level meeting with companies is scheduled for Tuesday to review the completed framework. The meeting is expected to explain the practical details of implementation, including how developers will submit models for evaluation and what information they will need to provide. Whether the meeting will result in any public disclosure remains unclear.
The fact that the companies were allowed to see a draft version of the framework while the public was not has drawn sharp criticism. Some observers have argued that the most powerful AI companies now know the secret thresholds, giving them an advantage over smaller competitors and open-source developers. Others have noted that the same companies also have commercial incentives to keep the details hidden, since public scrutiny of the evaluation process could complicate their release schedules.
A de facto gating mechanism
Because the benchmarks and thresholds are classified, the framework functions as a de facto gating mechanism that companies cannot publicly evaluate or challenge. The label "voluntary" applies in a narrow legal sense, but the practical effect is very different. A company that chooses not to submit its model for review could face reputational damage, investor uncertainty, or pressure from government agencies that also purchase or deploy AI systems.
Furthermore, the 30-day review period creates a hard timing constraint. Companies typically plan product launches months in advance. If the government can add an unexpected review window, that alone is enough to change release strategies. Even without a formal mandate, the threat of a delay is a powerful incentive to comply. The combination of voluntary language and unverifiable rules means that the government can apply its standards inconsistently, with limited accountability.
This is not the first time the US government has used soft power to shape AI development. In 2023, the Biden administration secured voluntary commitments from leading AI companies to submit models to external red-team testing and share safety information. Those commitments were also non-binding, but they became the industry norm within months. The current framework follows a similar pattern, but with an additional layer of secrecy that did not exist in the earlier agreements.
The Gold Eagle connection
The framework is not an isolated policy initiative. Earlier this month, the White House launched a program called Gold Eagle to coordinate AI-powered cyber defence across federal agencies. The model evaluation framework is the companion piece to Gold Eagle. Together, they are designed to address two sides of the same problem: identifying vulnerabilities and deciding which AI systems are powerful enough to require government oversight before they are released.
Gold Eagle is intended to use AI to find weaknesses in government networks, critical infrastructure, and perhaps the AI models themselves. The evaluation framework would then determine whether a model under development is subject to a 30-day preview. This pairing suggests that the government's primary concern with frontier AI is not general capability but cyber-related risk. The classified benchmarks are thought to be based on expertise from the National Security Agency and other intelligence agencies.
That lineage raises important questions. An evaluation framework built on classified intelligence criteria may be well suited to national security priorities, but it is not necessarily aligned with the broader goals of AI safety, such as avoiding bias, preventing dangerous misuse, or protecting individual privacy. The White House has not explained how the classified benchmarks relate to the range of risks outlined in the executive order, nor has it offered any mechanism for independent verification of the evaluation process.
What transparency promise?
Policymakers and AI safety advocates expected to see details of the framework once it was completed. They have not received them. The executive order promised that the administration would "strengthen transparency" around AI models, and the completion of the framework was supposed to be a milestone in that effort. Instead, the administration has offered a vague statement and a promise to discuss next steps with selected industry partners.
There is also a practical problem with keeping the thresholds secret. Without knowing what capabilities trigger review, developers of smaller models cannot plan ahead. A research lab might spend months training a model, only to discover at the end of the process that it falls under a classified threshold. The uncertainty is likely to have a chilling effect on innovation, particularly for companies that cannot afford to wait through an unexpected 30-day review.
Open-source developers face an even harder situation. They are not organized in a way that allows them to participate in a White House meeting or share a draft with legal counsel. If the government considers an open-source model to be covered by the framework, there is no clear process for requiring the developer to submit it. This could leave a blind spot at the very heart of the policy, or it could become a justification for more aggressive enforcement later.
The question that remains
The White House has said the framework is done. It has not said what is in it, who has seen it, or when it will begin affecting the AI industry. The only confirmed details are that the benchmarks are classified, the thresholds are classified, and a meeting with companies will take place on Tuesday. The question is whether a framework that nobody outside government can read, built on benchmarks nobody outside the intelligence community can see, qualifies as the transparency the executive order promised. The answer, so far, is no.