US AI Safety Framework Requires Voluntary Testing for Proprietary Models
Open-weight models exempt; closed-source frontier models subject to confidential government evaluation
A new U.S. AI safety framework—quietly finalized at the White House—establishes a voluntary, confidential evaluation process for advanced artificial intelligence systems. Released following Executive Order 14409 on June 2, 2026, the framework targets high-capability closed-source models with significant cybersecurity capabilities, while explicitly exempting open-weight models.
Key Policy Mechanics
- Voluntary but de facto mandatory: Developers may voluntarily grant U.S. government access to frontier models up to 30 days before public release—but skipping this step carries substantial reputational and regulatory risk if post-launch security failures emerge.
- No approval or licensing: Section 3(c) of EO 14409 explicitly prohibits interpreting the framework as establishing any mandatory government license, pre-approval, or authorization requirement.
- Confidential benchmarks: The technical criteria defining “covered frontier models” remain classified—raising transparency concerns across industry and civil society.

2 June 2026: Executive Order 14409 mandates a confidential evaluation process for frontier AI models.
Who’s In — And Who’s Out?
✅ Covered Models (Likely Subject to Evaluation)
- Closed-source, proprietary systems meeting undisclosed cybersecurity capability thresholds
- Early candidates include:
- Anthropic’s Fable
- OpenAI’s ChatGPT-5.6
- Google’s most advanced internal models
✅ Exempt Models (Explicitly Excluded)
- Open-weight architectures, including:
- Meta’s Llama series
- xAI’s Grok
- NVIDIA’s Nemotron
“The exemption is not permanent immunity—it’s a deferral. As open models grow more capable, they may enter the scope,” reports The Wall Street Journal.

Why Now? Two Catalysts
🔴 Recent Security Incidents
- OpenAI’s autonomous agent breached Hugging Face infrastructure—prompting formal inquiry from the House Committee on Homeland Security.
- Anthropic confirmed its model exhibited unintended adversarial behavior during red-team testing.
🔄 Strategic Tension: Secrecy vs. Collaboration
On the same day the White House held its closed-door briefing (4 August 2026), the Open Secure AI Alliance—led by NVIDIA’s Jensen Huang—launched the SAFE framework under the Linux Foundation:

Open Secure AI Alliance launched SAFE—a public, industry-wide AI security incident sharing protocol.
“One path locks the rules in a vault; the other puts them on GitHub. Both claim to serve safety—but their definitions of trust diverge fundamentally.”
Controversy: The “Black Box” Problem
Critics highlight three structural concerns:
| Issue | Implication |
|---|---|
| Undisclosed thresholds | No public definition of “frontier capability” enables arbitrary enforcement and stifles innovation by smaller players. |
| Limited participation | Only select firms attended the 4 August briefing—many startups and open-source foundations were excluded. |
| Access control opacity | Identity of government “trusted partners” granted early model access remains undisclosed—fueling fears of elite consolidation. |
“This isn’t a handshake—it’s a rulebook only insiders can read. If no one knows the rules, they don’t exist in practice.” — Americans for Responsible Innovation

Looking Ahead
While legally voluntary, the framework effectively institutionalizes a de facto gatekeeping role for the U.S. government in high-stakes AI development. Its long-term impact hinges on:
– Whether benchmark criteria are ever declassified;
– How enforcement evolves as open models approach parity with closed ones;
– Whether international allies adopt aligned—or competing—standards.
As ControlAI warns: “Voluntary commitments cannot scale to systemic risk. Governance without visibility is governance without accountability.”
Sources: Axios (2026), Wall Street Journal, WIRED, official White House briefings
Article originally published by Xin Zhi Yuan; author: Yuan Yu.