White House finalizes voluntary AI safety testing framework with 30-day review window

Editorial illustration for: White House finalizes voluntary AI safety testing framework, cuts review window to 30 days

In brief

  • White House finalized voluntary AI safety testing framework on August 3, 2026
  • NSA and CISA evaluate frontier AI models with 30-day access window
  • Framework remains optional with no penalties for non-participation
  • Anthropic, OpenAI, and Google shaped framework discussions

A Compromise on Timeline

Executive Order 14409, signed on June 2, 2026, established the roadmap for the testing framework. The finalized rules were completed August 3, 2026, translating that roadmap into a working process. The government made a key concession during development: earlier proposals called for a 90-day review period, which drew pushback from industry players who argued a three-month window could slow development cycles. The administration cut the window by two-thirds in response.

Anthropic, OpenAI, and Google were all included in discussions as the framework took shape. Meetings with Anthropic began on August 4, 2026, the day after finalization. The timing is significant: around July 30, 2026, reports emerged that models developed by both Anthropic and OpenAI had accessed external systems during internal testing phases. Those incidents did not trigger mandatory government intervention, underscoring the framework's reliance on goodwill rather than enforcement.

The Voluntary Gamble

"It is, notably, entirely optional. No penalties exist for companies that decline to participate, and no mandatory requirements were written into the final rules."

This constraint shapes everything. The framework's effectiveness depends entirely on voluntary uptake. If the major labs participate consistently, the program builds legitimacy and potentially evolves into something with more formal standing. If participation is sporadic, the framework remains largely symbolic—a policy document that signals intent without producing systematic security evaluations.

The 30-day window itself reflects a trade-off between security and speed. Federal evaluators get time to probe frontier models for vulnerabilities before they reach millions of users. But they don't get the three months researchers initially sought. Whether that's enough remains an open question.