A 60-Day Clock Ran Out at Midnight

On the morning of Tuesday, June 2, 2026, President Trump signed an executive order titled "Promoting Advanced Artificial Intelligence Innovation and Security." It got the number 14409. When it hit the Federal Register three days later it ran three pages. Inside those three pages, several clocks started at once. The longest one was 60 days, and it expires today, August 1.

Here's what that 60-day clock was supposed to produce. Two things. First, a classified benchmarking process for deciding which AI systems get designated a "covered frontier model." Second, a voluntary framework that lets developers hand the federal government a look at those models before anyone else sees them. The language in Section 3(b) is specific: developers may provide the government with access to covered frontier models "for a period of up to 30 days before they plan to release such models to other trusted partners," subject to confidentiality, cybersecurity, insider-risk, and intellectual-property protections.

So as of today, the most capable AI models built in the United States pass across a government desk before they reach the world. The National Security Agency and Commerce's Center for AI Standards and Innovation are the named reviewers, according to a July 27 report in The Information. That same report said the White House Office of the National Cyber Director had circulated a draft of the framework to OpenAI, Anthropic, and Google roughly two weeks earlier, and that the three companies made edits to it. Worth flagging: that reporting is single-sourced, and no agency or company has publicly confirmed it.

And then there's the sentence that carries the entire tension of this story. Also in Section 3: "Nothing in this section shall be construed to authorize the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models, including frontier models." The order goes out of its way to say this is not a permission regime.

Here's the deal, though. The threshold that decides who's covered is classified. The agency making the final call is the NSA. And the framework covers not just review but who else gets early access. In that setup, which lab is going to raise its hand and opt out? That's why the phrase you keep seeing attached to this thing is some version of voluntary on paper, mandatory in practice.

Who's Actually in the Room

Start with the White House, because the politics here are genuinely strange. This administration revoked Biden's AI executive order 14110 on day one in January 2025. Removing regulatory friction has been the stated theory of American AI leadership, and EO 14409 says so in its own text — the administration, it claims, "unleashed tremendous technological growth and economic investment in AI by slashing the bureaucratic constraints that the prior administration placed on America's AI developers and researchers." That same administration has now built a pre-release review pipeline for frontier models. That contradiction is the story's first turn.

Second party: CAISI. It used to be called the U.S. AI Safety Institute. In June 2025, Commerce Secretary Howard Lutnick restructured it into the Center for AI Standards and Innovation. "Safety" came out of the name and "Standards and Innovation" went in. But the scope of the work grew rather than shrank. CAISI now describes itself as "industry's primary point of contact within the U.S. government to facilitate testing and collaborative research related to harnessing and securing the potential of commercial AI systems."

Third: the labs on the receiving end. On May 5, 2026, CAISI announced new agreements with Google DeepMind, Microsoft, and xAI. It already had arrangements with OpenAI and Anthropic. That made five. CAISI Director Chris Fall framed it this way at the time: "Independent, rigorous measurement science is essential to understanding frontier AI and its national security implications." The agreements were renegotiated versions of earlier voluntary commitments, they support testing in classified environments, and — this is the part that matters technically — they cover evaluation of models with safeguards reduced or removed. CAISI wants to see the raw capability, not the product-shaped version.

Fourth, and maybe most interesting: the party that isn't there. Meta. Not among CAISI's five labs. Not on the list of companies reported to have received the draft framework. On April 8, 2026, Meta shipped Muse Spark, the first model out of Meta Superintelligence Labs — and it came with closed weights, fully proprietary. So the company that spent years as the standard-bearer for open models has locked down its own flagship while simultaneously staying off the government review table. Both moves at once.

And then there's the detail that defines the character of the whole arrangement: the NSA is holding the pen. The order directs the Secretary of the Treasury, the Secretary of War through the Director of the NSA, and the Secretary of Homeland Security through the Director of CISA to build the classified benchmarking process — with the NSA making the final determination, after consultation, on which models cross the threshold. Not a consumer protection agency. Not a competition regulator. Not even NIST, which houses CAISI. A signals intelligence agency is the gatekeeper.

What's Actually on Paper, and What Still Isn't

Break Section 3 into its parts. Subsection (a) orders a classified benchmarking process "to assess the advanced cyber capabilities of AI models" and to determine the threshold for "covered frontier model" designation. Subsection (b) orders the design of a voluntary framework with three functions: let developers work with the government to figure out whether their model qualifies; let them provide access up to 30 days before release; and let the government and the developer collaborate on choosing which "trusted partners" also get early access. Subsection (c) is the no-preclearance disclaimer.

That third function in (b) deserves more attention than it's gotten. It means a company doesn't decide alone who gets its model during the preview window — the government is a party to that decision. That is not review. That's shared control over the distribution sequence. The order calls it collaboration. Critics call it a gate. Both descriptions fit the text.

Plenty is still missing from the page. The Congressional Research Service analyzed the order in report IF13268 and landed on three gaps. One: the order never defines "covered frontier model." The classified benchmark does the defining, which means a company may not know in advance whether it's covered. Two: no funding mechanism is specified for any of the new work. Three: because the structure is voluntary, it "potentially creates coverage gaps if major developers decline to participate." CRS framed the choice for Congress plainly — keep the voluntary approach, or legislate mandatory federal pre-release testing.

Line the whole lineage up and the shape gets clearer.

Item White House Voluntary Commitments (Jul 2023) Executive Order 14110 (Oct 2023) EO 14409 framework (Aug 2026)
Legal basis Written pledges, no legal force Invoked the Defense Production Act Executive order, voluntary participation
Who's covered Signatory's own judgment Published compute threshold (10^26 operations) Classified benchmark, NSA decides
What government gets Promise to share red-team results Training plans and red-team reports Access to the model itself
Timing After the fact, ad hoc Reports before and after training Up to 30 days pre-release
Lead agency White House Commerce NSA + Commerce CAISI
Enforceable No Reporting obligation No (explicitly disclaimed)
Outcome Faded Revoked Jan 2025 Starts today

One more number worth knowing. According to analysis published by the Council on Foreign Relations, the draft Trump was scheduled to sign on May 21 set the review window at 90 days. He pulled back from signing that day, citing competition with China: "I don't want to do anything that's going to get in the way of that lead." Twelve days later the document reappeared with 90 days cut to 30. The compromise between national security review and shipping speed is sitting right there, measured in the 60 days that got deleted.

It also helps to know what CAISI has already been doing, because that's the practical weight behind the framework. CAISI and its predecessor conducted pre-deployment national security tests on OpenAI's o1, o3-mini, o3, o4-mini, GPT-4.5 and GPT-5; Anthropic's Sonnet 3.5, 3.7 and 4 and Opus 4; and xAI's Grok 3. Anthropic published its own account of the work on September 12, 2025. U.S. CAISI and UK AISI red teams got access to Constitutional Classifiers running on Claude Opus 4 and 4.1 before deployment, and they found five categories of weakness: prompt injections that used falsely claimed human-review annotations to slip past detection, universal jailbreaks built on sophisticated encoding, cipher-based attacks using character substitution, input/output obfuscation that split harmful strings into benign-looking fragments, and automated systems that iteratively refined their own attacks. Anthropic's own conclusion was that "giving government red-teamers deeper access to our systems enables more sophisticated vulnerability discovery." That's the case for the framework, made by one of the companies being reviewed.

What Each Side Is Getting Out of This

The administration gets two things. One is substantive. If NSA and CISA can characterize a frontier model's offensive cyber capability before it's public, defenders get lead time before that capability is loose. Matthew Ferren, writing in the CFR analysis, called this "a cybersecurity window of opportunity." The other benefit is political. An administration whose brand is deregulation needs an answer when it gets accused of ignoring AI risk, and a voluntary process that demonstrably functions is a much better answer than nothing.

The labs' calculus is subtler. On the surface this is pure cost. Handing an unreleased model to the government thirty days early means absorbing leak risk, schedule risk, and competitive exposure all at once. But there are three things on the other side of the ledger.

First, regulatory insulation. As the CFR piece put it, labs have reason to participate "if only to forestall more invasive regulation later." Every quarter a voluntary process runs without incident is a quarter the mandatory-testing bills lose momentum.

Second, a moat. The set of companies that can pass a government review conducted in classified environments — with the security infrastructure, cleared personnel, and legal machinery that requires — is very small. If this becomes the norm for shipping a frontier model, the qualification to be a frontier lab narrows to roughly the five companies already in the program. Nobody will say this out loud, and everybody has done the math.

Third, federal procurement. "Evaluated by CAISI" is a sales asset inside government buying processes. Worth keeping straight, though: a CAISI evaluation is a government national security review, not a public safety certification. Those are different claims, and the distinction matters when you see the phrase in a procurement document.

CAISI itself gets a reason to exist. An organization that dropped "safety" from its name needs to demonstrate national security value to survive budget pressure, and this framework institutionalizes exactly that. CAISI has been building the case with foreign model evaluations too. Its DeepSeek V4 Pro assessment, published May 1, 2026, tested across cybersecurity, software engineering, natural sciences, abstract reasoning, and mathematics using nine benchmarks, and concluded that the model is "the most capable PRC AI model evaluated by CAISI to date" while lagging "behind the frontier by about 8 months." On economics, it beat GPT-5.4 mini on five of seven benchmarks, at prices ranging from 53% cheaper to 41% more expensive. The methodological point is that CAISI deliberately used held-out benchmarks so it could compare its own measurements against the developer's published claims.

The NSA gets access, which is its own category of benefit. Very few institutions on earth can inspect a commercial frontier model with its safeguards stripped, before release. That's defensive intelligence. It's also a live, exclusive map of what the most capable American AI can actually do, held by a signals intelligence agency. Both readings are true. Which one weighs more is where reasonable people split.

This Has Been Tried Before — One Collapsed, One Got Deleted

The first comparison is the July 2023 White House voluntary commitments. Seven companies — Amazon, Anthropic, Google, Inflection, Meta, Microsoft, and OpenAI — signed pledges on safety, security, and trust at the White House, and eight more joined two months later. No legal force, no verification mechanism, no scheduled compliance review. You know how that went. Inflection was effectively hollowed out the following year, and no institution ever systematically published whether the commitments were kept. That's what pure voluntarism looks like when it dissolves: not a scandal, just silence.

The second comparison is the opposite extreme. Biden's EO 14110, signed October 2023, invoked the Defense Production Act to require developers of dual-use foundation models trained above a compute threshold (10^26 operations) to report training plans and red-team results to Commerce. It had teeth. The threshold was public, so companies knew whether they were covered. And it lasted fifteen months. It was revoked on Trump's first day in office, and there is almost no public record of what risks the reporting regime actually caught in the interim. That's the failure mode of regulation built on executive action alone: it doesn't get repealed, it gets erased.

The third comparison is older and, honestly, the most instructive. The Crypto Wars of the 1990s. The U.S. government classified strong encryption as a munition and controlled its export, effectively running a national-security pre-approval regime over a software capability. It did not hold. PGP's source code was printed as a book and exported that way. First Amendment litigation ground the controls down. By 2000 the restrictions had been substantially relaxed. The lesson was that when a technology can propagate as open source, a government gate binds domestic commercial actors and almost nobody else.

Each of those three throws a different question at today's framework. The 2023 commitments say a process with no obligation doesn't survive. EO 14110 says an obligation with no statute doesn't survive either. The Crypto Wars say that the moment weights are downloadable, a 30-day gate is only half a gate. EO 14409's framework doesn't cleanly escape any of the three. It's voluntary, it rests on an executive order, and it has essentially nothing to say about open-weight models.

What the Rest of the Industry Is Calculating Right Now

The clearest counter-move came from the open-weights camp, and the timing was not subtle. On July 24, exactly one week before the framework deadline, a three-page open letter titled "Open Weights and American AI Leadership" went out with 25 launch signatories — Nvidia, Microsoft, Meta, Mistral, Palantir, IBM, Andreessen Horowitz, Hugging Face, Mozilla, and the Linux Foundation among them. The central ask is to "keep the frontier plural by avoiding premature restrictions on open models that stifle competition or drive innovation overseas." It argues that "openness may be one of the most important paths to AI safety and security," defends distillation as a legitimate technique deserving "targeted legal and commercial frameworks rather than sweeping restrictions," and asks for expanded compute access for startups and researchers.

The real signal in that letter was who wasn't on it. OpenAI, Anthropic, and Google were all absent from the launch roster. The three companies reportedly sitting at the government review table and the coalition signing the open-weights letter split almost exactly along the same line. OpenAI added its name afterward, and by July 30 the signatory count had passed 230 organizations. If you want a map of how the industry is currently aligned, put the two documents side by side and read the names.

Meta's position is the most calculated one on the board. It's not in the CAISI agreements, not on the reported draft circulation list, and it is a founding signatory of the open-weights letter — while shipping Muse Spark with closed weights. Open-source rhetoric in the policy fight, proprietary strategy in the product. That combination lets Meta avoid the compliance burden of pre-release review while keeping the openness narrative, and right now it's the cheapest seat in the room. It also won't survive contact with a serious open-weights restriction, which may be exactly why the letter exists.

Chinese labs are running a different calculation entirely. CAISI has already published evaluations of DeepSeek V4 Pro, Z.ai's GLM-5.2 (completed July 17), and Kimi K3 (a joint assessment with the UK AISI, July 23). So the emerging structure is asymmetric by design: pre-release review for American models, post-release capability assessment for Chinese ones. American labs wait thirty days; Chinese labs don't wait at all. That asymmetry is the core of the competitiveness argument, and it's precisely why Trump balked on May 21. The number 30 is what that pressure produced.

Congress and the states are pushing in opposite directions from each other and from the White House. Senator Josh Hawley welcomed the order but said, "I would go farther. I think we ought to enact my legislation... that would make that sort of reporting and monitoring mandatory." Brendan Steinhauser of the Alliance for Secure AI argued that "Congress must now codify the White House's EO with legislative action." Meanwhile, House members introduced the Great American AI Act, which would preempt state laws on frontier AI development for three years. And Illinois just became the first state to require independent third-party audits of frontier AI models. Federal voluntarism, federal mandates, state regulation, and federal preemption are all being pushed at the same time, by different people, into the same space.

So What Actually Changes for You

If you're a developer — the immediate change is release cadence. Frontier model launches can now slide up to thirty days, and the preview access list is something the company negotiates with the government rather than decides alone. Next time you see "we're rolling this out to select partners first," you can no longer assume that list reflects purely commercial judgment. Practically: if your roadmap assumes a new frontier model version lands on a specific date, add a month of buffer and design a fallback path to the current generation.

If you work in policy or compliance — the direct scope here is narrow, maybe five companies. The knock-on effects are not. A classified definition of "covered frontier model" means you cannot cite that term in a procurement document or contract without defining it yourself, because there's no public definition to point at. And if Illinois-style third-party audit requirements collide with federal preemption attempts, multi-state compliance design gets genuinely messy. The concrete task this week is checking whether your vendor agreements have language covering model version changes, launch delays, and access restrictions arising from government review.

If you're an investor — read it two ways. First, this structure thickens the moat around the top handful of labs, because the security and legal capacity required to operate inside classified evaluation environments is itself a barrier to entry. Second, it degrades launch predictability. A 30-day gate plus negotiated preview lists is a variable that can move revenue recognition timing. And the whole thing rests on an executive order rather than a statute, which means it carries the same fragility that erased EO 14110 in fifteen months. The right frame isn't regulatory risk. It's regulatory instability risk, which is harder to price.

If you're a regular user — you'll barely notice. New models might arrive slightly later than they otherwise would. The thing worth holding onto is the underlying fact: the U.S. government now gets to look at the most capable American AI models with their safeguards removed, before you do. The case for that is defensive lead time on cyber threats. The concern is that a signals intelligence agency gets an exclusive, continuously updated map of commercial AI capability. Both of those are real, and which one ends up mattering more will be settled by several years of operating record that nobody has yet.

🥄 Three Things You're Probably Wondering

— If it's voluntary, can't a lab just say no? Legally, yes. The order explicitly disclaims any mandatory licensing, preclearance, or permitting requirement, and there's no penalty attached to declining. In practice it's harder than it sounds: the designation threshold is classified, the NSA makes the call, and federal procurement plus political exposure both sit on the table. The Congressional Research Service flagged the same tension from the other direction, warning that the voluntary structure "potentially creates coverage gaps if major developers decline to participate."

— Does something get published today? The order required agencies to design the framework within 60 days. It did not require them to publish it. A classified benchmarking process cannot be public by definition, and how much of the voluntary framework document becomes visible hasn't been established. Almost everything known beyond the executive order text traces back to single-source reporting from people who saw a draft, so the honest answer is to watch what actually gets confirmed in writing from here.

— Is Meta's absence a real problem? Too early to call. Meta isn't in the CAISI agreements or on the reported draft list, but that doesn't mean it stays out permanently. The deeper issue is structural: open-weight models and pre-release review don't fit together. Once weights are downloadable, a 30-day window stops meaning anything. That's the exact wall 1990s encryption export controls hit, and this framework hasn't answered the question yet.

Further Reading

Numbers are as of announcement and may change.