Member Login Become a Member
Advertisement

Allocating AI Priority: A Strategic Token Reserve?

  |  
09.24.2026 at 06:00am
Allocating AI Priority: A Strategic Token Reserve? Image

Abstract:

The current artificial intelligence market lacks a way to distinguish and prioritize strategically-beneficial uses of leading AI models. In times of war and crisis, adversaries could clog the queue of user requests, leading to delays in decision making and other suboptimal outcomes. To better manage AI resources, a priority rating scheme or a “strategic token reserve”  does have prior legal precedent and could provide a way to better manage these risks.


Somewhere in the United States tonight, a twelve-year-old boy might be using generative AI to make a video of a talking cherry fist-fighting with George Washington. Somewhere else, a local emergency manager might be trying to get the very same AI model to summarize four hundred pages of hazard mitigation filings before the hurricane season starts in earnest. Both requests enter the same queue, and both are served by the model on the same terms. But the emergency manager has no clue that Cherry v. George Washington might be ahead in the line for inference power.

This hypothetical demonstrates something concerning about the architecture of American artificial intelligence capacity. The market currently assumes that inference and compute are abundant enough so that the allocation of AI capabilities within society does not really matter. After all, the cost of the tokens themselves has fallen dramatically, even though overall spending on AI continues to rise. But what happens in a crisis when the demand for access to the best models far outstrips the need for AI capability? As more organizations and agencies (including critical infrastructure) integrate frontier models into their operations, policymakers should consider whether the market can effectively allocate model access during emergencies when demand may spike, given known supply side constraints with existing hardware. Currently, major AI providers appear to use rate limiting based on use of “tokens” or “requests” per interval of time as well as quota or tier-based pricing models.

Given these current allocation models, the United States should consider establishing either a strategic token reserve or using existing Defense Production Act authorities that allow the federal government to place priority-rated orders on goods and services. Doing so will enable government agencies to ensure that critical or emergency uses of frontier models receive sufficient priority during an incident, while avoiding broader disruptions to the AI marketplace.

Allocating AI as a Key Resource

To be clear, the problem raised by the Cherry  hypothetical is not a market “failure.” The existing market approach for AI inference allocates existing capability to whoever can pay the most, which is how any economic system ordinarily functions. Of course, the difficulty with this approach to pricing is that willingness to pay is often a poor proxy for strategic value. A hedge fund’s backtesting budget might dwarf the entire information technology spend of a public water utility, and yet, only one of these has a pressing need to identify vulnerabilities in their software that might lead to a loss of clean drinking water. But the current market setup will likely cater to the fund’s more extensive demand for tokens, and there might not be any price signal that conveys that one of these customers keeps fresh water flowing to the local municipality and the other does not.

There is good reason to believe that we cannot build our way out of this problem. While tech companies are building more data centers and bringing more compute power online every day, developing a prioritization framework would be good for economic and national security interest. Much like other strategic resources like semiconductor chips or petroleum, building additional inference capacity cannot be quickly scaled: Data centers can take years to permit, construct, and connect to the power grid. The total amount of inference capacity remains fixed in the short term, even though demand will likely spike as more enterprises develop novel use cases for AI models in their own operational contexts.

Because AI models are an informational resource and a decision-making tool, there are especially good reasons to set aside capacity or priority access in the event of a crisis. The governor facing a Category 5 hurricane or the CEO facing a cyber intrusion only has a few hours in which AI tools can actually make a difference. Once the decision window has passed, availability of the models is beside the point.

Priority ratings exist for exactly these conditions. A party bus does not need gasoline as much as a squad car, and under normal circumstances, allocation is not necessary. Which is why central planning of artificial intelligence availability is unlikely to be beneficial from a government or industry viewpoint. But some applications of AI capabilities are inevitably more socially-useful than others. Setting aside a strategically-useful amount of tokens (to oversimplify, tokens are the “currency” of AI models and correspond to small units of data process) could help ensure that inference will be available when it is most needed.

Adversaries Could “Drink Your AI Milkshake”

Beyond properly allocating model access during a crisis, having a token reserve can forestall an adversary from drawing down model availability. Consider a scenario where a threat actor wants to degrade decision-making capability during a particular crisis window. Ordinarily, a hostile entity could pursue options like cyber intrusions, sabotaging a data center, and disrupting power supplies. But these all require sufficiently good tradecraft to carry out, in addition to the usual risks of interdiction, attribution, and escalation.

But if an adversary opened commercial accounts to use leading frontier models, they could clog up these models with long-context, reasoning-heavy workloads that are designed to be computationally expensive. This would not be obvious junk queries a la denial of service attacks, or large floods of badly-formed input data that existing controls might be able to catch. Rather, this might look like complex traffic that is indistinguishable from bona fide enterprise-level customers.

This is not speculative. For example, the Open Worldwide Application Security Project elevated “unbounded consumption” to the top-ten list of risks for large language model applications. Furthermore, the National Institute of Standards and Technology includes energy-latency attacks in its report on adversarial machine learning taxonomy. Recent work by researchers at leading universities shows how carefully constructed and lengthy prompts to AI models can force the models into “pathologically long reasoning traces” that imposes outsized computational costs on model providers. Further, power providers are now factoring computational load requirements into their reliability and generation planning.

As a result, an adversary could possibly build a way to degrade model performance or availability only by infusing capital: no exploit to use, no vulnerability to discover, only money and patience to degrade the availability of the model, just to continually re-insert themselves in the queue while strategically-important tasks wait. And the model provider may not have any commercial incentive to probe at increased demand, especially if it is from a high-paying, high-volume set of customers.

From a legal perspective, this kind of adversarial move would fall into the grey zone. The adversary in this instance would have fully authorized, paid-for access. The Computer Fraud and Abuse Act does not apply when access has already been authorized, at most, the adversary user may face some terms of service problems. authorized access, obtained legitimately and paid for in full. This is a kind of hybrid attack that does not register as illegal, let alone hostile, constraining possible response options. In an environment where the Department of Justice has intervened in a Clean Air Act lawsuit trying to block a data center for xAI on grounds of national security, there is a need to properly allocate inference resources.

Existing Legal Solutions

As mentioned above, creating a strategic token reserve (like the Strategic Petroleum Reserve) or using priority-rating authorities in the Defense Production Act might solve the problem. Title I of the Defense Production Act authorizes the President to require companies to fulfill certain priority-rated contracts and to allocate materials when necessary for national defense, which includes critical infrastructure protection and restoration. Executive Order 13603 already distributes these kinds of priority-setting authorities to federal departments and agencies (of note, the federal government already has priority access for telecommunications). Using allocation authority, to set aside a strategically-beneficial number of tokens, could ensure immediate access for critical queries to frontier models: the cyber defenders find the APTs, while goofy videos wait in line.

There could be market nuances that have not yet been considered, or externalities that might make the cure worse than the disease. Luckily, the Defense Production Act offers a way for companies to enter into voluntary agreements, allowing them to discuss and coordinate action with the federal government while avoiding antitrust liability.

This might enable frontier model providers to coordinate on holding reserve capacity in the interests of national defense, which normally would raise several legal and policy considerations. Furthermore, these kinds of flexibilities would enable industry to develop solutions in collaboration with government, thereby supporting economic and security interests at the same time, avoiding more intrusive “redistribution” policy options. The AI Action Plan already contemplates arrangements to “codify priority access” for the Department of War during a national emergency, but having a strategic token reserve would enable other key players in the national security environment, such as private sector network defenders, to have sufficient access to the leading AI models during a crisis. Of course, any program would need to allocate or prioritize at machine-speed in order to be useful.

Building Capacity for the Future of AI

Building a Strategic Token Reserve and using existing legal authorities for allocating AI resources, could drive down national security risk and preserve America’s leading edge in this transformative technology. Extending priority access to other government agencies and critical infrastructure can help defend America from the mutating cyber threat environment, without disrupting the innovative AI market. As AI becomes increasingly important across all sectors of the economy, using these legal tools and building a reserve may be incredibly important in the middle of a crisis.

About The Author

  • Terence Check

    Terence Check is currently Deputy Chief Counsel (A) at the Cybersecurity and Infrastructure Security Agency and is a non-resident fellow at the U.S. Air Force Academy’s Institute for Future Conflict. He teaches national security and cybersecurity policy at the Ohio State University and Cleveland State University. All opinions and statements are his own and do not reflect the official views of any of these institutions.

    View all posts

Article Discussion:

0 0 votes
Article Rating
Subscribe
Notify of
0 Comments
Oldest
Newest Most Voted