IA · 21 July 2026 · 3 min read
The Invisible Token Crisis: US Army Runs Out of AI Budget While Google Designs New Silicon
In brief: In May 2026, the US Army CIO announced unlimited AI tokens via the Ask Sage platform. By mid-June, the pool was exhausted, forcing leadership to re-impose strict limits. This incident highlights the critical need for cost governance and architectural control in enterprise AI deployments.
by Team Mocchi's
The Illusion of "Unlimited Tokens" in the US Military
In May 2026, the Pentagon's tech leadership celebrated a major milestone: nearly half of the Department of Defense's (DOD) 3.5 million employees were actively using generative AI tools. Riding this wave of enthusiasm, the US Army CIO announced "unlimited tokens" for Ask Sage, an enterprise AI platform accredited for handling Controlled Unclassified Information. However, as revealed by a recent WIRED investigation, this illusion of abundance lasted less than a month.
By mid-June, employees at the Army’s Combat Capabilities Development Command (DEVCOM) received an urgent internal email: the token pool was completely exhausted, forcing leadership to re-impose strict usage limits. The push to drive AI adoption at all costs directly fueled this financial mishap. Employees had been automatically allocated 200,000 monthly tokens, with automatic top-ups. Those who did not use their quota were even prompted via automated emails to run more queries for mundane administrative tasks, such as reclassifying personnel descriptions. Consequently, the annual 100-million-token enterprise allocation vanished in just a few weeks.
Exploding Costs and the Need for Optimized Silicon
The DEVCOM incident is not an isolated blunder; it is the tip of an iceberg representing a broader economic sustainability crisis in generative AI. While token consumption might seem manageable for routine office work, it scales exponentially under operational demands. For instance, during the military exercise Operation Epic Fury, the Pentagon burned through approximately 20 billion tokens per day, according to defense sources cited by WIRED.
Every generated or processed token carries a precise computational, environmental, and financial cost. Integrating large language model (LLM) APIs without middle-tier architectural controls exposes organizations to unpredictable, runaway expenses. In response, the tech industry is shifting its focus from raw model scale to hardware efficiency and custom silicon design.
Google’s Countermove: The "Frozen v2" Chip
This intense pressure to lower inference costs is the driving force behind Alphabet's latest hardware venture. As reported by TechCrunch (citing details originally published by The Information), Google is secretly developing a custom server chip codenamed "Frozen v2."
Slated for a 2028 release, the chip aims to dramatically improve computational efficiency for the Gemini model family. Preliminary reports suggest that "Frozen v2" could be between 6 and 10 times more efficient than Google’s current Tensor Processing Units (TPUs), measured in tokens generated per unit of power.
While Google did not explicitly confirm the "Frozen v2" project, it told TechCrunch that its full-stack hardware-software co-design approach is critical to optimizing real-world workloads. This race for custom silicon—which includes OpenAI’s "Jalapeño" chip and Anthropic's partnership with Samsung—underlines a major industry shift: power efficiency and reducing the cost per token have become the ultimate competitive battlefield.
Mocchi's take
The US Army's experience proves that "unlimited tokens" in generative AI is a marketing illusion that can quickly derail any corporate budget. For Italian businesses embarking on their AI journey, the takeaway is clear: AI adoption requires rigorous architectural governance. Simply purchasing licenses or exposing API endpoints is a recipe for financial volatility. Companies must implement middleware layers—such as semantic caching, rate limiters, and token budget policies—and embrace hybrid architectures that leverage smaller, cost-effective open-source models for routine tasks. By carefully engineering token flows, businesses can harness the power of AI without exposing themselves to unpredictable financial strain.