The Token Factory: When AI Compute Becomes a Utility
At WAIC 2026, the 'Token Factory' emerged as one of the hottest concepts. From Jensen Huang's metaphor to real-world deployments, AI compute is shifting from hoarding GPUs to pay-as-you-go infrastructure.
If WAIC 2025 was about the large model arms race, WAIC 2026 belonged to a far more pragmatic concept: the Token Factory.
In June, NVIDIA CEO Jensen Huang likened modern AI data centers to "factories that produce tokens." Every token, he argued, could be transformed into code, answers, or design plans — directly generating commercial profit. The metaphor caught fire instantly.
A month later, at WAIC 2026 in Shanghai, "Token Factory" was everywhere. Moore Threads, Infinigence AI, StoneTech, and others all presented their takes on token production, optimization, and scheduling. Softtek prominently displayed the numbers for its Beijing No.1 Token Factory: a planned daily capacity of 1.4 trillion tokens. UCloud took the vision further, articulating a complete infrastructure stack spanning "Token Production → Distribution → Invocation → Governance → Operations."
The Real Pain Point: Sub-30% Utilization
The Token Factory buzz didn't come from nowhere. It emerged from a widespread embarrassment in China's AI infrastructure landscape.
According to the National Data Administration, China's daily AI token invocation volume surged from roughly 100 billion in early 2024 to over 140 trillion by March 2026 — a 1,400x increase in just eighteen months. Yet, during the same period, the average rack utilization rate at domestic intelligent computing centers hovered between 20% and 30%. Some enterprise facilities fell below 10%.
On one side, GPU rental prices remained sky-high due to chip shortages. On the other, expensive GPUs sat idle in server rooms. "Owning compute doesn't mean owning productivity" became the warning bell hanging over the entire industry.
The traditional compute rental model suffers from a structural flaw: enterprises purchase fixed hardware units. During peak periods, they're starved for capacity; during lulls, resources sit idle. Meanwhile, the most labor-intensive tasks — model deployment, performance tuning, energy management, daily operations — don't transfer with the rental contract. For small and medium-sized businesses, building in-house compute essentially locks up capital that should be fueling product development and market expansion.
How Token Factories Solve This
The core logic is simple: customers stop purchasing GPUs or renting racks. Instead, they buy the final output — tokens — produced by AI models processing information. It's the same principle as electricity: you don't build your own power plant; you pay for what you use.
Behind this simplicity lies a complete infrastructure chain: stable power supply, reliable networking, high-grade data centers, intelligent scheduling platforms, and 24/7 uninterrupted operations. A failure at any link means service interruption.
Take the Wuxiang Cloud Valley Intelligent Computing Center in Nanning, Guangxi, operated by Runjian Co., Ltd. (SZ: 002929). This Token Factory produces approximately 200 billion tokens per hour. Customers connect via API, invoke on demand, and pay per token. Through deep end-to-end optimization, its cluster utilization rate reaches roughly 57% — far above the industry average of under 30% — with first-token response times improved by over 30%.
Runjian is not a traditional tech company. It's an infrastructure operator spanning telecommunications, energy, and computing networks, with over two decades of experience in digital operations and maintenance. Its competitive advantage in the Token Factory race isn't "who has more GPUs," but cross-domain organizational capability. This highlights a key insight: Token Factories aren't exclusive to tech companies. Infrastructure players may actually hold the edge.
From Compute Hub to AI Productivity Store
Infinigence AI CEO Xia Lixue presented an even more ambitious framework at WAIC — the "Front Store, Back Factory, One Hub" strategy:
- Compute Distribution Hub: Aggregating scattered, heterogeneous compute resources, scheduling them elastically, and distributing on demand. Its cross-domain training system for heterogeneous clusters has already reached over 37,000 PFLOPS of compute across China, covering 16 chip types.
- Token Factory: The MaaS (Model-as-a-Service) platform layer, interfacing downward with chips and upward with models and applications.
- AI Productivity Store: Industry-facing solution layer, turning tokens into real business value.
This three-tier architecture reveals the full picture: the bottom layer handles resource aggregation and scheduling, the middle layer manages model inference and serving, and the top layer delivers industry-specific applications. Between them, agent swarms automate resource orchestration and optimize token production.
Infrastructure-ization: AI's "Grid Moment"
Looking back through computing history, each paradigm shift has been accompanied by infrastructure reinvention. The mainframe era centralized compute into single machines. The PC internet era distributed it to personal devices. The cloud era re-centralized it into data centers, but the unit of compute was the virtual machine or container.
Now, the foundational unit of the AI era is becoming the token.
This isn't merely a change in measurement units. When compute is accessed like a utility — metered, priced, and available on demand — the entire industry's cost structure and innovation threshold transform. Entrepreneurs no longer need to spend millions stockpiling GPUs. A startup can prototype an AI application MVP with a few hundred yuan worth of tokens.
Of course, Token Factories still face significant challenges: the cost of adapting heterogeneous chips, latency in cross-region compute scheduling, compatibility issues from rapid model iteration, and the standardization of token pricing. But the direction is clear. The next decade of AI infrastructure won't be about who stacks more GPUs. It will be about who can turn compute into productivity — reliably, efficiently, and sustainably.