The physical supply of data center GPUs grew only two and a half times while the demand for tokens expanded a hundredfold, creating a severe bottleneck in the 2026 market. As the midpoint of the year passes, the Chinese cloud computing sector has entered a complex economic phase widely characterized as the Ledger Maze. This term illustrates a paradoxical environment where, despite a massive surge in Artificial Intelligence and Model-as-a-Service adoption, neither the service providers nor their corporate users can clearly establish a path toward long-term financial viability. The industry has radically shifted its focus from traditional metrics, such as central processing unit core counts and storage bandwidth, to AI-specific indicators like token consumption, GPU allocation, and model call volumes. While these new performance indicators show triple-digit growth for industry giants like Alibaba Cloud and Baidu, the broader financial reality remains sobering and increasingly difficult for stakeholders to navigate. This transition reflects a fundamental change in how digital value is calculated, moving away from raw hardware capacity toward the intellectual output of large language models and their associated ecosystems.
The Paradox of Growth: Localized Explosions and Market Stagnation
Statistical Dissonance: Token Surges and Modest Revenue
The core of the current industry dilemma is a stark contradiction between localized AI growth and an overall market performance that appears unexpectedly flat. Data from late 2025 and early 2026 show that public cloud large model call volumes in China are skyrocketing, with projections suggesting they will reach 40,000 trillion tokens by year-end. For example, the Doubao model managed by ByteDance experienced a tenfold increase in daily average calls within a single twelve-month cycle, a feat that would normally signal a booming industry. However, report data from the China Academy of Information and Communications Technology indicates that the overall public cloud market growth has slowed to a modest range between 8% and 11%. This discrepancy suggests that while specific AI tools are seeing unprecedented usage, this activity is not translating into a commensurate expansion of the total economic pie for cloud vendors. Instead of creating a new frontier of wealth, the AI boom seems to be operating within the confines of a rigid financial ceiling that prevents the sector from achieving a macro-level breakthrough.
Furthermore, this statistical gap highlights a transition in how enterprises allocate their digital spending, often prioritizing experimental AI projects over reliable infrastructure. While Alibaba Cloud has reported that AI-related product revenue now exceeds 30% of its total external income, the underlying reality is that many of these gains are offset by the declining growth of traditional cloud storage and computing services. The localized frenzy surrounding large language models often masks the fact that many organizations are simply shifting their existing budgets around rather than injecting new capital into the ecosystem. This redistribution of funds creates an illusion of progress while the total revenue growth remains trapped in single or low double digits. For cloud providers, this means they are working harder and investing more to support resource-intensive AI workloads while receiving roughly the same total compensation from their clients as they did for less demanding legacy services.
Budget Cannibalization: The Shift from Legacy Cloud to AI
The phenomenon of budget cannibalization has become a defining characteristic of the 2026 cloud landscape as companies grapple with limited IT expenditures. Rather than expanding their total digital investments, many Chinese enterprises are cutting back on standard virtual machines and database services to fund their forays into generative AI and MaaS platforms. This zero-sum game within corporate budgets means that for every yuan spent on high-performance GPU clusters or token-based API calls, a yuan is often removed from the maintenance of traditional enterprise resource planning systems or secondary storage tiers. Consequently, the cloud market is witnessing a structural transformation where the high-growth AI segment is effectively eating the lunch of the stable, high-margin legacy business. This shift is particularly challenging for vendors who relied on the predictable, recurring revenue of standard cloud services to fund their long-term research and development efforts in more speculative fields.
Moreover, this cannibalization effect is exacerbated by the fact that AI workloads are significantly more expensive to support from an operational standpoint than traditional cloud tasks. While a legacy database might run efficiently on older hardware with minimal power consumption, a modern AI model requires constant access to high-end accelerators and sophisticated cooling systems. As clients migrate their spending toward these power-hungry applications, vendors find themselves in a precarious position where they are providing more complex and costly services for a budget that has not increased in proportion to the operational burden. This trend forces cloud providers to re-evaluate their service catalogs, as the high volume of AI activity often yields lower profit margins compared to the established cloud products they are replacing. The challenge for the remainder of the year will be to find a way to make AI a complementary revenue stream rather than a replacement for the profitable foundations of the cloud industry.
Operational Obstacles: Infrastructure Strain and Vendor Volatility
Capital Expenditure: The High Cost of Physical Hardware
For cloud vendors, the current AI revolution has evolved into an incredibly expensive endeavor marked by rigid asset depreciation and skyrocketing capital expenditures. Alibaba recently reported that its spending on property and equipment surged by 45% for the fiscal year, while competitors like Tencent and ByteDance have seen even more dramatic increases in their infrastructure budgets to maintain a competitive edge. The physical hardware required to support the relentless demand for generative tasks—specifically high-end GPUs—is becoming increasingly difficult and costly to procure due to global supply chain pressures and local demand. Despite billions of yuan in investment, the growth of physical computing resources simply cannot keep pace with the exponential demand for tokens, creating a persistent supply-demand gap that threatens to stall innovation. This imbalance places a massive financial strain on vendors who must commit enormous amounts of capital upfront for hardware that may become obsolete within a few years.
In addition to the high cost of acquisition, the rapid pace of hardware iteration means that the depreciation cycles for these expensive GPU clusters are shorter than ever before. Cloud providers are essentially in a race against time to monetize their hardware before the next generation of accelerators makes their current inventory less desirable. This creates a high-pressure environment where vendors must maintain near-constant utilization of their AI clusters to see any hope of a return on investment. The resulting financial pressure has led to a volatility in pricing models that makes long-term planning difficult for both the providers and their clients. Without a more stable hardware supply chain or a longer lifecycle for these assets, the cloud industry faces a future where capital expenditure requirements continue to outpace revenue growth, leaving thin margins for all but the most efficient players in the space.
Strategic Shifts: Engineering Services and Commercial Instability
The imbalance between infrastructure costs and market demand has led to significant commercial instability and a noticeable breakdown in trust between providers and their enterprise clients. There have been documented instances throughout 2026 where suppliers attempted to quintuple their prices immediately after signing contracts, citing the rising costs of GPU maintenance or the need to shift from flat-rate caps to volatile, consumption-based token billing. This unpredictability has made many large-scale enterprises wary of committing to long-term partnerships, as they fear being trapped by sudden price hikes or service limitations. To manage this chaos and help clients find tangible value in their AI investments, vendors are moving away from traditional, hands-off Software-as-a-Service models. Instead, they are increasingly deploying Forward Deployment Engineering teams, which involve sending specialized engineers to live and work within client organizations to implement and fine-tune AI solutions directly.
This pivot toward intensive engineering services indicates that even the most advanced cloud providers are still struggling to figure out how to make these models work efficiently at a commercial scale. By embedding their staff within client teams, vendors hope to shorten the distance between technical capability and business value, ensuring that AI implementations actually solve problems rather than just consuming resources. However, this model is inherently difficult to scale and adds yet another layer of operational expense to the vendor’s ledger. It transforms the cloud business from a high-margin, automated software play into a lower-margin, labor-intensive consulting business. While these engineering teams are essential for bridging the current capability gap, they also highlight the lack of standardized, “plug-and-play” AI solutions that can be easily monetized. Until the technology matures to the point where it can be deployed without such heavy human intervention, the profitability of the AI cloud will remain hindered by these high service costs.
Tactical Implementation: Client Frugality and the Search for Value
Strategic Frugality: Balancing Small Models and Testing Tiers
On the client side, enterprise users are finding it increasingly difficult to justify massive AI spending when the immediate return on investment remains elusive across many sectors. Large companies, such as Meituan, are spending billions of yuan on AI data procurement and model training, yet they often find that accuracy in critical applications—such as road network identification—remains stuck at suboptimal levels. To mitigate this financial risk, many enterprises have adopted a strategy of strategic frugality, carefully balancing their use of different model tiers. They frequently utilize smaller, cheaper language models ranging from 4B to 35B parameters for daily operations and routine tasks, reserving high-performing and expensive models strictly for testing and validation. This tiered approach allows them to explore the capability ceiling of the technology and stay competitive without draining their operational budgets on tasks that do not require maximum computational power.
This trend toward using “just enough” AI reflects a maturing market where buyers are no longer blinded by the hype of massive parameter counts. Instead, they are evaluating models based on a strict cost-to-performance ratio, often finding that a fine-tuned smaller model can outperform a generic large model for specific industrial tasks. This pragmatism is forcing cloud vendors to diversify their offerings, as the demand for mid-range, efficient models often exceeds the demand for the most powerful flagship systems. For the enterprises, this strategy serves as a buffer against the volatility of the token market, allowing them to maintain a consistent baseline of AI integration while scaling their usage of premium models only when a clear business case exists. As this trend continues, the success of a cloud provider may depend less on the sheer size of their primary model and more on the breadth and efficiency of their entire model library.
Risk Management: Avoiding Lock-in and Validating Returns
In sectors where the cost of an AI error is high, such as education or medical diagnostics, companies are often settling for break-even performance rather than spending exponentially more to reach near-perfection. For a firm like Onion Academy, a wrong answer provided by an AI tutor can lead to a direct financial loss through refunds or damaged reputation, making the pursuit of a 95% accuracy rate prohibitively expensive compared to a reliable 70% baseline. Consequently, these organizations are carefully managing their exposure by avoiding long-term, exclusive contracts with any single cloud vendor. By maintaining access to multiple providers simultaneously, such as Volcano Engine and DeepSeek, they can switch between them based on whoever offers the best price or performance at any given moment. This lack of contractual loyalty makes it nearly impossible for cloud providers to predict their future revenue streams accurately, further complicating the profitability puzzle.
The search for measurable value remained the primary focus for most corporate leadership teams as they closed out their mid-year reviews. They realized that while the technology successfully answered the question of whether complex tasks could be automated, the market had yet to prove that doing so was consistently worth the investment. Early signs of value validation began to emerge in specific areas like AI-driven customer service, where the cost of a token could be directly compared to the hourly wage of a human representative. However, for the majority of internal corporate processes that lacked a standard price tag, the negotiations over cost and ROI continued without a clear resolution. The industry moved into the second half of the year in a waiting game, anticipating the moment when the commercial flywheel would finally turn. This shift was expected only when enterprises could accurately quantify the savings or revenue generated by their AI investments, finally providing an exit from the ledger maze.
