As AI competition becomes increasingly geopolitical and macroeconomic in nature, it's important to understand the factors driving AI development in regions outside the US. It's already common to discuss how open frontier models are being driven by Chinese labs, but why is this the case? And where are the chokepoints and challenges that stymie frontier lab success and growth in China versus the US?
Here, we'll cover the ambitions of companies like Z.ai and Alibaba based on recent earnings calls. We'll also outline the challenges of chip development in China—the challenges are very different than in the US, as China has electricity supply but lacks access to bleeding edge chip development tools. This leads to business models that enable not just development of open source models, but different forms of monetization. State support and Chinese media culture (i.e., the same trends that drove TikTok's development) are also driving unique aspects of generative AI models.
Z.ai is one of the most successful frontier labs in China, and is now moving to operate its own data centers on Chinese-built chips; it may be the first frontier lab starting to do so in China. Z.ai announced in July[1] that it was building a 1-gigawatt data center, one of the largest for Chinese frontier labs. Z.ai's most recent GLM-5.3-Flash announcement specifically mentioned running all inference on Chinese-made chips[2].
This is a critical development for AI development in China: not only are open models tracking performance improvements (albeit still underperforming frontier labs like Anthropic and OpenAI), they are now starting to run their own data centers on chips made in China.
Z.ai's earnings call had interesting language about its next blockbuster model release, GLM-6.0. Z.ai founder Tang Jie mentioned that GLM-6.0 is being trained in an approach akin to recursive self-improvement (RSI), where the training algorithm can make decisions about its entire training pipeline[3]. This might be similar to how OpenAI's latest Astra model was developed as well[4].
It remains to be seen when the new model gets released, but there's certainly no shortage of ambition.
Alibaba's strategy is a two-pronged approach, in many ways similar to what Z.ai is doing: a frontier model that is competitive with most models and inference requiring best-in-class data center GPUs, combined with a “flash” model that is small enough to run on powerful consumer PCs and older hardware.
Its suite of models covers a huge parameter count: Qwen3.8 Max is a 2.4 trillion parameter model[5], Qwen3.8-Flash-Next is 125 billion, and Qwen3.8-27B is 27 billion[6]. Depending on the benchmarks, all three models are competitive with Claude Opus 4.6, released in February 2026[7].
The Qwen family of models is by far the most popular open source model family being run locally, as shown in Figure 1[8].
Alibaba is comparable to Amazon in its breadth of business from ecommerce through on-demand cloud computing services. It also has T-Head, its own chip design arm, again similar to Amazon and its Trainium strategy. As discussed in its earnings call[9]:
T-Head has established a full-stack portfolio of proprietary chips spanning GPUs, CPUs, and networking chips. As of early August, Zhenwu chips had served more than 650 customers. The supernode instance powered by T-Head's next-generation Zhenwu M890 AI chip, backed entirely by domestic supply chain, recently launched on Alibaba Cloud for commercial sale at scale. … Alibaba Cloud's Zhenwu M890 Supernode can efficiently run inference workloads for foundation models with more than 2 trillion parameters. Both Kimi K3 and Qwen 3.8-Max are already providing MaaS services to external customers through this supernode instance.
Like Z.ai, the language around a 100% domestically sourced chip supply chain is emphasized on earnings calls and corporate documents.
Alibaba's cloud revenue grew 45% this year. Its payback period for chips is 3 years, but the company mentions that even 7-year-old chips are still used for inference workloads:
In our data centers, A100 GPUs purchased in 2020 and V100 GPUs purchased in 2018 continue to be used by customers at close to full capacity.
This is the flash modeling approach in action—smaller models can be run on such hardware, ensuring it does not depreciate after 3 or 4 years.
Alibaba is planning to launch its next generation of T-Head chips in the second half of 2026:
In the second half of this year, T-head's second-generation domestic chip is gradually taping out with production to follow at a later stage. The chip will offer very strong compute performance and interconnect bandwidth, and we believe it will be fully capable of supporting large-scale model training.
Alibaba also recently launched Wan 3.0, a multimodal model that can generate video from images, documents, and prompts[10]. This is on the heels of MiniMax, which launched H3 on July 31[11]. ByteDance also has Seedance[12].
While Chinese models trail the performance of Western coding models, they tend to outperform on media generation. OpenAI's Sora is being discontinued in late September 2026[13]. Google's Gemini Omni 1.1 was released on August 27[14] and is the only model comparable to Chinese counterparts[15].
This is likely being driven by Chinese technology companies' business models. ByteDance created Douyin, the Chinese predecessor to TikTok, and ecommerce in China is highly dependent on short-form video content. This has since expanded into short-form entertainment of all sorts. The Economist[16] recently reported on 1-minute-long microdramas optimized for social media, and how they represent a $15B/year industry. Unsurprisingly, AI is now being used to generate this sort of content.
So far we've been looking at the actual software and modeling side of things, but you can't scale AI without having the proper chips and memory in place.
Most of the startups trying to provide Nvidia-like chips in China are doing so following Nvidia's own fabless chip design playbook. These include Cambricon, Hygon, Moore Threads, Biren Technology, Enflame, and MetaX. We won't provide a systematic review of each of these companies here, but rather simply say that while they all differ on their specific approaches to chip design, their success is ultimately enabled or hindered by access to chip fabrication. The chip designers are all supply-constrained: like Nvidia, there is huge demand for their chips, and the challenge is actually producing the chips to meet that demand. As long as they have access to chip supply, they can likely sell out their inventory.
As a brief overview of the various offerings and their market success, Bloomberg ran a survey across Chinese corporations and various chip designers, shown in Figure 2[17].
SMIC is the only company in China that is producing these bleeding edge chips. As such, judging the success of the fabless chip designers is a combination of (a) seeing if they can design next generation hardware, (b) confirming they have customers for their latest products, (c) understanding their relationship to SMIC and SMIC's capacity.
Cambricon is a case in point: (a) it is successfully designing chips, (b) nearly all of its Q1 2026 inventory has been bought by ByteDance[18], and (c) Beijing has supported Cambricon by ensuring SMIC set aside product capacity[18].
This begs the question: how is SMIC doing?
Chip development can be summarized as a three-step process of printing wafers, splitting wafers into dies, and then using dies to construct chips. SMIC shipped about 1.11 million 12-inch wafer equivalents in Q4 2025, or about 370,000 per month. It added about 50,000 12-inch equivalent wafer capacity in 2025, and expects to add 40,000 per month in 2026[19]—in other words, about 11% growth in wafer capacity in 2026. This is simply not enough to address demand, by far.
Worse still, yields—the error-free production rate—are a critical challenge. In December 2025, Bloomberg reported that SMIC's yield on Cambricon chips was about 20%[20], meaning that for every 5 chips' worth of wafers produced, only 1 is of a quality high enough to ship. By comparison, newer generation TSMC processes have 60% yields. This is a direct impact of US sanctions; newer Extreme Ultraviolet (EUV) ASML technologies help produce higher-quality wafers, but these are banned from export to China since 2019[21].
For example, adding 50,000 in monthly wafer capacity means that with 20% yields, you're effectively adding only 10,000 usable wafer-equivalents.
In short: SMIC is adding capacity, but chip production at the bleeding edge will be capacity-constrained for the foreseeable future.
The other part of the hardware equation is memory, and specifically High Bandwidth Memory. ChangXin Memory Technologies (CXMT) is a memory manufacturer working to catch up to SK Hynix, Micron, and Samsung. The company IPOed in late July[22], seeing its stock jump 466% that day and raising $8.6 billion[23]. The company saw revenue of $22.4 billion in the first half of 2026 and has ambitions to outcompete the memory stalwarts.
While there is huge execution risk in expanding to high bandwidth memory, The Information reported on August 31 that CXMT is now producing HBM3E, the type of memory used by major AI processors today[24], with Alibaba and Cambricon both testing the new memory in their processors this year.
While progress is seemingly good, the challenge with scaling memory production is that you require massive fabrication facilities that take years to build, just like with GPUs. At its current pace, SemiAnalysis[25] estimates CXMT will still be in fourth place in 2027, producing significantly fewer memory chips than competitors, as shown in Figure 3.
Without significant fab buildouts, CXMT won't be able to catch up on its current trajectory. Indeed, as with other supply-constrained memory manufacturers, much of the revenue growth in the space might be coming from price increases rather than increasing quantities of products sold.
If you can't run massive data centers and are limited by compute, another option is to license your technology. This is exactly what some Chinese frontier labs are doing—licensing their models to neoclouds via a revenue share approach. MiniMax H3, for example, is free to use commercially unless your platform generates $20 million in revenue, at which point a revenue sharing model is required[26]. MoonshotAI, the makers of Kimi K3, have a similar requirement[27]. If you can't build your own data center, you might as well release your code as per the above! Then you can have hyperscalers and neoclouds provide you with a delivery model.
The focus on open and local models is a strategic differentiator. If you lack access to the best AI chips, building frontier models that work on consumer hardware or older chips is the next best thing.
Next, frontier labs are more aggressively innovating across all modalities. China is the home of short-form video content such as TikTok (via ByteDance's Douyin), pioneered Key Opinion Leaders (KOLs; i.e., ecommerce influencers), and of course short-form social media-oriented microdramas. In Europe and North America, video content seems to revolve around traditional media, which is controversial due to talent displacement risks; the Writers Guild of America went on strike in 2023 to protest use of generative AI[28], and Sora 2 received very public backlash from Hollywood studios and talent agencies[29]. This is less of an issue in China—in other words, the odds are in favor of the frontier labs and big tech companies benefiting from such media.
As we discussed in The AI Trade is Becoming a Macro Trade, country-level industrial success is as much about public policy and state support as individual private company execution. China is doing a lot. As far as energy is concerned, China is executing aggressively on new solar power, nuclear power, and other energy production projects. It's planning to nearly double nuclear power generation from 62 GW in 2025 to 110 GW in 2030[30]. It is aiming to have over 50% of its electricity generation from non-fossil fuel sources by 2030[31].
China is coordinating a package between $300 billion and $600 billion to support the buildout of a national data center strategy[32] from 2025 to 2030. This includes requirements that 80% of the equipment be produced domestically[33]. The emphasis on domestic production across Alibaba and Z.ai's earnings announcements shows how important this is for the entire Chinese industry.
More important is the fact that the government is specifically forcing self-sufficiency. Back in January 2026, even when Nvidia was allowed to export H200 chips to China, it was the Chinese government that stepped in to prevent imports without its own permission[34][35], only allowing some companies to import H200 chips this past August[36].
This is one of the biggest indicators of how ambitious and important self-sufficiency is to China. Rather than simply depend on US and Western chips, the government is aggressively forcing its own industry to develop. While this might hurt progress in the short run (though, given the progress of Z.ai, MiniMax, Alibaba, etc. this might not even be the case), the long run could lead to a significantly more diversified and powerful AI industry in China.
China's AI industry is facing capacity constraints across the board, and this is true for SMIC, AI frontier labs, hyperscalers, and startups. China has the benefit of large amounts of energy and the ability to execute on electricity, data center, and network expansion, but the challenge of fabricating chips—and scaling SMIC's capabilities—remains.
US sanctions combined with SMIC constraints are leading to innovation on the modeling side. This is a big reason why frontier labs are developing models that run on high-end consumer PCs and have open source options that can be hosted, for a fee, by Western neoclouds. These capacity constraints will likely be addressed in the future—the question is when. Key signposts right now are looking at SMIC's capacity additions and yield. If either rise significantly, then there might be a structural shift in AI capacity.
Most interestingly, these capacity constraints, combined with the nuances of the Chinese economy such as a focus on short-form video content and ecommerce, are leading to interesting innovations, products, and research roadmaps different from those of Western model developers. It will be exciting to see how this progresses.
Join hundreds of AI, geopolitics, and economics experts. We'll send 1–2 emails per week, and we will never share your e-mail with anyone.