Chinese AI lab Zhipu claims its domestic chips match Nvidia on per-token cost

Key points
- Zhipu says it has deployed domestic chips for large-scale inference
- Per-token inference cost fell about 80% this year
- Revenue rose 399.7% and gross margin nearly halved
On Monday, August 31, Z.AI (2513.HK), the Beijing lab still widely known as Zhipu, published its first interim results announcement since listing in Hong Kong in January. In it the company said it had deployed domestic chips as a primary source of inference capacity, reaching what it described as 100,000-level deployment, with per-token inference cost down about 80% since the start of the year, as reported by Tencent News.
The wording is doing more work than the English coverage suggests. SemiAnalysis pointed out that the translated version reads "a cluster of," while the Chinese says only 100,000-level, with no claim about one cluster. That isn't a quibble. One hundred thousand accelerators wired into one cluster is a much harder engineering result than 100,000 cards deployed across a fleet. At least one major outlet rendered the disclosure in that stronger form. The South China Morning Post reported that the model ran entirely on "a cluster of 100,000 domestically produced chips," and noted that Zhipu hasn't disclosed who makes them.
Reports have pointed to Huawei Ascend, Moore Threads, and Hygon, and Chinese outlet Sina hedged that list with a "may come from." Which vendor it is changes what the disclosure means, so treat it as unconfirmed.
Advertisement
The half in numbers
The financials came in the same announcement. Zhipu reports in yuan, and revenue was about 954 million yuan for the half, up 399.7% year over year. Gross profit was about 252 million yuan, up 163.7%. The loss for the period was 2.07 billion yuan, against 2.35 billion yuan in the same half of 2025, a narrowing of about 12%. Research and development spending was 2.13 billion yuan, up 33.6%, which is more than twice what the company booked in revenue. Open platform and API revenue was about 825 million yuan, up 2,735.7%, and now makes up 86.5% of the total. The company put its MaaS platform at more than 5.8 million enterprise and developer users.
The growth came with sharply lower gross margin. Gross profit rose much more slowly than revenue, pushing the implied margin down from about 50% a year earlier to 26.4%. That could reflect rapid expansion of the lower-priced API business, but the announcement doesn't provide enough detail to isolate the cause.
What ties the two halves of the announcement together is the cost line. Zhipu open-sourced GLM-5.3-Flash on August 26 and said the online traffic during that launch ran on domestic chips, with end-to-end throughput three times a baseline on the same hardware, and per-token cost reaching parity with mainstream Nvidia (NVDA) GPUs. The three-times figure needs reading carefully. Zhipu described its own stack, an inference engine adapted from SGLang with prefill-decode separation, quantization, and interconnect work, measured against an earlier configuration on the same chips. It didn't identify the baseline software, the workload, or the measurement method, so the number describes its own improvement curve rather than a comparison anyone else can reproduce.
Advertisement
What the cost-parity claim means for Nvidia
US export controls seek to restrict Chinese access to advanced Nvidia hardware, and we have covered the rulemaking as it's tightened. Their practical effect depends partly on domestic substitutes remaining less capable or less economical. Putting a cost-parity claim in an exchange filing gives it more weight than a demonstration or a product announcement, although the performance comparison remains unaudited and independently unverified.
The document itself is strange for a financial filing. SemiAnalysis counted 19 of its 60 pages reading like a technical blog, plus a nine-page technical glossary, and noted that Zhipu defines the Pareto frontier of a state-of-the-art model as speed of iteration multiplied by intelligence index multiplied by per-task cost. Terms like GRPO and mHC appear in a document filed with an exchange.
Zhipu isn't listed in the United States and has no American depositary receipt, and several of the possible domestic chip suppliers trade only in mainland China. For investors limited to US-listed stocks, Nvidia is the clearest liquid read-through, although Zhipu's disclosure represents a competitive risk rather than direct exposure. We last wrote about Nvidia on a separate SemiAnalysis teardown of its rack costs.
Four disclosures would make the investment case easier to evaluate: the chip vendor named in a filing rather than inferred by reporters, serving economics stated in a form an outside party can compare, a gross margin recovering from 26.4%, and loss and operating cash burn shown as a share of revenue rather than only in absolute yuan.
Advertisement