Advertisement

Chinese AI lab Zhipu claims its domestic chips match Nvidia on per-token cost

Chinese AI lab Zhipu claims its domestic chips match Nvidia on per-token cost

Key points

  • Zhipu says it has deployed domestic chips for large-scale inference
  • Per-token inference cost fell about 80% this year
  • Revenue rose 399.7% and gross margin nearly halved

On Monday, August 31, Z.AI (2513.HK), the Beijing lab still widely known as Zhipu, published its first interim results announcement since listing in Hong Kong in January. In it the company said it had deployed domestic chips as a primary source of inference capacity, reaching what it described as 100,000-level deployment, with per-token inference cost down about 80% since the start of the year, as reported by Tencent News.

The wording is doing more work than the English coverage suggests. SemiAnalysis pointed out that the translated version reads "a cluster of," while the Chinese says only 100,000-level, with no claim about one cluster. That isn't a quibble. One hundred thousand accelerators wired into one cluster is a much harder engineering result than 100,000 cards deployed across a fleet. At least one major outlet rendered the disclosure in that stronger form. The South China Morning Post reported that the model ran entirely on "a cluster of 100,000 domestically produced chips," and noted that Zhipu hasn't disclosed who makes them.

Reports have pointed to Huawei Ascend, Moore Threads, and Hygon, and Chinese outlet Sina hedged that list with a "may come from." Which vendor it is changes what the disclosure means, so treat it as unconfirmed.

Advertisement

The half in numbers

The financials came in the same announcement. Zhipu reports in yuan, and revenue was about 954 million yuan for the half, up 399.7% year over year. Gross profit was about 252 million yuan, up 163.7%. The loss for the period was 2.07 billion yuan, against 2.35 billion yuan in the same half of 2025, a narrowing of about 12%. Research and development spending was 2.13 billion yuan, up 33.6%, which is more than twice what the company booked in revenue. Open platform and API revenue was about 825 million yuan, up 2,735.7%, and now makes up 86.5% of the total. The company put its MaaS platform at more than 5.8 million enterprise and developer users.

The growth came with sharply lower gross margin. Gross profit rose much more slowly than revenue, pushing the implied margin down from about 50% a year earlier to 26.4%. That could reflect rapid expansion of the lower-priced API business, but the announcement doesn't provide enough detail to isolate the cause.

What ties the two halves of the announcement together is the cost line. Zhipu open-sourced GLM-5.3-Flash on August 26 and said the online traffic during that launch ran on domestic chips, with end-to-end throughput three times a baseline on the same hardware, and per-token cost reaching parity with mainstream Nvidia (NVDA) GPUs. The three-times figure needs reading carefully. Zhipu described its own stack, an inference engine adapted from SGLang with prefill-decode separation, quantization, and interconnect work, measured against an earlier configuration on the same chips. It didn't identify the baseline software, the workload, or the measurement method, so the number describes its own improvement curve rather than a comparison anyone else can reproduce.

Advertisement

What the cost-parity claim means for Nvidia

US export controls seek to restrict Chinese access to advanced Nvidia hardware, and we have covered the rulemaking as it's tightened. Their practical effect depends partly on domestic substitutes remaining less capable or less economical. Putting a cost-parity claim in an exchange filing gives it more weight than a demonstration or a product announcement, although the performance comparison remains unaudited and independently unverified.

The document itself is strange for a financial filing. SemiAnalysis counted 19 of its 60 pages reading like a technical blog, plus a nine-page technical glossary, and noted that Zhipu defines the Pareto frontier of a state-of-the-art model as speed of iteration multiplied by intelligence index multiplied by per-task cost. Terms like GRPO and mHC appear in a document filed with an exchange.

Zhipu isn't listed in the United States and has no American depositary receipt, and several of the possible domestic chip suppliers trade only in mainland China. For investors limited to US-listed stocks, Nvidia is the clearest liquid read-through, although Zhipu's disclosure represents a competitive risk rather than direct exposure. We last wrote about Nvidia on a separate SemiAnalysis teardown of its rack costs.

Four disclosures would make the investment case easier to evaluate: the chip vendor named in a filing rather than inferred by reporters, serving economics stated in a form an outside party can compare, a gross margin recovering from 26.4%, and loss and operating cash burn shown as a share of revenue rather than only in absolute yuan.

Advertisement

Frequently asked questions

What did Zhipu say about domestic chips in its interim report?

Z.AI (2513.HK), the Beijing lab known as Zhipu, said in its first interim results announcement since its January 2026 Hong Kong listing that it had brought domestic chips into its mainstay inference compute and reached what it described as 100,000-level deployment, and that per-token inference cost had fallen about 80% since the start of the year. The results were published on August 31, 2026. The claim is the company's own and has not been independently verified.

Did Zhipu run 100,000 chips in a single cluster?

That is not what the Chinese version says. SemiAnalysis noted that the English translation reads as a cluster, while the Chinese text says only 100,000-level, with no claim about one cluster. The distinction matters: 100,000 accelerators wired into a single cluster is a considerably harder engineering result than 100,000 cards deployed across a fleet, and English-language reports have carried the stronger version, including the South China Morning Post, which said the model ran entirely on a cluster of 100,000 domestically produced chips.

Which Chinese chips is Zhipu using?

Zhipu has not publicly disclosed the supplier, according to the South China Morning Post. Reports have pointed to Huawei Ascend, Moore Threads and Hygon, and Chinese outlet Sina hedged that list with a phrase meaning they may come from those vendors, so none has been confirmed. Which vendor it is changes what the disclosure implies about domestic accelerator capability.

What were Zhipu's first-half 2026 results?

Revenue was about 954 million yuan, up 399.7% year over year, with gross profit of about 252 million yuan. The loss for the period was 2.07 billion yuan, against 2.35 billion yuan a year earlier, a narrowing of about 12%. Research and development spending was 2.13 billion yuan, more than twice the revenue booked in the half. Gross profit rose far more slowly than revenue, so the implied gross margin fell from about 50% a year earlier to 26.4%. Open platform and API revenue was about 825 million yuan, up 2,735.7%, and 86.5% of the total. Zhipu is not listed in the United States and has no American depositary receipt. This is general information, not investment advice.

More on NVDA

Dennis Singleton
Dennis Singleton

Dennis Singleton has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.