415.tech
AI & tech, from the frontlines of Silicon Valley
Zhipu AI's GLM-5.3-Flash ran 62 trillion tokens on 100,000 domestic Chinese chips

Zhipu AI's GLM-5.3-Flash ran 62 trillion tokens on 100,000 domestic Chinese chips

Zhipu AI revealed that Ox Alpha, the anonymous model that topped OpenRouter's coding rankings with 10.3 trillion tokens — nearly 31% of the platform's weekly volume — is GLM-5.3-Flash, served entirely from a cluster of 100,000 domestically produced chips. The stealth trial is the concrete data point China's chip-independence push has lacked: global-scale inference, no Nvidia, with usage measured before anyone knew the provenance. Zhipu shares closed up more than 12% at HK$1,160 in Hong Kong.

Source: scmp.com

Post on XEmail

The deployment marks a significant test of China's ability to handle large-scale global inference workloads on home-grown hardware, as Beijing seeks to reduce reliance on advanced processors from US market leader Nvidia amid tight export controls.

SCMP reporting

Why this matters

  • → Proves China can run global-scale AI inference on domestic chips without Nvidia
  • → 62 trillion tokens in stealth trial demonstrates feasible chip independence path
  • → Competitive performance ranked #1 for coding, validating alternative hardware viability
China's chip independence bet
Also in this edition