
Zhipu AI's GLM-5.3-Flash ran 62 trillion tokens on 100,000 domestic Chinese chips
Zhipu AI revealed that Ox Alpha, the anonymous model that topped OpenRouter's coding rankings with 10.3 trillion tokens — nearly 31% of the platform's weekly volume — is GLM-5.3-Flash, served entirely from a cluster of 100,000 domestically produced chips. The stealth trial is the concrete data point China's chip-independence push has lacked: global-scale inference, no Nvidia, with usage measured before anyone knew the provenance. Zhipu shares closed up more than 12% at HK$1,160 in Hong Kong.
Source: scmp.com ↗
The deployment marks a significant test of China's ability to handle large-scale global inference workloads on home-grown hardware, as Beijing seeks to reduce reliance on advanced processors from US market leader Nvidia amid tight export controls.
Why this matters
- → Proves China can run global-scale AI inference on domestic chips without Nvidia
- → 62 trillion tokens in stealth trial demonstrates feasible chip independence path
- → Competitive performance ranked #1 for coding, validating alternative hardware viability