RELEASED
2026-09-18
Zhipu launches GLM-5.3-FlashX at 200 tokens/second, running on ~100k domestic chips
A speed-focused iteration of GLM-5.3-Flash (320B total / 18B active) claiming 200 tokens/second — 5x the prior Flash — optimized by a self-hosted 'Infra Agent'; inference reportedly runs wholly on domestically produced accelerators. Priced at 2.5x the original Flash rate; Zhipu's stock gained roughly 7% on the news. Throughput and hardware claims are lab-reported and not yet independently reproduced.