Ox Alpha revealed: @Zai_org’s GLM-5.3-Flash, the first native multimodal model in the GLM-5 series. Ox Alpha was the biggest model ever on OpenRouter, processing over 20 trillion tokens in 6 days. Continue to use the model now: https://t.co/avBrUW8BfZ

GLM-5.3-Flash is built for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture preserves accurate long-context capabilities while reducing compute overhead. We expect many providers to onboard this model throughout the week.

1M-token context. 131K max output. `max` reasoning by default. Launch pricing via @Zai_org is 50% off through Sep 9 at 16:00 UTC: $0.075/M input $0.25/M output $0.015/M cached input Then $0.15/M, $0.50/M, and $0.03/M.
The volume is the signal: a cheap 1M- model absorbed frontier-scale traffic in under a week. The discounted launch pricing has a hard end date on September 9.
Checking sign-in…
Loading comments…