Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: https://t.co/tzOmB7gdZP Available now across all official platforms: Weights: https://t.co/9LRMahY9Wa API: https://t.co/VcaQnzYmS9 Coding Plan: https://t.co/Nk8Y98HNhU ZCode: https://t.co/Peepqv4XSx Chat: https://t.co/WCqWT0qCQb AutoClaw:

Standard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03
On the https://t.co/w8mHB85z4n Code Bench, which measures real-world coding performance, GLM-5.3-Flash clearly outperforms GLM-5.2 at every effort level and performs on par with Claude Opus 4.8.

Architectural enhancements, combined with an optimized pre-training corpus, enable GLM-5.3-Flash to deliver greater intelligence with less compute.

Open MIT weights plus $0.15 input and $0.50 output per million puts a 1M- model within reach for self-hosting or cheap API routing, with vendor benchmarks claiming parity with far pricier closed models.
Checking sign-in…
Loading comments…