
Model size is still the default shorthand for capability, and this lays out why that shorthand now misleads when picking a model to build on. The GLM-5.3 detail also shows long-horizon RL environments, not scale, driving the current jump in engineering task performance.
Sign in to comment.
Loading comments…