August 15, 2026AgentsCodingOpen Source

GLM-5.3 went from 4.6 to 28.3 on Terminal-Bench without a new base model

Zhipu shipped GLM-5.3 on August 14 and it took the top of Hacker News with a thousand points. The headline number is Terminal-Bench 3.0 going from 4.6 to 28.3. That is a six-fold jump on the benchmark that measures whether a model can actually finish work inside a real terminal, and it came out of the same base weights as GLM-5.2. No new pretraining run. Just extended post-training with more executable environments, more long-horizon tasks, and more RL compute.

That detail is the story. Everyone assumes the next capability jump costs another pretraining run and another data center. Zhipu got a 50 percent internal coding improvement, DeepSWE v1.1 from 46.2 to 66.9, and SWE-Marathon v1.1 from 19.4 to 42.5, by pointing reinforcement learning at environments the model has to survive rather than problems it has to answer. If that generalizes, the frontier is a lot cheaper to move than the capex announcements suggest, and a lot of compute deals signed this year were priced on the wrong assumption.

The cyber numbers are the uncomfortable part. GLM-5.3 hits 84.5 percent on CyberGym, edging past Mythos 5 and GPT-5.6 Sol. Zhipu says the gap to closed frontier models persists on deep exploitation tasks like ExploitBench, which is the honest caveat, but a model that is about to have downloadable weights leading a vulnerability benchmark is a different conversation from a hosted API leading it. Zhipu is holding the weights for roughly two weeks pending safety evaluation. Read that as what it is: a Chinese lab running a staged release for cyber capability, which is exactly the pattern the Western labs adopted, arriving in China without anyone legislating it.

Practical status right now. The model is live only through the GLM Coding Plan at 18, 80, and 168 dollars a month for Lite, Pro, and Max, with a 1M-context route as glm-5.3[1m] and reasoning effort at low, high, or max. Standard per-token API pricing is not published. No model card, no license, no local deployment recipe, and every benchmark above was run by the vendor with no independent reproduction yet.

So the sober read is: strongest open-weights coding claim on the table, on paper, with the weights not on the table. If you already pay for the Coding Plan and have repository tasks you can measure, run it today and see whether Terminal-Bench transfers to your repo. If you need a license or local control, the interesting date is two weeks out. Details at z.ai/blog/glm-5.3.
← Previous
Ops Log: August 14, 2026
Next β†’
Qwen3.8-27B: a 27B model claiming 73 on Terminal-Bench and 84.3 on OSWorld
← Back to all articles

Comments

Loading...
>_