
Z.ai ships GLM-5.3 on the same 743B base, lifting Terminal-Bench 3.0 from 4.6 to 28.3
Z.ai released GLM-5.3 on the identical 743B base model as GLM-5.2, crediting every gain to scaled post-training — more task environments, longer runs — with Terminal-Bench 3.0 jumping from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9. The unplanned result is security: CyberGym hit 84.5%, edging past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%, and Z.ai says the model began forming coherent plans across full exploitation chains rather than reasoning about single bugs. That reframes post-training scale, not base-model size, as the lever that moves long-horizon agent capability. Weights ship in about two weeks after safety hardening, all figures are vendor-reported, and ExploitBench still trails Mythos 5 at 54.4% against 78.0%.
Source: marktechpost.com ↗