[AINews] Death of Params: Z.ai CEO Jie Tang on GLM 5.3 and the new Post-training Scaling Law
Z.ai CEO Jie Tang pointed out that the number of parameters is no longer the only criterion for measuring the significance of a model. He emphasized that it is necessary to consider factors such as data volume, computational resources, and operating conditions comprehensively. The significant improvement of GLM-5.3 stems from reinforcement learning conducted in long-horizon environments. Its training environment encompasses real-world engineering and scientific research workflows, requiring the model to independently complete the entire process from diagnosing bottlenecks to delivering performance optimizations. To support this process, the team built a synthetic environment and a reward signal pipeline. They used a research agent to collect real task patterns and generate multi-step dependent environments, ensuring the reliability of the reward signals through a verifier without reference solutions. Jie Tang noted that advanced skills, such as detecting software vulnerabilities, rely on long causal reasoning rather than simple parameter memory. He also proposed five scaling adjustments, including MoE sparsity.