brief research
Zero-RL training scales to 1 trillion params
A new paper trains a 1-trillion-parameter model with reinforcement learning and no human-labeled data, reporting emergent self-verification on math benchmarks.
The model, Ring-2.5-1T-Zero, uses clipped importance sampling and mixed-precision control to stabilize training at scale, and reports improved sample efficiency and structured reasoning across seven math benchmarks versus smaller zero-RL runs.
sources 1 cited
1 arxiv.org Ring-Zero: Scaling Zero RL to a Trillion Parameters