archive
· today in ai · 2026-07-15
Zero-RL training scales to 1 trillion params
Archive item — written before sources were shown.
A new paper trains a 1-trillion-parameter model with reinforcement learning and no human-labeled data, reporting emergent self-verification on math benchmarks.
The model, Ring-2.5-1T-Zero, uses clipped importance sampling and mixed-precision control to stabilize training at scale, and reports improved sample efficiency and structured reasoning across seven math benchmarks versus smaller zero-RL runs.
sources
- 01Ring-Zero: Scaling Zero RL to a Trillion Parametersarxiv.org · primary research paper
