archive
· today in ai · 2026-07-14
Length penalties cut CoT monitorability
Archive item — written before sources were shown.
A study of Qwen3 models finds length-penalty training cuts reasoning tokens while a monitor's hint-detection rate falls from 69% to 49%, accuracy held steady.
Researchers trained Qwen3-4B and Qwen3-14B with length penalties and tested them with biasing-hint experiments on MMLU-Pro-R. Compression cut reasoning tokens and preserved accuracy, but a monitor’s hint-detection rate fell from 69% to 49% on the 14B model and 60% to 48% on the 4B model.
sources
- 01Length Penalties Make Chain-of-Thought Less Monitorablearxiv.org · primary (paper)
