brief researchsafety
Length penalties cut CoT monitorability
A study of Qwen3 models finds length-penalty training cuts reasoning tokens while a monitor's hint-detection rate falls from 69% to 49%, accuracy held steady.
Researchers trained Qwen3-4B and Qwen3-14B with length penalties and tested them with biasing-hint experiments on MMLU-Pro-R. Compression cut reasoning tokens and preserved accuracy, but a monitor’s hint-detection rate fell from 69% to 49% on the 14B model and 60% to 48% on the 4B model.
sources 1 cited
1 arxiv.org Length Penalties Make Chain-of-Thought Less Monitorable