2 papers across 2 sessions
We show that limiting a model's confidence during training can improve test-time scaling in mathematical reasoning.