Does RL Incentivize Reasoning in LLMs Beyond the Base Model?

spwa4 · 2025-04-22T12:24:35 1745324675

I don't like papers that ask a question in the title, so here's the answer:

"RL boosts sampling efficiency but reduces the reasoning capacity boundary."

Perhaps better to put it like this: Given one, or few attempts, RL trained models beat non-RL models. Given many attempts, non-RL models come up with better answers.

yorwba · 2025-04-22T11:58:03 1745323083

They write "We manually inspect CoT validity to ensure correct answers stem from valid reasoning, not lucky guesses." but the example answer they show at the end only gets the correct number due to two errors canceling out. The model calculates 195+367+562+900 and gets 1924 instead of 2024, and also turns -437 - 2*234 into -805 instead of -905, but in total 1924-805 = 2024-905 = 1119 and from there the remaining steps are correct again.

It would be interesting to know how much of the sampling efficiency improvement from reinforcement learning is due to being better at basic arithmetic (something which could also be achieved by giving the model access to a calculator tool) and how much is due to choosing the correct approach for solving the problem more often.

（评论） (comments)

（评论）
(comments)