New Reinforcement Learning Approaches Aim to Improve Language Model Reasoning

Researchers propose novel RL techniques to enhance compositional reasoning and address training instabilities in large language models.

New Reinforcement Learning Approaches Aim to Improve Language Model Reasoning

Researchers have published several papers introducing novel reinforcement learning techniques to improve reasoning capabilities in large language models.

According to arxiv.org, a paper titled “Tropical Reinforcement Learning” proposes changing the mathematical foundation of RL by using maximum operations instead of sum operations for probabilities. The approach “enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts,” according to the abstract. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), their TROPIC algorithm “outperforms the strongest on-policy baselines by up to 16 percentage points.”

Separately, arxiv.org reports on “Segment-wise On-Policy Distillation (Seg-OPD),” which trains models to revise intermediate reasoning steps. The method “consistently outperforms the compared state-of-the-art baselines in reasoning accuracy with an average relative improvement of 5.22% across diverse models and tasks,” according to the paper.

Another arxiv.org paper examines failure modes in on-policy distillation (OPD), finding that “OPD can also collapse into excessively long and repetitive generation.” The researchers explain this through implicit rewards: “the teacher implicitly rewards student behaviors, even those it rarely exhibits itself.” They found that masking problematic responses and using supervised fine-tuning initialization can mitigate these issues.

All papers were published on October 5, 2026.