Researchers Propose New Methods for Reinforcement Learning in Large Language Models

Three new papers present techniques for improving LLM training through reinforcement learning, addressing critic-free methods and inference-time scaling.

Researchers Propose New Methods for Reinforcement Learning in Large Language Models

Multiple research teams have published new approaches to improving large language model training and inference through reinforcement learning techniques.

According to a paper on arxiv.org, researchers proposed “Follow the Winners” (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to reinforcement fine-tuning (RFT). The method replaces group rollouts with an ordinal filter on replay-buffer samples, designed for agents acting in stateful environments where repeated rollouts are impractical. The paper states that FTW “matches GRPO and PPO on Sokoban and Search-R1 baselines,” demonstrating a trade-off from a value model or group rollouts to CPU memory. The work was presented as a poster at NeurIPS 2026.

Separately, arxiv.org published research introducing “Cognitive Relative Policy Optimization” (CRPO), a reinforcement learning framework for mental health assessment. According to the paper, CRPO extends group relative policy optimization by incorporating stage-dependent uncertainty modeling and achieved “an average improvement of 10.4 percentage points in weighted F1-score over the best reinforcement learning baseline” across 8 mental health datasets.

A third paper on arxiv.org presented an analytically tractable model for inference-time scaling, examining Bayesian linear regression with reward-weighted sampling in LLM-as-a-judge scenarios. The research, published at the International Conference on Machine Learning 2026, theoretically demonstrates that generalization error in “best-of-k” scenarios “decays as Θ(1/k²)” when using the teacher as reward.