Researchers have published several papers examining different approaches to improving large language model performance and efficiency.
According to arxiv.org, a new framework enables collaboration between small language models (SLMs) and large language models (LLMs) by treating it as an information acquisition problem rather than a computation allocation problem. The SLM serves as the primary reasoner and “selectively queries a black-box LLM advisor only when needed, issuing targeted queries rather than delegating the reasoning process itself,” the paper states. The approach uses a three-stage RLVR framework and demonstrates improved performance-cost tradeoffs on mathematical reasoning and coding tasks.
In asynchronous training research, arxiv.org reports on a group mass capping GRPO (GMC-GRPO) method that addresses stale rollouts from earlier policies. The paper establishes that “compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(\epsilon^{-4})$ to $O(\epsilon^{-2})$.” Experiments using Qwen3 models showed improved robustness to stale rollouts.
Another arxiv.org paper introduces Cancellation-Aware Response Masking (CARM), which addresses off-policy issues in LLM reinforcement learning. The method “takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling.” CARM improved mathematical reasoning scores by up to 3.13 percentage points and code generation pass rates by 2.88 points.
Finally, arxiv.org research found that tool use through code synthesis prevents accuracy collapse in out-of-distribution reasoning tasks, particularly as problem depth increases.