Skip to content
EN

Reinforcement learning

Rewards, policies, environments and results, for language models and beyond.

79 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Reinforcement learning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. LED uses intermediate-layer diversity to restore exploration

    The paper reports that RL post-training reduces entropy in the final-layer distribution while intermediate layers retain higher entropy. Its training-free Latent Exploration Decoding method uses intermediate-layer distributions to improve sampling in reasoning models.

    Engineers can evaluate a decoding approach aimed at improving pass@n without further training.

  2. Single-Rollout Asynchronous Optimization for Agentic RL

    The linked paper examines asynchronous RL for post-training language models, where models are updated as rollouts arrive. The post says GLM 5.2 PPO uses async RL improvements including clipping adjustments, and also links to VAPO.

    Engineers working on agentic RL can review approaches to asynchronous training and stability.

  3. Privileged Self-Distillation Can Degrade Long Reasoning Traces

    The paper studies self-distillation with privileged information, such as a math solution, in thinking models. It reports degradation on long reasoning traces across five Qwen3 and OLMo models evaluated on AIME.

    It highlights a potential trade-off when using privileged information to improve reasoning models.

  4. RubricEM decomposes research-agent policies with rubrics

    RubricEM trains research agents to create rubrics before working, then use them to plan, search, review, and answer. A judge scores trajectory stages, and judged attempts become lessons for future tasks.

    Stagewise rubric rewards and reusable lessons offer an approach to training agents beyond single final-score feedback.

  5. EfficientRollout speeds up RL rollout generation for LLMs

    The paper presents system-aware self-speculative decoding for RL rollouts, using a quantized copy of the model, a roofline-based activation rule, and adaptive draft lengths. It reports up to 19.6% faster rollout generation and 12.7% faster training steps.

    It describes ways to reduce rollout latency without changing the model being trained.

  6. Two-Phase Distillation for Multi-Task Agentic LLMs

    The paper studies consolidating separately trained, task-specific RL experts into a multi-task model through distillation, instead of training one model on mixed tasks. The post describes using off-policy distillation for initialization and on-policy distillation for refinement.

    The approach offers an alternative to mixed-task training and examines a limitation of off-policy distillation in multi-task settings.

  7. DreamSmooth smooths sparse rewards in model-based RL

    The ICLR paper proposes temporally smoothing rewards across neighboring steps, using kernels such as Gaussian or EMA, to make reward modeling easier. The post says it evaluates the method on RoboDesk and Shadow Hand.

    Reward smoothing may help engineers train model-based RL agents when rewards are sparse in time.

  8. AutoDecompiler Uses RL for Feedback-Driven Decompilation

    The post describes AutoDecompiler as an RL-optimized model for multi-turn decompilation that uses feedback directly, unlike prior one-shot approaches.

    It may interest engineers exploring reinforcement learning for iterative code generation and decompilation.

  9. TMAX trains terminal agents with diverse Dockerized RL environments

    The TMAX post describes 14.6k Dockerized RL environments and an outcome-only DPPO recipe for training terminal agents. It reports a 9B model reaching 27% on Terminal-Bench 2.0.

    The recipe and environment design may help engineers train terminal agents for multi-turn shell tasks.

  10. A Proposal for Training-Free Reasoning with Verified Examples

    The post proposes keeping an LLM frozen and retrieving verified reasoning from an external pool instead of updating weights. It uses SymbCoT-style symbolization for checking logical structure; whether this can approximate RL gains remains open.

    It outlines a possible alternative to weight updates, while highlighting verification and deployment as unresolved challenges.

  11. Infrastructure optimizations for RL at trillion-parameter scale

    A deep dive into prime-rl 0.6.0 training trillion-parameter MoE models, covering FP8, wide expert parallelism, P/D disaggregation, router replay, and 3-D parallelism.

    Engineers can review training and inference techniques for scaling RL workloads on large MoE models.

  12. Post claims OPD is more efficient than RL

    The author claims OPD uses compute and samples more efficiently than RL, which they say finds high-reward reasoning traces. The linked PDF is titled “llm_distillation.pdf.”

    The claim concerns the compute and sample efficiency of approaches to training language models.

  13. A proposed test for reward-hacking mitigations

    The post contrasts blocking suspicious tool calls and returning dummy information with penalizing a CoT monitor, which it says can lead to obfuscation. It asks whether a head-to-head test in the same environment has measured this difference.

    The comparison could help engineers evaluate whether interventions stop reward hacking or merely make it harder to detect.

  14. Why RL Scaling Laws Differ from Pretraining

    The post contrasts pretraining and RL scaling laws, noting that RL compute includes sampling and policy updates and may be measured in FLOPs or GPU hours. It also distinguishes within-run and across-run extrapolation.

    It highlights compute and evaluation choices that complicate comparisons and extrapolation in RL experiments.

  15. Overview of RL methods for reasoning LLMs

    A blog post surveys RL methods including REINFORCE, PPO, RLHF, GRPO, RLOO, Dr. GRPO, DAPO, CISPO, MaxRL, DPPO, and ScaleRL.

    Useful as a single starting point for engineers comparing RL methods used with language models.

  16. Running verl RLHF Training on AMD GPUs with ROCm 7.0

    An AMD ROCm blog describes deploying verl for RLHF training on AMD GPUs, with ROCm optimization and Docker scripts. It reports throughput and convergence results.

    Useful for engineers evaluating AMD GPU infrastructure and deployment options for verl-based RLHF.

  17. CogRouter adapts reasoning depth at each agent step

    CogRouter uses four hierarchical cognitive levels to adjust reasoning depth for LLM agents. Its training combines supervised fine-tuning and policy optimization; the post reports an 82.3% benchmark success rate for a 7B model.

    Step-level reasoning control may help engineers trade off agent performance and token use.

  18. Self-distillation for continual learning without reward functions

    The post describes a method that uses a model conditioned on a demonstration as a teacher for the same model generating text without that demonstration. The student is trained to match the teacher’s token distributions on its generated text.

    This approach may help engineers train models on new tasks while reducing catastrophic forgetting, without defining a reward function.

  19. Online RL for HPC Code Generation with Machine Benchmarks

    The article describes using online reinforcement learning with real-machine benchmark rewards to improve LLMs’ HPC code generation. It notes that generated code’s runtime performance is not guaranteed.

    Engineers can see an approach to training code-generation models using measured runtime performance.

  20. MaxRL uses likelihood-based training for binary-reward RL

    The post describes MaxRL, which uses additional rollouts to approximate maximum-likelihood training with non-differentiable sampling. It says the method addresses underweighting of hard prompts and reports up to 20× test-time efficiency versus GRPO.

    The approach may matter to engineers evaluating reward optimization and compute scaling for RL systems.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor