Skip to content
EN

Reinforcement learning

Rewards, policies, environments and results, for language models and beyond.

79 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Reinforcement learning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. SA-RAG Uses Spreading Activation for Knowledge-Graph Retrieval

    SA-RAG applies spreading activation to knowledge-graph-based RAG: activation spreads from query-matched entities through weighted connections, helping retrieve documents linked to highly activated entities.

    It offers an alternative to LLM-guided iterative retrieval for finding connected evidence in multi-hop questions.

  2. Stabilizing Reinforcement Learning with LLMs

    A Qwen team paper examines instability in reinforcement learning for large language models. The post highlights training–inference and old–new policy gaps, and mentions importance sampling and routing replay.

    The paper may help engineers understand practices for stabilizing reinforcement learning in large language models.

  3. Olmo 3’s SFT, DPO, and RLVR training pipeline

    Olmo 3 applies SFT on curated reasoning traces, followed by DPO preference tuning and RLVR using Group Relative Policy Optimization. The post says the combined SFT and DPO setup is a better starting point for RL than SFT alone or direct RL from the base model.

    The pipeline offers a concrete example of how data curation and preference tuning can shape reasoning-model RL training.

  4. RLoop Uses Iterative Policy Initialization for RL

    RLoop alternates reinforcement-learning exploration with RFT consolidation from filtered expert trajectories. The post reports gains over RL on Qwen-2.5-7B-Math with DAPO-17K, including about 9% higher Avg@32 and over 15% higher Pass@32.

    The approach offers a policy-initialization strategy for reducing overfitting and forgetting in RLVR pipelines.

  5. Why RL Training Can Be More Fragile Than SFT

    The post argues that SFT and RL can share a loss form but differ in practice because RL depends on noisy feedback and more complex rollout, log-probability, and reward-shaping infrastructure. It says these factors can destabilize training.

    It highlights data-quality and infrastructure risks engineers should monitor in RL training systems.

  6. RLAC uses an adversarial critic for free-form generation

    RLAC trains a generator and adversarial critic with an external validator, checking the most likely failure rather than everything. The post reports higher factuality with 5.7× fewer calls and better code scores using 9% of the data and 97.5% fewer test cases.

    The approach may reduce validation costs while improving factuality and code scores.

  7. Supervised Reinforcement Learning for Step-Wise Reasoning

    The post introduces a framework called Supervised Reinforcement Learning, using expert trajectories to teach small language models to reason through difficult problems. The linked paper is titled “Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning.”

    The paper may be relevant to engineers exploring trajectory-based training methods for language-model reasoning.

  8. Scaling Latent Reasoning with Looped Language Models

    The paper scales looped language models to 2.6 billion parameters and trains them on more than 7 trillion tokens. The author reports performance on par with state-of-the-art language models two to three times larger.

    It offers engineers a result on scaling language models through a looped architecture.

  9. Socratic-Zero co-evolves agents to generate reasoning data

    Socratic-Zero uses Teacher, Solver, and Generator agents to create reasoning data from 100 seed questions. The post reports that Socratic-Solver-8B improves by 20.2 points over prior data-synthesis methods across seven math reasoning benchmarks.

    The paper describes an approach to generating reasoning curricula without human-labeled data.

  10. RE-Searcher uses goals and self-reflection for robust search

    RE-Searcher is a goal-driven, self-reflective search agent for language models. It sets a search goal and checks whether retrieved evidence meets it, aiming to resist misleading cues in noisy search environments.

    Its goal-setting and evidence checks offer an approach to making search agents more resilient to noisy results.

  11. DeepMind studies autonomous discovery of reinforcement learning algorithms

    A Nature paper titled “Discovering state-of-the-art reinforcement learning algorithms” reports that AI can autonomously discover better RL algorithms. The study was led by David Silver.

    Algorithm discovery could offer engineers a way to develop new reinforcement learning methods.

  12. EditScore uses reward models for RL in image editing

    EditScore is a family of 7B–72B reward models for evaluating and guiding complex image edits, built on EditReward-Bench. The post says it enables RL training and outperforms GPT-5 on the benchmark.

    Domain-specific reward models may help engineers train systems for complex image-editing tasks.

  13. KL-Regularized RL and Mode Collapse

    The paper studies reverse- and forward-KL regularization in reinforcement learning. It shows mathematically and empirically that the usual mode-seeking versus mass-covering intuition does not necessarily transfer to RL.

    Engineers using KL regularization in language-model RL may need to reassess assumptions about diversity.

  14. Meta-learning discovers reinforcement learning update rules

    The paper describes a method that uses meta-learning across agents and environments to discover RL update rules. Its DiscoRL rule reports state-of-the-art benchmark performance, including on Atari and ProcGen.

    Engineers can see how learned update rules may transfer across environments without being tuned for each one.

  15. BAPO stabilizes off-policy reinforcement learning for LLMs

    BAPO is a method for stabilizing off-policy reinforcement learning for large language models, using balanced policy optimization with adaptive clipping. The work addresses settings including partial rollout and experience reuse.

    Engineers working on LLM reinforcement learning can review the method and its code for off-policy training settings.

  16. Ring-1T combines trillion-parameter MoE with three-stage RL training

    The post describes Ring-1T, a 1T-parameter MoE reasoning model with about 50B active parameters per token. Its training uses long-CoT SFT, verifiable-reward reasoning RL, and general RLHF, alongside IcePop, C3PO++, and ASystem.

    The training phases and systems components offer concrete approaches to scaling RL for large reasoning models.

  17. A Study of RL Design Choices for Agentic Reasoning

    The paper systematically investigates reinforcement learning for agentic reasoning in LLMs across data, algorithms, and reasoning modes. The preview highlights a comparison of stitched synthetic trajectories with real end-to-end tool-use trajectories.

    It examines design choices that can inform how engineers train LLM agents to reason and use tools.

  18. LaSeR uses last-token self-rewarding for LLM reinforcement learning

    Tencent introduces LaSeR, an algorithm that aligns last-token scores with true rewards for LLM reasoning and self-rewarding. The post says it uses one extra token inference.

    The approach may be relevant to engineers exploring lower-cost reward optimization for language models.

  19. SkyRL tx v0.0.2 supports training multi-LoRA models

    SkyRL tx is an open backend for the Thinky Machines Tinker API. Version v0.0.2 supports training multi-LoRA models.

    Engineers can explore an open backend for training multi-LoRA models through the Tinker API.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor