Skip to content
EN

Reinforcement learning

Rewards, policies, environments and results, for language models and beyond.

79 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Reinforcement learning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. LENS jointly trains reasoning and image segmentation with RL

    LENS is a reinforcement-learning framework for text-prompted image segmentation that jointly optimizes reasoning and segmentation. The paper says supervised fine-tuning often omits explicit chain-of-thought at test time, limiting generalization to unseen prompts and domains.

    Engineers working on vision systems can assess an approach that trains reasoning and segmentation together.

  2. Unsloth guide to reinforcement learning with GRPO

    Unsloth documentation presents a beginner-to-advanced reinforcement learning guide, including how to train a DeepSeek-R1 reasoning model with GRPO.

    Engineers can use it as a guide to applying GRPO to reasoning-model training with Unsloth.

  3. Xiaomi MiMo releases RL environments and training code

    A post points to Xiaomi MiMo’s GitHub repository for RL environments and training code, plus a MiMo-V2.6 RL dataset on Hugging Face.

    The repository and dataset may help engineers explore or build RL workflows for language models.

  4. SmolDataEnvs: 5,000 verifiable RL tasks

    SmolDataEnvs is an open-source collection of 5,000 verifiable reinforcement-learning environment tasks for hill-climbing small models in code and data science. The release includes environments, evaluations, and training.

    Engineers can use the tasks and accompanying evaluations to train and assess small models.

  5. DeepSeek Elastic Compute reports sandbox capacity at production scale

    A post quotes a DeepSeek Elastic Compute paper: one production-scale DSec unit spans about 160 nodes and supports over 380,000 concurrent sandboxes, with more than 5,000 creations per second.

    The reported figures offer a concrete reference point for large-scale sandbox infrastructure.

  6. CodeMidas builds coding RL environments from source code

    CodeMidas presents an agentic pipeline that turns implemented functionality in open-source codebases into executable RL environments. It targets diverse tasks with reliable verifiers.

    Engineers building coding agents can use source code as a scalable source of RL tasks.

  7. CodeMidas Generates Code-Based RL Training Tasks

    Xiaomi’s CodeMidas uses agents to extract executable tasks from open-source codebases rather than relying on issues or commits. It generated 5,545 tasks across 23 languages.

    The approach offers engineers a way to build RL training environments directly from source code.

  8. MiMo Pro improves its DeepSWE score after RL runs

    The post says MiMo Pro's DeepSWE score rose from 58.41 to 72.57 after RL runs. It compares the result with a score of 74 attributed to Astra, Gemini 3.8 Flash, and Opus 5.

    The reported scores offer a benchmark comparison for evaluating RL runs on DeepSWE.

  9. Self-Rewarding Language Models use LLM-as-a-Judge feedback

    The paper studies using a language model itself, prompted as an LLM-as-a-Judge, to provide feedback during training. It contrasts this approach with human-preference reward models that remain frozen during training.

    Engineers can assess an approach to generating training feedback without relying solely on separate, frozen reward models.

  10. Never Give Up shifts RL compute toward harder LLM problems

    The paper reports that RL improves LLM performance more on problems the model already solves easily than on harder problems. Its adaptive “Never Give Up” method samples until success to direct compute toward harder problems.

    Engineers can assess an adaptive sampling strategy for allocating RL training compute across problems of different difficulty.

  11. RULER uses LLM-ranked trajectories for GRPO rewards

    The post describes RULER in OpenPipe ART: engineers specify evaluation criteria in plain English, an LLM ranks agent trajectories, and ART turns the comparisons into rewards for GRPO training. It reports using the setup with a Qwen3 1.4B agent playing 2048.

    It offers an approach to evaluating multi-step agent behavior without hand-coding a set of scoring rules.

  12. Visual introduction to contraction operators in RL

    An article in an intuitive RL series about contraction operators. The author describes a visual-first approach that aims to explain ideas with precision and few words.

    It offers engineers a visual resource for learning a foundational RL concept.

  13. Writeup on Policy Gradient Derivations and Variants

    The author describes an article covering foundational reinforcement learning papers, deriving the policy gradient algorithm, and discussing its variants.

    It may help engineers review the derivation behind policy gradient methods and how they have evolved.

  14. Sampling methods for optimizing objectives without RL

    The post introduces a series on methods beyond reinforcement learning for optimizing evaluable objectives. Part I surveys sampling methods.

    It gives engineers an overview of alternative approaches to objective optimization.

  15. Practical resources on GRPO and OPD

    The post shares two resources for practical RL-style post-training beyond SFT. One covers techniques for making vanilla GRPO work at scale; the other is about OPD.

    The GRPO guide may help engineers address practical challenges in scaling RL training.

  16. A Guide to Reinforcement Learning for LLMs

    A guide covering RL fundamentals, policy gradients, REINFORCE, and LLM-specific formulations such as token-level MDPs, completion-level bandits, and outcome versus process rewards.

    It connects core RL concepts to methods and setups used to train language models.

  17. SQL-based RL replay buffer with weighted top-k sampling

    Replayhouse uses a table as an RL replay buffer and implements Efraimidis-Spirakis weighted top-k sampling in SQL, allowing draws to be filtered with WHERE. The post reports sampling 8,192 rows from 50 million in about 1.1 seconds on a laptop.

    A table-based buffer with SQL filtering may simplify replay-buffer sampling workflows.

  18. WANDR: An RL Environment for Evaluating Agentic Search

    Perplexity says it is open-sourcing an evaluation and RL environment for agentic search, synthesized from production traces with weak human supervision. The company says it uses these environments internally to train models.

    Engineers can examine an environment built from production traces for evaluating and training agentic search systems.

  19. Policy Gradients Part 1: The REINFORCE Estimator

    The first post in a series derives unbiased policy gradients with the REINFORCE estimator without differentiating through the environment, and analyzes variance scaling with trajectory length.

    Useful for understanding a foundational policy-gradient method and its variance properties.

  20. A proposed account of how RL builds on SFT

    The post describes a paper’s hypothesis that SFT solutions contain tangled useful components, while reward-guided variation during RL helps separate them into reusable skills and routing rules for new problems.

    The hypothesis offers a way to think about how RL may improve generalization beyond SFT examples.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor